AI detection and journal submission
Publishers screen submissions now, and the register that gets flagged is the one journals ask you to write in.
Submitting a paper means it will probably be scored, by the publisher's screening, by a reviewer running a check out of curiosity, or by an editor following up on a reviewer's remark. None of those necessarily reach you, and none of them come with the error rate attached.
Journals are screening, and the tools are the ordinary ones
Publishers have added AI screening to submission workflows, generally using the same detectors available to everyone else rather than anything special. Which means the same error rates apply, and the same weaknesses.
Disclosure requirements are now common too, and they vary considerably between publishers. What a given journal asks you to declare is a matter for that journal's own policy, and those are worth reading rather than guessing at. We are not the right people to summarise them.
The register problem
This is the part worth understanding, because it is structural rather than bad luck.
Detectors measure how regular text is. Journals ask for regular text. Field conventions constrain vocabulary, structure is prescribed, hedging is expected, and every round of co-author and copy-editor attention smooths a little more variation out.
Measured across our corpus, the standard deviation of sentence length within a document has a median of 8.7 words for pre-2020 arXiv writing against 6.4 for generated text. Academic prose already sits closer to the generated end than casual writing does, before anyone opens a language model. Full version in AI detection in academic papers.
The consequence is uncomfortable and worth stating plainly: the better you follow the conventions a journal asks for, the more machine-like your paper looks to a classifier.
Which parts of a submission carry the risk
Methods is the highest-risk section in most papers. Formulaic by design, frequently reusing phrasing from your group's earlier work, deliberately narrow vocabulary.
The structured abstract is short and templated, and short text is where every detector is least reliable.
Related work falls into a repeated shape quickly.
Cover letters are short, formal and follow a conventional structure, which is the worst combination available. They are also rarely checked, but worth knowing.
Two things that raise a score with no model involved
Language editing services. Many journals recommend them, and many authors writing in a second language use them. Editing smooths prose, and smoothing strips out exactly the variation detectors read as human. The service improves your paper and raises your score at the same time. See does Grammarly get flagged as AI.
Writing in English as a second language. Published research keeps finding elevated false positive rates here, and a large share of submitted research is written by non-native speakers. See why non-native English writing gets flagged.
Both mean a screening score correlates with something other than how a paper was produced.
Checking before you submit
Run the full manuscript rather than the abstract. More text means more signal and a more reliable number.
Score the prose, not the .tex source. Detectors that read your preamble and equation environments as text return numbers that are about your markup. See why LaTeX breaks AI detectors.
Look at which sections carry the signal rather than the document percentage. A paper flagged in its methods is showing the property every methods section has. A paper flagged evenly throughout is telling you something about the writing as a whole, and those want different responses.
Ours is free and unlimited with no account, takes .tex, .pdf and .docx, and scores each paragraph separately. We publish our false positive rate: 0.8% at the public setting, 0.4% stricter, measured on exactly this register.
We also ran it across 12,750 full arXiv papers to see the trend, in how much of arXiv reads as AI written.