unslop

AI detection in academic papers

Academic writing scores higher than almost any other kind, whoever wrote it. Detectors measure regularity, and papers are regular on purpose.

4 min read

Academic prose is the hardest register we know of for AI detection, and not because of anything to do with AI. Detectors measure how regular text is. Academic writing is deliberately, professionally regular.

The overlap is structural, not accidental

Detectors read predictability, rhythm and word distribution, covered in how AI detectors work. Text that is predictable, evenly paced and formally worded scores as machine written.

That is also a fair description of a competent methods section.

Field conventions constrain the vocabulary. Every discipline has an agreed way of saying things, and departing from it is a fault rather than a flourish. Constrained vocabulary is predictable vocabulary.

The structure is prescribed. Introduction, methods, results, discussion. Structured abstracts go further and prescribe the sentences themselves.

Hedging is expected. Claims get qualified, and qualified sentences run longer and more uniform than plain assertions.

Everything gets smoothed. Supervisors, co-authors, reviewers, copy editors. Each pass removes irregularity, and irregularity is precisely what a detector reads as human.

Our own corpus shows the size of it. Across pre-2020 arXiv, the standard deviation of sentence length within a document has a median of 8.7 words. Generated text sits at 6.4. Academic writing already lives closer to the generated end than casual writing does, which leaves the detector a narrower gap to work with before anyone touches a language model.

Some sections are far riskier than others

This is not evenly spread through a paper.

Methods is the worst offender in most papers. Formulaic by design, often reusing phrasing from the group's earlier work, describing procedures in a deliberately narrow vocabulary.

Abstracts, structured ones especially. Short and templated, and short text is where every detector is least reliable.

Related work falls into a repeated shape fast. One citation per sentence, same rhythm throughout.

Introduction and discussion are usually the most varied and the safest.

Which is why a single document percentage is a poor instrument for a paper. A paper flagged mostly in its methods is displaying the property every methods section has. A paper flagged evenly throughout is telling you something about the writing as a whole. Only per-passage output separates those two, and they call for completely different reads.

Length helps you, formatting hurts you

More text means more signal, so a full paper is a far better input than an abstract. Every detector is more reliable per document than per sentence, by a wide margin.

Formatting cuts the other way. Equations, tables, reference lists and LaTeX commands are not prose, and a detector that scores them anyway is measuring your preamble. See why LaTeX breaks AI detectors.

Non-native English

A large share of published research is written by people whose first language is not English, and detectors flag that writing at higher rates. Same cause again: learned English gets applied more consistently and reaches for the common construction more often.

For a journal or a department this is worth sitting with, because it means a detection score correlates with something other than how the paper was produced. See why non-native English writing gets flagged.

What we built for this

We built the detector on academic writing specifically, because the register is the whole problem. The human side is pre-2020 arXiv, the largest clean source of definitely human academic prose in existence. The generated side is matched to it: the same papers, rewritten, so the two differ in how they were produced rather than in what they are about.

On held-out data from that corpus, our public setting flags 0.8% of genuine academic writing and catches 88.3% of generated academic text. On the strict setting the wrong-flag rate drops to 0.4%, which is 99.6% specificity on real academic prose.

Bar chart comparing wrong flags on genuine human academic writing. unslop 0.4 percent of documents against Turnitin 4 percent, a ten-fold difference.
Bar chart comparing wrong flags on genuine human academic writing. unslop 0.4 percent of documents against Turnitin 4 percent, a ten-fold difference.

We also ran it across 12,750 full papers to see what the trend looks like, which is written up in how much of arXiv reads as AI written. The method is in how the unslop AI detector works.

Run a paper through it. Free, no account.

Check any text with our detector, free and unlimited →