unslop

How accurate is Turnitin's AI detection?

Turnitin publishes more numbers than most, and read together they say something different from the headline.

3 min read

Turnitin is the detector most people meet, usually without choosing it, and it publishes enough about its own behaviour to answer this properly rather than by impression.

The four published numbers

Wrong about human writing 4% of the time. This is the figure that matters, because it is the error that arrives attached to somebody's work.

Under 1% at the document level. Same tool, larger unit. The gap between this and the 4% is the whole story of how much text a classifier needs before it is reliable.

Below 20% is less reliable still. Turnitin states that documents scoring under 20% show a higher incidence of wrong flags, because a low percentage is built from a small flagged region.

Wrong flags cluster. Their own analysis found 54% of falsely flagged sentences sit directly beside a sentence scored as AI written, and 26% two sentences away. Four in five land next to a genuine detection rather than scattering.

What 4% means at the scale it runs

The rate is knowable. Which specific document it applies to is not, and no version of the tool will say.

A department processing 10,000 submissions a term expects a large number of documents where the indicator is wrong about somebody's own writing. That is not a malfunction, it is the published rate doing what published rates do at volume.

Bar chart comparing wrong flags on genuine human academic writing. unslop 0.4 percent of documents against Turnitin 4 percent, a ten-fold difference.
Bar chart comparing wrong flags on genuine human academic writing. unslop 0.4 percent of documents against Turnitin 4 percent, a ten-fold difference.

Ours is 0.4%, which is 99.6% specificity, measured across 15,900 documents of real academic writing. Ten times fewer wrong flags on the error that lands on a person.

What the indicator is measuring

Not a match. The AI indicator finds nothing and points at nothing, unlike the similarity score next to it on the same report, which points at a source you can open. See how Turnitin AI detection works.

It reads regularity: predictable word choice, and sentence lengths that stay close to each other. Our corpus puts human academic writing at a within-document sentence-length standard deviation of 8.7 words and generated text at 6.4.

Two overlapping histograms of within-document sentence length variation. Human writing centres on a standard deviation of 8.7 words, generated text on 6.4.
Two overlapping histograms of within-document sentence length variation. Human writing centres on a standard deviation of 8.7 words, generated text on 6.4.

Where the 4% concentrates

Wrong flags are not spread evenly across a cohort. They gather in writing that is regular for ordinary reasons.

Formal and conventional prose, which is what students are asked to produce.

Methods, procedures and literature reviews, formulaic by design.

Heavily edited work, because smoothing removes variation.

Writing by non-native English speakers, repeatedly measured at elevated rates in published evaluations. See why non-native English writing gets flagged.

So the published average understates what those groups actually meet. See AI detection false positives, by type of writing.

The practical problem with all of it

You never see the number. Turnitin is licensed to institutions, the report is generated for staff, and there is no student version. See can I use the Turnitin AI detector myself.

Which leaves one useful action, and it happens before submission: score the document yourself, per paragraph, and fix the passages that read flat.

Ours is free and unlimited with no account, takes .pdf, .docx and .tex, has no word cap, and never stores your text.

Check any text with our detector, free and unlimited →