unslop

Turnitin's AI false positive rate

What Turnitin publishes about how often its AI detector is wrong about human writing, and how that compares to a detector calibrated for a low false positive rate.

2 min read

Turnitin publishes two figures for how often its AI detector is wrong about human writing: under 1% of documents, and around 4% of sentences. Both come from Turnitin.

At institutional volume those rates stop being abstract. Ten thousand submissions a term at 1% is around a hundred documents where the indicator is wrong. The rate is knowable. Which specific document it applies to is not.

Why the rate is the number that matters

Detectors advertise accuracy. Accuracy blends two errors that have nothing in common: missing generated text, and flagging writing somebody wrote themselves. Only one of those lands on a person.

A detector can post an excellent accuracy figure and still produce a large absolute number of wrong flags, because the volume is large. The rate on human writing is the number worth asking about, and it is the one most tools do not lead with.

Where wrong flags land

They are not spread evenly. Turnitin's own analysis of falsely flagged human sentences found 54% sit directly beside a sentence the model scored as AI written, and another 26% sit two sentences away.

Four in five occur next to a genuine detection. Mixed documents, where drafted and written passages alternate, are the hardest case for every detector, and they are becoming the normal case as drafting tools get built into ordinary writing software.

Some writing sits nearer the boundary whoever produced it: regular sentence construction, a formal register, a narrow vocabulary range, repeated technical structures. Research has repeatedly found higher false positive rates for non-native English writers for those reasons. See why non-native English writing gets flagged.

What we measure

We build a detector, so this is disclosure, not an impartial review.

Our measured operating points: 0.4 percent of human academic documents flagged catches 85 percent of AI academic text, rising to 1.3 percent flagged for 91.6 percent caught.
Our measured operating points: 0.4 percent of human academic documents flagged catches 85 percent of AI academic text, rising to 1.3 percent flagged for 91.6 percent caught.

We calibrate backwards from the error that matters. Rather than picking the threshold with the best headline accuracy, we fix a tolerable false positive rate and accept whatever detection rate follows. The threshold is the score only 1% of known human documents exceed.

Measured on held-out pre-LLM academic writing:

operating pointwrong about human writingAI academic text caught
stricter0.4% of documents85.0%
public default0.8% of documents88.3%

Our detection rate varies sharply by which model produced the text, from 67.7% on the hardest family to 94.0% on the easiest. We publish that alongside the false positive rate rather than behind it, in how the unslop AI detector works.

Run any text through it. Free, unlimited, no account.

Check any text with our detector, free and unlimited →