unslop

AI detectors for teachers: what the number can carry

A detection score is a probability with a published error rate. Here is what that means for a class of thirty, and for a cohort where many students write in a second language.

3 min read

If you teach, you are probably not choosing a detector. One is already running inside the system where work gets submitted, and it produces a number whether or not anyone asked. See what AI detectors do teachers use.

What is worth your time is knowing what that number is, because the error rate is published and the arithmetic at class scale is not intuitive.

The arithmetic at your scale

Turnitin publishes a false positive rate of under 1% of documents. That sounds negligible until you multiply it.

A department processing 10,000 submissions a term expects roughly 100 documents where the indicator is wrong about human writing. At the sentence level, where Turnitin's own figure is around 4%, the count is far higher.

The rate is knowable. Which specific document it applies to is not, and no version of this tool will tell you. That is the whole difficulty in one sentence.

Wrong flags are not randomly distributed

Two patterns matter for reading a report.

They cluster. Turnitin's analysis found 54% of falsely flagged human sentences sit directly next to a sentence scored as AI written, and another 26% two sentences away. Four in five land beside a genuine detection. A submission with one assisted paragraph will show elevated scores in the human writing around it.

They concentrate in particular kinds of writing. Detectors measure regularity, and several kinds of ordinary student writing are regular: formal register, template-shaped assignments, technical or scientific prose, and anything heavily edited. The mechanism is in why AI detectors flag human writing.

The cohort effect worth knowing about

Published research keeps finding higher false positive rates for non-native English speakers. Learned English is applied more consistently and reaches for the common construction more often, and consistency is exactly what these tools read as machine-like. See why non-native English writing gets flagged.

This means a vendor's published average understates what you will see in a cohort with many second-language writers. The rate is not uniform across your students, and the direction of the error is predictable.

Reading a report rather than a number

Four things tell you more than the percentage does.

How much text was scored. Short submissions and assignments heavy with tables, equations or references may contain far less qualifying prose than the page count suggests.

Whether the score is above or below 20%. Turnitin's own reliability claim changes there, and they state that documents below 20% show a higher incidence of false positives.

Whether flags cluster or spread. Concentrated in one section points at that section. Spread evenly points at a property of the writing, which for a careful or non-native writer is frequently just how they write.

What a second tool says. Detectors disagree for structural reasons, and wide disagreement usually means the text is genuinely borderline rather than that one tool is wrong. See why AI detectors disagree.

What the indicator is

A probability from a classifier, not a match against a source document. Nothing was found, so there is nothing to click through to, which is the key difference from the similarity score sitting next to it on the same report.

Turnitin's own documentation states the indicator should not be used as the sole basis for action. What follows from a score is governed by your institution's procedures, which vary widely and which we are not the right people to describe.

A second measurement

Ours is free and unlimited with no account, takes .pdf, .docx and .tex, and scores each paragraph separately rather than returning one number for the document. We publish the false positive rate: 0.8% at the public setting and 0.4% at a stricter one, measured on pre-LLM academic writing, a corpus of 15,900 documents built for exactly this measurement.

The method is in how the unslop AI detector works.

Check any text with our detector, free and unlimited →