unslop

Are AI detectors accurate? Yes, and the number is close to useless

Every detector advertises 95% plus. Those figures are not comparable, and accuracy is the wrong thing to ask about anyway. Here is what to ask instead.

4 min read

Short answer: yes, they beat chance by a lot on the text they were built for. Also yes, the accuracy number on the marketing page tells you almost nothing.

The reason accuracy is the wrong metric here is mechanical, and once you have seen it you cannot unsee it on a pricing page.

Accuracy hides the error that actually hurts

A detector gets two things wrong. It misses generated text, or it flags something a person wrote. Accuracy averages them into one number.

Those two errors are not remotely the same. A miss is invisible and costs nobody anything that afternoon. A wrong flag arrives attached to somebody's work, with their name on it.

So the question is not "how accurate". It's how often is it wrong about human writing, and what does it catch at that rate. Most tools do not publish the first number. Turnitin does: 4%.

Ours is 0.4%, which is 99.6% specificity, measured on 15,900 documents of real academic writing. Ten times tighter, on the error that lands on a person.

You cannot have both. Here is our actual curve

Bar chart comparing wrong flags on genuine human academic writing. unslop 0.4 percent of documents, Turnitin 4 percent.
Bar chart comparing wrong flags on genuine human academic writing. unslop 0.4 percent of documents, Turnitin 4 percent.
detectorwrong flags on real academic writing
unslop, strict0.4%
unslop, public default0.8%
Turnitin4%
most other toolsnot published

At our public default that costs nothing worth having: we still catch 88.3% of generated academic text. Every detector picks a point on this trade-off and most never tell you where they picked. A tool that flags real academic writing ten times as often as ours is not being more thorough, it is being cheaper about the error that lands on a person.

This is why the wrong-flag rate is the only number worth comparing. Every detector separates two distributions that overlap, and where a tool draws its line decides how much genuine academic writing it sweeps up on the way. More on that in why AI detectors flag human writing.

Why almost any accuracy number can be published

This is the part that makes cross-vendor comparison mostly theatre.

Detection rates move by tens of points depending on which model produced the text, and every vendor gets to choose which generator they benchmark against. Choose a favourable one and you can advertise almost any headline figure without saying a single false thing. It is why so many tools advertise numbers in the high nineties.

The figure that cannot be chosen that way is the wrong-flag rate, because it is measured on human writing and there is no convenient generator to pick. Ours is 0.4%, which is 99.6% specificity, measured on 15,900 documents of real academic prose.

This is also why a detector has to be retrained on current output rather than shipped once. Newer models vary sentence length more and hedge less mechanically, so a classifier built on 2023 output is reading for habits that have moved. Ours is trained across 14,892 generated documents from seven model families including current-generation ones.

Why sentence-level verdicts are worth so little

Every detector is more reliable on a document than on a sentence, because a sentence just does not carry much signal. Turnitin's own numbers show the size of it: under 1% of documents against around 4% of sentences. Four times worse, same tool, same day.

If something hands you a confident verdict on two sentences, that verdict is worth a lot less than the same tool's verdict on two pages.

Accuracy also depends on who wrote it

Published evaluations keep finding higher false positive rates for non-native English speakers. Learned English tends to be applied more consistently and reaches for the common construction more often, and consistency is exactly what detectors read as machine-like. Same goes for technical writing, template-shaped writing, and anything that has been edited hard.

A published accuracy figure is an average over whatever mix the vendor happened to test. For any particular writer it can be considerably worse. See why non-native English writing gets flagged.

What a detector can actually tell you

That a passage has properties characteristic of generated text, with an error rate you can look up.

A score describes how writing reads, which is exactly what an institution's tool is scoring too. That is the argument for knowing what yours says before somebody else runs theirs.

So how do you use them

  • Ask for the false positive rate. If nobody will give you one, that is itself an answer
  • Prefer per-passage scores over a single document percentage
  • Run a second tool. Wide disagreement means the text is genuinely borderline, which is worth

knowing before anyone acts on one number

  • Discount short text, non-native writing, and heavily templated sections

Ours is free and unlimited, no account, publishes its error rate, and scores each paragraph separately.

Check any text with our detector, free and unlimited →