unslop

Can AI detectors be wrong about your essay?

Yes, at a rate the vendors publish. The interesting part is that the errors are not random, so you can tell in advance whether your writing sits in the exposed group.

3 min read

Detectors are classifiers. Classifiers have error rates. The error rate that matters is the one measured on human writing, because that is the error that arrives with somebody's name on it.

Turnitin publishes theirs: 4%.

Ours is 0.4%, which is 99.6% specificity, measured across 15,900 documents of real academic writing.

Bar chart comparing wrong flags on genuine human academic writing. unslop 0.4 percent of documents against Turnitin 4 percent, a ten-fold difference.
Bar chart comparing wrong flags on genuine human academic writing. unslop 0.4 percent of documents against Turnitin 4 percent, a ten-fold difference.

The errors have a shape

This is the useful part. Wrong flags are not scattered evenly across a cohort, they concentrate in writing with specific measurable properties.

Predictable phrasing. Detectors read how surprising your word choice is. Formal, conventional prose is unsurprising by construction. Anything written to a field convention lands here.

Even sentence rhythm. Human academic writing in our corpus varies sentence length with a standard deviation of 8.7 words, generated text with 6.4. Writing at a steady rhythm reads as polished to a person and machine-made to a classifier.

Two overlapping histograms of within-document sentence length variation. Human writing centres on a standard deviation of 8.7 words, generated text on 6.4.
Two overlapping histograms of within-document sentence length variation. Human writing centres on a standard deviation of 8.7 words, generated text on 6.4.

Connective openings. "Moreover", "Furthermore", "Additionally" open 2.4% of human sentences in our corpus and 5.8% of generated ones.

Heavy editing. Every revision pass smooths, and smoothing strips out variation. The more carefully edited a piece is, the closer it sits to the generated end.

English as a second language. Published evaluations keep finding elevated wrong-flag rates, because learned English applies its rules more consistently. See why non-native English writing gets flagged.

If several of those describe your writing, you are in the exposed group, and that is knowable before anybody scores you.

Where wrong flags land inside a document

Turnitin's analysis found 54% of falsely flagged sentences sit directly beside a sentence scored as AI written, and 26% two sentences away. Four in five cluster around a genuine detection rather than scattering.

They also state that documents scoring below 20% carry a higher incidence of wrong flags, because a small flagged region is the thinnest evidence a report can rest on.

Why two tools disagree on the same essay

Different training data, different thresholds, different segmentation, different handling of references and equations. Two detectors can compute nearly the same internal score and report different verdicts because of where each drew its line.

Wide disagreement is informative rather than embarrassing: it usually means the text sits near a boundary. See why AI detectors disagree.

What to do with this

Find out where your writing sits before somebody else does. That is a few minutes of work and it converts an unknown into a specific list of paragraphs.

If some of them come back flat, the fixes are mechanical and leave your argument untouched. Vary sentence length. Cut connective openings. Keep the detail you were about to generalise. See what makes text read as machine written.

Ours is free and unlimited with no account, takes .pdf, .docx and .tex, scores every paragraph separately, publishes its wrong-flag rate, and never stores your text.

Check any text with our detector, free and unlimited →