Why AI detectors are unreliable on short text
Every detector is several times worse per sentence than per document. It is not a flaw in any particular tool, it is how much evidence a sentence contains.
Paste a paragraph into a detector and you get a confident-looking number. Paste two sentences and you get an equally confident-looking number, which is worth considerably less. Nothing in the interface tells you that, so it is worth knowing where the line is.
The size of the effect, from the vendors
Turnitin publishes both figures, which makes the comparison easy: under 1% of documents, and around 4% of sentences.
Same tool, same model, same afternoon. Four times the error rate, purely because a sentence carries less evidence than a document.
Turnitin also states that documents scoring below 20% show a higher incidence of false positives, which is the same effect again: a low score is built from a small amount of flagged text, and small amounts of text are where classifiers are least reliable.
Why more text helps this much
Detectors read statistical properties, and statistics need samples.
Sentence length variation is the strongest single signal in most detectors, and it is a property of a sequence of sentences. On one sentence there is no variation to measure. On three there is barely any. It takes a paragraph before the number means anything, and a document before it is stable.
Predictability averages out. A single unusual sentence can be genuinely surprising or genuinely flat by chance. Over a thousand words the noise cancels.
Word distribution needs words. Whether a document over-selects certain phrasing is a question about frequencies, and frequencies on forty words are close to meaningless.
Our own detector ranks text from roughly 75 words upward, but the threshold it uses is calibrated on whole documents. That is why we score per paragraph and report the document number as the authoritative one, rather than putting a verdict on individual sentences.
Where this bites in practice
Sentence-level highlighting. Many tools mark individual sentences as AI written. Those marks carry the thinnest evidence anywhere on the report, and they are also the most visually convincing, which is an unfortunate combination.
Abstracts. A structured abstract is short and templated, which is the worst pairing available. A score on an abstract alone should be treated as much softer than a score on the paper.
Short assignments. A 300-word response gives a detector far less to work with than a 3,000 word essay, and the error rate is correspondingly higher.
Free tools with word caps. A cap that forces you to check a document in fragments does not just inconvenience you, it degrades the result, because you are asking for several unreliable small answers instead of one reliable large one. See free AI detectors.
Reading a short-text score
- Treat anything under a few hundred words as indicative rather than informative
- Prefer a document score over the sum of sentence scores
- Be especially cautious with sentence-level highlighting on a mixed document, where Turnitin's
own analysis found 54% of falsely flagged sentences sit directly beside a genuine detection
- Run the whole thing rather than an extract, wherever the tool allows it
Ours takes full documents with no cap and scores each paragraph separately while reporting the document figure as the one that carries the calibration. Free, no account.