Documents that are part human, part AI
The most common real case is the one a single document percentage describes worst, and there is a published number showing exactly how far it drifts.
Almost nobody generates a whole document and submits it. The realistic case is a draft somebody wrote, with a paragraph reworked by a model, or an outline that was generated and then filled in, or a section pasted in and edited afterwards.
Most detectors are built and evaluated on the clean case, whole documents on one side or the other, and hand back a single percentage. On a mixed document that one number describes nothing in particular, and there is a published figure showing how far it drifts.
The clustering effect
Turnitin's own analysis of falsely flagged sentences found that 54% sit directly beside a sentence scored as AI written, and another 26% two sentences away.
Four in five wrong flags land next to a genuine detection.
That is not a coincidence, it is the mechanism showing through. Detectors segment text and score regions, and the boundary between a generated passage and a human one is blurry to a classifier reading a sliding window. Signal bleeds outward. The human sentences immediately around a generated passage get pulled up with it.
The practical consequence: in a mixed document, the human writing around an assisted passage is the writing most likely to be flagged wrongly. Somebody who used a model for one paragraph of ten will not see one paragraph flagged. They will see a smear.
Why a document percentage is the wrong output here
A single number for a mixed document is an average over two different populations, and the average describes neither.
A 30% score can mean 30% of the text is strongly flagged and the rest is clean, or the whole document sits at moderate suspicion throughout. Those are completely different situations and the document number cannot separate them.
Turnitin also states that documents scoring below 20% carry a higher incidence of false positives, which is the same problem again: a low percentage is built from a small flagged region, and a small region is the thinnest evidence a report can rest on.
What to look at instead
Per-passage scores, and where they sit. Concentration in one section is a different finding from an even spread. Concentrated points at that section. Even points at a property of the whole writing, which for careful, formal or non-native writing is often just how the person writes. See why AI detectors flag human writing.
Whether the flagged region has clean edges. Given the clustering above, the flagged region in a mixed document will usually be wider than the assisted passage that caused it. Treat the boundaries as soft.
How long the flagged region is. A flag over 60 words is worth much less than a flag over 600, for the ordinary reason that short text carries less evidence. See why detectors are unreliable on short text.
Editing changes the picture, in both directions
A generated passage that somebody rewrote heavily may score lower than the human paragraphs around it, because editing introduces variation.
And human text that was edited hard scores higher, because editing also smooths. Every round of revision, every co-author pass, every language service pulls a document toward regularity, and regularity is the signal. See does Grammarly get flagged as AI.
So in a document with mixed provenance and mixed editing history, passage scores track how each passage reads rather than where it came from. That is exactly why a single document percentage is the wrong output, and why we report every paragraph separately instead.
Checking a mixed document
Run the whole thing, not the section you are worried about. More text means a more reliable number, and the surrounding passages are what let you see whether a flag is concentrated or smeared.
Ours is free and unlimited with no account and scores each paragraph separately alongside the document figure, precisely because one number for a mixed document hides the thing worth seeing. The method is in how the unslop AI detector works.