What AI percentage is too high?
There is no threshold that means anything on its own. What the number is made of matters far more than what it is, and the vendors publish enough to show why.
The question everybody asks first, and the honest shape of the answer is that a percentage is not a scale of guilt. It is the share of qualifying prose a classifier scored above its own internal line, and two documents with the same number can be completely different situations.
Here is what actually changes how a number should be read.
Below 20% carries less than it looks like
Turnitin states directly that documents scoring under 20% show a higher incidence of wrong flags.
The reason is arithmetic. A low percentage is built from a small amount of flagged text, and a small amount of text is the thinnest evidence a classifier can work from. Everything gets more reliable with length. See why AI detectors are unreliable on short text.
So a 15% is not "a bit of AI". It is a small flagged region, which is the regime where the tool is least sure of itself.
Spread matters more than size
A 30% concentrated in one section and a 30% smeared evenly across the document are different findings with the same number.
Concentrated points at that section, and the first question is what kind of section it is. A methods description or a procedure is formulaic by design, and formulaic reads as machine-made.
Even points at a property of the writing as a whole, which for a careful, formal or second-language writer is frequently just how they write. See why AI detectors flag human writing.
A document percentage cannot separate those two, which is the argument for reading per-paragraph scores instead of one figure.
Flagged regions are wider than what caused them
Turnitin's own analysis found 54% of falsely flagged sentences sit directly beside a sentence scored as AI written, and 26% two sentences away.
Four in five wrong flags cluster around a genuine one. So in any document with mixed origins, the highlighted region overstates the assisted part, and the boundaries should be treated as soft. See documents that are part human, part AI.
The number the percentage should always be read against
The rate at which the tool is wrong about human writing. Turnitin publishes 4%.

Ours is 0.4%, which is 99.6% specificity, measured on 15,900 documents of real academic writing. Ten times fewer wrong flags on the error that arrives attached to somebody's work.
A percentage without a wrong-flag rate beside it is an unlabelled number.
How to read your own
- How much prose was scored. Tables, equations, references and short assignments leave far
less qualifying text than the page count suggests
- Where it sits. Concentrated or spread
- What kind of section carries it. Formulaic sections score high for structural reasons
- What a second tool says. Wide disagreement means the text is near a boundary rather than
that either tool is broken. See why AI detectors disagree
Ours is free and unlimited with no account, scores every paragraph separately rather than handing back one figure, and never stores your text.