What percentage of AI is acceptable?
There is no number that means anything on its own, and the vendors' own figures explain why. What a percentage is made of matters far more than its size.
The hope behind this question is that some threshold exists below which a score is fine. It does not, and the reason is worth understanding, because it changes what you look at instead.
What any institution treats as acceptable is set by its own policy, which varies enormously and which we are not the right people to describe. What we can describe is the measurement, since that part is the same everywhere.
A percentage is not a scale
The number is the share of qualifying prose a classifier scored above its own internal threshold. It is not a confidence level, not a share of your ideas, and not a quantity of anything you did.
Two documents with the same percentage can be completely different situations, which is exactly why no single cutoff works.
Below 20% carries less, not more
Turnitin states directly that documents scoring under 20% show a higher incidence of wrong flags.
The reason is arithmetic. A low percentage is built from a small flagged region, and a small region is the thinnest evidence a classifier can produce. Everything gets more reliable with length. See why detectors are unreliable on short text.
So a 15% is not "a little bit of AI". It is a small flagged region, which is the regime where the tool is least sure of itself.
Spread matters more than size
A 30% concentrated in one section and a 30% smeared evenly are different findings with the same number.
Concentrated points at that section, and the first question is what kind of section it is. Methods, procedures and literature reviews are formulaic by design and score high for structural reasons.
Even points at a property of the writing as a whole, which for a careful, formal or second-language writer is frequently just how they write. See why AI detectors flag human writing.
A document percentage cannot separate those. Per-paragraph scores can, which is the argument for reading them instead.
The flagged region overstates its cause
Turnitin's own analysis found 54% of falsely flagged sentences sit directly beside a sentence scored as AI written, and 26% two sentences away.
Four in five wrong flags cluster around a genuine one, so in any mixed document the highlighted region is wider than whatever produced it and the boundaries are soft. See documents that are part human, part AI.
The number every percentage should be read against
How often the tool is wrong about human writing. Turnitin publishes 4%.

Ours is 0.4%, which is 99.6% specificity, measured on 15,900 documents of real academic writing. A percentage without a wrong-flag rate beside it is an unlabelled number.
A more useful question than "what is acceptable"
What does my own writing normally score? Run three or four pieces you wrote before any of this was a question. If they all land in the same range, that range is your writing, and a new document inside it is not telling you anything about that document.
Half an hour, once, and every score afterwards becomes a comparison rather than a verdict. See find your own baseline score.
Ours is free and unlimited with no account, has no word cap, scores every paragraph separately, and never stores your text.