unslop

What a Turnitin AI score means

It is not a confidence level. A 40% does not mean the report is 40% sure. It means something quite different, and the difference matters.

3 min read

The percentage on a Turnitin AI writing report is an estimate of how much of the document the classifier thinks was machine generated. It is not how confident the tool is.

So a 40% is a claim about roughly two fifths of your text. Not a claim that the report is moderately sure about all of it. People read it the second way constantly, and it changes what they think the number supports.

What the percentage counts

Turnitin cuts the submission into passages, scores each one, and reports what proportion landed above its internal threshold.

  • 0% means no passage crossed the line
  • 40% means about two fifths did
  • 100% means effectively all of it did

Text too short to score gets excluded, which is why the percentage sometimes seems to describe less material than you actually submitted.

A low score tells you less than a high one

This comes from Turnitin rather than from us: documents scoring below 20% show a higher incidence of false positives.

The scale is not symmetric in reliability. A 65% is a stronger statement than a 15% is. And the reason is mechanical rather than mysterious. A low score is built from a small amount of flagged text, and small amounts of text are where every classifier is least reliable. Turnitin's own error rate runs around 4% of sentences against under 1% of documents, for exactly this reason.

Look at whether the flags bunch together

If you check which passages got marked, they often sit next to each other rather than scattering through the document.

That is not a coincidence. Turnitin's analysis found 54% of falsely flagged human sentences sit directly beside a sentence the model scored as AI written, and another 26% two sentences away.

Which means two documents can produce the same percentage and be telling you different things. A 30% concentrated in one continuous section behaves quite differently from a 30% distributed evenly through every paragraph. The second pattern is more consistent with the classifier finding your writing generally regular than with it finding a specific generated passage.

What the score cannot separate

The indicator estimates whether text resembles model output. It has no way of knowing how the text came to look that way, and several very different situations land in the same place:

  • text a model generated
  • text you drafted and a model edited
  • text a model drafted and you edited
  • text you wrote entirely yourself in a formal, regular register
  • text you wrote in a language that is not your first

None of these are separable from the words alone. The text carries no record of how it was made. That is a limit of the method rather than a fault in this particular tool, and it applies to ours too. See why AI detectors flag human writing.

Four questions that get you most of the way

How much text was actually scored? Short submissions, or documents thick with tables, equations and references, can contain far less qualifying prose than the page count suggests.

Is it above or below 20%? Turnitin's own reliability claim changes at that line.

Are the flags clustered or spread? Clustering points at a region. Even distribution points at a property of the writing.

What does a second tool say? Detectors disagree for structural reasons, so a second score tells you whether the first reflects your text or that tool's calibration. Ours is free and unlimited, and we publish our false positive rate so you can weigh it: 0.8% at the public setting, measured on pre-LLM academic writing.

Check any text with our detector, free and unlimited →