unslop

Is Grammarly's AI detector accurate?

Grammarly both edits your writing and scores it. Those two functions pull in opposite directions, which is the most useful thing to understand about reading a number from it.

3 min read

Grammarly is a writing assistant first, and its AI detection is a feature within that. That matters because the assistant's suggestions change the properties the detector measures.

The accuracy question, and the one underneath it

A detector makes two errors. It misses generated text, or it flags something a person wrote. An accuracy figure averages them, so it can be raised by improving either one, and only the second error ever lands on somebody.

The number to ask for is how often it is wrong about human writing. Turnitin publishes 4%. Ours is 0.4%, which is 99.6% specificity, measured on 15,900 documents of real academic writing.

Bar chart comparing wrong flags on genuine human academic writing. unslop 0.4 percent of documents against Turnitin 4 percent, a ten-fold difference.
Bar chart comparing wrong flags on genuine human academic writing. unslop 0.4 percent of documents against Turnitin 4 percent, a ten-fold difference.

Whatever headline figure any tool advertises, check whether it is a blended accuracy percentage or a wrong-flag rate. They are different claims and only one describes your own exposure.

The tension specific to Grammarly

Accepting editing suggestions makes prose more regular. Sentences get tightened toward a comfortable length, constructions get standardised, hedges get cleaned up. That is the product working.

Regularity is what detectors measure. Our corpus puts human academic writing at a within-document sentence-length standard deviation of 8.7 words and generated text at 6.4, and every editing pass moves a document toward the lower figure.

Two overlapping histograms of within-document sentence length variation. Human writing centres on a standard deviation of 8.7 words, generated text on 6.4.
Two overlapping histograms of within-document sentence length variation. Human writing centres on a standard deviation of 8.7 words, generated text on 6.4.

So heavily edited human writing scores higher than a rough draft of the same piece. This is the most common way people meet a wrong flag, and it is a separate question from how good any one detector is. See does Grammarly get flagged as AI and why good writing gets flagged more.

What no detector's accuracy figure covers

Short text. Every detector is far more reliable per document than per sentence. Turnitin's own numbers: 4% wrong on sentences against under 1% on documents. See why detectors are unreliable on short text.

Register. Formal, conventional and technical prose scores high because it is regular, not because of how it was produced. See why AI detectors flag human writing.

Who wrote it. Published evaluations keep finding elevated wrong-flag rates for non-native English writers. A published average is an average over whatever mix the vendor tested. See why non-native English writing gets flagged.

Getting a second number

Detectors disagree for structural reasons, and wide disagreement means the text sits near a boundary rather than that one tool is broken. See why AI detectors disagree.

Ours is free and unlimited with no account, takes .pdf, .docx and .tex, has no word cap, scores every paragraph separately, and never stores your text.

Check any text with our detector, free and unlimited →