unslop

Is QuillBot's AI detector accurate?

QuillBot sells a paraphraser and gives away a detector. Worth knowing what that combination means before you read a number from either one.

3 min read

QuillBot is a paraphrasing tool first. The AI detector is a free feature alongside it, which is a common arrangement and it shapes what the free check is for.

The question "accurate" does not answer

A detector makes two different errors and they are not remotely equivalent.

It misses generated text, which is invisible and costs nobody anything that afternoon. Or it flags text a person wrote, which arrives attached to their work with their name on it.

An accuracy figure averages those together, so it can be raised by getting better at either one. A tool that is very willing to flag scores well on accuracy against generated samples and produces a large pile of wrong flags on real writing, and the single headline number hides that completely.

So the question worth asking of QuillBot, or of anything else, is how often is it wrong about human writing, and what does it catch at that rate.

Turnitin publishes 4%. Ours is 0.4%, which is 99.6% specificity, measured across 15,900 documents of real academic writing.

Bar chart comparing wrong flags on genuine human academic writing. unslop 0.4 percent of documents against Turnitin 4 percent, a ten-fold difference.
Bar chart comparing wrong flags on genuine human academic writing. unslop 0.4 percent of documents against Turnitin 4 percent, a ten-fold difference.

When you read any vendor's figure, check which of the two it is. A blended accuracy percentage and a wrong-flag rate are different claims and only one of them tells you your own risk.

The part specific to QuillBot

The same company sells the paraphraser and the detector, and the two work against each other in a way worth understanding.

Paraphrasing smooths. It applies one transformation across a document, replacing whatever unevenness the writing had with the tool's own consistency. Detectors read exactly that evenness: our corpus puts human academic writing at a within-document sentence-length standard deviation of 8.7 words and generated text at 6.4.

Two overlapping histograms of within-document sentence length variation. Human writing centres on a standard deviation of 8.7 words, generated text on 6.4.
Two overlapping histograms of within-document sentence length variation. Human writing centres on a standard deviation of 8.7 words, generated text on 6.4.

So running your own writing through a paraphraser can raise the score a detector gives it. See does QuillBot get flagged and paraphrasing to avoid AI detection.

What any free detector is reliable about

Three limits apply to all of them, QuillBot included, and they matter more than the brand.

Length. Every detector is far more reliable on a document than on a paragraph, and on a paragraph than on a sentence. Turnitin's own figures show the size of it: 4% wrong on sentences against under 1% on documents. A word cap that forces you to check an essay in fragments degrades the answer rather than merely inconveniencing you. See why detectors are unreliable on short text.

Register. Formal, conventional, heavily edited prose scores high whoever wrote it. See why AI detectors flag human writing.

Who wrote it. Published evaluations keep finding elevated wrong-flag rates for non-native English writers. See why non-native English writing gets flagged.

Running a second check

Detectors disagree for structural reasons, so a second measurement is worth having, and wide disagreement tells you the text sits near a boundary. See why AI detectors disagree.

Ours is free and unlimited with no account, takes .pdf, .docx and .tex, has no word cap, scores every paragraph separately rather than handing back one document number, and never stores your text.

Check any text with our detector, free and unlimited →