unslop

Can teachers tell if you used ChatGPT?

Two separate questions get mixed together here: what the software reports, and what a person notices. They fail in different places.

3 min read

The software runs automatically and produces a number. The person reads your work and forms an impression. Neither of them sees how the document was made, and they go wrong in different directions, which is worth separating.

What the software reports

A probability that your writing has properties characteristic of generated text. Predictable word choice, and sentences that stay close to the same length.

It is wrong about human writing at a rate Turnitin publishes: 4%.

Bar chart comparing wrong flags on genuine human academic writing. unslop 0.4 percent of documents against Turnitin 4 percent, a ten-fold difference.
Bar chart comparing wrong flags on genuine human academic writing. unslop 0.4 percent of documents against Turnitin 4 percent, a ten-fold difference.

Ours is 0.4%, which is 99.6% specificity, measured on 15,900 documents of real academic writing.

What a person notices

Different signals entirely, and often more reliable ones.

A gap against your earlier work. Somebody who has read three of your assignments has a model of how you write. A sudden shift in register is visible without any tool.

Confident wrongness. Generated prose states things smoothly whether or not they are correct. A fluent paragraph containing a claim nobody in the field makes reads oddly to somebody who knows the field.

Citations that do not resolve. The most common tell there is, and the easiest to check.

Generic specificity. "Various studies have shown" where a person would name two. Detail is the thing generated text reliably thins out.

Nothing from the seminar. Work that engages with none of what was discussed in the room.

Notice that none of these are what the classifier measures. A document can read completely natural to a person and score high, or read strangely to a person and score clean.

Where each one goes wrong

The software flags careful writing. Formal register is predictable, even sentence rhythm looks machine-made, heavy editing removes variation, and writing in a second language shows elevated rates in published evaluations. Our corpus puts human sentence-length variation at a standard deviation of 8.7 words against 6.4 for generated text, and every editing pass pulls a document toward the lower figure. See why AI detectors flag human writing.

The person goes wrong on unfamiliarity. A student who has improved, changed register for a new assignment type, or worked with a writing centre reads differently from last term for ordinary reasons.

What this means if you are submitting

The two failure modes stack. Careful, formal, well-edited writing is the profile most likely to score high, and it is also the profile least likely to look suspicious to a reader, so the number arrives without anything to explain it.

That is the case worth pre-empting, and the only step available is to look at your own document the way the classifier does before anybody else does.

Ours is free and unlimited with no account, takes .pdf and .docx, has no word cap, scores every paragraph separately so you can see which ones read flat, and never stores your text.

Check any text with our detector, free and unlimited →