unslop

ChatGPT prompts to avoid AI detection

Prompts move the score much less than the internet suggests, and the reason is structural. Here is what a prompt can and cannot reach.

3 min read

Every list of these prompts contains the same instructions: write naturally, add personality, vary your style, write like a human. They are worth understanding rather than collecting, because whether a prompt does anything at all depends on which property it touches.

Why most of them do nothing

Detectors read two things: how predictable the word choice is, and how much the sentence length varies across the document.

The second one is a product of how a model samples text, not of what you asked it to write about. A model told to "write naturally" still produces sentences clustered around a comfortable length, because that is what its sampling does. The instruction reaches the content and not the rhythm.

Our corpus puts human academic writing at a within-document sentence-length standard deviation of 8.7 words and generated text at 6.4. Instructions about tone leave that number roughly where it was.

Two overlapping histograms of within-document sentence length variation. Human writing centres on a standard deviation of 8.7 words, generated text on 6.4.
Two overlapping histograms of within-document sentence length variation. Human writing centres on a standard deviation of 8.7 words, generated text on 6.4.

The instructions that reach something real

A prompt only helps when it names a countable property rather than a mood.

"Vary sentence length. Include at least two sentences under six words and two over thirty." This is checkable and it targets the strongest signal there is. Vague versions of it do nothing.

"Do not open any sentence with Moreover, Furthermore, Additionally, or Overall." Connective openings appear in 2.4% of human sentences and 5.8% of generated ones, so removing them moves a measurable quantity.

Bar chart comparing how often sentences open with a connective such as moreover or furthermore. Human writing 2.4 percent, generated writing 5.8 percent.
Bar chart comparing how often sentences open with a connective such as moreover or furthermore. Human writing 2.4 percent, generated writing 5.8 percent.

"Name specific figures, methods and examples. Do not write 'various' or 'significant'." Generalisation is the habit that reads as generated, and it is one a prompt can genuinely suppress.

"Vary paragraph length. Do not end paragraphs by restating the opening sentence." Targets the paragraph template.

Notice the pattern: every effective instruction is a rule you could verify by counting, and every ineffective one is an adjective.

What still will not work

"Write like a human." Names no property. The model has no dial for this.

"Write at a Flesch-Kincaid grade level of 8." Readability is not what detectors measure. Text can be simple and perfectly even.

Asking for typos or grammatical errors. Correctness is not scored. This costs a reader and buys nothing.

Asking for unusual punctuation or unicode substitutions. Not measured, and trivially stripped before scoring.

The limit worth knowing

A prompt shapes text at generation time. It cannot inspect the result, so it cannot tell you whether it worked, and the same prompt produces different structural properties on different runs.

That is why the reliable workflow is generate, measure, then edit the specific passages that came back flat, rather than hunting for a prompt that solves it in one shot. The edits are mechanical and take minutes. See how to avoid AI detection.

Measuring instead of guessing

Score per paragraph rather than as one document number, because a prompt tends to affect some passages and not others, and an average hides which.

Ours is free and unlimited with no account, has no word cap so you can compare several attempts, scores every paragraph separately, and never stores your text. Wrong-flag rate on genuine human writing 0.4%, published, against Turnitin's 4%.

Check any text with our detector, free and unlimited →