Why non-native English writing gets flagged more often
The measured effect, the linguistic properties behind it, and why it follows directly from what detectors measure rather than from any rule about who wrote the text.
Detectors flag writing by non-native English speakers at higher rates than writing by native speakers. This has been found repeatedly in published evaluations, and it is not a quirk of one tool.
The cause is mechanical, and once you see what detectors measure it is close to inevitable.
What detectors actually measure
Three signals carry most of the weight, covered in full in how AI detectors work:
- predictability, how expected each word is given the previous ones
- rhythm, how much sentence length and structure vary
- vocabulary distribution, which words and phrases appear and how often
Text that is predictable, regular and formal scores as machine written. That description is about properties of text. It says nothing about who produced it.
Why learned English lands in that profile
Four properties of writing in a second language push in exactly that direction.
Narrower vocabulary range. A learner has a large but more constrained working vocabulary, and reaches for the common word more often. Common words are predictable words.
More regular sentence construction. Learned grammar tends to be applied consistently. Native writers break their own patterns constantly, often ungrammatically, and that irregularity is a large part of what detectors read as human.
Fewer idioms and less colloquial variation. Idiom is where language is least predictable, and it is the last thing acquired.
Careful, formal register. Writing in a second language for an academic or professional audience usually means writing carefully. Careful writing is smoother, and smoothness reads as generated.
Every one of these is also a description of good, clear second-language writing. The signal detectors use to identify machine text overlaps with the signal of a competent non-native writer, and no threshold separates them cleanly because they are not separable properties.
The same effect elsewhere
This is one instance of a general pattern. Any writing that is regular and formal sits closer to the line:
- technical and scientific prose, especially methods sections
- text written to a template, such as structured abstracts or grant sections
- heavily edited writing, since editing smooths variation
- translated text, for the same reasons as learned English
We cover the general case in why AI detectors flag human writing.
What this means for a score
A detection score on non-native English writing carries a higher prior probability of being wrong than the same score on native writing. The tool's published false positive rate is an average over its test set, and if that set was mostly native-speaker text, the rate for non-native writing is higher than the number advertised.
That is worth knowing when weighing a single score, and it is a reason to ask what a detector was measured on rather than only what it measured.
Our own numbers come from academic writing, much of which is by non-native English speakers, since a large share of pre-2020 arXiv is. We publish the false positive rate we measure, 0.8% at our public setting and 0.4% at a stricter one, along with where the detector is weakest, in how the unslop AI detector works. We have not published a separate breakdown by author language, and until we do, treat the overall figure as an average rather than a guarantee for any particular group.
Reducing the risk
If your writing is regularly flagged, the changes that help are the same ones that make prose read better, and none of them require changing what you are saying:
- Vary sentence length deliberately. The strongest single signal, and the easiest to
change. Follow a long sentence with a short one.
- Cut leading connectives. "Moreover", "Furthermore", "Additionally" raise the score and
rarely carry meaning.
- Keep a draft history. Version history is a record of process; finished text is not.
- Check before submitting. Ours is free and unlimited and scores each paragraph
separately, so you can see which passages carry the signal rather than only a document number.