Is Copyleaks' AI detector accurate?
Copyleaks came from plagiarism detection, and the habit of reading an AI score the way you read a similarity score is the mistake worth avoiding.
Copyleaks is a plagiarism-detection company that added AI detection. Both scores can appear on the same report, and they work nothing alike.
Two scores, two completely different things
Similarity matches your text against a database. It is a lookup, so every result comes with a source you can open and check. A high similarity score points at something specific.
AI detection matches nothing. It is a classifier's probability that your writing has the statistical properties of generated text. Nothing was found, so there is nothing to click through to. It comes with an error rate rather than a citation.
Reading the second like the first is how a probability turns into a certainty in somebody's head. See how Turnitin AI detection works, which covers the same confusion in the tool most people meet first.
The number that decides whether a detector is any good
Accuracy blends two errors: missing generated text, and flagging text a person wrote. Only the second one lands on somebody, so the figure to ask for is the wrong-flag rate on human writing.
Turnitin publishes 4%. Ours is 0.4%, which is 99.6% specificity, measured across 15,900 documents of real academic writing.

When any vendor advertises a percentage, check which of the two claims it is. A blended accuracy number can be raised by getting better at either error, and only one of them is your problem.
What moves a score, on any tool
The properties are the same everywhere because they belong to the text rather than the vendor.
Even sentence rhythm. The strongest signal. Human academic writing in our corpus varies sentence length with a standard deviation of 8.7 words, generated text with 6.4.

Predictable word choice. Formal and technical registers are predictable by construction.
Heavy editing. Smoothing removes variation, so a well-revised draft scores higher than a rough one.
Writing in a second language. Repeatedly measured at elevated rates in published evaluations. See why non-native English writing gets flagged.
Length changes the answer more than the brand does
Turnitin's own figures make the point: 4% wrong on sentences against under 1% on documents. Four times worse at the smaller unit, same tool, same day.
So a sentence-level highlight from any detector carries the thinnest evidence on the report, and a verdict on a full document carries the most. Run the whole thing. See why detectors are unreliable on short text.
A second measurement
Ours is free and unlimited with no account, takes .pdf, .docx and .tex, has no word cap, scores every paragraph separately, and never stores your text.