unslop

Group work and AI detection

One submission, several writers, one number. The way detectors handle mixed documents means the flag rarely lands where the cause is.

3 min read

A group submission is a document with several authors, different drafting habits, and usually a final merge pass by whoever volunteered. That is the exact document shape detectors handle least precisely, and it produces a specific failure worth understanding before you submit one.

The flag spreads past its cause

Turnitin's own analysis found 54% of falsely flagged sentences sit directly beside a sentence scored as AI written, and another 26% two sentences away.

Four in five wrong flags land next to a genuine detection. In a group document, that means if any one section draws a flag, the sections written by other people around it get pulled up with it. The highlighted region is wider than whatever caused it, and the boundaries are soft.

The people whose writing gets swept up have no way of knowing that from the number alone. See documents that are part human, part AI.

The merge pass raises the score

Somebody harmonises the voice at the end. That is good editing and it works against you here.

Merging smooths. It evens out sentence rhythm across sections written by different people, which is precisely the variation a classifier reads as human. Our corpus puts human academic writing at a within-document sentence-length standard deviation of 8.7 words against 6.4 for generated text, and a harmonising pass moves a document toward the lower figure.

Two overlapping histograms of within-document sentence length variation. Human writing centres on a standard deviation of 8.7 words, generated text on 6.4.
Two overlapping histograms of within-document sentence length variation. Human writing centres on a standard deviation of 8.7 words, generated text on 6.4.

A group document that reads as one voice scores higher than the same document with four voices left in it.

A single number is meaningless here

The document percentage averages over sections written by different people with different habits, at least one of whom writes in a second language in most groups.

A 25% could be one section at 90% and the rest clean, or every section sitting at moderate suspicion. Those are entirely different situations and the document figure cannot tell them apart. Per-section scores can.

What to do before submitting

Score the assembled document per paragraph, not per contributor. The merged version is what gets submitted, so it is the version to check.

Look at which sections sit highest. Usually it is the formulaic ones: the methods description, the background section, the part somebody assembled from sources. Those score high for structural reasons. See AI detection false positives, by type of writing.

Leave some unevenness in. A little variation in rhythm between sections is closer to how human documents actually read. Harmonising to a single flat voice is what pushes the number up.

Check after the merge, not before. The score on four separate drafts tells you little about the document that will be submitted.

Keep the record

Group documents usually have excellent provenance and nobody thinks to keep it. Shared drive version history, the chat where sections were divided, individual drafts before the merge, and the comment threads. All of it is dated and all of it shows the document growing. See how to show you wrote your essay.

Ours is free and unlimited with no account, takes .pdf and .docx, has no word cap so the whole assembled document goes through at once, scores every paragraph separately, and never stores your text. Wrong-flag rate 0.4%, published, against Turnitin's 4%.

Check any text with our detector, free and unlimited →