What is a detector trying to recognise?

A detector does not find a hidden author signature. It estimates how closely word, sentence and rhythm patterns resemble training examples. The output is an estimate rather than a directly observed fact.

A score depends on training data, thresholds, length, language, topic and model version. Two detectors can therefore disagree sharply about the same paragraph.

What did the Popular Science test find?

Popular Science tried five tools on two human and two samples per tool. Pangram, Grammarly and GPTZero classified all four correctly; Scribbr missed both samples and Copyleaks caught one of two.

Four examples show behaviour, not general accuracy. Different languages, prompts, lengths or editing processes can change the result.

Laptop showing a page about responsible AI writing
One tool or platform cannot reliably establish the full writing process on its own.
Aerps.com / Unsplash · Sources ↗ · Image terms ↗

Why do errors happen?

Short, formal school answers offer few clues and can be falsely flagged. Rewriting, translation and personal examples can also weaken patterns a detector seeks.

Many documents are co-produced by people and tools. A single label cannot reveal which parts the author understands or how the work developed.

Bias differs across systems

An 2026 study evaluated 16 systems and found that several were more likely to label English-language learner essays as machine generated.

A tool must be tested on the language, population and task where it will be used. Results from long English articles should not automatically judge short Serbian school essays.

A practical review process

Ask the author to explain the argument, sources and work sequence. Review notes, earlier versions and revision history, and verify the citations.

Use detectors only as supporting information, save the full report and give the author a chance to respond. Disagreement is a reason for review rather than punishment.

A student writes notes beside a laptop
Notes, drafts and revision history reveal the writing process more reliably than a single detector score.
Reuben Hu / Unsplash · Sources ↗ · Image terms ↗

Documenting the process

An idea, source notes, outline, first draft and later revisions create a useful work trail. Version history helps even when no was used.

Students should not surrender private conversations or an entire device. Schools should state permitted assistance and review rules in advance.

Conclusion: signal, never verdict

Detectors and experienced human reviewers can perform well on carefully defined tasks, but success does not automatically transfer to every language and model.

A detector may open a question but cannot close it alone. Understanding, sources, conversation and a trace of the writing process matter more than one score.

Key terms

— human writing incorrectly labelled as generated. — how closely a displayed probability matches the actual frequency of correct results.

Sources