Skip to content

Research

How it works

How our detector reaches a verdict, and the cases it gets wrong.

What it reads

Our detector is a language model trained on academic papers and student essays. Some were written by people and some by AI models.

It reads English academic and student writing. It is not built or tested for other languages or other kinds of writing, such as fiction, email or business documents.

How a reading is made

You paste a paper. Our detector reads it sentence by sentence and in longer passages, and marks the passages it reads as AI-written.

From those marks it works out one number: the share of the paper it reads as written or edited by an AI model.

Treat the marks as a reading, not a finding. We have not measured how exactly they fall on the sentences an AI model wrote, on the detector that is live.

From that share to a verdict

Each check ends in one of three verdicts.

  • Human: the AI-written share is at or below the line your setting draws.
  • Mixed: part of the paper reads as written or edited by an AI model.
  • AI: at least 80% of the paper reads that way.

What a flag means

When this site says a paper was flagged, it got Mixed or AI.

A flag from our detector is a reason to look closer. It is not proof that anyone cheated.

The three settings

The setting decides where the line between Human and Mixed sits. Standard is the default, and a check that names no setting gets it.

  • Accusation-safe, above 15%. Only flags a paper when a substantial part of it reads as machine-written. In our tests it removed no false flags, because the default had none, and it called 10 more AI-edited papers human-written.
  • Standard, above 6%. Flags a paper when more than a small part of it reads as machine-written.
  • Sensitive, above 2%. Flags a paper at the first small sign of machine-written text. It flagged 2 of 1,928 human-written documents. Use it to prompt a second look, not a conversation about misconduct.

What a setting cannot change

Changing the setting never turns an AI verdict into Human. It only moves the line for papers that are partly machine-written. An AI verdict needs at least 80%, which is above every setting’s line.

When it does not answer

Under 50 words, our detector does not score the text. It tells you that instead of guessing. From 50 to 149 words the reading is labelled weak. Machine-written text under 300 words has not been measured.

If the detector is not answering, a check also comes back "Not scored". Nothing else scores the text in the meantime.

Where it gets things wrong

  • Light polish. When an AI model lightly polishes a human paper, our detector mostly misses it. Of 39 polished papers, it flagged 3. No setting fixes this.
  • AI-edited papers. Even at the default setting, it called 38 of 140 partly AI-edited papers human-written. Some had only a small AI-written share, so the record of the edits agreed they were Human.
  • Paid humanizer tools. They can hide machine writing from our detector. At one tool’s most expensive setting, none of 12 papers was called AI and 4 came back human-written.
  • Heavily rewritten text. On heavily rewritten machine text, our detector and the commercial detector we compared it with often disagree: same verdict on 31 of 51 documents. That measures disagreement, not who is right.
  • Recent writing. Every human-written document we tested was written before 2022, so we know it is human. Writing from after AI tools became common has not been tested, and neither has writing done with grammar checkers or translation tools.

How the live detector was chosen

The detector that is live was chosen from five training runs by a rule written before the results.

Its default line moved from 10% to 6% on 2026-09-15. Any figure dated before that day was measured at 10%.

We run and grade our own tests. No outside party has checked these results yet.

The model card lists each result with its date and source.

Check a paper yourself.

Free with a verified school email, up to 1,000 papers a month.