Using it well
How to read a verdict
Read a result in this order: first whether our detector answered, then the passages it marked, and the verdict word last.
The Human team · Sep 15, 2026
A result from our detector has a verdict word at the top: Human, Mixed or AI. Most people read that word and stop. It sums up everything under it, and on its own it tells you the least.
Check that it answered
Under 50 words our detector does not score the text. You get "Not scored" and no share, because it tells you that instead of guessing. From 50 to 149 words it gives a reading and labels it weak. Machine-written text under 300 words has not been measured, so give a short answer less weight than a full essay.
If the detector is not answering, you also get "Not scored", with a retry. Nothing else scores the text in the meantime.
Read the marked passages
Our detector marks the passages it reads as AI-written. A Mixed verdict with one marked paragraph in twelve pages is a different conversation from a Mixed verdict with half the pages marked. Only the marks show that difference. They are a reading, not a finding: we have not measured how exactly they line up with the sentences an AI model wrote.
The percentage is the share of the paper our detector reads as written or edited by an AI model. It is not the chance that a student cheated.
Know what the three words mean
- Human: the AI-written share is at or below the line your setting draws. Standard, the default, draws it at 6%.
- Mixed: part of the paper reads as written or edited by an AI model.
- AI: at least 80% of the paper reads that way.
When our site says a paper was flagged, it got Mixed or AI. Changing the setting never turns an AI verdict into Human. It only moves the line for papers that are partly AI-written.
Then weigh it
A flag from our detector is a reason to look closer. It is not proof that anyone cheated.
Our detector flagged none of 1,928 human-written papers and essays it was never trained on. The 95% ceiling on its false-flag rate is 0.155%: for a school checking 10,000 papers, no more than about 16 wrongly flagged papers at 95% confidence. That holds only for English academic and student writing like the test set, all written before 2022.
It also misses things. When an AI model lightly polishes a human paper, our detector mostly misses it: of 39 polished papers, it flagged 3. Keep both facts in mind when a colleague asks how much to trust a result.
Our detector, run v1_8_s2_seed303, standard setting (6%). False flags measured 2026-09-15, source docs/reports/unseen_human_1200_v1_8_2026-09-15.md. Polished papers measured 2026-09-15, source docs/reports/sensitivity_points_2026-09-14.md.