Skip to content

Research

Model card

The detector that is live, what it was tested on, what it measured, and its limits in full.

The detector that is live

The service runs v1_8_s2_seed303. The detector’s own health check is the authority on that, and the status page reads it live, so the two can be checked against each other.

Every number on this card belongs to that run. Every reading the product shows carries the line "Our detector, run v1_8_s2_seed303". A reading without a run name is not one of ours.

It was chosen from five training runs by a rule written before the results. Its default line moved from 10% to 6% on 2026-09-15. Figures dated before that day were measured at 10%.

Intended use

Checking English academic and student writing, and marking the passages that read as AI-written so a person can look closer.

A flag from our detector is a reason to look closer. It is not proof that anyone cheated. It is not intended as the only basis for an academic, disciplinary, employment or legal decision about a person, or for automated penalties of any kind.

Out of scope

  • Text under 50 words, which it does not score.
  • Languages other than English.
  • Kinds of writing it was not trained or tested on, such as fiction, email, social posts, code and business documents.
  • Plagiarism, images and web addresses. It checks none of them.

What it was trained and tested on

The human-written training text included 98 essays from three published collections of English-learner essays, next to 1,396 papers by university students.

Every human-written document in training and testing was written before 2022, so it is known to be human. The test sets are published research collections. None is a set of real classroom submissions.

The 481 fully machine-written test documents were written by AI models used in training, or rewritten by our own tools: 236 came from the same eight AI models used in training, and 245 were AI text rewritten by our own rewriting tools, not by paid humanizer services. They are long documents.

Of the 1,928 human-written test documents, 1,200 had never been scored at all. The other 728 had been used to report earlier results, never to train or to set the line. None was used in training.

Measured results

MeasurementResult95% ceilingDateSource
Human-written papers and essays it was never trained on, wrongly flagged0 of 1,9280.155%2026-09-15docs/reports/unseen_human_1200_v1_8_2026-09-15.md; PRD.md §8.27, §8.28
The same ceiling for a school checking 10,000 papersnone observedabout 16 papers2026-09-15PRD.md §8.27
Fully machine-written documents missed, at every setting0 of 4810.62%2026-09-13docs/reports/v1_8_seed2_score_2026-09-13.md; docs/reports/sensitivity_points_2026-09-14.md §4
Learner essays kept out of training as a fairness test, wrongly flagged, at every setting0 of 1,1950.25%2026-09-13, recounted 2026-09-15docs/reports/v1_8_seed2_score_2026-09-13.md
Learner essays it had never scored, wrongly flagged0 of 6000.50%2026-09-15docs/reports/unseen_human_1200_v1_8_2026-09-15.md §2, §5
Partly AI-edited papers flagged102 of 140not applicable2026-09-15docs/reports/sensitivity_points_2026-09-14.md §4; PRD.md §8.28
Partly AI-edited papers given the verdict the record of the edits expects124 of 140not applicable2026-09-15PRD.md §8.28
Partly AI-edited papers called human-written38 of 140not applicable2026-09-15docs/reports/sensitivity_points_2026-09-14.md
Lightly polished papers flagged3 of 39not applicable2026-09-15docs/reports/sensitivity_points_2026-09-14.md §4
Paid-humanizer papers called human-written19 of 176not applicable2026-09-15PRD.md §8.24, §8.28; docs/reports/v1_8_ship_decision_2026-09-13.md §3b
Paid-humanizer papers at one tool’s most expensive setting0 of 12 called AI, 4 human-writtennot applicable2026-09-15docs/reports/v1_8_ship_decision_2026-09-13.md §3b
Same verdict as the commercial detector we compared it with, heavily rewritten machine text31 of 51not applicable2026-09-15docs/reports/b5_run_compare_v1_8_2026-09-15.json

Our detector, run v1_8_s2_seed303, standard setting (6%) unless a row says otherwise. The 95% ceiling is the highest error rate still believable at 95% confidence when we observe zero errors. It is the number to plan against, not the zero.

The three settings

Accusation-safeStandard (default)Sensitive
Flags a paper when more than this share reads as machine-written15%6%2%
Human documents wrongly flagged, of 1,9280 (ceiling 0.155%)0 (ceiling 0.155%)2 (0.10%, ceiling 0.28%)
Fully machine-written documents missed, of 481000
Partly AI-edited papers flagged, of 14092102103
Expected verdict on edited papers, of 140114124125
Rewriting-tool outputs called human-written, of 70855
Paid-humanizer papers called human-written, of 176271915
What it costsCalls 10 more edited papers human-written than the default, every one a paper where a model added a paragraph, plus 3 more rewriting-tool outputs and 8 more humanizer papers. Removes no false flags, because the default has none.Nothing extra measured. The closest human document sits comfortably below its line.Buys back 1 edited paper and 4 humanizer papers, and costs 2 false flags in 1,928.

Our detector, run v1_8_s2_seed303, measured 2026-09-15. Sources: PRD.md §8.28; docs/reports/sensitivity_points_2026-09-14.md §2, §4. The 2 false flags at the sensitive setting were academic papers, not learner essays. Changing the setting never turns an AI verdict into Human.

Limits

  • Zero observed is not zero forever. The ceiling, not the count, is the number to write into a policy.
  • The false-flag figures hold only for English academic and student writing like the test set. At the sensitive setting the ceiling is 0.28%, about 28 in 10,000.
  • Every human-written document we tested was written before 2022. Writing from after AI tools became common has not been tested.
  • No grammar-checking product was tested, so we cannot say whether a paper written with one will be flagged or not.
  • When an AI model lightly polishes a human paper, our detector mostly misses it. Of 39 polished papers, it flagged 3. No setting fixes this. On an earlier run, most of these papers read exactly 0% machine content, and a line cannot move below zero.
  • Even at the default setting, our detector called 38 of 140 partly AI-edited papers human-written. On the 16 papers where it disagreed with the record of the edits, the commercial detector we compared it with got all 16 wrong too.
  • The expected verdict on edited papers comes from a labelling rule we wrote. Across our five training runs, the 124 of 140 is the best of five and a statistical tie with the others.
  • Paid humanizer tools can hide machine writing from our detector. At one tool’s most expensive setting, none of 12 papers was called AI and 4 came back human-written. Of all 176 paid-humanizer papers, 59 were called AI, 98 Mixed and 19 Human. Two services were tested.
  • On heavily rewritten machine text, our detector and the commercial detector we compared it with often disagree: same verdict on 31 of 51 documents. Neither is ground truth there.
  • It does not answer under 50 words, and readings under 150 words are weak. Machine-written text under 300 words has not been measured.
  • Evidence on AI models it was not trained on is thin for this run: 183 documents from two models that carry a small amount of training text, all in one subject area.
  • One learner collection served as the fairness test. Two more were planned and never obtained. A count of zero false flags on learner essays is not proof of fairness.
  • The marked passages are a reading. Their placement accuracy has not been measured on this run, and no number on a result is the probability that a person used AI.
  • It is not built or tested for languages other than English or for kinds of writing other than academic and student writing.
  • We run and grade our own tests. No outside party has checked these results yet.
  • The first check after a quiet spell can take longer. There is no load test.

Ethical considerations

A published error rate helps a school write a policy and answer an appeal. It also says plainly that an error will eventually happen to someone.

A flag from our detector is a reason to look closer. It is not proof that anyone cheated. It does not show intent.

Ask us how pasted text is handled: hello@olive.is.

Check a paper yourself.

Free with a verified school email, up to 1,000 papers a month.