Research
Model card
The detector that is live, what it was tested on, what it measured, and its limits in full.
The detector that is live
The service runs v1_8_s2_seed303. The detector’s own health check is the authority on that, and the status page reads it live, so the two can be checked against each other.
Every number on this card belongs to that run. Every reading the product shows carries the line "Our detector, run v1_8_s2_seed303". A reading without a run name is not one of ours.
It was chosen from five training runs by a rule written before the results. Its default line moved from 10% to 6% on 2026-09-15. Figures dated before that day were measured at 10%.
Intended use
Checking English academic and student writing, and marking the passages that read as AI-written so a person can look closer.
A flag from our detector is a reason to look closer. It is not proof that anyone cheated. It is not intended as the only basis for an academic, disciplinary, employment or legal decision about a person, or for automated penalties of any kind.
Out of scope
- Text under 50 words, which it does not score.
- Languages other than English.
- Kinds of writing it was not trained or tested on, such as fiction, email, social posts, code and business documents.
- Plagiarism, images and web addresses. It checks none of them.
What it was trained and tested on
The human-written training text included 98 essays from three published collections of English-learner essays, next to 1,396 papers by university students.
Every human-written document in training and testing was written before 2022, so it is known to be human. The test sets are published research collections. None is a set of real classroom submissions.
The 481 fully machine-written test documents were written by AI models used in training, or rewritten by our own tools: 236 came from the same eight AI models used in training, and 245 were AI text rewritten by our own rewriting tools, not by paid humanizer services. They are long documents.
Of the 1,928 human-written test documents, 1,200 had never been scored at all. The other 728 had been used to report earlier results, never to train or to set the line. None was used in training.
Measured results
| Measurement | Result | 95% ceiling | Date | Source |
|---|---|---|---|---|
| Human-written papers and essays it was never trained on, wrongly flagged | 0 of 1,928 | 0.155% | 2026-09-15 | docs/reports/unseen_human_1200_v1_8_2026-09-15.md; PRD.md §8.27, §8.28 |
| The same ceiling for a school checking 10,000 papers | none observed | about 16 papers | 2026-09-15 | PRD.md §8.27 |
| Fully machine-written documents missed, at every setting | 0 of 481 | 0.62% | 2026-09-13 | docs/reports/v1_8_seed2_score_2026-09-13.md; docs/reports/sensitivity_points_2026-09-14.md §4 |
| Learner essays kept out of training as a fairness test, wrongly flagged, at every setting | 0 of 1,195 | 0.25% | 2026-09-13, recounted 2026-09-15 | docs/reports/v1_8_seed2_score_2026-09-13.md |
| Learner essays it had never scored, wrongly flagged | 0 of 600 | 0.50% | 2026-09-15 | docs/reports/unseen_human_1200_v1_8_2026-09-15.md §2, §5 |
| Partly AI-edited papers flagged | 102 of 140 | not applicable | 2026-09-15 | docs/reports/sensitivity_points_2026-09-14.md §4; PRD.md §8.28 |
| Partly AI-edited papers given the verdict the record of the edits expects | 124 of 140 | not applicable | 2026-09-15 | PRD.md §8.28 |
| Partly AI-edited papers called human-written | 38 of 140 | not applicable | 2026-09-15 | docs/reports/sensitivity_points_2026-09-14.md |
| Lightly polished papers flagged | 3 of 39 | not applicable | 2026-09-15 | docs/reports/sensitivity_points_2026-09-14.md §4 |
| Paid-humanizer papers called human-written | 19 of 176 | not applicable | 2026-09-15 | PRD.md §8.24, §8.28; docs/reports/v1_8_ship_decision_2026-09-13.md §3b |
| Paid-humanizer papers at one tool’s most expensive setting | 0 of 12 called AI, 4 human-written | not applicable | 2026-09-15 | docs/reports/v1_8_ship_decision_2026-09-13.md §3b |
| Same verdict as the commercial detector we compared it with, heavily rewritten machine text | 31 of 51 | not applicable | 2026-09-15 | docs/reports/b5_run_compare_v1_8_2026-09-15.json |
Our detector, run v1_8_s2_seed303, standard setting (6%) unless a row says otherwise. The 95% ceiling is the highest error rate still believable at 95% confidence when we observe zero errors. It is the number to plan against, not the zero.
The three settings
| Accusation-safe | Standard (default) | Sensitive | |
|---|---|---|---|
| Flags a paper when more than this share reads as machine-written | 15% | 6% | 2% |
| Human documents wrongly flagged, of 1,928 | 0 (ceiling 0.155%) | 0 (ceiling 0.155%) | 2 (0.10%, ceiling 0.28%) |
| Fully machine-written documents missed, of 481 | 0 | 0 | 0 |
| Partly AI-edited papers flagged, of 140 | 92 | 102 | 103 |
| Expected verdict on edited papers, of 140 | 114 | 124 | 125 |
| Rewriting-tool outputs called human-written, of 70 | 8 | 5 | 5 |
| Paid-humanizer papers called human-written, of 176 | 27 | 19 | 15 |
| What it costs | Calls 10 more edited papers human-written than the default, every one a paper where a model added a paragraph, plus 3 more rewriting-tool outputs and 8 more humanizer papers. Removes no false flags, because the default has none. | Nothing extra measured. The closest human document sits comfortably below its line. | Buys back 1 edited paper and 4 humanizer papers, and costs 2 false flags in 1,928. |
Our detector, run v1_8_s2_seed303, measured 2026-09-15. Sources: PRD.md §8.28; docs/reports/sensitivity_points_2026-09-14.md §2, §4. The 2 false flags at the sensitive setting were academic papers, not learner essays. Changing the setting never turns an AI verdict into Human.
Limits
- Zero observed is not zero forever. The ceiling, not the count, is the number to write into a policy.
- The false-flag figures hold only for English academic and student writing like the test set. At the sensitive setting the ceiling is 0.28%, about 28 in 10,000.
- Every human-written document we tested was written before 2022. Writing from after AI tools became common has not been tested.
- No grammar-checking product was tested, so we cannot say whether a paper written with one will be flagged or not.
- When an AI model lightly polishes a human paper, our detector mostly misses it. Of 39 polished papers, it flagged 3. No setting fixes this. On an earlier run, most of these papers read exactly 0% machine content, and a line cannot move below zero.
- Even at the default setting, our detector called 38 of 140 partly AI-edited papers human-written. On the 16 papers where it disagreed with the record of the edits, the commercial detector we compared it with got all 16 wrong too.
- The expected verdict on edited papers comes from a labelling rule we wrote. Across our five training runs, the 124 of 140 is the best of five and a statistical tie with the others.
- Paid humanizer tools can hide machine writing from our detector. At one tool’s most expensive setting, none of 12 papers was called AI and 4 came back human-written. Of all 176 paid-humanizer papers, 59 were called AI, 98 Mixed and 19 Human. Two services were tested.
- On heavily rewritten machine text, our detector and the commercial detector we compared it with often disagree: same verdict on 31 of 51 documents. Neither is ground truth there.
- It does not answer under 50 words, and readings under 150 words are weak. Machine-written text under 300 words has not been measured.
- Evidence on AI models it was not trained on is thin for this run: 183 documents from two models that carry a small amount of training text, all in one subject area.
- One learner collection served as the fairness test. Two more were planned and never obtained. A count of zero false flags on learner essays is not proof of fairness.
- The marked passages are a reading. Their placement accuracy has not been measured on this run, and no number on a result is the probability that a person used AI.
- It is not built or tested for languages other than English or for kinds of writing other than academic and student writing.
- We run and grade our own tests. No outside party has checked these results yet.
- The first check after a quiet spell can take longer. There is no load test.
Ethical considerations
A published error rate helps a school write a policy and answer an appeal. It also says plainly that an error will eventually happen to someone.
A flag from our detector is a reason to look closer. It is not proof that anyone cheated. It does not show intent.
Ask us how pasted text is handled: hello@olive.is.
Check a paper yourself.
Free with a verified school email, up to 1,000 papers a month.