How we test
About Human
Our detector marks the passages in a paper that it reads as AI-written. Before a new version goes live, it has to pass a test on essays by people learning English, essays it never trained on.
Why we started here
Detectors and people learning English
A 2023 study by Liang and colleagues tested seven AI detectors of that time on 91 TOEFL essays written by people learning English. On average the detectors wrongly called 61% of those essays AI-written, while scoring 88 essays by US eighth graders, native speakers, as human almost every time. None of those detectors was ours, so the study says nothing about how ours performs.
Liang et al., GPT detectors are biased against non-native English writers. Patterns 4(7), published 2023-07-10 (preprint arXiv:2304.02819, 2023-04-18). Read 2026-09-15.
Fairness testing
Starting with learner essays
98
Learner essays in training, beside 1,396 papers by university students
0 of 1,195
Learner essays flagged in the fairness test, at every setting. 95% ceiling 0.25%
0 of 600
Learner essays flagged that our detector had never scored. 95% ceiling 0.50%
Our detector, run v1_8_s2_seed303. Fairness test read on 2026-09-13 and recounted at all three settings on 2026-09-15; the 600 read on 2026-09-15. Sources: docs/reports/v1_8_seed2_score_2026-09-13.md, docs/reports/unseen_human_1200_v1_8_2026-09-15.md.
Learner writing in training
We put learner writing into training on purpose: 98 essays from three published collections of English-learner essays, next to 1,396 papers by university students. That is a small share of the training set.
A test a release must pass
We kept a separate set of learner essays out of training and used it as a pass or fail test for each release. No writer in that set wrote a training essay. Our detector flagged none of the 1,195 essays, at every setting. The 95% ceiling on that, the highest error rate still believable, is 0.25%.
- The essays come from one university English language program.
- They were written between 2005 and 2012.
Essays it had never read
We then ran 600 learner essays our detector had never scored. It read every one as human-written. We also ran 180 learner essays through our detector and a commercial detector on the same text. Neither flagged any.
- The 600 come from three collections whose other essays were partly used in training, written between 2001 and 2019.
- The 180 are part of the 1,195 above. A small set, and a tie.
What this does not prove
Across both tests, no learner essay was flagged. That is a count on 1,795 essays. It is not proof of fairness, and we do not claim our detector is fairer than any other.
- One collection served as the fairness test. Two more were planned and never obtained, so we have no results on them.
- Only written essays were tested, all written between 2001 and 2019.
- Learners who now write alongside AI tools, grammar checkers or translation tools have not been tested.
Settings for caution
A school picks how much of a paper must read as AI-written before it is flagged. Accusation-safe (flags only above 15%) removed no false flags in our tests, because the default had none. It raises how much AI text it takes before a student is asked to explain their work, and it called 10 more AI-edited papers human-written. No setting turns an AI verdict into Human.
- Accusation-safe, above 15%: only flags a paper when a substantial part of it reads as machine-written.
- Standard, above 6%, the default: flags a paper when more than a small part of it reads as machine-written.
- Sensitive, above 2%: flags at the first small sign. It flagged 2 of 1,928 human-written documents. Use it to prompt a second look.
Where it falls short
When an AI model lightly polishes a human paper, our detector mostly misses it: of 39 polished papers, it flagged 3. Paid humanizer tools can hide machine writing from it. On heavily rewritten machine text, it and the commercial detector we compared it with often disagree: same verdict on 31 of 51 documents. It reads English academic and student writing only. It does not score text under 50 words, and readings under 150 words are weak. Even at the default setting, it called 38 of 140 partly AI-edited papers human-written. On some of those, the AI had changed so little that our own labelling rule agreed. Every human-written document we tested was written before 2022, so writing from after AI tools became common has not been tested.
Read the model cardWhat a flag is for
A flag from our detector is a reason to look closer. It is not proof that anyone cheated. Read the marked passages next to the student’s drafts and earlier work, then talk to the student. Don’t use a flag as the only reason for a penalty. We run and grade our own tests, and no outside party has checked these results yet.
Talk to us
If your school is deciding whether to use our detector, or a student has asked you about a result, write to hello@olive.is.
Check a paper yourself.
Free with a verified school email, up to 1,000 papers a month.