DetectorsFor administrators
A Detector Covers the AI Models Someone Measured It On, and No Others
A detector covers the AI models someone measured it against, and no others. Human measured fourteen models in September 2026: 4,084 documents from 310 source papers. The part of that result speaking to writing the detector never learned from is narrower: 900 documents from three models kept out of training, plus 419 more, from papers nothing in the project had touched, scored once after the model was frozen. Three or four of 176 documents from paid rewriting services read further from machine-written than under the model being replaced.
Bill Nguyen & HumanUpdated
What Changed on 21 September 2026?
Human replaced the detector model behind human.olive.is with a retrained one on 21 September 2026. The three sensitivity settings did not move: accusation-safe flags a paper above 15 percent machine content, standard above 6 percent as the default, and sensitive above 2 percent. Below 50 words the detector still declines to answer. What changed is the evidence underneath the settings, and one row moved the wrong way.
The settings behave much as before: at the default, 102 of 140 part-edited papers are still flagged and 123 of them get the verdict the record of the edits expects; at the sensitive setting, 2 of 728 human-written papers are flagged. The row that moved the wrong way is in the losses below.
Fourteen AI models from three vendors each wrote fresh documents to the outline and length of a human paper the detector had never trained on: 4,084 documents, drawn from 310 distinct source papers, across 24 federal field-of-study families. The documents cluster rather than sitting one to a paper: each 300-document set comes from between 216 and 249 distinct source papers, and Fable's 184 from 152. Thirteen models contributed 300 accepted documents each; Anthropic Fable 5.1 contributed 184, accepted from 255 attempts drawn from a smaller pool, because a standing content rule keeps part of the corpus away from that model. 66 of its 71 rejections were the length gate.
Every document in those sets is machine-written by construction, so AI is the correct verdict, a Human verdict an outright miss and a Mixed verdict a partial one. The retrained detector returned no outright misses and two partial ones across all 4,084. The model it replaces returned one outright miss and three partial ones under one of the two reading modes, and one and four under the other, on the same documents. Four documents separate them, one of which was the only outright human-written call either model made anywhere in the set. That is Human's own finding on its own test sets, not an outside audit.
That pooled figure is about recognition rather than reach. Three of the fourteen models were fixed in advance as carrying no writing in any training or development window; for the other eleven the count that would settle it, how much of that model's writing sits in the fitting data, is not published beside their rows. What speaks to writing the detector never learned from is the 900 documents below, and the 419 that follow them.
Count the Documents Behind Every Model Row
A row saying a detector caught a model is worth the number of documents behind it, and then worth a second look at how many different source papers those documents came from. With no misses in 300 documents the most the true miss rate can be at 95 percent confidence is 0.99 percent; pooled over 4,084 it is 0.073 percent. The arithmetic is one minus 0.05 to the power one over the count.
Run the same arithmetic on source papers and every ceiling rises. The 4,084 documents were written from 310 distinct papers, and a single 300-document set sits on as few as 216, which reads 0.99 percent counting documents and 1.38 percent counting papers. The per-document ceiling is the optimistic reading and the per-paper ceiling the conservative one. Both belong in a procurement file, and this programme's own rule for a complete row requires both.
| Set | Documents | Source papers | Called human-written | Called partly machine | Ceiling per document | Ceiling per source paper |
|---|---|---|---|---|---|---|
| Fourteen model sets pooled | 4,084 | 310 | 0 | 2 | 0.073 percent | 0.96 percent |
| Three models with no text in training | 900 | 242 | 0 | 0 | 0.332 percent | 1.23 percent |
| Reserve set, scored once after the model was frozen | 419 | 150 | 0 | 0 | 0.712 percent | 1.98 percent |
Those are Human's own measurements on its own test sets, September 2026, and every row holds under both reading modes. A count of zero is not a rate of zero. Both ceilings belong in a procurement file, and the arithmetic that turns a ceiling into papers works the same way on any vendor's number.
One more figure belongs beside the headline: the retrained model is not the cleanest row on record. An earlier draw of the same recipe, never a shipping candidate, scored no outright misses and one partial one on the same 4,084 documents.
Which Rows Show a Detector Handling Writing It Never Learned From?
Two of them, and not the pooled figure. A row from a model whose writing is in the training mix shows the detector recognising text like what it learned from, a different claim. Three models were named in advance, before any result existed, as the gate: models with no text in any training or development window. That pre-registered trio is where the evidence about unfamiliar writing sits.
Three is the size of that gate rather than a census of the fourteen. The other eleven rows are certified neither way, because the number that would settle each of them, the count of that model's writing in the fitting data, is not published beside the row.
The trio is Anthropic's Opus 4.8 and Google's Gemini 3.7 Flash and Gemini 3.5 Flash. Pooled, their row reads no outright misses on 900 documents from 242 source papers, across 24 field families and two vendors, under both reading modes, with a ceiling of 0.332 percent counting documents and 1.23 percent counting papers. The honest note beside it: the model being replaced scores that row identically, so this gate could not have told the two apart.
The second row is the reserve. Those same three models wrote 419 further documents from 150 source papers overlapping the fourteen coverage sets on zero, the training side on zero and the sealed sets on zero. They were scored once, after the checkpoint was chosen on development data and the model frozen; nothing in the reserve fed that pick. No misses and no partial calls: ceiling 0.712 percent counting documents, 1.98 percent counting papers. Two limits travel with it: the reserve spans 10 field families rather than 24, and it is 419 documents of a planned 450, a shortfall recorded before any score existed.
So the sentence the reports support is narrower than a marketing page would like, and the narrowness is the point. Evaluated is not the same as never trained on. Nineteen model identifiers are registered for this programme and fourteen have a measured row. One vendor's column is thin: of six registered OpenAI identifiers only one has a row; GPT-6 Astra is registered and was accepted on five writing probes of five on 18 September 2026 but has no row yet; and GPT-5.4 was refused by its own vendor through this account on 20 September 2026, one Sunday short of being recorded as retired. The one OpenAI model with a row is announced for retirement next month, on the vendor's own changelog, which this project has not verified. A claim that every current model is covered would be false about that column.
Read the Losses Before Reading the Gains
A rule written before the results returned one word: do not ship. Ten of its thirteen clause rows passed and three failed. Those three failures are two clauses, one disagreement with a second detector adjudicated against Human and a per-document tolerance on commercially rewritten text that fails under each of the two reading modes, and the founder overrode both on 21 September 2026 with the failing rows in front of him.
Both readings stay on file, and the three failures are not a seed's bad luck: every one of four draws of this recipe fails the same rows. The founder then changed the rule for the next candidate. A row now passes either by regressing nothing or by improving measurably, at least three documents net and at least two better for every one worse, with four things no improvement can buy off: a single wrongly flagged human paper, a single missed machine-written document, a single human-written call on the untrained-on gate, one on the reserve. Under the new rule this model passes all thirteen rows and the candidate rejected two days earlier is still refused. The change was made after this result, and the rule document says so itself.
- Paid humanizer services, the largest measured weakness. On 176 sealed documents from two commercial humanizer services, the retrained detector still reads 19 as human-written under one reading mode and 20 under the other. Against the model it replaces it reads 10 closer to machine-written, and 3 further away under one mode and 4 under the other; documents not called AI move from 117 to 110. Every draw of this recipe is worse on at least two of the 176, and a humanizer figure quoted without its reading mode is not reproducible. Independent work reports the same shape: fewer than 4 percent of AI-labelled rewrites were still flagged after a humanizing pass in a 2026 study of 642 abstracts 4. What a paid rewriting pass does to a draft is why that row is hard.
- One edited paper over-called. On 140 papers where a model edited part of a human draft, the edit record expects Mixed on one document and the retrained detector calls it AI, where the previous model called it Mixed. Every draw makes the same call.
- One held-out human paper. On 549 human-written papers kept out of training, the retrained detector returns one partly-machine verdict where the previous model returned none, on the same paper in every draw. The shipping rule does not gate that row; it is published because it moved the wrong way.
What passed, on the release test batch under one mode: no wrongly flagged papers among 600 human-written ones, ceiling 0.498 percent; no misses among 220 machine-written ones, ceiling 1.353 percent; and 123 of 140 edited papers given the verdict the edit record expects, against 121 for the model it replaces on the same instrument. On 70 sealed documents from a rewriting pass both models call 5 human-written, and four draws call 2, 3, 5 and 7 against a ceiling of 6, so the tie is a coin landing well. None of 1,195 learner essays kept out of training was flagged.
This is an estimate from our detector. Treat a flag as a reason to look closer, not as a finding.
What Does a Daily Model Check Promise, and What Does It Not?
Since 17 September 2026 a job has compared three vendors' model lists with the models already measured. One is a live listing call; the other two are lists the vendors' own command-line tools carry locally. On 21 September it saw 57 identifiers, 35 from Anthropic, 15 from Google and 7 from OpenAI, of which 20 are spellings it counts as already registered: 8 Anthropic, 7 Google, 5 OpenAI.
That 20 is a tally of spellings in vendor lists, not the size of the register, which holds 19 identifiers and one alias. Five dated records exist, 17 to 21 September 2026. The first was a rehearsal that did not call the Google list and the second was written late in the day, so four mornings have run the check as described. The record is days long, not months.
Two of its behaviours matter more than the schedule. It writes could not read rather than clean whenever something it depends on cannot be compared. On 19 September all three lists were read, and the word was written because the Anthropic tool's version moved mid-record, so that night's list no longer described the same baseline. On 20 September every list read cleanly and the word came from a probe that returned nothing. Neither night was a missing list, and neither was reported as a failure. And before recording a model as retired it waits for the vendor's own refusal, its 400 or 404 class, on two consecutive Sunday re-probes.
On a genuinely new release the job does less than its name suggests, on purpose. It writes the identifier into the morning's record, notifies a person and files the model as a question for a person to answer. Registering the model is a person's work on a person's authority: the job never edits the registry or the tests, never trains and never publishes a measured row. The plan gives it one optional hand-off: a check for that model's writing in the training data, a short writing probe and a check that the endpoint answers, writing nowhere but a probes directory and filing the registration itself as a task for a person. As installed on 21 September 2026 that hand-off is switched off, and the record says so by carrying no hand-off command. A published row needs 300 accepted documents across six field-of-study families or more; the promise is a day to notice, not a day to cover.
That gap is what a procurement question should test, whichever vendor is answering. A shared benchmark of twelve detectors found accuracy dropping against adversarial edits, sampling changes and generative models the detectors had not been built for 1. Nature reported in August 2026 that the accuracy rates in circulation are the companies' own internal figures rather than externally checked ones 2. An outside test in July 2026 found one commercial detector produced no false positives across 495 human-written passages while missing 30 of 297 AI passages imitating a particular author 3. The clauses that put those questions into an RFP are the practical version of the same test.
Common questions
Does a detector work on an AI model released after the detector was built?
Nobody knows until someone measures it on that model. A detector reads statistical habits in text, and a model it has never met may not share them. That is why coverage evidence is measured per model rather than as one accuracy figure. Human measured fourteen models in September 2026 and reports the document count and the confidence ceiling behind each row. The rows that speak to an unfamiliar model are the ones from models kept out of training: 900 documents, written from 242 source papers, across three such models, none missed, ceiling 0.332 percent counting documents and 1.23 percent counting papers. A shared benchmark found detector accuracy dropping against generative models the detectors were not built for 1.
Does a zero in a coverage table mean the detector never misses that model?
No. A zero is a count of observed misses in a finite test, not a rate. With no misses in a set of documents, the most the true miss rate can be at 95 percent confidence is one minus 0.05 raised to the power one over the number of documents. That is 0.99 percent for a 300-document row and 0.073 percent pooled across 4,084. Then ask which papers those documents were written from, because the same arithmetic on source papers is the conservative reading: the 4,084 sit on 310 papers, which allows 0.96 percent. Both ceilings are figures to plan against. The sets are also academic writing rather than classroom submissions.
Did the retraining improve results on models the detector has never seen?
Not measurably. On the pre-registered set of three models with no text in any training or development window, 900 documents in all, both the retired model and the retrained one return no outright misses and no partial ones, under both reading modes. That gate could not have told them apart. The gains that can be measured are on models whose writing is in the training mix. A separate reserve of 419 documents, scored once after the checkpoint was frozen, confirms the untrained-on result rather than improving it: no misses, no partial calls, ceiling 0.712 percent counting documents and 1.98 percent counting the 150 source papers behind them.
What got worse in this release?
Three things, all published. On 176 sealed documents from two commercial humanizer services the retrained detector reads 3 documents further from machine-written under one reading mode and 4 under the other, while reading 10 closer under each; documents not called AI move from 117 to 110. On one of 140 edited papers it returns AI where the record of the edits expects Mixed and the previous model returned Mixed. And on 549 human-written papers kept out of training it returns one partly-machine verdict where the previous model returned none. A rule written before the results failed on the humanizer row and on one disagreement with a second detector; the founder overrode both, and then changed the rule.
Does a daily check of vendor model lists mean a new model is covered the day it ships?
No. The check notices; it does not measure. Each morning it compares three vendors' model lists with the models already measured, writes what it saw into a dated record, and files anything new as a question for a person. It does not register the model, never trains and never publishes a measured row; registration, and the check that none of the model's writing is in the training data, are a person's work on a person's authority. A published row needs 300 accepted documents across at least six field-of-study families, which takes days. It also reports could not read rather than clean when something it depends on cannot be compared, which is what it wrote on 19 and 20 September 2026.
What should a school ask a detector vendor about model coverage?
Four questions, and they work on any vendor. Which AI models has the detector been measured against, named one by one? How many documents stand behind each row, and what confidence ceiling does that count buy? Which of those models had writing in the training data, and which were deliberately kept out? And what happens between a vendor releasing a model and a measured row existing for it? Nature reported in 2026 that the accuracy rates in circulation are largely the companies' own internal figures 2, so a row without a document count and a training-status column is an assertion rather than evidence.
References
- 1.RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors Association for Computational Linguistics (ACL 2024), 2024. aclanthology.orgSupports the gap between advertised accuracy and measured robustness: 12 detectors fooled by adversarial attacks, sampling variations, repetition penalties and unseen generative models.
- 2.AI-detection tools have made huge leaps forward — how good are they? Nature (Naddaf and Van Noorden), 2026. nature.comSupports the point that the current accuracy rates in circulation are the firms' own internal figures and are not externally verified.
- 3.AI detectors rarely flag human writing, but sometimes miss AI text imitating real authors Epoch AI (Jaeho Lee), 2026. epoch.aiOn author-imitation text published 15 July 2026, one commercial detector missed 30 of 297 passages and produced no false positives on 495 human-written passages; a Mixed verdict is counted there as a non-detection.
- 4.Why AI Detection Fails for Academic Integrity Karr, Khvatskii, Hua and Chawla, University of Notre Dame (arXiv), 2026. arxiv.orgSupports the humanizer figure quoted: fewer than 4 percent of AI-labelled rewrites still flagged after humanization, measured on a cohort of 642 abstracts of 25 to 500 words.
4 sources, numbered by first appearance.
General guidance for teachers, administrators and students. What holds at one institution, on one assignment, may not transfer to another.
Human reports how much of a document reads as machine-written. It does not report a probability that a person used AI, it does not check for plagiarism, and no number it produces stands for a student's honesty. This is an estimate from our detector. Treat a flag as a reason to look closer, not as a finding.