PangramFor teachers
Pangram Reports 97.67 Percent on Humanized Text; Outside Tests Split
Pangram says yes, and dates the claim. On its own test set, Pangram 4's technical report reads humanized text as AI-generated 97.67 percent of the time, and as Mixed or AI-generated 98.83 percent [1]. Per-tool rates in that report run between 92.78 and 99.70 percent [1], so which humanizer ran changes the answer. The outside record splits. A 2025 economics working paper reported near-zero error rates on humanizer output, while a May 2026 magazine test, run two months before Pangram 4 shipped, found one humanizer that Pangram called human every time.
Bill Nguyen & HumanUpdated
What Does Pangram Publish About Humanized Text?
Pangram publishes a catch rate for humanized text and dates it, and the two 2026 figures name the model version as well. An August 2025 update reported above 90 percent on all the notable humanizers Pangram tested 2. Pangram 3.3, released 13 May 2026, claimed twice as many commercially humanized texts caught as the version before it 4. The Pangram 4 technical report gives a rate: 97.67 percent 1.
Every one of those figures is Pangram's own, measured on Pangram's own set.
The average is the number that travels worst. In the same report, the share of humanized text read as AI-generated runs from 92.78 to 99.70 percent across thirteen commercial systems, and the share read as Mixed or AI-generated from 95.28 to 100 percent 1.
Seven points separate the best-caught system from the worst. At the bottom of that range, roughly one humanized document in fourteen reads as something other than AI-generated, and one in twenty-one reads as neither Mixed nor AI. That is a different Tuesday from the one the headline figure describes.
Pangram 4 also scores the processing itself. Next to its provenance labels it carries a humanizer score, and marks a segment as humanized once that score passes a threshold set at 0.91 on release 3. One check yields two readings: whether the writing reads as machine-written, and whether it reads as rewritten to look otherwise. On text already read as AI, the humanizer flag fires between 91.52 and 99.39 percent of the time across the same thirteen systems 1.
A high humanizer score describes the rewriting, not the first draft. The pattern it reads is the subject of what humanized text looks like.
Has Anyone Outside the Company Tested It on Humanized Text?
Three outside results name Pangram specifically, and they do not agree. A 2025 National Bureau of Economic Research working paper evaluated four detectors and reported near-zero error rates for Pangram. Those rates held across models and threshold rules, on passages of 50 words or fewer, and on humanizer output 5. The Atlantic ran ChatGPT and Claude output through Walter Writes AI in May 2026, and reported that Pangram called the result human-written every time 6.
The date on the second result is the part that falls off when it gets quoted.
May 2026 is before Pangram 4 shipped on 29 July 2026 3, so the article describes an earlier model meeting one commercial humanizer, not the version a school licenses this term. That doesn't make the finding wrong. It makes it a measurement with an expiry date, which is what the vendor's July table is as well.
The third result names the model version it tested, which is what makes it readable at all. A 2026 University of Notre Dame paper scored Pangram 3.2 and GPTZero on abstracts humanized with Undetectable AI v11, and fewer than 4 percent of AI-labelled rewrites stayed flagged across the two detectors 8. The 3.3 release note says some text from the latest OpenAI and Anthropic releases had been labelled human by the model 3.3 replaced 4. The cohort matters too. Those were 642 abstracts of 25 to 500 words 8, short academic prose and not term papers. The NBER paper's robustness on passages that short was measured on a different model and a different set 5.
Nature's 2026 survey of the field states the position plainly. Only the accuracy rates the firms announce from their own internal testing can be current, and those are not externally verified 7. Set the vendor's July table beside a humanizer's undated promise and the comparison is between two unaudited documents, one of which at least says what it measured and when.
Why Do the Humanizer Sales Pages Say the Opposite?
The one humanizer page cited here sells a promise rather than a measurement, and publishes nothing behind it. Rephrasy's homepage guarantees a 100 percent pass rate on all major AI detectors, with no sample size, no date, no detector version and no method beside it 9. The promise names no detector, so reading it as an answer to Pangram's figures is an inference. Pangram's claims can be argued with because they name what was measured; a bare pass rate can't.
Undetectable AI's ethics page states that a single AI text detection score alone is not enough to rule out cheating or deception, and that the company has never condoned cheating 10. That is a statement about a clean reading: passing a detector clears nothing.
Human's own in-house measurement, as of September 2026, found that whether a humanized document is caught depends on the humanizer's price tier more than on which humanizer it is. This is an estimate from our detector. Treat a flag as a reason to look closer, not as a finding. That measurement is Human's own held-out testing rather than an independent audit, and no figure is published with it.
The gap between the free tools and the subscriptions is the subject of free humanizer versus paid tier.
Read the Model Version and the Date Before Reading the Number
Each side of this argument publishes a snapshot of a moving target. Pangram shipped 3.3 on 13 May 2026 and Pangram 4 on 29 July 2026, eleven weeks apart 43. The 3.3 announcement says what prompted it: some text from the latest OpenAI and Anthropic releases had been labelled human 4. A humanizer benchmark from 2025 therefore describes a model that has been replaced since.
| Claim | Source | Date |
|---|---|---|
| 97.67 percent of humanized text read as AI 1 | Pangram Labs | July 2026 |
| Near-zero errors on humanizer output 5 | NBER working paper 5 | 2025 |
| Human-written every time on one humanizer 6 | The Atlantic | May 2026 |
| 100 percent pass rate on all major AI detectors 9 | Rephrasy | no date given |
Two questions settle most humanizer-versus-detector claims. Which detector version was tested, and on what date. A humanizer page naming neither is marketing rather than measurement, and a vendor table naming both is still a vendor table.
For a school the consequence is a procurement question: a licence bought on a figure measured in 2025 rests on a model that has since been replaced, and the renewal question is what the current version was measured on, and by whom.
The wider accuracy picture, including what Pangram reports on ordinary human writing, sits in is Pangram accurate.
What a Teacher Can Do With One Suspect Paragraph
Treat the paragraph as a reason to ask rather than a conclusion, which is what the vendor recommends first. Pangram's model card records a non-zero error rate and states that false accusations of AI use can cause reputational damage and emotional harm 3. Its chief executive told The Atlantic the tool should never be the ending arbiter, only the starting point for a more thorough investigation 6.
On 642 short abstracts scored with Pangram 3.2 and GPTZero, honest, guideline-compliant AI editing carried a higher sanction risk than humanizer-assisted evasion 8. If that carries to coursework, which the paper did not test, a rule policed by a detector rewards concealment.
Nature relays a figure from Pangram's own technical study: when the firm used consumer AI to substantially modify human-written student essays, the model still labelled the results fully human 41 percent of the time 7. A real draft reworked by a chatbot is the case that figure describes.
What survives scrutiny is process evidence. The draft history and the document's version log say when the words appeared, and five minutes of conversation about the argument on page two says whether the student can hold it.
A paragraph scoring as humanized is a reason to ask, and the asking is set out in how to talk to a student about AI.
Common questions
Does Pangram label humanized text as AI, or only as suspicious?
Both, and they're separate readings. Pangram 4 assigns provenance labels to stretches of a document, and its technical report says humanized text reads as AI-generated 97.67 percent of the time and as Mixed or AI-generated 98.83 percent 1. A second score covers the processing: a segment is marked humanized once its humanizer score passes a threshold set at 0.91 on release 3. A high humanizer score is a statement about the text having been rewritten, not about who wrote the first draft.
Which humanizers does Pangram catch least often?
Pangram's August 2025 post does name tools. The lowest rates in that table were Undetectable AI at 90.3 percent, TwainGPT at 92.7, Just Done at 93.5 and humanizeai.io at 93.8, with ten tools at 100 percent 2. Those are Pangram's own figures on an undescribed set, measured on an earlier model the post does not name. The Pangram 4 technical report names no products. It labels thirteen systems Commercial A through M, and the share of humanized text read as AI-generated runs from 92.78 to 99.70 percent 1.
Why do independent tests of the same detector disagree?
The tests differ in model, humanizer and month, and no source says which difference explains the split. A 2025 National Bureau of Economic Research working paper found near-zero error rates that held on humanizer output 5. The Atlantic found the opposite on one tool in May 2026 6, two months before Pangram 4 shipped 3. Nature's 2026 survey concluded that only the firms' internal figures can be current and that none of them are externally verified 7.
Do humanizer companies say their tools are meant for schoolwork?
The two pages cited here do not answer that directly. Rephrasy guarantees a 100 percent pass rate on all major AI detectors 9. A different company, Undetectable AI, states on its ethics page that it has never condoned cheating, and that a single AI text detection score alone is not enough to rule out cheating or deception 10.
Does a humanized paper that a detector reads as human-written clear the student?
No, and the reverse holds too: a flag does not convict. A detector reports a reading on text. A 2026 University of Notre Dame paper found, on 642 abstracts of 25 to 500 words scored with Pangram 3.2 and GPTZero, that honest, guideline-compliant AI editing carried a higher sanction risk than humanizer-assisted evasion 8. A defensible decision leans on process evidence: the draft history and what the student can explain about the argument.
References
- 1.Pangram 4 Technical Report Pangram Labs and University of Maryland (arXiv), 2026. arxiv.orgThe 97.67 percent AI-generated and 98.83 percent Mixed-or-AI figures on humanized text; Table 14's per-system AI recall of 92.78 to 99.70 percent and Mixed-or-AI recall of 95.28 to 100 percent across thirteen commercial systems; and the auxiliary humanizer head's separate 91.52 to 99.39 percent true-positive range.
- 2.How well does Pangram perform on humanizers? (Updated August 2025) Pangram Labs, 2025. pangram.comPangram's August 2025 claim of above 90 percent on all the notable humanizers it tested, a vendor measurement rather than an audit.
- 3.Pangram 4 Model Card Pangram Labs, 2026. pangram.comThe 29 July 2026 release date, the humanizer score and its 0.91 threshold at release, and the statement that the model has a non-zero error rate and that false accusations cause real harm.
- 4.Meet Pangram 3.3 Pangram Labs, 2026. pangram.comThe 13 May 2026 release, the claim of twice as many commercially humanized texts caught, and the admission that some text from the latest OpenAI and Anthropic releases had been labelled human.
- 5.Artificial Writing and Automated Detection (Working Paper 34223) National Bureau of Economic Research (Jabarian and Imas), 2025. nber.orgThe independent evaluation of four detectors reporting near-zero error rates for Pangram that held across models, threshold rules, passages of 50 words or fewer, and humanizer tools.
- 6.America Has a Pangram Problem The Atlantic (Matteo Wong), 2026. theatlantic.comThe reporter's May 2026 test in which Pangram called Walter Writes AI output human-written every time, and the chief executive's statement that Pangram should never be the ending arbiter.
- 7.AI-detection tools have made huge leaps forward — how good are they? Nature (Naddaf and Van Noorden), 2026. nature.comThat only the firms' own internal accuracy rates can be current and none are externally verified, and the figure it relays from Pangram's own technical study: 41 percent of substantially modified student essays still labelled fully human.
- 8.Why AI Detection Fails for Academic Integrity Karr, Khvatskii, Hua and Chawla, University of Notre Dame (arXiv), 2026. arxiv.orgFewer than 4 percent of AI-labelled rewrites still flagged after Undetectable AI v11 humanization, scored on Pangram 3.2 and GPTZero; the cohort of 642 abstracts of 25 to 500 words; and the higher sanction risk carried by honest, guideline-compliant AI editing.
- 9.Rephrasy Rephrasy, 2026. rephrasy.aiThe homepage promise of a 100 percent pass rate on all major AI detectors, published with no sample, date or detector version.
- 10.Ethics Undetectable AI, 2026. undetectable.aiThe vendor's own statements that it has never condoned cheating and that a single AI text detection score alone is not enough to rule out cheating or deception.
10 sources, numbered by first appearance.
General guidance for teachers, administrators and students. What holds at one institution, on one assignment, may not transfer to another.
Human reports how much of a document reads as machine-written. It does not report a probability that a person used AI, it does not check for plagiarism, and no number it produces stands for a student's honesty. This is an estimate from our detector. Treat a flag as a reason to look closer, not as a finding.