Skip to content

HumanizersFor teachers

Undetectable AI Promises a Human Score; Pangram Labs' Own August 2025 Table Reports 90.3 Percent Caught, Its Lowest of Twenty Tools

Undetectable AI sells a rewrite scored as human, and the two published tests of it disagree. Pangram Labs' own August 2025 table reports catching 90.3 percent of Undetectable AI rewrites, the lowest of the twenty tools it lists [1]. So about one rewrite in ten got through. A 2026 Notre Dame study scored humanized abstracts on Pangram 3.2 and found fewer than 4 percent of the AI-labelled rewrites still flagged [4]. The tests differ in text, length, date and detector version, and neither settles what a particular student did.

Bill Nguyen & HumanUpdated

Does Undetectable AI Work Against a School's Detector?

Sometimes, and the published evidence splits by detector, by version and by passage length. Undetectable AI's home page sells a rewrite that will "Score as human written on AI detectors and improve readability" 2. On its 27 August 2025 results page, Pangram Labs reported catching 90.3 percent of that product's rewrites 1.

Ten of the twenty tools on that table sit at 100.0 percent, and the lowest figure on the list belongs to Undetectable AI 1. The page does not say how many samples it ran per tool 1. Read the other way, close to one rewrite in ten did not come back flagged on Pangram's own test.

Two limits travel with the table. It describes one detector, and the school marking a paper next Tuesday may be running another. It also describes a replaced model: Pangram 4 succeeded Pangram 3.3.2 on 29 July 2026 5.

A detection rate is a count on a test set, and Pangram's own model card says as much from the other end. It acknowledges a non-zero error rate, and it states that false accusations of AI use can lead to serious consequences, including reputational damage and emotional trauma 5. A percentage published about a corpus is not a probability about one student's essay.

What Did the One Outside Study Actually Test?

Published abstracts, not essays. A 2026 University of Notre Dame paper retained 642 abstracts of 25 to 500 words, then humanized every variant with Undetectable AI 4. Of the rewrites the detector had labelled AI beforehand, fewer than 4 percent stayed flagged 4. That is a false negative rate above 96 percent.

The paper's word filter allowed abstracts from 25 words, below the 50-word floor Pangram 4's model card sets, and it does not say how many were that short 45.

The scoring ran on Pangram 3.2 4. Pangram 4 replaced Pangram 3.3.2 on 29 July 2026, and the newer model carries a separate humanizer score that calls a segment humanized only once that score clears a threshold set at 0.91 5. The study measured the version available in its own window, which is all any study can do. A year separates two honest measurements of the same product, and only the Notre Dame paper names the detector version it ran 4; Pangram's page does not say 1. The fuller record of what the detector company publishes sits in what Pangram publishes about humanized text.

One finding in that paper matters more in a department meeting than any detection rate. Light AI editing of the kind a guideline permits, the paper's proxy for honest assistance, was flagged more often than fully humanized text 4. Students may be flagged while drafting their own ideas and using AI only for clarity or grammar, yet others who generate inorganic drafts can evade detection after humanization 4.

What Does a School See When a Paper Has Been Rewritten?

A percentage, from whichever checker the institution licenses, with no field naming the tool that produced the text. Where that checker is Turnitin, its own page says the AI checker is meant to surface text spinners and AI bypassers, also called humanizers 7.

Pangram's model card describes segment labels and a humanizer flag 5, and Turnitin's page names no product 7. Neither describes a rewrite history. The arithmetic under a verdict is worth knowing before a conversation starts. Pangram labels each stretch of a document Human, AI-Assisted or AI-Generated. It then reads the document as AI only when at least 80 percent of its characters carry the AI label, and as Human at 90 percent human characters 5. A rewrite that moves a handful of segments can move the verdict, when the document sits near one of those lines, without touching the argument.

The market behind these rewrites is not a niche one. NBC News reported that Turnitin keeps a list of 150 tools that adjust text so a detector will not flag it, some charging as much as $50 for a subscription 6. The same report cited an academic integrity company that counted 43 humanizer sites drawing 33.9 million visits in a single month 6. The general version of the question sits here: whether AI humanizers work.

Why Does the Answer Change Every Few Months?

Because both sides ship, and the pages that sell the rewrite carry no version. The August 2025 table describes a detector generation since replaced and names no version of its own; the home-page promise carries none either 12. The one test that recorded versions is the Notre Dame paper, which names Pangram 3.2 and Undetectable AI v11 4. A detection figure is a reading of one tool against one model on one date.

A screenshot from last spring doesn't describe the model running this term.

The seller says so too. Undetectable AI's ethics page states that the company has never condoned cheating and never will. The same page 3 also states that "A single AI text detection score alone is NOT enough to rule out cheating or deception." A vendor selling rewrites and a researcher studying false positives land on the same sentence.

Whether a rewrite gets past a checker and whether it is allowed are separate questions. The University of Southern California's academic integrity office states that work created by a generative AI tool and presented as a student's own counts as plagiarism, whether it was copied word for word or paraphrased 8. Write the tool, the version and the date beside any score that lands in a file. A hearing six months from now may have no way to reconstruct which model produced a number.

Check the Document, Not the Product Page

Run the whole paper rather than the paragraph that reads oddly. Run it at its real length, on whatever version the school licenses today. Pangram's own page quotes a study finding its miss rate rises a bit on shorter passages while staying low 1, and its model card refuses anything under 50 words 5. Then ask the questions a percentage cannot answer, starting with where the draft came from.

The draft history answers a question a second detector run cannot, where the text came from, and it costs one email: reading a Google Docs version history.

An oral follow-up is the other place to look, and it works whether or not a rewriting tool was ever involved: the questions worth asking after a flag.

A flag describes text. A finding about a person needs evidence about that person.

For the paper already on the desk rather than the product category behind it, Human at human.olive.is reads up to 1,000 words pasted in, three times a day, with no account and nothing stored. It reports how much of that text reads as machine-written or machine-edited, sentence by sentence, with a document verdict of Human, Mixed or AI. This is an estimate from our detector. Treat a flag as a reason to look closer, not as a finding. It is tuned for college-level academic writing, and below 50 words it declines to answer.

One in-house finding on rewriting tools, measured on Human's own held-out sets rather than by an independent audit, names no figure: whether a humanized document is caught depends on the humanizer's price tier more than on which humanizer it is. The student's side of that question is free humanizer tiers against paid ones.

Common questions

Does a rewritten paper come back clean on every detector?

No, and the two published results disagree by corpus, by detector version and by year. Pangram's own August 2025 table reports catching 90.3 percent of Undetectable AI rewrites, the lowest detection rate of the twenty tools listed there 1. A 2026 Notre Dame paper humanized published abstracts with the same product. Scoring with Pangram 3.2, it reported fewer than 4 percent of the AI-labelled rewrites still flagged 4. Both were measured on corpora rather than on a student's essay, and neither describes the detector a particular school runs this term.

Can a teacher tell which tool rewrote a paper?

Not from the checker. A report gives a share of the document, a date and a submission, with no field naming a rewriting tool. Turnitin's page says its checker surfaces text spinners and bypassers; it does not claim to name which product was used 7. The evidence that names a process sits outside the detector, in the draft history and in a conversation about the argument.

Is using a rewriting tool against the rules even when nothing is flagged?

Policy decides that, and the score does not. The University of Southern California's academic integrity office states that work created by a generative AI tool and presented as a student's own counts as plagiarism, whether copied verbatim, near-verbatim or paraphrased 8. Undetectable AI's own ethics page states that the company has never condoned cheating and lists cheating in school as unacceptable use 3. Detection and permission are two different questions, and only one of them changes with a software release.

Why did an outside study find almost every rewrite getting through?

Because that is what the paper measured. The 2026 Notre Dame paper retained 642 published abstracts of 25 to 500 words and scored them with Pangram 3.2 4. Its word filter allowed abstracts from 25 words, below the 50-word floor Pangram 4's model card sets, and it does not say how many were that short 45. Its authors say the result may not generalize to other detector versions or to classroom genres 4, and Pangram 4 has since replaced Pangram 3.3.2 5. Whether a full essay on a current version behaves the same way has not been measured.

What should a teacher do with a flag on a paragraph that reads as rewritten?

Check the whole document rather than the paragraph, since a study quoted on Pangram's own page found slightly more misses on shorter passages 1, and Pangram's model card refuses text under 50 words 5. Record the checker and its version next to the score, because the model card names Pangram 4 as the successor to Pangram 3.3.2 5. Then ask for the draft history and talk the student through the argument. A label describes text; the finding needs evidence about the student.

References

  1. 1.How well does Pangram perform on humanizers? (Updated August 2025) Pangram Labs, 2025. pangram.comThe per-tool table dated 27 August 2025: twenty tools, ten of them at 100.0 percent, Undetectable AI the lowest detection rate at 90.3 percent, with no sample count per tool stated on the page.
  2. 2.Undetectable AI Undetectable AI, 2026. undetectable.aiThe home-page promise that a rewrite will score as human written on AI detectors and improve readability, carrying no detector version or date.
  3. 3.Ethics Undetectable AI, 2026. undetectable.aiThe same vendor's statement that it has never condoned cheating, that cheating in school is unacceptable use, and that a single AI text detection score alone is not enough to rule out cheating or deception.
  4. 4.Why AI Detection Fails for Academic Integrity Karr, Khvatskii, Hua and Chawla, University of Notre Dame (arXiv), 2026. arxiv.orgThe 642-abstract cohort of 25 to 500 words, fewer than 4 percent of AI-labelled rewrites still flagged after Undetectable AI humanization, and the finding that honest AI editing carries a higher sanction risk than humanizer-assisted evasion.
  5. 5.Pangram 4 Model Card Pangram Labs, 2026. pangram.comThe 29 July 2026 release succeeding Pangram 3.3.2, the 50-word intended-use floor, the three segment labels and the 90 and 80 percent character thresholds behind a document verdict, the 0.91 humanizer threshold, and the stated non-zero error rate and harm of false accusations.
  6. 6.To avoid accusations of AI cheating, college students are turning to AI NBC News, 2026. nbcnews.comTurnitin's list of 150 tools charging as much as $50 for a subscription, and Cursive's count of 43 humanizer sites with 33.9 million visits in one month.
  7. 7.AI Checker Solutions: Ensure Academic Integrity Turnitin, 2026. turnitin.comTurnitin's own description of its AI content checker as surfacing text spinners and AI bypassers, also called humanizers, in student submissions.
  8. 8.Addressing Generative AI University of Southern California, Office of Academic Integrity, 2026. academicintegrity.usc.eduThe statement that work authored by a generative AI tool but represented as the student's own is plagiarism, whether copied verbatim, near-verbatim or paraphrased.

8 sources, numbered by first appearance.

General guidance for teachers, administrators and students. What holds at one institution, on one assignment, may not transfer to another.

Human reports how much of a document reads as machine-written. It does not report a probability that a person used AI, it does not check for plagiarism, and no number it produces stands for a student's honesty. This is an estimate from our detector. Treat a flag as a reason to look closer, not as a finding.

Human

Check a paper

Paste a paper and see which sentences read as machine-written, with a verdict of Human, Mixed or AI. Nothing is stored and no account is needed.