Skip to content

HumanizersFor students

Whether an AI Humanizer Works Depends on the Detector, the Tool and the Year

Humanizers sometimes work, and never in general: each published result is tied to one detector release. A 2023 study of a research paraphraser, not a commercial humanizer, cut DetectGPT from 70.3 percent accuracy to 4.6 percent [1]; how far that carries to commercial tools is an assumption. Pangram's own 2026 technical report reads humanized text as AI-generated 97.67 percent of the time [2]. A Notre Dame test of Pangram 3.2 and GPTZero, on short abstracts, left fewer than 4 percent of rewrites first labelled AI still flagged after one humanizer ran [3].

Bill Nguyen & HumanUpdated

Do AI Humanizers Still Beat AI Detectors?

Yes against some detectors in some outside tests, and against the detectors one 2023 study tested it was not close. DIPPER paraphrasing dropped DetectGPT from 70.3 percent accuracy to 4.6 percent at a fixed 1 percent false positive rate, without much change to what the text meant 1. Watermarking did not stop the same paraphraser. Neither did GPTZero, nor OpenAI's own classifier 1. DIPPER is a research paraphraser though, so how far that carries to commercial humanizers is an assumption.

Pangram 4 adds a head trained to recognise humanized text, and its own report gives very different figures 2. Pangram 4 reads humanized text as AI-generated 97.67 percent of the time, and as Mixed or AI-generated 98.83 percent 2. The per-tool figure travels further: AI recall across the 13 commercial humanizers it evaluated runs from 92.78 to 99.70 percent 2. Those are the vendor's measurements of its own detector rather than an audit.

TestDetectorResult
DIPPER, 2023 1DetectGPT70.3% to 4.6% accuracy
Pangram 4 report 2Pangram 497.67% read as AI
Per-tool AI recall 2Pangram 492.78% to 99.70%
Notre Dame, 2026 3Pangram 3.2, GPTZerounder 4% of AI-labelled rewrites still flagged
Reporter's account, 2026 4Pangramhuman-written in every trial reported, count not given

The rows differ in detector, release, text length and humanizer, and no cited test varies one of those alone.

Two outside results point the other way. A reporter for The Atlantic wrote in May 2026 that Pangram called every output he pasted from one commercial humanizer human-written, naming no Pangram release and giving no count 4. The Notre Dame study names its detectors, its set and its lengths 3.

Fewer than 4 percent of the rewrites those two detectors had first labelled AI stayed flagged after one commercial humanizer ran 3. That reading and the 97.67 percent figure differ in more than one way. Pangram 3.2, one of the two detectors the Notre Dame authors scored with, is an earlier release than the one behind the 97.67 percent figure, GPTZero is a different product again, and the Notre Dame set is short abstracts run through a single humanizer 3.

Every figure above belongs to one test, and none is a property of humanizers in general.

What Does a Humanizer Change in a Sentence?

Wording, by two documented routes. Pangram Labs defines the category as tools that rewrite AI-generated text to evade detectors, and names synonym swapping as one technique 5. Another is inserting nonsensical phrases, on the theory that a detector reads gibberish as unlikely to have come from a machine 5. The Atlantic's reporter described the same thing as anodyne rewording, a swapped transition clause and introduced grammatical oddities 4.

A swapped word changes the register of a sentence without changing its job, and a paragraph can read as slightly wrong for four sentences running before anything in it looks deliberate. A grader reads the prose before any tool does. Odd word choice in a paper whose earlier drafts read plainly is the kind of thing a teacher raises in conversation, and that conversation needs no detector at all.

What a humanizer does to writing works through the sentence-level edits, and what humanized text looks like collects the artefacts they leave behind.

Which Tool and Which Price Tier Change the Answer?

Both do. Pangram's per-tool AI recall across the commercial humanizers it evaluated spans 92.78 to 99.70 percent, a spread of about seven points on the vendor's own table 2. Human's in-house testing, a separate measurement on a different detector, finds that whether a humanized document is caught depends on the humanizer's price tier more than on which humanizer it is.

That is Human's own finding on its own held-out sets, not an outside audit, and it is stated without a figure. This is an estimate from our detector. Treat a flag as a reason to look closer, not as a finding.

The market underneath these tools is large 6. Turnitin told NBC News it keeps a list of 150 tools that adjust text so a detector will not flag it, some charging as much as 50 dollars for a subscription 6. Cursive counted 43 humanizer sites drawing a combined 33.9 million visits in October 2025, in the same report 6.

Forum threads are where this gets confusing. A student who reports that a humanizer went unflagged and a student who reports being caught by the same tool can both be telling the truth. Tier, detector release and passage length may all differ between them.

The price spread inside that market is where free humanizer vs paid tier picks the question up.

Does an Unflagged Paper Mean the Work Is in the Clear?

No. One published policy shows why. The University of Southern California's Office of Academic Integrity, as of September 2026, counts material created by a generative tool and represented as a student's own work as plagiarism, whether it was copied verbatim, near-verbatim or paraphrased 7. A humanizer produces the paraphrased case, and the policy already names it. Other institutions write their own rules, and a student checks the one that governs the course.

A clean reading does not touch that. Light editing of an abstract, the Notre Dame authors' proxy for guideline-compliant AI assistance, was flagged at 38 to 80 percent depending on the detector, while humanizer output fell below 4 percent 3. Their summary of the asymmetry reads: "Honest AI-editing results in a higher sanction risk than humanizer-assisted evasion" 3.

Those figures describe 642 abstracts of 25 to 500 words, so they cover short academic passages and not full papers 3.

A reading describes text. It says nothing about who wrote it.

A paper that comes back clean has been described as reading like human writing, and used a humanizer and still got flagged covers the case that lands the other way.

Read the Reading Before Trusting Either Claim

Check what a number was taken on before treating it as evidence. Ask of any pass rate, a humanizer's or a detector's: which detector release, how many texts, what date. The Notre Dame figures name their two detectors, a cohort of 642 abstracts and a 25 to 500 word range 3, which is why anyone can argue with them at all.

Short text is where a reading gets weak in both directions. Human, at human.olive.is, declines to score anything under 50 words, and treats a reading on 50 to 149 words as weak, because a short passage gives a detector less to read. A humanized paragraph and an honest one both get thinner evidence at that length.

Human reads a text sentence by sentence and reports how much of it reads as machine-written or machine-edited, with one document verdict of Human, Mixed or AI. No setting turns an AI verdict into a Human one; an AI verdict needs at least 80% of the document to read as machine-written. Those are in-house measurements on held-out sets rather than an outside audit. This is an estimate from our detector. Treat a flag as a reason to look closer, not as a finding.

One question a teacher can settle without a detector is process. Drafts and revision history answer more than any percentage on either side of this, and so does a short conversation about the argument, which can teachers tell if a humanizer was used describes in practice.

Common questions

Does a humanizer remove plagiarism from a paper?

No. A humanizer rewrites wording; it does not add a citation or change where the material came from. The University of Southern California's integrity office, as of September 2026, counts work created by a generative tool and presented as a student's own as plagiarism, and paraphrasing it does not change that 7. Rewriting the sentences produces the paraphrased case the policy already names.

Why do two detectors disagree about the same humanized paragraph?

The published results come from different releases, different texts and different humanizers. Notre Dame found fewer than 4 percent of AI-labelled rewrites still flagged with Pangram 3.2 and GPTZero 3, while Pangram's report on Pangram 4 has humanized text read as AI 97.67 percent of the time 2. Pangram 4 adds a head trained to recognise humanized text 2, so a result on one release is not evidence about another. Length adds to the spread, since short text gives any detector less to read.

Do free humanizers work as well as paid ones?

Human's in-house testing finds that whether a humanized document is caught depends on the humanizer's price tier more than on which humanizer it is, stated without a figure because it is an internal measurement rather than an independent audit. A tier is not a brand, so two students naming the same tool can land on different sides of a flag. Turnitin told NBC News its list of 150 such tools includes some charging as much as 50 dollars for a subscription 6.

Can a teacher tell a humanizer was used without running a detector?

Sometimes the prose raises a question; no cited study measures how often. Synonym swaps and inserted nonsensical phrases are documented techniques 5. Odd wording is a reason to ask about process, never a finding. Drafts, version history and a conversation about the argument answer that question without any tool.

Does lightly editing a student's own writing with AI trigger a flag?

It can. In the Notre Dame cohort, light AI editing of an abstract, the authors' proxy for guideline-compliant assistance, was flagged at 38 to 80 percent depending on the detector, against under 4 percent for humanizer output 3. That is 642 abstracts of 25 to 500 words and not full papers, so it describes short passages, and on those passages the permitted edit drew the flag more often than the concealed rewrite 3.

References

  1. 1.Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense Krishna, Song, Karpinska, Wieting and Iyyer (arXiv), 2023. arxiv.orgDIPPER paraphrasing dropping DetectGPT from 70.3 percent to 4.6 percent accuracy at a fixed 1 percent false positive rate, and evading watermarking, GPTZero and OpenAI's classifier.
  2. 2.Pangram 4 Technical Report Pangram Labs and University of Maryland (arXiv), 2026. arxiv.orgPangram's own figures: humanized text read as AI-generated 97.67 percent of the time and as Mixed or AI-generated 98.83 percent; Table 14's per-system AI recall running from 92.78 percent (Commercial H) to 99.70 percent (Commercial F) across 13 commercial humanizers; and, in a separate column, the auxiliary humanizer head's per-system true positive rate of 91.52 to 99.39 percent for telling humanized AI from unmodified AI.
  3. 3.Why AI Detection Fails for Academic Integrity Karr, Khvatskii, Hua and Chawla, University of Notre Dame (arXiv), 2026. arxiv.orgScored with Pangram 3.2 and GPTZero on 642 abstracts of 25 to 500 words, with a single commercial humanizer: fewer than 4 percent of rewrites first labelled AI still flagged, light AI editing used as a proxy for guideline-compliant help flagged at 38 to 80 percent depending on the detector, and the stated sanction asymmetry.
  4. 4.America Has a Pangram Problem The Atlantic (Matteo Wong), 2026. theatlantic.comA reporter running ChatGPT and Claude output through the humanizer Walter Writes AI in May 2026 and finding that Pangram invariably called the twice-baked article human-written; also his description of the rewriting as anodyne rewording, a swapped transition clause and introduced grammatical oddities. Page read 2026-09-18.
  5. 5.What is a humanizer? Pangram Labs, 2025. pangram.comPangram Labs defining a humanizer as a tool that rewrites AI-generated text to evade detectors, and naming inserted nonsensical phrases as a documented technique.
  6. 6.To avoid accusations of AI cheating, college students are turning to AI NBC News, 2026. nbcnews.comTurnitin's list of 150 tools charging as much as 50 dollars a subscription, and Cursive's count of 43 humanizer sites with 33.9 million visits in one month.
  7. 7.Academic Integrity & Generative AI University of Southern California, Office of Academic Integrity, 2026. academicintegrity.usc.eduThe statement that material created by a generative tool and represented as a student's own work is plagiarism whether copied verbatim, near-verbatim or paraphrased. The page carries no date; read 2026-09-18.

7 sources, numbered by first appearance.

General guidance for teachers, administrators and students. What holds at one institution, on one assignment, may not transfer to another.

Human reports how much of a document reads as machine-written. It does not report a probability that a person used AI, it does not check for plagiarism, and no number it produces stands for a student's honesty. This is an estimate from our detector. Treat a flag as a reason to look closer, not as a finding.

Human

How to read a result

How a verdict is made, what the caveat means, the false-flag arithmetic, and why short and edited text is harder to read.