Skip to content

HumanizersFor teachers

Humanized AI Text Reads Odd More Often Than Robotic: Swapped Words, Inserted Nonsense, Stray Typos

A synonym-swap pass reads like a thesaurus accident rather than a chatbot, though a paraphrase can still carry the machine's habits. Pangram Labs, which audits humanizer tools, prints the shape: "I require to obtain my vehicle repaired" where the ordinary sentence said "I need to get my car fixed" [1]. Two more artefacts can travel with it: a nonsense phrase some humanizers drop mid-paragraph, and typos a pass can leave in the prose [1]. Each one is a reason to open the draft history. Not a finding.

Bill Nguyen & HumanUpdated

What Does Humanized AI Text Actually Look Like?

Three artefacts, and none of them reads like a chatbot on its own. First, a word sitting one step off the ordinary choice. Pangram Labs, which audits these tools, prints the pair: the plain sentence is "I need to get my car fixed because the engine is making a strange noise", and its rewrite is "I require to obtain my vehicle repaired because the motor is creating a peculiar sound" 1.

Meaning survives that. Register doesn't.

The second artefact means nothing at all. Humanizers sometimes add nonsensical phrases to a piece of writing in the hope that a detector will read the gibberish as unlikely to have come from a machine 1. The example Pangram prints is a fragment of code and the page range 18 to 23, dropped between two sound sentences 1.

The third is plain damage. A pass can leave typos and broken grammar behind, a misspelling where the draft had none 1. Surface errors beside an untouched structure are a reason to open the version history. Nothing more than that.

Picture the three together on one marked page. A paragraph carries two words nobody in the seminar would choose and one clause that resolves to nothing, and a misspelling sits inside a sentence that is otherwise doing careful comparative work. Read a full paragraph before deciding anything. A single sentence carries little of this.

The register is the tell. A mistake on its own is not.

Why Can a Humanized Paper Read Worse Than the Draft It Came From?

When it does, the rewrite has optimised for a score and spent clarity to get there. Pangram Labs says as much about its own synonym example: swapping words for synonyms damages the text and lowers its clarity and fluency, sometimes to comedic effect 1. The sentence still parses. What it has stopped doing is sounding like anybody in particular, and that is a different failure from sounding like a machine.

How much of that damage shows varies from pass to pass. A 2023 stress test of recursive paraphrasing reported the method cutting detection rates sharply while degrading text quality only slightly in many cases 3.

Price is the variable Human's own measurement found to matter. Human, tuned for college-level academic writing at human.olive.is, reports one in-house finding here, measured September 2026 on its own held-out sets rather than by an independent audit: whether a humanized document is caught depends on the humanizer's price tier more than on which humanizer it is. This is an estimate from our detector. Treat a flag as a reason to look closer, not as a finding.

Price tiers aren't hypothetical. NBC News reported that Turnitin keeps a list of 150 tools, some charging as much as $50 for a subscription, that adjust text so a detector will not flag it 8.

What a pass does to a student's sentences, clause by clause, is what-a-humanizer-does-to-writing. The artefacts here are what survives that pass and reaches the marking pile.

Which Tells Did Expert Readers Name When They Were Tested?

Vocabulary, first and by a wide margin. In an ACL 2025 study, five annotators who frequently use language models for writing tasks read 300 articles and wrote down why they judged each one. Vocabulary clues appeared in 53.1 percent of those explanations, sentence structure in 35.9 percent, originality in 23.7 percent and formatting in 15.0 percent 2.

The structural tells are concrete enough to teach. The researchers' definition of the structural clue names a high frequency of "not only ... but also" and consistently listing three items; one annotator put it as the "it's not just this, it's this" comparison, alongside listings of specifically three ideas 2.

One explanation described writing filled with the same language it used to describe everything: inspirational, stunning, essential 2. The researchers' own definition of the originality clue reads like a line off a marking rubric: machine writing is straightforward, safe, and lacking in surprises or humor 2.

Their judgement held up under a humanizer.

The researchers built one themselves by prompting o1-Pro with instructions drawn from the annotators' earlier comments, then put a batch of 60 articles, 30 of them humanized, in front of the same panel. The expert majority vote was perfect on all 60 2.

Across the study as a whole, that same majority vote missed 1 of 300 articles 2.

Whether the skill survives a real marking pile, with a name attached to every paper, is can-teachers-tell-humanizer-used.

Why Is Reading Odd Not Evidence of Anything?

Because detectors misread plain prose in the same direction. A 2023 Stanford study simplified the word choices in essays written by US eighth graders and watched the average misclassification rate across the detectors it tested climb from 5.19 percent to 56.65 percent 4. Nothing about the authorship changed. The vocabulary did.

The reverse held too. Rewriting human TOEFL essays with richer vocabulary cut the average false positive rate from 61.22 percent to 11.77 percent 4. The authors read that as a sign the detectors may be penalising a limited range of expression rather than reading authorship 4.

Bloomberg Businessweek put names to who pays for that. Its reporting, drawn from students, academics and AI developers, says the students most exposed to a wrong accusation are likely those who write in a more generic manner, including neurodivergent students and students working in a second language 7. A humanizer vendor makes the same point from the other side. Undetectable AI's ethics page states that a single AI text detection score alone is not enough to rule out cheating or deception 6.

The incentive runs backwards, too. A 2026 Notre Dame team took 642 published abstracts of 25 to 500 words, had a model rewrite each one, and ran the rewrites through a humanizer. Light, guideline-compliant edits to the same abstracts were flagged 38 to 80 percent of the time, depending on the detector and the field, while fewer than 4 percent of the humanized rewrites were 5. That is the finding that should worry a department: guideline-compliant AI editing carried a higher sanction risk than humanizer-assisted evasion 5. The authors say classroom writing may behave differently 5.

How that lands on a cohort writing in a second language is ai-detectors-non-native-english-writers-policy.

Check the Paragraph Against the Student's Own Drafts

Open the version history before opening a detector. A paragraph that reads wrong is a question about how it was written, and the draft record can show whether the passage grew across several saves or arrived whole in one. Odd wording earns that look. It earns nothing further.

Read the whole document rather than the sentence that snagged. The Notre Dame cohort ran as short as 25 words 5, and Human itself reports a weak reading under 150 words and none under 50.

Compare the paper against something the student already handed in. A graded essay from week three costs nothing and outlives every version number a detector ships.

The order those checks belong in is tell-if-student-used-chatgpt.

If a reading would help after that, take one. Human reads a document sentence by sentence and reports how much of it reads as machine-written or machine-edited, with a document verdict of Human, Mixed or AI. This is an estimate from our detector. Treat a flag as a reason to look closer, not as a finding. Paste as much of the paper as fits, up to 1,000 words. A lone short paragraph gives the detector little to read, and under 50 words Human declines rather than guessing.

Common questions

Does humanized text always contain nonsense phrases?

No. Inserted nonsense is one technique among several that Pangram Labs documents, alongside synonym swapping and general quality damage 1. A 2023 stress test of recursive paraphrasing reported the method cutting detection rates sharply while degrading text quality only slightly in many cases 3. A careful rewrite can leave a paper that reads well. Absent artefacts don't show that nothing happened, in the same way that present ones don't show that something did.

Can an odd-sounding paper simply be a second-language writer's work?

Yes, and that is the most expensive mistake available here. A 2023 Stanford study found that rewriting human TOEFL essays with richer vocabulary cut the average false positive rate from 61.22 percent to 11.77 percent, which the authors read as a sign the detectors may be penalising limited linguistic range rather than authorship 4. Bloomberg Businessweek, drawing on students, academics and AI developers, reported that the students most exposed to a wrong accusation are likely those who write in a more generic manner, including neurodivergent students and students working in a second language 7. Odd register is a reason to ask, never a reason to conclude.

What did expert readers say gave machine writing away?

Vocabulary, mostly. In an ACL 2025 study, clues about word choice appeared in 53.1 percent of the explanations five expert annotators wrote, ahead of sentence structure at 35.9 percent, originality at 23.7 percent and formatting at 15.0 percent 2. The researchers' definition of the structural clue names a high frequency of "not only ... but also" and consistently listing three items. Their definition of the originality clue calls machine writing straightforward, safe, and lacking in surprises or humor 2.

Why would a humanized essay contain typos?

Because the pass damaged the surface. Pangram Labs prints a degraded rewrite carrying misspellings and grammatical errors 1. A humanized paper can arrive with typos that no earlier draft from the same student shows, and that gap is the question worth asking. Surface errors beside an untouched structure are a reason to open the version history rather than a rule about the person who handed the paper in.

Is an odd paragraph enough to open a misconduct case?

No. The sources that address it point the other way, including a vendor selling evasion: Undetectable AI's own ethics page states that a single AI text detection score alone is not enough to rule out cheating or deception 6. A reading describes text. A case is about a person. Bring evidence about that person: a draft history, a comparison with earlier work, a conversation in which the student walks through the argument.

How many humanizer tools are a teacher actually up against?

Enough that naming one is beside the point. NBC News reported that Turnitin keeps a list of 150 tools, some charging as much as $50 for a subscription, that adjust text so a detector will not flag it 8. An academic integrity company counted 43 humanizer sites drawing a combined 33.9 million visits in a single month 8. The artefacts described here come from the technique rather than the brand, which is why they outlast any particular list.

References

  1. 1.What is a humanizer? Pangram Labs, 2025. pangram.comThe 27 January 2025 definition post: the synonym-swap pair about a car and its engine, the statement that swapping words damages clarity and fluency sometimes to comedic effect, the nonsensical-phrase technique with its code-fragment example, and the degraded rewrite carrying misspellings and grammatical errors.
  2. 2.People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text Russell, Karpinska and Iyyer (arXiv / ACL 2025), 2025. arxiv.orgFive expert annotators over 300 articles: the explanation categories at 53.1 percent vocabulary, 35.9 percent sentence structure, 23.7 percent originality and 15.0 percent formatting; the named patterns; the o1-Pro humanizer built from the annotators' own comments, with a perfect expert majority vote on a 60-article batch of which 30 were humanized and 1 of 300 missed overall.
  3. 3.Can AI-Generated Text be Reliably Detected? Sadasivan, Kumar, Balasubramanian, Wang and Feizi (arXiv / Transactions on Machine Learning Research), 2023. arxiv.orgThe study's own summary of the trade-off: recursive paraphrasing significantly reduces detection rates while only slightly degrading text quality in many cases.
  4. 4.GPT detectors are biased against non-native English writers arXiv (Stanford University; published in Patterns), 2023. arxiv.orgSimplifying the word choices in US eighth-grade essays raised the average misclassification rate from 5.19 percent to 56.65 percent; enriching the vocabulary of human TOEFL essays cut the average false positive rate from 61.22 percent to 11.77 percent.
  5. 5.Why AI Detection Fails for Academic Integrity Karr, Khvatskii, Hua and Chawla, University of Notre Dame (arXiv), 2026. arxiv.orgThe cohort of 642 published abstracts of 25 to 500 words, fewer than 4 percent of the humanized machine rewrites still flagged against 38 to 80 percent of the light abstract-only edits, and the finding that honest AI editing carries a higher sanction risk than humanizer-assisted evasion.
  6. 6.Ethics Undetectable AI, 2026. undetectable.aiThe humanizer vendor's own statement that a single AI text detection score alone is not enough to rule out cheating or deception.
  7. 7.AI Detectors Falsely Accuse Students of Cheating-With Big Consequences Bloomberg Businessweek, 2024. cs.utexas.eduThe reporting that students most susceptible to inaccurate accusations are those who write in a more generic manner, including neurodivergent students and students who speak English as a second language.
  8. 8.To avoid accusations of AI cheating, college students are turning to AI NBC News, 2026. nbcnews.comTurnitin's list of 150 tools charging as much as $50 for a subscription, and Cursive's count of 43 humanizer sites drawing 33.9 million visits in one month.

8 sources, numbered by first appearance.

General guidance for teachers, administrators and students. What holds at one institution, on one assignment, may not transfer to another.

Human reports how much of a document reads as machine-written. It does not report a probability that a person used AI, it does not check for plagiarism, and no number it produces stands for a student's honesty. This is an estimate from our detector. Treat a flag as a reason to look closer, not as a finding.

Human

Check a paper

Paste a paper and see which sentences read as machine-written, with a verdict of Human, Mixed or AI. Nothing is stored and no account is needed.