WritingFor students
Humanizing AI Text by Hand Means Rewriting From Notes, Not Swapping Words
Humanizing AI text by hand means rewriting it, not swapping words. The moves that change how a paragraph reads are structural. Vary sentence length. Cut the opener that repeats down the page, drop the hedges stacked on every claim, break the three-item list habit, and then write each claim again from notes instead of from the draft on the screen. A detector reading afterwards is a receipt on the prose. A rewrite of a chatbot draft changes how it reads and nothing about where it came from.
Bill Nguyen & HumanUpdated
What Does Humanizing AI Text by Hand Involve?
It involves rewriting from notes instead of editing the sentence on the screen. A hand pass has three stages, and the middle one happens with the draft closed. Read the draft once and mark every sentence the writer could not defend out loud. Close the file. Write those claims again from memory, and only then reopen the draft and keep the facts and citations it was carrying.
What comes back is shorter and less even, and it reads as one person's.
The rewrite has to happen away from the draft because of what a detector is built to read. Pangram Labs, which sells one, trains on roughly a million human-written documents 1. For each of those it generates an AI-written mirror matching the original on as many axes as possible, and Pangram says the aim of that design is a model that classifies, in its own phrase, solely based on specific characteristics of LLM writing 1. Which characteristics those are, the page does not say.
The habits this page counts, sentence rhythm and hedging, are the ones a hand pass can reach and a synonym pass cannot. A word swap can still move a reading, as the Stanford study below shows, but the rhythm and the argument stay exactly where they were.
The second reason to close the file is slower and matters more. A paragraph rewritten from notes can be defended in an oral follow-up. A paragraph copied from a chatbot and reworded cannot, whatever any tool reports about either.
The tool version of this job is a different product with different consequences: what a humanizer does to writing.
Why Does Chasing a Detector Score Produce Worse Writing?
Because the score answers to surface form, and surface form is cheap to move in either direction. A 2023 Stanford study had ChatGPT enrich the vocabulary of human TOEFL essays. The average false positive rate across seven detectors fell from 61.22 percent to 11.77 percent 2. Then the same team had ChatGPT simplify the word choices in essays by eighth graders in the United States, and misclassification climbed from 5.19 percent to 56.65 percent 2.
In both experiments the argument stayed the same and a model changed the words.
The paraphrase literature says the same thing from the other side. A 2023 model called DIPPER dropped DetectGPT's accuracy from 70.3 percent to 4.6 percent at a fixed 1 percent false positive rate. Its authors record that the meaning of the input was not appreciably modified 3.
A number that falls that far while the argument stays put is not measuring the argument.
The prompt version of the same move costs even less. In the same Stanford study, one follow-up instruction asked the model to rewrite its own output in more literary language, and detection of machine-written college essays dropped from 100 percent to 13 percent 2.
A practical trap sits underneath the arithmetic. All seven detectors in that study flagged the same 18 of the 91 human-written TOEFL essays, and 89 of the 91 were flagged by at least one 2. Seven instruments that disagree that much 2 leave a student who edits toward one tool's reading optimising against a tool that may not be the one a course runs. That cycle has its own page: the detector and humanizer loop.
One narrow diagnostic survives all of this. A sentence-level reading points at sentences. It does not rank them, and it does not say what is wrong with them.
Which Habits Give a Machine Draft Away?
Even sentence length, a repeated opener, stacked hedges and boosters, ideas grouped in threes, and a connective at the head of every paragraph. All of them are countable on paper in about ten minutes. Counting beats rereading, because a writer rereading their own prose hears the intended sentence instead of the one on the page.
Start with the margin count: write the word count of each sentence beside it for one page. A narrow band is the tell. The hand fix is one deliberately short sentence in every paragraph.
| What to count | What to look for | The hand fix |
|---|---|---|
| Words per sentence, in the margin | A narrow band | One sentence under eight words per paragraph |
| The first two words of every sentence | One opener repeated | Start on the subject of the claim instead |
| Hedges and boosters per paragraph | Typically, generally, arguably on one side; significantly, crucially, notably on the other | Delete the word, or replace it with the figure |
| Items per list | Threes | Two, or four, or a sentence carrying none |
| Paragraph openers | However, Additionally, Moreover, Ultimately | Open on the noun the paragraph is about |
Students resist the hedge count, because hedging feels careful. At volume it reads as evasive and flattens the writer's voice.
Boosters do the same damage upward: significantly, crucially, notably, vitally. Cut both lists and see what is left of the claim. Then say the claim plainly.
The three-item habit is hard to unlearn. Threes have a rhythm, and a reader hears it long before anyone names it. Two examples and a sentence on why those two are enough reads better, and it is easier to defend.
The adjacent question, for an essay that was never machine-written in the first place, is how to stop an essay sounding like AI.
Rewrite From Notes, Not From the Sentence on the Screen
Open a blank file and write the paragraph again from what the writer knows. Only then compare the two versions. A benchmark of 6,500 texts, spanning human-written and machine-generated material along with expert-edited versions, found that expert editing evades machine-text detection, while the same material edited by another language model is still unlikely to read as human-written 4.
That result is about machine text edited afterwards. The Stanford result above is about human text a model reworded. The two point the same way: a detector reads surface form, and either kind of editor can move it. What a person rewriting from notes changes, and a model does not, is whether the writer can defend the paragraph afterwards.
The same finding carries a warning. If careful human editing can move a reading, a low reading proves nothing about who wrote the paper. Neither does a high one. A flag settles nothing on its own, and the point cuts against a rewrite offered as evidence as hard as it cuts against an accusation.
One unfairness is worth naming before choosing a route. A 2026 University of Notre Dame study found that light, guideline-compliant AI editing of abstracts was flagged more often than machine-written text run through a humanizer 5. The cohort was 642 abstracts of 25 to 500 words, so the finding describes short academic passages, not whole papers 5.
Editing within the rules draws more flags than hiding a machine draft does.
The remedy is a better policy and a disclosed draft, not a hidden one, which is why the disclosure section below is specific about wording. The chatbot-shortcut version of the same request has its own page: make this sound better.
When Is a Hand Rewrite the Wrong Answer?
When the draft was machine-written and the course does not allow it. Rewriting that draft changes how it reads and nothing else. A paper a detector now calls Human is still a paper the student did not write. The honest routes are narrower: disclose the use and cite the tool, or write the piece from notes and hand in that.
Disclosure has settled wording already.
MLA's March 2023 guidance asked writers to acknowledge functional uses of an AI tool, editing prose among them, in a note or in the text, even where nothing generated by the tool is quoted 6. MLA revised that post in August 2025 and now advises acknowledging substantive uses of AI 8, so the syllabus and the current MLA page decide what counts. The writing center at the University of Tennessee, Knoxville goes further for drafts in progress. It tells students to point out any part of a draft containing generated output, even before the citations are in 7.
An integrity meeting costs more than the sentence, and the sentence is what a student brings to it.
A worked version of that sentence sits at an AI disclosure statement, and the syllabus governs whichever way it is written.
The other wrong answer is rewriting in the dark. A writer working from notes needs to know which sentences still carry the machine draft's shape before the pass starts, and Human at human.olive.is marks them: it reads a paper sentence by sentence and reports how much of it reads as machine-written or machine-edited, with a document verdict of Human, Mixed or AI. Below 50 words it declines to answer. Between 50 and 149 words a reading is weak, because the detector has little to read.
Its published accuracy figure is an in-house measurement on held-out sets rather than an independent audit: 0 false flags in 1,928 human documents it had never seen; the statistical ceiling on that is 0.155%, measured September 2026. This is an estimate from our detector. Treat a flag as a reason to look closer, not as a finding. The reading a student takes after a rewrite is a receipt on the second draft, and it goes into any integrity meeting as a reading, never as proof.
The loop a student can run today is check, revise, check again: paste up to 1,000 words at human.olive.is, three times a day, with nothing stored and no account; rewrite the marked sentences from notes; read the same passage a second time. A longer paper goes through in parts. A writing tool in private preview reads a draft with the same detector, colouring each sentence by what the detector reads it as, keeping quotations, numbers, citations and anything the writer marks as their own untouched, and helping rewrite the rest in the writer's own words, and it never tells a person that their own writing is not theirs.
Common questions
Does rewriting AI text by hand count as cheating?
The syllabus decides, and the rewrite does not change what is being submitted. Where a course bans generative drafting, a hand-rewritten chatbot draft is still a chatbot draft. The rewrite only makes it harder to describe honestly later. Where a course allows AI with acknowledgment, MLA's 2023 guidance asked for a note covering functional uses such as editing 6, and its August 2025 revision advises acknowledging substantive uses of AI 8; the current MLA page decides what counts. Read the course policy first and decide what to write second. The reverse order is where the trouble starts.
Will a hand rewrite lower a detector score?
Sometimes, and no one can promise it. In a benchmark of 6,500 texts, human-written and machine-generated alongside expert-edited material, expert editing did evade detection, while text edited by another model did not read as human 4. That is a result on a research set, and it predicts nothing about one essay on one instrument. Two detectors disagree on the same paper often enough that editing toward a number means chasing a moving target. Rewrite for the reader who grades the paper. Any reading is a receipt on the prose, nothing more.
Which words make writing read as machine-made?
No word list settles it, and betting on one backfires. The 2023 Stanford study moved detector verdicts in both directions by having ChatGPT change vocabulary alone. Richer word choices cut false positives from 61.22 percent to 11.77 percent, and simpler ones raised misclassification from 5.19 percent to 56.65 percent 2. Hedges such as typically, generally and largely are worth cutting anyway, because they weaken the claims a marker is trying to grade. Cut them for the argument, not for the score.
Is asking a chatbot to sound more human the same thing?
No, and the measured difference runs the other way from the sales pitch. In the expert-editing benchmark, text revised by another language model remained unlikely to be recognised as human-written. Expert human revision did evade detection 4. A second machine pass also adds a second thing to disclose. That request has its own page, make this sound better, and it is worth reading before pasting a paragraph into a chat window.
What if a student's own writing keeps reading as machine-made?
That is a documented failure of the instruments, not a fault in the writing. The 2023 Stanford study suggests detectors may penalise writers with a limited range of expression; its evidence is essays by non-native English writers, and plain prose is the kind it found at risk 2. In that study 89 of the 91 TOEFL essays by real people were flagged by at least one of the seven detectors 2. Keep the draft history and save dated versions. Raise it with the instructor before submission. Rewriting to look less plain costs the clarity that made the essay readable.
References
- 1.How AI Detection Works Pangram Labs, 2026. pangram.comPangram's description of mirror prompts: for each human example an AI-generated example is produced matching the original on as many axes as possible, so the model classifies solely based on specific characteristics of LLM writing.
- 2.GPT detectors are biased against non-native English writers arXiv (Stanford University; published in Patterns), 2023. arxiv.orgThe vocabulary interventions in both directions: having ChatGPT enrich human TOEFL essays cut the average false positive rate from 61.22 percent to 11.77 percent, having ChatGPT simplify eighth-grade essays raised misclassification from 5.19 percent to 56.65 percent, and the stated explanation that detectors penalise writers with limited linguistic expressions.
- 3.Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense Krishna, Song, Karpinska, Wieting and Iyyer (arXiv), 2023. arxiv.orgDIPPER dropping DetectGPT's accuracy from 70.3 percent to 4.6 percent at a constant 1 percent false positive rate, without appreciably modifying the semantics of the input.
- 4.Beemo: Benchmark of Expert-edited Machine-generated Outputs arXiv (NAACL 2025), 2024. arxiv.orgThe 6,500-text benchmark of human-written, machine-generated and expert-edited texts, and its finding that expert-based editing evades machine-generated-text detection while model-edited texts are unlikely to be recognised as human-written.
- 5.Why AI Detection Fails for Academic Integrity Karr, Khvatskii, Hua and Chawla, University of Notre Dame (arXiv), 2026. arxiv.orgThe finding that guideline-compliant light AI editing results in a higher sanction risk than humanizer-assisted evasion.
- 6.How Do I Cite Generative AI in MLA Style? MLA Style Center, 2023. style.mla.orgMLA's instruction to acknowledge all functional uses of an AI tool, including editing prose and translating words, in a note, in the text or in another suitable location.
- 7.GenAI Tools and Writing: Information and Guidelines for Students Judith Anderson Herbert Writing Center, University of Tennessee, Knoxville, 2026. writingcenter.utk.eduThe guidance that any part of a draft containing generated output should be pointed out, even in a draft where citations have not yet been added.
- 8.Beyond Citation: Describing AI Use in Your Work MLA Style Center, 2025. style.mla.orgMLA's August 2025 post advising authors to acknowledge substantive uses of AI, published as the updated guidance the 2023 citation post now points to.
8 sources, numbered by first appearance.
General guidance for teachers, administrators and students. What holds at one institution, on one assignment, may not transfer to another.
Human reports how much of a document reads as machine-written. It does not report a probability that a person used AI, it does not check for plagiarism, and no number it produces stands for a student's honesty. This is an estimate from our detector. Treat a flag as a reason to look closer, not as a finding.