Blog / Research
What changes when AI text is humanized: 138 pairs, measured
Every study of machine-written prose compares human text to AI text. There is a third category almost nobody has measured: the same text before and after a rewrite. We measured 138 pairs of ours, and published the cases where it barely moved.
Read this first, because it changes what the numbers mean. The texts below are not our customers' texts. They are a corpus we wrote ourselves, put through the same production endpoint a paying user hits. We did it that way because our own Terms do not permit the other version — the detail is in the method, and it is the second finding of this piece, not a disclaimer at the bottom.
The one question nobody has measured
Every study of machine-written prose compares two categories. Human text on one side, AI text on the other — em-dash counts, burstiness, perplexity, the vocabulary that models over-reach for. There are dozens of these. They are the basis of every detector pitch and every "how to spot AI writing" listicle.
There is a third category and almost no public measurement of it: the humanized pair — one text, before and after a rewrite that was asked to make it read like a person wrote it. Not AI vs human. AI vs the-same-AI-text-rewritten. Nobody outside the small set of companies that run these rewrites holds that pair, and as far as we can find, none of them has published the distributions.
So we measured ours. 138 English before/after pairs across two corpora, scored with the same word lists and thresholds that power our public AI Tells Scanner — no new regexes invented for the occasion.
Three things are worth knowing before the tables:
- The median AI-tell density goes from roughly 12–14 events per 1,000 words to 0.
- For about a third of the texts there was nothing there to move in the first place.
- Sentence rhythm gets burstier in roughly two thirds of texts and flatter in roughly three in ten — a failure direction large enough that an average would hide it.
No AI detector was run. Not one. What that excludes is spelled out below.
Method
N. 138 English pairs, in two corpora scored separately because they disagree on magnitude and we would rather publish the disagreement than average it away.
| Corpus A | Corpus B | |
|---|---|---|
| Rewritten | 2026-09-14 | 2026-09-24 |
| Records in file | 76 | 80 |
| Excluded — not a completed rewrite | 0 | 8 (rejected by our own quality rails) |
| Exact-duplicate inputs removed | 0 | 0 |
| Pairs | 76 | 72 |
| English pairs (headline set) | 66 | 72 |
| Language split | en 66 · es 5 · fr 5 | en 72 |
| Median input length (words, all pairs) | 77.5 | 49.5 |
| Length range (words, all pairs) | 18 – 861 | 26 – 247 |
Window. Both corpora were rewritten in September 2026, ten days apart, by the production endpoint at intensity=medium, multi-pass. The model and pipeline config may differ between the two runs and from today's production. That is a real limitation, not a hedge.
Exclusions, all of them. Eight corpus-B records never completed a rewrite — they were refused by our own quality rails, so they are out of the headline. (A 10% rejection rate on a corpus we wrote is itself a product signal; we are not filing it as data cleaning.) Four other bench runs were rejected outright because output == input on every row — input-echo artefacts, not rewrites. Two further runs had real pairs but ran on a local build rather than production, so they are out of scope. The ten Spanish and French pairs are excluded from every headline, because every slop and cliché list in the scanner is English: they score 0.0 both before and after, which means "not measured," not "clean."
Aggregate only. No text is reproduced anywhere in this piece or in the published raw data. No per-item or per-account figures. Every number is a distribution — median, interquartile range, and the share of pairs moving each direction.
Why the corpus is synthetic, verbatim. Before computing a single statistic we read our own live policies. The decisive line, from Terms §6:
"We process your text solely to provide the humanization service and do not use your content to train our models."
solely is an exhaustive-purpose limitation, and computing style statistics for a public report is not providing the service to the person who submitted the text. The only research-shaped permission in the privacy policy is for "aggregate, anonymized usage statistics (never individual content)" — and our method requires reading individual content to produce each per-pair measurement, even though only aggregates get published. So the verdict was recorded as NOT PERMITTED → synthetic corpus used, and from the source report:
"This is not first-party user data. It is our own synthetic corpus run through the production endpoint, used because our own Terms and Privacy Policy do not permit the first-party version. […] Real user text differs in genre mix, length, language, prompt style and messiness. Treat every figure as 'what the production model does to text we wrote,' not 'what it does to what customers submit.'"
No production query against the stored input or output columns was run at any point. The thing this substitution costs us is specific and worth naming: there is no api / playground / UI split in what follows, because that dimension only exists in the table we did not query.
What changed
Rates are per 1,000 words. down / flat / up is the share of pairs moving that direction, with a small dead band (0.1 per 1,000 words; 0.01 for the CV) so that 4.01 → 4.00 counts as flat rather than as an improvement.
Composite AI-tell density — the one that moves
The composite is the plain sum of four scanner rate checks: slop vocabulary, cliché phrases, em dashes, three-item lists. It is our own construction, unweighted and untrained — a different weighting would move the failure shares later in this piece.
| median in | IQR in | median out | IQR out | Δ median | down | flat | up | |
|---|---|---|---|---|---|---|---|---|
| Corpus A (n=66) | 12.26 | 0 – 16.05 | 0 | 0 – 6.51 | −4.51 | 68.2% | 30.3% | 1.5% |
| Corpus B (n=72) | 13.68 | 0 – 44.40 | 0 | 0 – 17.63 | −2.09 | 51.4% | 40.3% | 8.3% |
Read the IQRs, not just the medians. The lower quartile is 0 in both corpora going in — a median that moves 12.26 → 0 is partly a statement about how concentrated the tells were, and not only about how hard the rewrite worked.
Cliché phrases — the single most-moved measure in the study
| median in | IQR in | median out | IQR out | Δ median | down | flat | up | |
|---|---|---|---|---|---|---|---|---|
| Corpus A (n=66) | 2.15 | 0 – 11.88 | 0 | 0 – 0 | −1.60 | 59.1% | 40.9% | 0.0% |
| Corpus B (n=72) | 0 | 0 – 1.15 | 0 | 0 – 0 | 0 | 22.2% | 73.6% | 4.2% |
59.1% of corpus-A pairs shed cliché density and none gained any. This is the cleanest directional result anywhere in the data, and it is also the narrowest: corpus B's inputs barely contained clichés, so there was nothing to remove.
Slop vocabulary
| median in | IQR in | median out | IQR out | down | flat | up | |
|---|---|---|---|---|---|---|---|
| Corpus A (n=66) | 0 | 0 – 0 | 0 | 0 – 0 | 18.2% | 78.8% | 3.0% |
| Corpus B (n=72) | 0 | 0 – 30.80 | 0 | 0 – 0 | 33.3% | 61.1% | 5.6% |
Corpus B's upper quartile falls from 30.80 to 0 — the slop reduction is real, but it lives in the upper tail of texts that had slop to lose: 33.3% of pairs moved down and 61.1% had nothing to move.
Sentence rhythm (coefficient of variation) — up is the intended direction
| median in | IQR in | median out | IQR out | Δ median | up | flat | down | |
|---|---|---|---|---|---|---|---|---|
| Corpus A (n=39) | 0.40 | 0.29 – 0.52 | 0.47 | 0.36 – 0.55 | +0.04 | 69.2% | 5.1% | 25.6% |
| Corpus B (n=30) | 0.33 | 0.24 – 0.43 | 0.41 | 0.31 – 0.47 | +0.03 | 60.0% | 10.0% | 30.0% |
Note the n. This is measured on 39 of 66 and 30 of 72 pairs, because the scanner refuses to judge rhythm below five sentences. The CV figures describe the longer half of the corpus and nothing else. The down column gets its own section below.
Length — humanized text comes back slightly longer
| median word count in | median out | median length Δ | IQR of Δ | longer | shorter | |
|---|---|---|---|---|---|---|
| Corpus A (n=66) | 74 | 78 | +3.23% | −3.49% – +10.02% | 62.1% | 33.3% |
| Corpus B (n=72) | 49.5 | 55 | +5.03% | −2.00% – +14.27% | 68.1% | 30.6% |
If you are budgeting word count, API spend, or a fixed-length slot, plan for the rewrite to add a few percent. Mean sentence length barely moves either way (A: 17.05 → 17.80 words; B: 15.50 → 16.25), with pairs split roughly in half on direction.
Paragraph structure — never touched, 138 of 138
| pairs | paragraph count changed | |
|---|---|---|
| Corpus A | 66 | 0 |
| Corpus B | 72 | 0 |
A clean negative result and the most practically useful one in the set: the rewrite works inside the shape you give it. Opener repetition, measurable on the 18 corpus-A pairs with three or more paragraphs, also did not move (2 in, 2 out); corpus B has no multi-paragraph items, so it is unmeasurable there.
What barely changed, or got worse
Same prominence as the section above, because these are results.
| Corpus A | Corpus B | |
|---|---|---|
| English pairs scored | 66 | 72 |
| Inputs with a composite of 0 — nothing for these checks to fix | 20 (30.3%) | 28 (38.9%) |
| Pairs where a % change in the composite is even definable | 46 | 44 |
| Barely moved (|Δ| / composite in < 5%) | 2 of 46 — 4.3% | 5 of 44 — 11.4% |
| Got worse (composite rose) | 1 of 66 — 1.5% | 7 of 72 — 9.7% |
| Zero-tell inputs that acquired tells | 1 | 2 |
| Got less bursty (sentence rhythm flattened) | 11 of 39 — 28.2% | 9 of 30 — 30.0% |
The honest headline is the range, not the friendlier end: the statistics got worse for 1.5% to 9.7% of texts, and barely moved for another 4% to 11% — and for about three in ten, the rewrite flattened the sentence rhythm it was supposed to vary.
Then there is the result that matters most and flatters us least.
| Corpus A | Corpus B | |
|---|---|---|
| Texts the scanner flagged on ≥1 pattern, before | 24 | 27 |
| Texts the scanner still flags on ≥1 pattern, after | 19 | 18 |
| Cleared | 5 of 24 | 9 of 27 |
On our own scanner's conservative thresholds, most text that read as machine-written still reads that way afterwards. The density improves for the majority of pairs; the verdict usually does not flip. Those two sentences are both true at the same time, and any version of this report that published only the first one would be dishonest.
Two more measures we are reporting as untested rather than unchanged, because the corpus did not contain enough of them to test:
- Em dashes. 98.5% of corpus-A pairs and 90.3% of corpus-B pairs had no em-dash change, because there were almost no em dashes going in. We cannot claim the model removes them. The most-repeated tell in the entire genre, and we have nothing on it.
- Contractions. 86.4% and 65.3% flat, same reason. We cannot claim the model adds them.
What this does and does not say
No AI detector was run. Not GPTZero, not Originality, not Turnitin, nothing. Every number above is a style measurement taken with our own published word lists.
Style change is not a detection outcome. The two are not known to be the same axis and we did not test the link. Nothing in this report supports a claim about detection, "bypassing," or "undetectability" — if you arrived looking for evidence that rewriting text defeats a detector, this study does not contain it, and we would rather say so plainly than let the composite table imply it.
This is not user data. See the method. It is our own corpus through the production endpoint, because our policies do not permit the first-party version. Real submissions differ in genre, length, language, prompt style and messiness.
n is small — 66 and 72 English pairs — and the two corpora disagree on magnitude. Where they disagree we have published both.
The composite is ours. Nobody else uses it. It is a plain sum of four published per-1,000-word rates, and a different weighting would move the "barely moved" and "got worse" shares. The inputs are listed in the appendix so you can re-weight them.
Non-English is not measured at all.
How to use it
If you write. Three findings are directly actionable. Clichés are what a rewrite reliably removes, so they are also the thing worth removing yourself first. Paragraph structure survives untouched, so the shape you hand over is the shape you get back. And output runs 3–5% longer at the median — budget for it. ToHuman is the rewrite tool these pairs came from, and the honest scope of what it rewrites is prose: body copy, email, documentation, and for anything code-shaped, the comments, docstrings, READMEs, commit messages and PR descriptions around the code — never the code itself. We measured that boundary separately and wrote up where it holds and where it breaks in can you humanize AI-generated code.
If you edit. The verdict table is the one to keep. A tool can cut measured tell density by most of its range and still leave text that trips the same pattern checks — 5 of 24 and 9 of 27 cleared. Treat a rewrite as a first pass, not a finish line, and re-read for the specific thing you were worried about. Our comparison of the tools in this category, including what each one actually claims, is at best AI humanizer tools 2026; the AI Tells Scanner runs the same checks used here, in your browser, on your own text.
If you research this. The pair corpus is the missing third category in the AI-vs-human literature and we would like more of it to exist. The scanner source (public/ai_tells_scanner.js) is public, so the measurement is reproducible against any corpus of pairs: 90 slop patterns, 196 cliché phrases, 4 cliché regexes, plus em-dash, sentence-uniformity, triad and opener-repetition checks with fixed thresholds. If you are generating pairs at scale programmatically, the humanizer API is the same endpoint these runs used. Two gaps are cheap to close and we would welcome someone closing them first: a testset that actually contains em dashes and contractions, and a detector arm on the same pairs.
Credit line. If you cite these numbers:
ToHuman, "What changes when AI text is humanized: 138 pairs, measured" (October 2026).
https://tohuman.io/blog/what-changes-when-ai-text-is-humanized-2026
Corrections and replications are welcome — including ones that contradict this.
Method appendix
Scoring code. Every AI-tell measure comes from the live AI Tells Scanner (public/ai_tells_scanner.js), required unmodified into Node via its module.exports. The published numbers therefore come from exactly the code that powers /ai-tells-scanner: 90 slop-vocabulary patterns, 196 cliché phrases, 4 cliché regex patterns, and the scanner's own em-dash, sentence-uniformity, triad and opener-repetition checks, with its own thresholds and its own sentence, paragraph and word splitters. No word list or regex was re-invented for this report.
Composite. slop_1k + cliche_1k + emdash_1k + triads_1k — the plain sum of the scanner's four rate-based checks in events per 1,000 words. Unweighted, untrained. The scanner's other two checks are excluded deliberately: sentence-length CV is a ratio that should move up, and opener repetition is a share of paragraphs. Both are reported separately above.
Contraction rate. Locally defined: an allow-list of roughly 60 contracted forms (don't, it's, we're, …) per 1,000 words. A generic \w+'\w+ pattern would score every possessive as a contraction and would inflate formal academic prose most, so possessive 's is excluded.
Direction shares. Dead band of 0.1 per 1,000 words for rates, 0.01 for the CV. "Barely moved" = |Δcomposite| / composite_in < 5%, among pairs with composite_in > 0. "Got worse" = composite_out > composite_in.
Burstiness n. The scanner will not judge rhythm under five sentences (THRESHOLDS.minSentences), so the CV rows cover 39 of 66 (A) and 30 of 72 (B) pairs.
Genre mix. Corpus A: academic 34, blog 16, multilingual 10, personal 9, business 7; tiered 56 paragraph-length / 20 document-length. Corpus B: long-control 20, blog-short 11, email-short 11, business-note 8, essay-para 7, human-sourced controls 7, other 8.
Source. 100% bench-synthetic in both corpora, rewritten by the production RunPod endpoint. Full statistics report and machine-readable output: marketing/reports/humanization-corpus-stats-2026-10.md and its .raw.json.