How to test an AI humanizer
Twenty minutes, no cost beyond the free tiers, and at the end you have your own numbers instead of somebody's marketing page. This is the method we use, written out so you can run it.
1. Fix your test texts first
Pick three pieces of text and never change them again. Ours are a 400-word argumentative essay, a 300-word product description and a 250-word technical explainer, all model-generated. Fixing them is the whole trick: if the text changes between tools, the comparison is worthless, and most published comparisons quietly do exactly that.
Use text in the same genre as your real work. A humanizer that handles marketing copy well can be poor on technical writing, because the thing it removes - repetition, uniform sentence length, connective phrases - is sometimes the thing technical writing needs.
2. Get a baseline
Before touching any humanizer, run all three texts through your detectors and write the numbers down. Without a baseline you cannot tell a good humanizer from a text the detectors were never going to flag anyway.
3. Use more than one detector
A single detector is a single opinion, and humanizer vendors know which one their tool beats. Our panel is GPTZero, Originality.ai, Copyleaks, Winston AI, Sapling, ZeroGPT - free tiers on all of them are enough for this. Record every score, not the best one.
4. Humanize, then re-run
Paste the same text, take the default settings, and put the output back through the same detectors on the same day. Note the date. You are measuring one tool on one day against detectors that will change.
5. The step everyone skips: read the output
Compare it to the original, sentence by sentence, and check three things:
- Numbers and names. Humanizers rewrite them surprisingly often. A changed figure is a broken document.
- Hedges and caveats. "May reduce" turning into "reduces" is a meaning change that reads as an improvement.
- Idiom. Synonym substitution produces phrases nobody says. "Conduct an investigation of the matter" is what a low score on readability looks like.
A tool that drops detector scores to zero and mangles two facts has not helped you. It has produced a document you now have to proofread against the original, which costs more time than writing it yourself.
6. Write down the date
Your result is true for the day you measured it. Detectors are retrained, humanizers are updated on their own schedule, and a number without a date will still look current in six months when it is badly wrong. Re-run the test before you renew any subscription.
What this method cannot tell you
- Whether a specific institution's specific detector will flag your specific document.
- Whether using the tool is permitted where you are. See academic integrity.
- Whether the vendor keeps your text. That is a reading exercise, not a testing one - see the buyer's checklist.