Do not index
How do you prove your writing is human when the detector says it is not? A Substack writer just demonstrated why that is the wrong fight to pick. Dr Sam Illingworth wrote a post entirely himself, ran it through Pangram, the AI detector Substack now partners with, and it scored 100% AI. Then he ran the same post through a free humanizer tool, changed nothing else, and the new version scored 100% human. His own summary: "According to Substack and Pangram the version I am actually publishing scored 100% AI. The same argument, same ideas, one free tool in between." So here is my answer. Stop trying to win the detection game. It is broken in both directions, and the only test that still works is specificity.
The post about the experiment pulled 545 likes and 284 comments within days, which tells you how many writers are living some version of this problem. It is not an edge case. The same piece cites a Stanford study that found detectors flagged 61% of essays by non-native English speakers as AI. Careful, polished, second-language prose is exactly what the statistical models mistake for machine output.
I am writing this for ghostwriters charging $5k to $30k per month whose clients have started asking whether the work will get flagged, for agency owners between $200k and $2M in revenue shipping content across dozens of accounts, and for founders writing their own material who have already been accused once and felt how little recourse exists.
This is not for anyone hunting a tool to slip AI slop past detectors. That arms race has no finish line and both sides of it are selling you ammunition. Skip this if your content problem is volume rather than trust. And if your differentiator is still the phrase "100% human-written," this article will not change your model, but the next false positive will.
Why AI detectors fail in both directions
Detectors do not measure origin. They measure statistical patterns, and patterns can be faked in both directions. Professional prose that has been edited toward clarity, consistent rhythm, and clean structure starts to look machine-generated, which is how a fully human post scores 100% AI. Meanwhile a humanizer tool scrambles the surface pattern without touching the substance, which is how laundered output scores 100% human. The net effect is the worst possible combination for anyone who sells writing. Real writers get falsely accused with no appeal process, while actual slop passes with a receipt claiming it is human. A verdict that punishes the careful and clears the lazy is not a quality system. It is a lottery with your reputation as the stake.
Once you accept that, the question changes. Clients and readers were never actually asking "did a model type this." They were asking "is there a real person with real experience behind this." That question has an answer no detector can fake, and it is the basis of the only test I trust.
The Specificity Sort
I run every draft through what I call the Specificity Sort. Take the piece line by line and sort each claim into two piles. Pile one holds the lines a smart stranger could have written from a search, general truths, safe observations, advice with no fingerprints on it. Pile two holds the lines that could only come from someone who was in the room, a number from a live project, a client objection quoted from an actual call, a decision with a date and a cost attached. The sort is the only detector that matters, because readers run it instinctively even when the software fails. Nobody questions whether a post was typed by a model when it contains details no model could know.
The target ratio is the whole game. A post that is 80% pile one gets skimmed, doubted, and forgotten regardless of what any detector says about it. Flip the ratio and the authorship question evaporates. In client work this changes the brief entirely. The scarce input is not writing time, it is extracted experience, which means the interview matters more than the draft and the follow-up question matters more than the thesaurus.
For agencies and ghostwriters, the sort belongs inside your delivery process, not just your taste. Every piece that ships under a client's name either builds their credibility reserve or quietly drains it, and a flagging incident drains it fast. This is the same logic behind the quality control system that prevents client churn, because trust failures compound long before they show up in a cancellation email.
The strategic implication reaches past this news cycle. Detection will stay broken, accusation will stay cheap, and trust will keep migrating from tools to track records. The writers and agencies that hold value through that shift are the ones with an archive full of pile-two material, work so specific it authenticates itself. The real business risk was never being flagged by Pangram. It is building a body of work that passes every detector and still sounds like nobody lived it.
