I recently began testing some of my articles with Substack's new AI detector. The results were unsettling.
Some articles I had written, using only a spelling and grammar checker, were judged to be substantially AI-generated. Collaborative pieces, in which the ideas and writing were mine but used AI for research, verification, and editing, were judged even more harshly. After a while, I found myself looking in the mirror for transistors behind my eyes.
Was it just me?
I decided to run a few experiments. I fed various AI detectors passages by George Orwell, Bertrand Russell and Somerset Maugham. All three were identified, at least in part, as AI-generated. So were excerpts from Tolstoy and Hemingway.
This was reassuring in one sense. I probably was not a machine. But it raised a more troubling question: What exactly were these programs detecting—and was that the thing we actually wanted to measure?
The classic writers the programs distrusted tended to share certain qualities. Their prose was clear, their arguments orderly, their sentences polished. They could begin with something particular and reveal what it meant, or reduce a complicated thought to a sentence that stayed in the reader’s mind.
Those are not defects. They are qualities generations of editors and writing teachers have tried to encourage. They are also qualities artificial intelligence has learned to imitate.
That is the central problem. A detector cannot look into the history of a sentence. It cannot know who had the idea, chose the evidence, made the judgments or struggled through the revisions. It can only decide whether the finished prose resembles patterns associated with machine-generated writing.
