Writing Quality After AI
The opportunity to write better is right in front of us
As you may have heard, Substack shipped an AI detector last week. If it flags you, you’re not warned. You find out, and then you can file an appeal. Pangram estimates how much of a post, Note, reply, or comment was written by a human and how much was written with artificial intelligence.
Almost immediately, the predictable debate began (again):
Is the detector accurate?
Will human writers be falsely accused?
Is undisclosed assistance dishonest?
What about translation, accessibility tools, grammar correction, dictation, or a sentence rewritten after a conversation with a machine?
These are all legitimate questions. None of them touches the most important one: “why are we using authorship as a proxy for quality in the first place?”
My friend Limited Edition Jonathan wrote a fantastic piece about this on Substack (of course!). And I’m livestreaming with him there later today for this week’s Signals & Subtractions. Subscribers are notified, subscriptions are free.
The comfort of a binary
“Human or AI?” is a seductively simple framing. Two categories, a percentage, and the glow of certainty. Quality is harder. It requires us to say what a piece of communication is supposed to accomplish, what evidence would show that it accomplished it, and what kinds of failure matter.
A human can write thousands of empty words. A model can help compress a difficult idea into a paragraph that the reader finally understands. A human can fabricate a source. A model can surface a contradiction. A machine can generate polished nonsense. A person can write very sincere nonsense. Origin doesn’t settle the value question.
This doesn’t mean provenance is irrelevant. Readers may care whether a writer personally formed the sentences, just as they may care how any valued object was made. A catalogue can tell you that an apple is a Northern Spy, a grower can tell you that it is organic, and a photograph can show you its surface. None tells you whether it is crisp, bruised, sweet, or fit for a pie. Provenance can still be part of the value proposition. It can label labor, method, responsibility, and relationship.
But provenance and quality are two different variables.
Substack’s own interface demonstrates the distinction. Its Pangram integration assigns a statistical estimate to the text. Its separate “How I make this” field lets a publisher describe the process. One is classification. The other is testimony. Neither tells the reader whether the work is any good. The detector asks where the sentences probably came from. The reader’s actual problem is whether they deserve attention and meet expectations.
We never finished defining writing
Writing is old, familiar, and so deeply embedded that we rarely ask what it is for. We teach its forms, police its syntax, and sort it into poetry, persuasion, instruction, journalism, fiction, definition, criticism, and correspondence. But these categories are historical interfaces, not natural partitions. A poem can instruct. A definition can persuade. A joke can carry a theory. A technical specification can shame its implementer. A novel can provide more usable political knowledge than a policy paper. The genres are useful because human attention, institutions, and markets require shelves. The shelves should not be mistaken for the world.
Beneath them, writing is a technology for producing effects in minds across distance and time. It can preserve a claim, transfer a model, coordinate action, alter emotion, expose a relationship, or make an experience available for reconstruction. Usually it does several at once.
The quality of writing is therefore not one property. It is a vector.
An apple can be crisp but sour, beautiful but mealy, excellent for pie and disappointing eaten from the hand. A piece of writing may be accurate but incomprehensible. Clear but trivial. Dense with concepts but impossible to retain. Memorable but false. Elegant but evasive. Useful to an expert and disastrous for a novice. Responsible in its uncertainty but weak in its evidence. No single score can represent all of this without hiding value judgments inside its weighting.
But these dimensions aren’t just matters of subjective taste. Many are measurable.
Measurement does not remove judgment. It makes the observable dimensions and the values used to weight them available for argument.
From applause to witnessable quality
Historically, “good writing” has often collapsed two different statements:
“I like this.” That’s taste. You have every right to your opinion, and you don’t have to explain it. Assuming you really do like it, it’s undeniably true.
“This communication demonstrably does something well.” Now this... this can be examined!
If an instructional text is meant to improve performance, test whether readers can perform the task. If an explanation is meant to produce understanding, measure what readers correctly infer afterward. If an argument is meant to justify a conclusion, identify its claims, evidence, warrants, counterarguments, and unsupported leaps. If a report is meant to inform a decision, record whether its factual claims are traceable and its uncertainty is calibrated. If a definition is meant to distinguish cases, test its boundary conditions.
These are external effects that another person or an agentic system can witness.
The internal structure of a text also yields useful signals. Researchers measure propositional idea density: roughly, the number of assertions, relations, or qualifying ideas per unit of language. Discourse theories map how statements support, contrast with, elaborate, or condition one another. Cohesion systems examine whether references resolve and concepts recur coherently. Argument-quality research separates relevance, sufficiency, acceptability, and cogency. Plain-language standards evaluate whether readers can find, understand, and use what they need.
None of these measures is “quality.” Together they begin to form an assay.
Consider a quality profile with dimensions grouped into four families:
Text-internal: conceptual yield, semantic compression, coherence and entailment
Source relation: claim traceability, uncertainty calibration, novelty relative to a declared context
Reader effect: reader effort, inferential usefulness, retention and transfer
Fit and use: actionability, adaptation to audience, resistance to misinterpretation
Raw density can reward jargon, compression, and unreadability. A better measure might ask how many recoverable concepts are transferred per unit of reader effort, then identify the means of transfer: explanation, analogy, comparison, narrative, humor, emotional appeal, demonstration, authority, repetition, or social pressure.
That is no longer a style detector. It is an account of communicative work.
Writing has a new reader
For most of writing’s history, its intended final destination was a human nervous system. That’s no longer true. Large language models encounter the world principally through representational inputs, with text as their most mature and interoperable medium.
A model does not experience writing as a person does. It doesn’t move its eyes across a page, hear a remembered parental voice, become tired after one sitting, or feel a metaphor land in the body. Even multimodal models generally bring what they perceive into structures that can be related to language. Text has become the lingua franca of models, tools, agents, interfaces, specifications, and protocols.
Humans invented writing as an external memory and coordination technology. Machines, arriving late to the party, have become voracious consumers of the textual environment. We built writing around the constraints of human hands, eyes, institutions, and cultures. Now another kind of reader inhabits it at scale. We had to learn to read, but machines had to read to learn.
The inhuman reader makes some qualities newly visible. Models can compare definitions across thousands of documents, identify unsupported transitions, construct claim graphs, test whether instructions determine a unique action, and estimate how much semantic structure survives paraphrase. They can instrument aspects of quality that were previously too expensive to inspect everywhere.
They also introduce new failure modes. A model grader may prefer prose that resembles its own conventions. It may confuse fluency with truth or familiar argument forms with sound ones. Any serious quality system must therefore produce inspectable evidence, not a mysterious verdict.
The right output is not “82 percent good.” It is a set of claims that can be challenged, such as:
these are the propositions detected
these are their dependencies
these lack support
these terms shift meaning
these instructions produce divergent actions
these readers failed to recover this concept
these sources do or do not entail the claims attributed to them
Quality becomes witnessable when the measurement leaves a trail.
The big opportunity hidden inside the Pangram debate
The Substack controversy is more than another argument about AI detection. It is a moment when a platform is teaching millions of readers to look for a machine-generated score beside a piece of writing.
That interface could normalize suspicion about origin. Or it could become the first crude step toward something more valuable: routine inspection of communicative quality.
Imagine that instead of asking only “Did AI write this?”, a reader could ask:
What are the principal claims, and what evidence supports them?
Which claims are factual, inferential, normative, or rhetorical?
What concepts are introduced, and how efficiently are they made recoverable?
Where does the piece rely on analogy, humor, authority, identity, or emotional pressure?
What prior knowledge does it assume?
What remains ambiguous, and what would falsify its conclusions?
How faithfully could another reader or model reconstruct its meaning?
What did it add that was not already present in its sources?
No answer would substitute for reading. The point would be to make the work of the text legible.
This would also help people become better writers, not by teaching them to appear more human to a classifier, but by showing where communication succeeds or fails. A writer could discover that a paragraph contains six claims and one support relation, that an analogy carries more argumentative weight than the evidence, or that readers retain the anecdote but lose the mechanism.
Such feedback would not flatten style. Properly designed, it would separate style from function. A sentence could remain eccentric, lyrical, culturally specific, or funny while its communicative effects became more observable.
The decisive design principle is that no agent, platform, or institution should stand in as the source of quality. Humans and models can both produce evidence about it. Neither is a proxy for it.
Beyond the authorship panic
The internet’s immediate problem is not that machines can arrange words. It is that the cost of producing plausible language has collapsed while our methods for evaluating meaning remain primitive.
Authorship detection responds by trying to restore scarcity. It marks some language as costly, personal, and human. That may be useful for communities whose promise is a relationship with a particular person. But scarcity is not quality. Labor is not truth. Personality is not coherence. Humanity is not a guarantee.
The deeper opportunity is a culture in which communication can be evaluated by what it contains, what it does, and what others can verify about it. An origin label can name the variety. It cannot tell us whether the fruit is worth eating.
Pangram asks whether a text resembles the statistical traces of a human author. The question for the next generation of tools is whether the text carries meaning well.
That question is harder. It is also finally the right one.
Meta notes
Obvious first question to the AI-assisted label up top on this post: “How was this article written?” I would like my answer to always be: “Very well, thank you!”
But for this particular case, I’ll pull back the curtain all the way:
It started with me dictating into Wispr Flow into Perplexity. I dictated full passages and asked for research in return that sounded like that. I never assume my ideas are my own, and I’m always suspicious when I can’t find any prior art. I was also wrestling with some of the central concepts about what language is for LLMs and asking for pushback around how to express that.
I took the Perplexity research and brought it to Claude Code. Why? Because I hadn’t fleshed out the idea fully yet, I knew that my context layer, which operates very well in Claude Code, had most every blog post that I’ve ever written to draw upon, and would allow me to draw upon my own references because it can find them for me faster and better than I can now.
I had Claude Code make some changes, which I opened and edited extensively (by hand) in Obsidian (a markdown text editor that can connect most modern LLMs via MCP). But I’ve been doing the majority of my writing in simple text editors for over a dozen years now, and mostly in Markdown. Lucky for me, that’s also the native default for AI and super token efficient.
Once I had the idea a bit more focused and had pruned away many of the sprawly bits, I took the current draft to ChatGPT/Codex along with a handoff document with some specific context from my original intent and the references and links that I had not yet worked in. It rendered the whole thing fresh, gaining about 50% more words in the process. Words that I had not written or spoken, but that were very much in line with what I was trying to write. More importantly, the words were better sequenced into a different structure than I had before. Some of my repetition was gone. Other new repetitions were introduced that I had to strip back out. But mostly I just had it provide critique. And Claude did again too for typos and such.
I continued to work in Obsidian, which has been my primary writing tool for five years. Every bit of formatting and linking that’s here I did very quickly in Obsidian, while I was swapping words, reconstructing sentences, and moving and dropping entire paragraphs without having to use the usual copy/paste commands. Seems silly maybe but it works for me. I also continued to dictate in Wispr Flow as well, and whispered to myself to make sure the words still flowed. Polished until I thought I was finished writing.
I iterated with Claude a few different concepts for what would work as an image. I had it write a structured prompt for me to take to Gemini to generate, because I was sure I wanted the telltale little Gemini star in the lower right on this one. The first idea was too difficult for Gemini to render well, which took a couple attempts to determine. The fallback creative idea worked, which was a structured prompt that I iterated three times to add my own idea on top.
Seeing the resulting image inspired me to tighten the entire piece around that central analogy. Codex made a final pass, trimming about 20% of the main essay, breaking out one conceptual tangent into its own piece rather than cluttering this one, and introducing a few lines that anchor it conceptually to the image. As the writer, I winced at some of the changes like I would with any editor, but realized it was a better piece as a result. I made a few small tweaks, checked the links on last time, and published.
Does everything I write that’s AI-assisted look like this pattern? Absolutely not. Sometimes it’s more involved. Oftentimes it’s far easier. I think I’ve experimented with every possible variant and workflow by now. If it sounds like a lot of work at this point, you don’t write very much, do you? Or if you do, you’re a far better writer than I am. I’m a fast and trained writer, but good usually takes me some time. After all, I’m crafting something special for you, and I think about you a lot when I’m writing. I’m trying to make something that’s actually good...
...with or without AI.
Research notes
Substack’s “How can I detect AI on Substack?”
Pangram’s “How does Pangram work?”
Sam Rogers writes and builds about AI, learning, and the systems we use to decide what deserves trust. He publishes at sam-rogers.com and Signals & Subtractions.



I published an new episode on the same subject with @limitededitionjonathan. We argued the underlying question anyway, and he found the hole in it.
He had fed a real Monet to ChatGPT and Claude, told each one it was AI generated, and asked how close it came. Both called it soulless. Which is a problem for what I've argued here: if models carry the same bias people do, then an evidence layer built on models inherits the bias, and my proposal has the same defect I said the detector had.
I don't have a clean answer to that yet. Best I have is that evidence you can inspect and contest fails differently than a score you can only accept, but "fails differently" is not "doesn't fail."
So, a question I'd love to see answered: if a platform could show you one thing about a post before you read it, and it wasn't allowed to be who or what typed it, what would you want it to be?
The episode, with his side of it: https://sigsub.show/episodes/ep-005/