Graphite Study Finds Persistent AI Writing Tells, From 'This Matters' to Corrective Framing
Graphite study finds AI models still have distinct writing tells, such as 'this matters' and corrective framing.
Graphite researchers began with a corpus of 10,000 articles published before the release of ChatGPT to serve as a human-generated control group. They then had different AI models rewrite those articles from summaries, an approach intended to reduce source bias. With matching samples from humans and each model, the researchers compared how often certain words and phrases appeared and examined broader patterns in sentence construction.
Claude Opus 5.5 showed a set of habits that separate it from human writers. Its biggest tell is the word “dependable,” which appears 23 times more often than in human samples. The model now avoids the “it’s not X, it’s Y” construction, according to Graphite, but it still tends to say something “is more than an X, it’s a Y.” Above all, Opus 5.5 frequently explains why a subject matters, using “this matters” 116 times more often than human writing and “why X matters” 92 times more often.
OpenAI’s Astra has a different set of tip-offs. It likes to describe “another dimension” of whatever it is discussing, and it tends to hedge claims by saying an action “may provide” or “can provide” a particular benefit. Its biggest tell is what Graphite calls corrective framing, in which a topic is defined as “not simply X” or offered as an alternative, “rather than relying on X.” Those constructions were more than 100 times more common in Astra-generated prose than in human writing.
The study also found that frontier labs have responded to criticism that models overuse em-dashes. In Graphite’s samples, Opus 5.5 used the punctuation mark 99 percent less often than Opus 5. Astra used it 88 percent less than human samples, while Gemini 3.1 Pro had almost completely eliminated the em-dash from its writing.
The overall pattern differs by model family, according to Greg Druck, Graphite’s chief AI officer. “It turns out that Claude models are actually getting closer to the human word distribution over time,” Druck told TechCrunch. “And for the GPT models, it’s getting further away.”
Druck said the total number of tells is mostly holding steady even as specific ones change. “It’s not like the tells are decreasing,” he told TechCrunch. “They are managing to remove the most well-known tells, but other ones pop up. And every model version has its own.”
The persistence of these habits comes as the labs have focused on human-like writing styles. In the Opus 5.5 release, Anthropic said the model “communicates more naturally than prior models,” adding that early users “found its writing clearer and easier to follow.” OpenAI made similar claims when releasing the GPT-6 versions of Sol and Luna, saying users could “expect to see more clarity, less jargon, [and] fewer odd turns of phrase.”
Druck is skeptical about how much the labs can do to eliminate telltale constructions or phrases. “A general hypothesis I have is that the labs are less able to control some of these things than you might expect,” he said. “These are giant models with billions of parameters. They have some finite number of tests they can run, and things slip through.”