Unmasking AI: The Subtle Linguistic Quirks That Give Models Away
Unmasking AI: The Subtle Linguistic Quirks That Give Models Away
Think AI writing is seamless? Think again. New research reveals specific words and phrases that betray even the most advanced LLMs. Learn how to spot them!
As AI-generated content floods our digital landscape, the quest to differentiate human prose from machine output has never been more pressing. While early tells, like excessive em-dashes or a penchant for the word “delve,” have largely been eradicated, modern AI models still leave behind subtle linguistic fingerprints.
Recent research by marketing firm Graphite shines a light on these evolving habits, analyzing how frontier models craft their sentences. They’ve uncovered over 13,000 phrases that appear at least twice as often in AI content compared to human-written samples—their definition of a “tell.” It’s a fascinating insight into the ever-adapting nature of AI language generation.
Take Claude Opus 5.5, for instance.
Its most prominent tell is the word “dependable,” popping up a remarkable 23 times more often than in human text.
Opus also loves to emphasize importance, using the phrase “this matters” an astonishing 116 times more frequently, and “why X matters” 92 times more.
It also has a soft spot for the construction “is more than an X, it's a Y.”
OpenAI's Astra, on the other hand, presents its own unique quirks. This model frequently describes “another dimension” of its subject matter and tends to temper claims with phrases like “may provide” or “can provide” a benefit.
Astra's biggest giveaway is what Graphite terms the “corrective framing,” where a topic is defined as “not simply X” or presented as an alternative, “rather than relying on X.” These constructions were found to be over 100 times more common in Astra's prose than in human writing.
Interestingly, AI labs have clearly responded to previous criticisms regarding overused punctuation. Opus 5.5 now uses em-dashes 99% less often than its predecessor, Opus 5. Astra has reduced its usage by 88% compared to human samples, and Gemini 3.1 Pro has almost entirely removed the em-dash from its writing.
The Persistent Nature of AI Quirks
Despite these efforts to refine models, the overall number of linguistic tells isn't decreasing; they're merely evolving. As Graphite's chief AI officer, Greg Druck, explains,
They are managing to remove the most well-known tells, but other ones pop up. And every model version has its own.
This constant evolution makes AI detection a perpetual cat-and-mouse game.
It's surprising that these tells persist, especially given the labs' focus on producing human-like writing. Anthropic boasted that Opus 5.5 “communicates more naturally than prior models,” and OpenAI made similar claims for its GPT-6 versions.
Yet, Druck remains skeptical about how much control labs truly have over these nuances.
A general hypothesis I have is that the labs are less able to control some of these things than you might expect,
Druck says. “These are giant models with billions of parameters.
They have some finite number of tests they can run, and things slip through.”
What Does This Mean for the Future?
Understanding these evolving AI writing patterns is crucial for anyone interacting with digital content. It's not about outright detection, but about recognizing the subtle linguistic signatures that distinguish machine-generated text. As models continue to advance, so too will our methods for spotting their distinctive voice, or lack thereof.