Claude Opus 5.5 Writing Tells Changed, but They Did Not Disappear
Claude Opus 5.5 writing tells survived Anthropic’s style overhaul, despite a 99% reduction in the model’s use of em dashes. The strongest signal is now “dependable,” which appeared 23 times more often in model-written articles than comparable human work.
That finding comes from Graphite’s analysis of articles covering 9,974 matched topics. The researchers identified 2,548 words, phrases, and sentence patterns that Opus 5.5 used at least twice as often as human writers.
The result complicates Anthropic’s claim that Opus 5.5 communicates more naturally. Its vocabulary moved measurably closer to human writing, yet thousands of statistical differences remained. The model did not lose its linguistic fingerprint. It traded several familiar habits for a new collection of preferences.
OpenAI’s GPT-6 Astra produced another distinct pattern. It qualified claims more often and leaned on corrective framing, while Opus favored superlatives and declarations of importance. The comparison turns AI writing detection into a moving-target problem. Each model version changes the evidence that editors, teachers, and readers think they recognize.
A 23-fold preference exposed the new Opus pattern
The latest Claude Opus 5.5 writing tells are ordinary words made unusual by repetition.
Graphite published its Opus 5.5 update after Anthropic released the model on September 22, 2026. The research extended an earlier analysis across ten AI models and a human control collection.
The researchers used 9,974 topics for which every model and the human corpus had a corresponding article. Human articles were published before ChatGPT’s release, reducing the chance that those samples contained unmarked generative AI output.
Each model received a summary and generated a new article about the same subject. This approach reduced topic differences, although it could not remove every effect caused by prompting or source selection.
Researchers then compared length-normalized word and phrase frequencies. They also studied frames, which are recurring language patterns with gaps that can contain up to three words.
Under Graphite’s definition, a tell appears at least twice as often in AI writing as in the human collection. Frequency filters exclude expressions that appear in only a handful of documents.
The resulting AI writing study found 2,548 Opus 5.5 tells. That was only 4% below the 2,666 recorded for Opus 5.
“Dependable” led the model’s evaluative adjectives at 23 times the human rate. “Steady” appeared 11 times as often, while “thoughtful” appeared nine times as often. “Meaningful” reached eight times the human rate.
Those words are not inherently artificial. A human writer can reasonably call a service dependable or describe a response as thoughtful. Their value lies in aggregate behavior, not isolated appearances.
The transition language was also distinctive. “Looking ahead the” appeared at 40 times the human rate. “Adds another layer” reached 27 times, and “what comes next” reached 24 times.
Opus 5.5 also retained a preference for contrast. The pattern “is more than a _ it” appeared 98 times as often as it did in human articles. “Rather than simply” appeared at 32 times the human rate.
Its strongest category involved flagging importance. “This matters” appeared 116 times as often as in the human collection. “Why _ matters” appeared 92 times as often.
These figures do not mean every Opus article contains those phrases. They describe relative frequency across a large, controlled corpus. A rare phrase can produce a large ratio when humans almost never use it.
That distinction matters for anyone tempted to treat a single expression as proof. The study maps population-level habits. It does not provide a reliable authorship verdict for one paragraph, email, or student essay.
The most defensible conclusion is narrower. Opus 5.5 repeatedly reaches for certain evaluative and organizational phrases when it converts summaries into articles. Those preferences remain visible even after a major style revision.
The familiar AI writing tells are already obsolete
The em dash stopped being a useful shortcut because model behavior changed faster than popular detection advice.
For several years, readers treated em dashes as an informal marker of generated text. The punctuation became associated with Claude because earlier versions used it frequently.
Opus 5.5 largely abandoned that habit. Graphite measured 0.015 em dashes per 1,000 words, compared with 2.92 for Opus 5. That represents a decline of about 99%.
A separate evaluation produced a similar directional result. An Arize researcher generated 57,000 words with Opus 5.5 and found only two em dashes. His Opus 5 sample produced 12.9 per 1,000 words.
That smaller independent writing test used 20 frozen briefs across five genres and five subject areas. The researcher manually marked 153 examples before building an automated evaluator.
The dataset was much smaller than Graphite’s corpus, but it directly tested Anthropic’s revised model against its predecessor. It found that other recognizable “Claudisms” fell by roughly half rather than disappearing.
The change creates an uncomfortable reversal. Human writers now risk suspicion for punctuation that current frontier models rarely use.
A model developer can suppress an obvious habit through post-training, system instructions, or generated-output preferences. Users can also request different punctuation with one sentence in a prompt.
That makes surface clues fragile. “Delve,” em dashes, neat three-part lists, and polished conclusions can describe a style. None establishes authorship.
The same problem affects commercial AI detectors. A classifier trained on yesterday’s models can lose accuracy after a provider changes output behavior. The underlying model may retain the same name while its style shifts through an update.
False positives carry real consequences. A teacher might accuse a student who naturally writes formal prose. An editor might flatten a contributor’s distinctive punctuation to avoid automated suspicion.
False negatives are just as easy to produce. A user can request irregular sentence lengths, remove favored phrases, or ask a second model to rewrite the first output. Human editing can erase obvious patterns without changing the text’s origin.
The practical response is to evaluate provenance, factual support, and revision history together. Stylistic evidence can prompt closer review, but it should not decide the case alone.
For professional teams, that means preserving source notes and drafts. A searchable personal knowledge base can help connect claims to their evidence and record how a document developed.
That workflow addresses the actual quality risk. Undisclosed AI use can matter, but unsupported claims and missing accountability usually create the greater operational problem.
Graphite’s findings therefore weaken the case for simple detection rules. They strengthen the case for process evidence, editorial review, and transparent policies.
Anthropic improved the style without erasing the fingerprint
Opus 5.5 became statistically closer to human writing while retaining almost as many measurable tells as Opus 5.
Anthropic introduced Opus 5.5 with a direct communication claim. The company said the model writes more naturally, places important information first, and avoids more jargon and idiosyncratic phrasing.
Early testers reportedly found the prose clearer and easier to follow. Anthropic also presented better communication as a safety benefit because readers could inspect the model’s work more easily.
Graphite’s results partly support that position. The researchers measured word-distribution divergence using Jensen-Shannon divergence, a statistic that compares two probability distributions.
A score of zero would indicate identical word use. A score of one would indicate no overlap.
Opus 5 scored 0.064 against human writing. Opus 5.5 fell to 0.052, a 19% reduction. That means its overall word distribution moved closer to the human corpus.
The model also improved on “mannered prose,” which Anthropic defines as language that substitutes flourish or metaphor for direct statements. Graphite reported a score of 10.57 for Opus 5.5, down 37% from Opus 5’s 16.75.
Human articles scored 6.65. GPT-6 Astra scored 7.91. Opus therefore improved substantially while remaining more mannered than either comparison group.
The mannered-prose result needs careful interpretation. Graphite used Claude Opus 5 itself to score a subsample of 1,000 matched topics on a scale from zero to 100.
The evaluator did not know which model produced each article. Even so, its judgment reflects another model’s interpretation rather than a mechanical measurement.
The tell count creates the central tension. Opus 5.5 had better overall vocabulary similarity, but its 2,548 tells were only 4% fewer than Opus 5’s total.
Those results measure different properties. Distributional similarity asks how the entire vocabulary compares with human writing. A tell count asks how many specific patterns cross a frequency threshold.
A model can improve on the first measure without eliminating the second. Broad vocabulary choices can become more human while narrow habits remain heavily concentrated.
Anthropic’s Opus 5.5 announcement also highlighted stronger knowledge-work performance. In one internal research task, 16 of 18 reports cleared a strict quality threshold across different effort settings.
Anthropic said an invented number or quotation would cause a report to fail. Earlier Opus and Fable models did not clear that bar in any attempt, according to the company.
Those internal results are relevant because natural prose and reliable research are separate dimensions. A repetitive phrase does not make a report factually wrong. Human-sounding language does not make a claim accurate.
The company’s broader argument concerns usability. Clearer writing can reduce review time even when an editor can still identify recurring stylistic patterns.
Graphite’s research does not disprove that improvement. It places a boundary around it. Opus 5.5 appears more human in aggregate, but “more human” is not equivalent to indistinguishable.
Opus and Astra leave different linguistic fingerprints
The important competition is not human prose against one fixed AI style, but shifting model fingerprints against stable detection assumptions.
GPT-6 Astra did not simply perform better or worse than Opus 5.5. It showed a different collection of habits.
Astra qualified claims more often. “May provide” appeared 18 times as frequently in Astra as in Opus 5.5. “Not necessarily” appeared 17 times as frequently.
The phrase “does not establish” showed an even larger difference, reaching 275 times the Opus 5.5 rate. These patterns suggest Astra often protects a claim with explicit limits.
Opus 5.5 leaned harder on superlatives. Compared with Astra, it used “the most popular” 45 times as often and “the most powerful” 24 times as often.
“Perhaps the most” appeared 29 times as frequently. The frame “one of the best _ about” reached 79 times the Astra rate.
That contrast affects tone. Astra can sound cautious or corrective. Opus can sound more evaluative and confident, especially when assigning importance.
Graphite’s CEO Ethan Smith described the pattern as difference rather than steady progress. Each new model can become better along one dimension while developing another recognizable cadence.
The reported industry coverage reached the same conclusion. A model comparison noted that the study evaluated ten models on the same 9,974 topics.
Opus 5.5’s word distribution moved toward humans. Astra moved away from them, increasing from 0.101 for GPT-5.6 Sol to 0.109.
Yet Astra produced less mannered prose than Opus. That prevents a simple ranking based on one score.
A user might prefer Astra’s directness while disliking its hedging. Another might prefer Opus’s flow while removing unnecessary superlatives during editing.
The findings also complicate model attribution. A passage can contain traits associated with several systems because users combine tools in one workflow.
One model might perform research, another might draft, and a third might edit. A human could then revise the result before publication.
Prompts add another layer. Explicitly banning a phrase can alter its frequency. Asking for cautious scientific language can make Opus resemble Astra along one dimension.
Fine-tuned systems and application-level instructions create further variation. The output visible to a user may reflect the provider’s model, the application’s system prompt, and the user’s request.
This is why the phrase “AI style” has become too broad. Frontier models no longer share one stable voice, even when they repeat some familiar structures.
The competitive pressure falls on model labs and AI detection vendors alike. Labs want prose that adapts to user intent. Detection vendors need systems that survive rapid model changes.
Editors face another pressure. They must distinguish between harmless stylistic repetition and defects that damage accuracy, clarity, or trust.
A sentence that says “this matters” may be tedious, but it is not automatically false. An elegant sentence with an invented source is far more serious.
The useful lesson is not to hunt for one replacement tell. It is to recognize that generated prose carries changing statistical signatures, none of which should substitute for editorial judgment.
The study cannot identify who wrote a single passage
Graphite measured corpus-level differences, not a forensic test for individual documents or authors.
The study’s scale makes its aggregate findings persuasive. Its design also imposes limits on how those findings should be used.
First, the experiment focused on article generation from summaries. It does not automatically describe emails, fiction, legal analysis, classroom essays, or casual conversation.
Different tasks change a model’s language. A technical report may naturally contain more qualifications. Marketing copy may naturally contain more superlatives.
Second, the human control articles predated ChatGPT. That protects the baseline from obvious contamination, but it introduces a time difference.
Editorial conventions change. Search optimization changes. Publishers alter headline structures, paragraph lengths, and preferred punctuation.
Some differences attributed to models might partly reflect the contemporary writing environment. Graphite’s matched-topic design reduces content bias, but it cannot erase every historical difference.
Third, frequency ratios can exaggerate rare behavior. If humans almost never use a phrase, a modest number of model uses can generate a striking multiple.
The researchers applied minimum-frequency requirements to address that problem. Words needed to appear in at least 500 model articles for one ranked view, while other analyses used their own thresholds.
Even with those filters, a ratio needs its base rate. “This matters” at 116 times the human rate sounds decisive, but it does not show how often the phrase appeared per article.
Fourth, the results reflect specific model versions and experimental settings. Provider updates, inference settings, system prompts, and user instructions can all change output.
Fifth, the study came from Graphite, a marketing and growth company rather than an independent academic laboratory. Its researchers released the underlying tells data and raw articles, which supports external scrutiny.
Independent replication remains valuable. Researchers should test other genres, languages, prompt designs, and human baselines. They should also report absolute frequencies beside relative ratios.
The early Arize test provides one useful comparison. It supported the sharp reduction in em dashes while finding that broader stylistic habits remained.
Still, its 20 briefs cannot validate every result from nearly 10,000 aligned topics. The two projects used different designs and answered different questions.
The strongest external reporting has kept those boundaries visible. The original writing-tells coverage described the model habits without claiming that any phrase proves AI authorship.
Readers should apply the same restraint. A suspicious phrase can motivate a question. It cannot answer that question by itself.
For publishers, a more reliable review asks whether claims have sources, quotations match originals, and conclusions follow from evidence. It also asks who accepts responsibility for the final text.
For schools, authorship policies should define acceptable assistance before disputes occur. Draft history, citations, and oral follow-up provide stronger evidence than punctuation or vocabulary alone.
For businesses, disclosure rules should match the risk. A routine internal summary does not require the same controls as medical guidance, financial analysis, or public reporting.
The study reveals tendencies, not guilt. Treating it as a detector would turn careful descriptive research into an unreliable enforcement tool.
Three signals will show whether the new tells last
The next test is whether Opus 5.5’s current fingerprint survives replication, user prompting, and the next model update.
The first signal is independent reproduction of the Graphite results. Researchers need to generate comparable corpora with different prompts, genres, and human samples.
If “dependable,” “this matters,” and the contrast frames remain elevated, the case for a persistent Opus 5.5 signature strengthens. If they collapse under minor prompt changes, they are weak attribution signals.
The second signal is how quickly Anthropic changes the model’s language. The company clearly responded to criticism of Opus 5’s communication style.
A later update could reduce current tells as sharply as Opus 5.5 reduced em dashes. That would confirm that visible stylistic traits are adjustable product behaviors.
The third signal is whether detection vendors publish version-specific evaluations. A detector should be tested on current models, mixed human-AI workflows, and edited outputs.
Accuracy claims also need false-positive rates across different writing communities. Non-native English writers, technical specialists, and people with formal styles should not become collateral damage.
The more useful future may involve provenance rather than linguistic guesswork. Signed generation records, preserved drafts, and clear disclosure could establish process without pretending that prose alone reveals its origin.
For writers, the immediate action is simple. Search drafts for repeated evaluative words, canned contrasts, unnecessary superlatives, and sentences that announce importance instead of demonstrating it.
Then verify every factual claim. Remove unsupported confidence. Keep unusual punctuation when it serves the sentence rather than deleting it to avoid suspicion.
Claude Opus 5.5 writing tells provide a valuable editorial checklist, but they are not a verdict. Will publishers use the findings to improve review, or turn another temporary pattern into an unreliable shortcut?



