NEWS

Claude Opus 5.5 Cuts Em Dashes by 95%, But Responses Keep Getting Longer

Independent analysis shows Claude Opus 5.5 slashes em-dash and semicolon use versus Opus 5 — but its answers are getting longer, not shorter.

Dylan H.

News Desk

September 26, 2026
8 min read
Claude Opus 5.5 Cuts Em Dashes by 95%, But Responses Keep Getting Longer

Claude Opus 5.5's writing looks less like an AI wrote it — mostly

Anthropic released Claude Opus 5.5 on September 22, 2026, and independent analysis published this week shows the model has quietly changed one of the most recognizable "tells" of AI-generated text: the em dash. According to a BleepingComputer report citing data from the AI benchmarking tool Arena, Opus 5.5 uses roughly 95% fewer em dashes than its predecessor, Claude Opus 5 — a drop from 15.2 per 1,000 words to just 0.8. Semicolon usage fell in similar fashion, from 6.10 to 1.64 per 1,000 words. The catch: while sentences got shorter, overall responses got longer, rising from an average of 453 words to 481 words — making Opus 5.5 the most verbose Opus model measured to date.

For a cybersecurity and IT audience, the story is bigger than punctuation trivia. Stylistic fingerprints like em-dash frequency, semicolon use, and sentence cadence have become an informal shorthand — used by editors, content moderators, and even some phishing-detection heuristics — for flagging AI-generated text. If a leading model is deliberately (or incidentally) shedding those tells while writing more, not less, the reliability of "AI-detection by vibe" keeps eroding.


Details

AttributeValue
ModelClaude Opus 5.5
DeveloperAnthropic
Release dateSeptember 22, 2026
Predecessor comparedClaude Opus 5
Analysis sourceArena (AI benchmarking tool), citing high-reasoning Text Arena responses collected August–September 2026
Em dash frequency15.2 → 0.8 per 1,000 words (≈95% reduction)
Semicolon frequency6.10 → 1.64 per 1,000 words
Avg. sentence length12.14 → 10.03 words
Avg. response length453 → 481 words (increase)
API pricing$4 / 1M input tokens, $20 / 1M output tokens (down from $5 / $25 on Opus 5)
Independent benchmark#1 of 212 models on Artificial Analysis's Intelligence Index (score 58); 95th of 212 on verbosity

What the analysis found

The em-dash and semicolon drop

Arena's dataset, built from high-reasoning Text Arena responses gathered across August and September 2026, found that 10 of 12 tracked writing measures moved in what it characterized as a "better" direction for Opus 5.5 relative to Opus 5. The headline figure — a roughly 95% cut in em-dash usage — lines up with a long-running complaint among readers and editors that em dashes had become a near-universal giveaway of LLM-authored prose. The parallel decline in semicolon use suggests a broader shift away from complex, clause-heavy sentence construction, reinforced by the drop in average sentence length from 12.14 to 10.03 words.

The verbosity paradox

Despite writing shorter, simpler sentences, Opus 5.5 produces longer overall answers than Opus 5 — 481 words on average versus 453. That finding is corroborated by Artificial Analysis, an independent benchmarking firm, which ranked Opus 5.5 first of 212 models on its Intelligence Index (a score of 58) but 95th of 212 on verbosity, measuring 260 million output tokens generated across its Index run against a median of 88 million for other models. This sits awkwardly next to Anthropic's own marketing claims that Opus 5.5 uses fewer tokens and generates output more than 30% faster than Opus 5 — a case where a vendor's efficiency framing and an independent verbosity measurement point in different directions, even though both can be technically true depending on task type and prompt.

Why the style shifted at all

Context for the change came a day after the BleepingComputer report's dataset window closed: on September 23, 2026, Anthropic engineer Jackson Kernion, who works on Claude's fine-tuning, said publicly that the dense, jargon-heavy "Claudeish" prose style many users had complained about since Claude Opus 4.6 was not a bug or a deliberately "nerfed" model, but a side effect of reinforcement-learning post-training that increasingly optimized for math, code, and technical explanations aimed at other AI systems rather than human readers. Kernion said he "hadn't been as happy about a model's writing since Opus 4.6," and described Opus 5.5's style adjustments as a partial fix to that reward-design tradeoff rather than a full resolution — framing it as, in his words, "a hard problem to solve." Anthropic's own release material for Opus 5.5 describes the model as putting "the most important information up front," being "less likely to use jargon or idiosyncratic phrases," and following user-supplied writing rules more consistently.

What's still detectable

Not every AI writing tell has vanished. A separate style analysis cited in coverage of the release cautions that em dashes on Opus 5.5 are merely "less common than on Opus 5, but worth searching for before you publish," and flags other lingering habits: "flourish words" such as "load-bearing" and "quietly," plus closing sign-offs that offer to expand further on the content. In other words, the stylistic fingerprint has changed shape rather than disappeared.

Impact Assessment

Impact AreaDescription
Content authenticity / AI detectionPunctuation- and cadence-based heuristics for spotting AI-written text (including in phishing and disinformation triage) lose reliability as models like Opus 5.5 shed those tells while other markers ("flourish words," sign-offs) persist in weaker form.
Enterprise writing workflowsTeams using Claude for drafting, documentation, or customer-facing copy may see cleaner prose out of the box, reducing manual "de-AI-ify" editing passes — but longer average responses could increase review time and token spend.
API cost planningThe advertised $4/$20 per-million-token pricing (down from $5/$25) is a real per-token savings, but Artificial Analysis's verbosity findings suggest actual cost-per-task may not fall as much as the headline price cut implies if response length grows.
Vendor benchmark trustThe gap between Anthropic's "fewer tokens, 30% faster" claim and Artificial Analysis's independent verbosity ranking (95th of 212) is a reminder that vendor-reported efficiency metrics and third-party measurements can diverge — evaluate both before re-architecting cost-sensitive pipelines.
Social engineering / phishing contentReduced reliance on stylistic AI markers in generated text is a double-edged trend: it improves legitimate content quality, but the same shift removes a signal defenders have informally used to spot machine-generated lures and scam copy.

Recommendations

For security and content-moderation teams

  • Do not rely on punctuation-based heuristics (em-dash or semicolon frequency, sentence length) as a standalone signal for flagging AI-generated phishing, spam, or disinformation content — treat them as weak, decaying indicators rather than reliable detectors.
  • Where AI-content detection matters (compliance, academic integrity, brand-abuse monitoring), prioritize model-agnostic detection approaches (provenance metadata, C2PA-style content credentials, behavioral/network signals) over stylistic fingerprinting, since per-model writing style is now a moving target the vendor actively tunes.
  • Track vendor writing-style changes as part of routine threat-model reviews for any AI-assisted content-screening tool your organization operates.

For engineering and platform teams integrating Claude

  • If migrating from Opus 5 to Opus 5.5, re-baseline token-per-task costs empirically rather than assuming the $5→$4 / $25→$20 price cut translates directly into lower total spend; the Artificial Analysis verbosity data suggests output length can offset per-token savings on some workloads.
  • Re-test prompts that explicitly instructed Opus 5 to avoid em dashes or "sound less like AI" — those workarounds may now be redundant or could interact unpredictably with Opus 5.5's already-adjusted defaults.
  • Set explicit length constraints in prompts or output_config where response brevity matters, given the model's documented tendency toward longer answers.

For general users and editors

  • Continue a manual review pass before publishing AI-assisted content; residual tells (jargon-adjacent "flourish words," offer-to-expand sign-offs, occasional em dashes) still surface per independent style reviews.
  • Treat "sounds more human" as a quality-of-life improvement, not a guarantee of factual accuracy — writing-style changes are orthogonal to the model's reasoning or citation reliability.

Key Takeaways

  1. Independent analysis (Arena, via BleepingComputer) found Claude Opus 5.5 uses about 95% fewer em dashes than Opus 5 (15.2 → 0.8 per 1,000 words) and far fewer semicolons (6.10 → 1.64 per 1,000 words).
  2. Sentences got shorter on average (12.14 → 10.03 words), but overall responses got longer (453 → 481 words), making Opus 5.5 the most verbose Opus model tracked so far.
  3. Artificial Analysis independently ranked Opus 5.5 #1 of 212 models on intelligence but 95th of 212 on verbosity, which sits in tension with Anthropic's own "fewer tokens, 30% faster" marketing claims.
  4. Anthropic engineer Jackson Kernion attributed the earlier "Claudeish" writing style — dating back to Opus 4.6 — to RL post-training that over-optimized for machine-legible technical output at the expense of natural, human-readable prose.
  5. Some AI writing tells persist even in Opus 5.5, including occasional em dashes, stock "flourish words," and offer-to-expand closing lines — style detection heuristics should be treated as unreliable and decaying, not eliminated.
  6. Opus 5.5 also shipped at lower list pricing ($4/$20 per million input/output tokens, down from $5/$25), though real-world cost impact depends on the verbosity tradeoff documented above.

Sources