INDEX 179 / NEWS · 12 MIN

Sonnet 5.5 Just Ate Opus. Anthropic Did It on Purpose.

Sonnet 5.5 scores 70.6% on Terminal-Bench — beating Opus 5.5's 66.4% — at half the price. Here's why Anthropic is cannibalizing its own premium tier.

CL

ComputeLeap Team

Share

Sonnet 5.5 Just Ate Opus. Anthropic Did It on Purpose.

Two AI chips side by side — a smaller bright chip outshining a larger dimmer one, representing Sonnet 5.5 outperforming Opus 5.5

Anthropic released Claude Sonnet 5.5 on September 28 — six days after Opus 5.5 — at the same $2/$10 per million token price as its predecessor. That alone would be unremarkable. What isn't: Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, the agentic coding evaluation that developers actually care about. Opus 5.5 scores 66.4%. The cheaper model just outperformed the flagship on the benchmark that tracks real-world terminal usage.

@PhotonCap reacting to Sonnet 5.5 release — what did he say? just Anthropic released Sonnet 5.5

View original post on X →

This isn't a rounding error. It's a 4.2-percentage-point gap on a benchmark where both models were already competing against GPT-6 Sol — which OpenAI launched five days earlier at exactly the same $2/$10 price point. The mid-tier model price is now a commodity. The question is what that means for everyone building on these APIs.

The Numbers That Matter

Here's the full benchmark picture, sourced from Anthropic's own announcement:

BenchmarkSonnet 5.5Opus 5.5Sonnet 5Gap
Terminal-Bench 4.070.6%66.4%10.3%Sonnet leads by 4.2pp
CursorBench 4.055.5%57.8%34.1%Opus leads by 2.3pp
FrontierCode 1.146.2%54.4%42.4%Opus leads by 8.2pp
GDPval-AA v2.11,8441,8461,449Near-identical
AA Briefcase v1.11,8111,8221,359Near-identical
Humanity's Last Exam64.5%67.7%54.9%Opus leads by 3.2pp

Two things jump out. First, the generation-over-generation leap from Sonnet 5 to Sonnet 5.5 is enormous — a 60-point jump on Terminal-Bench alone. Second, Sonnet 5.5 and Opus 5.5 are separated by single-digit margins on nearly every benchmark. On knowledge work (GDPval), the gap is literally 2 points out of 1,846.

The pricing couldn't be more different:

Sonnet 5.5Opus 5.5Ratio
Input$2/MTok$4/MTok2x
Output$10/MTok$20/MTok2x
Cache reads$0.20/MTok$0.20/MTok1x
Cache writes$2.50/MTok$5/MTok2x

Opus costs exactly double per token — for single-digit benchmark differences. That's the headline. But it's not the whole story.

@LuminaBench — Claude Sonnet 5.5 beats Astra and Fable 5.1 on benchmarks

View original post on X →

CodeRabbit's real-world code review benchmarks add texture to the numbers. On 44 actual pull requests, Sonnet 5.5 averaged 6 minutes 33 seconds per review versus Sonnet 5's 13 minutes 31 seconds — a 2x speed improvement that directly translates to cost savings in agentic workflows.

The Effort-Level Twist (And Why the Headline Lies)

WARNING

Before you drop Opus from your stack, read the fine print. Anthropic's own benchmarks use different effort levels for each model — and effort level changes everything about cost-per-task.

Here's the detail that most coverage is missing: Sonnet 5.5's 70.6% on Terminal-Bench was measured at Max effort. Opus 5.5's 66.4% was measured at Xhigh effort — one tier lower. Anthropic doesn't publish what Opus would score at Max effort, but the pattern from other benchmarks suggests it would close or erase the gap.

More importantly, as analyst Andrea Saez noted, the effort dial flips the cost story entirely:

  • At Max effort, Sonnet 5.5 costs approximately $7.60 per task — because it burns more tokens and tool calls to reach its peak score
  • At Xhigh effort, Opus 5.5 matches Sonnet 5.5's best score of 56 (on the Artificial Analysis Intelligence Index) for just $3.46 per task
Andrea Saez Substack analysis — Claude 5.5 is brilliant, confusing, and very hungry

View original post on Substack →

That's right: for the hardest problems, Opus at a lower effort setting can be cheaper than Sonnet cranked to maximum. The per-token price is half, but the per-task price isn't — because Sonnet needs more tokens to reach the same result.

This is the contrarian read that most of the "Sonnet killed Opus" takes are missing. As Saez puts it: "Picking your effort level matters more than picking your model."

What the Community Is Saying

The Hacker News thread hit 874 points and 604 comments within hours — and the conversation quickly moved past the benchmarks into the structural implications.

Hacker News thread discussing Sonnet 5.5 — 874 points, 604 comments

View on Hacker News →

User datadrivenangel cut straight to the cannibalization question: "Opus 5.5 on Low seems smarter, cheaper, and faster than Sonnet on Medium, so what's the point?" It's a fair question, and one Anthropic's own pricing table invites.

The most substantive pushback came from inopinatus, who argued that LLMs still struggle with foundational design decisions — they produce "superficially plausible but profoundly ill-considered" data structure choices, then waste tokens addressing symptoms. The implication: benchmark improvements don't help if the model is faster at making the wrong architectural call.

INFO

User mattm identified a deeper issue with model value measurement: LLMs lack a "sense of importance," treating all code changes equally and missing the architectural problems a human would recognize immediately. This matters because cost-per-task measurements assume every task is worth automating — and they don't account for the cost of automating the wrong thing.

On the other side, satvikpendem countered directly: "LLMs these days write better architected and produced code than most programmers." Whether that's true or aspirational, the fact that the debate is happening at all tells you where the Overton window has moved.

The Broader Pricing War

Sonnet 5.5 didn't launch in a vacuum. Here's what happened in the seven days before it shipped:

  • Sept 22: Anthropic releases Opus 5.5 with a 20% price cut — $4/$20, down from Opus 5's pricing
  • Sept 23: OpenAI launches GPT-6 Sol at $2/$10 — exactly matching what would become Sonnet 5.5's price
  • Sept 28: Anthropic ships Sonnet 5.5 at $2/$10

Two companies. One week. Three frontier model launches. Identical mid-tier pricing. Fortune called it what it is: a price war, noting that "the releases came within hours of each other" despite "both companies' recent calls for an AI slowdown."

@MilkRoadAI — Dario wants to slow down AI but Anthropic keeps shipping

View original post on X →

The irony wasn't lost on anyone. Dario Amodei spent the preceding weeks calling for the industry to slow frontier development. Then Anthropic shipped its second frontier model in six days. As Milk Road AI noted: "Dario wants to slow down AI but Anthropic keeps shipping."

But the real pressure isn't coming from OpenAI. It's coming from below. DeepSeek's models now deliver comparable performance at 10--30x lower cost on equivalent tasks. Gemini 3.8 Flash sits at $0.75/$3.75 introductory pricing. The floor is dropping fast.

INFO

Ara Kharazian, Lead Economist at Ramp, told Fortune: "AI bulls assume that there will be highly performant models that provide more and more value, therefore they should be more expensive. But that is not how normal technology makes it to market." The AI pricing trajectory is following the same deflationary curve as cloud compute, storage, and bandwidth before it.

This is the context that makes Sonnet 5.5 significant beyond its benchmarks. Anthropic isn't just releasing a better mid-tier model — it's acknowledging that the premium pricing moat has eroded. When your $2 model beats your $4 model on a key benchmark, and your competitor's $2 model matches yours, the differentiation has to come from somewhere else.

The Real Moat: Cost Per Task, Not Cost Per Token

Anthropic's positioning of Sonnet 5.5 centers on a subtle but important shift. They're not advertising lower token prices (they kept them identical to Sonnet 5). They're advertising lower total cost to complete a job.

The argument: Sonnet 5.5 generates output 30% faster and uses fewer total tokens and tool calls per task. If a coding task that took Sonnet 5 ten tool calls now takes Sonnet 5.5 seven, the effective cost drops 30% even at the same per-token rate. As VentureBeat reported, the competitive frame has shifted "from per-token pricing to cost per completed job."

This reframes the competitive landscape entirely. At the sticker-price level, everyone has converged on $2/$10 for mid-tier models. The battle is now over efficiency — which model wastes fewer tokens getting to the right answer?

This trend didn't start with Sonnet 5.5. We covered the early signals when Fable 5.1's real strategy was a 75% cache price cut, and the broader dynamics in AI's $700B subsidy clock. Sonnet 5.5 is the logical next step: when per-token pricing converges, the model that completes work in fewer steps wins on total cost.

The Cannibalization Is a Feature

Here's our thesis: Anthropic is intentionally commoditizing its own mid-tier to win on distribution.

The AI model market in late 2026 looks increasingly like the cloud computing market circa 2015. Baseline capabilities have converged. Token prices are racing toward marginal cost. The labs that win will be the ones with the deepest developer adoption and the stickiest platform — not the ones with the highest benchmark scores.

By making Sonnet 5.5 nearly as good as Opus 5.5 at half the price, Anthropic is betting that:

  1. Volume beats margin. More developers building on Sonnet 5.5 means more lock-in to Anthropic's API ecosystem, even at lower per-unit revenue.
  2. Opus becomes the upsell. The pitch shifts from "use Opus because it's the best" to "start on Sonnet, escalate to Opus when the task demands it" — a classic freemium-to-premium funnel.
  3. Efficiency is the new benchmark. When token prices converge, the model that completes a task in fewer steps wins on total cost. Anthropic is competing on token economics, not token prices.

This is the same playbook AWS ran with EC2 reserved instances and spot pricing — offer a cheap default tier that captures volume, then monetize the premium tier through capability differentiation. Opus keeps its edge on sustained judgment, complex reasoning, and multi-step planning. Sonnet captures the volume.

What This Means for You

TIP

The practical playbook for teams evaluating Sonnet 5.5 vs Opus 5.5.

1. Benchmark on YOUR workload, not Anthropic's. Terminal-Bench and CursorBench measure different things. Sonnet 5.5 wins on agentic terminal work; Opus 5.5 wins on ambiguous multi-file coding. Which one matches your use case?

2. Test effort levels before picking a model. At medium effort, Sonnet 5.5 is dramatically cheaper. At max effort, it can cost more than Opus at xhigh. Run your most expensive tasks at both Sonnet-medium and Opus-low before deciding.

3. Measure cost per completed task, not cost per token. If Sonnet 5.5 solves your problem in 7 tool calls instead of Sonnet 5's 10, the 30% cost reduction is real even though the price hasn't changed. Track total API spend per successful outcome.

4. Watch for Haiku 5.5. Anthropic announced Haiku 5.5 is coming "in the weeks ahead." If it follows the same pattern — near-Sonnet performance at Haiku pricing — the commoditization accelerates further down the stack.

5. Don't ignore the cache pricing anomaly. Both Sonnet 5.5 and Opus 5.5 charge $0.20 per million cached tokens — identical. For agentic workflows that reread large contexts, the Sonnet pricing advantage is smaller than the 2x sticker price suggests. Factor your cache hit rate into the calculation.

The Bottom Line

Sonnet 5.5 is the strongest evidence yet that AI model quality is commoditizing faster than anyone — including the labs themselves — expected. When your $2 model beats your $4 model on a coding benchmark, and your competitor's $2 model has the exact same price tag, the game has changed.

The moat isn't in the model anymore. It's in the ecosystem — the developer tools, the API reliability, the context window handling, the agentic framework maturity. It's in the Fable 5.1 cache pricing plays and the effort-level economics and the thousand small decisions that determine whether a developer's first project on your platform becomes their hundredth.

Anthropic knows this. That's why Sonnet 5.5 eating Opus isn't a bug — it's the strategy.


Claude Sonnet 5.5 (claude-sonnet-5-5) is available now via the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry at $2/MTok input and $10/MTok output.

AUTHOR
CL

ComputeLeap Team

The ComputeLeap editorial team covers AI tools, agents, and products — helping readers discover and use artificial intelligence to work smarter.

DISCUSSION

Join the discussion

Have thoughts on this article? Discuss it on your favorite platform:

NEWSLETTER

The ComputeLeap Weekly

Get a weekly digest of the best AI infra writing — Claude Code, agent frameworks, deployment patterns. No fluff.

WEEKLY. UNSUBSCRIBE ANYTIME.