Claude Sonnet 5.5 Tops Agentic Coding Tests at Sonnet 5 Pricing, but Max Effort Costs More
Sonnet 5.5 tops agentic coding tests at Sonnet 5 pricing, until Max effort flips the math.

Anthropic released Claude Sonnet 5.5 on September 28, 2026 — the evening before OpenAI's annual developer conference — positioning its mid-tier model as the stronger option for agentic coding tasks at the same list price as its predecessor, Claude Sonnet 5. The new model scored 70.6% on Terminal-Bench 4.0, the primary benchmark for autonomous command-line task completion, surpassing Claude Opus 5.5's 66.4% on the same test. That headline result contains an important footnote from Anthropic's own announcement: the two models were not tested at the same effort level, and the cost structure at maximum computational load inverts in ways developers need to understand before routing work to the cheaper option.
The release lands in a compressed competitive window. Anthropic launched Claude Opus 5.5 on September 22 as the first model in the Claude 5.5 family; Sonnet 5.5 is the second. The new model keeps Sonnet 5's list price of $2 per million input tokens and $10 per million output tokens — approximately half Opus 5.5's $4 input and $20 output pricing — and arrives timed to put benchmark data in developers' hands the night before OpenAI DevDay 2026, which is scheduled for September 29 in San Francisco.
Read more: Anthropic's Claude Code appears to route to Opus 5.2 as Model 2 runs its research lab
What Claude Sonnet 5.5 Is and What Changed
Sonnet 5.5 is Anthropic's second Claude 5.5 model. Anthropic describes it as more than 30% faster than Sonnet 5 and capable of reducing per-task inference costs by up to 30% at comparable quality levels — a cost reduction driven not by lower token prices, which are unchanged, but by needing fewer tokens and tool calls to complete the same work.
The model targets what Anthropic calls everyday professional work: multi-file coding, document synthesis, UI generation, long-form business workflows, and agentic pipelines where a model must execute sequences of actions across tools and environments without human supervision at each step. Anthropic positions Sonnet 5.5 as the workhorse of the Claude 5.5 pair — the model developers run constantly — while Opus 5.5 handles the highest-complexity sustained tasks. A third Claude 5.5 model, Haiku 5.5, is expected in the coming weeks.
Beyond raw performance numbers, Sonnet 5.5 carries two capabilities that mark genuine generational advances: a near-quadrupling of visual chart comprehension accuracy and first-time completion of Pokémon Red using only raw game screenshots, without any image-processing tool assistance. Both point to the same underlying shift — the model can now read complex, dense visual information that earlier Claude versions could not reliably parse.
How the Effort System Controls Both Performance and Cost
The most consequential thing to understand about Claude Sonnet 5.5 is not its benchmark scores but how its effort system works — and why it creates an unexpected cost pattern at maximum settings.
Anthropic's API exposes an effort parameter for Claude models with five tiers: Low, Medium, High, Xhigh, and Max. The parameter controls how many tokens the model allocates to internal extended thinking before producing a final answer. At Low effort, the model reasons briefly and responds quickly. At Max effort, the model's internal chain-of-thought faces essentially no token budget cap and can reason for as long as it determines useful.
The default tier is Medium in Claude.ai consumer apps and in Claude Code; the Claude Platform API defaults to High. Developers who do not explicitly configure effort run at whichever tier their access point defaults to. This distinction matters because the benchmark numbers in Anthropic's announcement were not all collected at the same effort tier — Terminal-Bench 4.0 results were run at Max effort for Sonnet 5.5 and at Xhigh for Opus 5.5, a meaningful difference in computational budget that the headline comparison obscures. Anthropic discloses this in a footnote to its benchmark table.
The deeper issue emerges when developers push Sonnet 5.5 to Max effort on complex coding tasks. Independent testing found that the model generated an average of approximately 193,000 output tokens per task at Max effort — the highest per-task token output that evaluator had recorded, exceeding Opus 5.5's Max-effort figure by roughly 60%. Because output tokens determine inference cost and Sonnet 5.5's per-token price is lower than Opus 5.5's, the relationship between tier and expense is not intuitive: at Max effort, independent testing found a Sonnet 5.5 task costs approximately $7.60 versus $5.98 for the equivalent Opus 5.5 task. The "cheaper model" becomes the more expensive one per task when both are running at maximum load.
The reason is behavioral rather than architectural. At Max effort, Sonnet 5.5 aggressively dispatches multi-agent subagents for code review in a way Opus 5.5 does not. These subagents each consume tokens, dramatically inflating total per-task output. On FrontierCode 1.1 — a benchmark that measures whether code changes would pass a human merge review — this behavior actually hurts performance: Sonnet 5.5 scored 52.1% at Xhigh effort but only 46.2% at Max, because the additional subagents frequently made out-of-scope edits to files not relevant to the task and triggered timeout failures on longer runs. Anthropic itself explains this inversion in its footnotes: "At Max effort, Sonnet 5.5 more often ran Claude Code's code-review skill, which splits the review across many subagents, and in two cases this led to a timeout or to extra edits beyond the task's scope."
The practical guidance is clear from the data: Low to Medium effort is appropriate for bug fixes, document generation, and single-file coding; High or Xhigh suits multi-file engineering tasks and complex document synthesis; Max effort should be reserved for long-horizon agentic work on genuinely ambiguous problems, and even then, Opus 5.5 at High or Xhigh may produce a better cost-to-quality ratio than Sonnet 5.5 at Max.
Benchmark Results With Context
On Terminal-Bench 4.0, which measures autonomous execution of multi-step command-line tasks including environment setup, debugging cycles, and long-horizon execution without human checkpoints, Sonnet 5.5 reached 70.6%. Its predecessor Sonnet 5's score on the same test was approximately 10.3%, making this roughly a sevenfold improvement in one generation. The comparison with Opus 5.5's 66.4% requires the effort-level caveat noted above: the scores were collected under different conditions — Max for Sonnet 5.5, Xhigh for Opus 5.5 — and should be read as broadly comparable rather than strictly equivalent.
On FrontierCode 1.1, which simulates realistic pull-request review by penalizing both incorrect fixes and out-of-scope changes to unrelated files, Sonnet 5.5 scored 52.1% at Xhigh effort, above GPT-6 Sol's 49.3% on the same benchmark. Opus 5.5 scores 54.4% on FrontierCode — a two-point lead over Sonnet 5.5 at Xhigh. The effort-level inversion at Max (46.2%) is specific to this benchmark's scoring structure, which punishes the kind of expansive editing that Max-effort subagent dispatch tends to produce.
CursorBench 4.0, which draws its task set from real developer coding sessions in the Cursor environment rather than synthetic scenarios, placed Sonnet 5.5 at 55.5% — a gain of more than 21 percentage points over Sonnet 5's 34.1%, and within about two points of Opus 5.5's 57.8%. GDPval-AA v2.1, which uses Elo scoring across 44 professional occupations and nine industries to estimate real-world work quality, gave Sonnet 5.5 a score of 1,844 versus Opus 5.5's 1,846. The two-point gap on an Elo scale is statistical noise; the models are effectively tied on general knowledge work.
Humanity's Last Exam with tools — a multidisciplinary reasoning benchmark designed to challenge frontier models — reached 64.5%, an improvement of roughly ten percentage points versus Sonnet 5's 54.9%. That represents the clearest evidence of across-the-board reasoning quality improvement beyond coding-specific gains.
After launch, Artificial Analysis's independent Intelligence Index placed Sonnet 5.5 second overall, behind Opus 5.5, on its composite scoring across multiple capability dimensions. That ranking positions Sonnet 5.5 as the leading mid-tier model available by this measure, above GPT-6 Sol at the same price level.
Visual Understanding: The Chartography Leap and Pokémon Red
The capability improvement with the widest implications beyond coding is Sonnet 5.5's visual comprehension. On Chartography, a benchmark that tests whether a model can accurately interpret complex charts, graphs, and data visualizations from screenshots without access to underlying data or image-processing tools, Sonnet 5.5 scored 61.6%, up from Sonnet 5's 15.6%. That is a 3.95× improvement in a single generation — the largest single-generation jump on this benchmark for any Claude model. Opus 5.5 scores 64.4% on the same benchmark, placing Sonnet 5.5 close behind its more expensive sibling.
The practical consequence became clear through Anthropic's internal Pokémon Red playthrough, which the company has used as an ongoing long-horizon visual agent test. The game presents a particularly demanding visual reading challenge: sprite-based map elements, small inventory icons, and trigger objects that must be identified from low-resolution pixel art. Prior Claude models consistently failed at specific bottlenecks. Earlier Opus versions spent extended periods trapped in the game because they could not identify and collect a key card or locate a floor switch whose pixel representation was too small to reliably parse. In May 2026, a Claude Opus model became the first to complete the game — but only with two tool augmentations: a screenshot-zoom tool and a memory snapshot tool that preserved visual state across sessions.
Sonnet 5.5 completed the full game without either tool, relying only on raw screenshots. This matters not as a gaming accomplishment but as evidence of genuine visual detail recognition at a resolution and complexity level where previous mid-tier models consistently failed. The same capability applies to reading dense financial charts, engineering diagrams, scanned documents with complex layouts, and other real-world visual material where exact detail interpretation is required.
OSWorld 2.1, which tests a model's ability to operate a desktop computer by interpreting screenshots and issuing click and keyboard actions, placed Sonnet 5.5 at 80.1% versus Opus 5.5's 81.8% — a gap of less than two percentage points, suggesting the two models are nearly equivalent on computer-use tasks.
Read more: Claude Code adds AGENTS.md support after 13 months and a TypeScript hook system
Competitive Position the Day Before OpenAI DevDay
Anthropic's timing is deliberate. Releasing Sonnet 5.5 the day before OpenAI's annual developer conference ensures that developers arriving at DevDay have had overnight access to the new model and its benchmark data. The price alignment with GPT-6 Sol at the mid-tier is equally pointed: developers evaluating the two at the same cost now have Artificial Analysis Intelligence Index data showing Sonnet 5.5 ranked above GPT-6 Sol.
Anthropic's own benchmark table did not include a publicly reported GPT-6 Sol Terminal-Bench score, noting that OpenAI did not report GPT-6 Sol performance on that test. On FrontierCode 1.1, Sonnet 5.5 at Xhigh (52.1%) leads GPT-6 Sol (49.3%). Whether OpenAI DevDay 2026 on September 29 produces model announcements that shift this competitive standing remains to be seen.
Anti-Distillation Classifiers and the Security of Model Reasoning
A feature that received less attention than the benchmark results is Sonnet 5.5's anti-distillation protection — the first time this mechanism has appeared in a Sonnet-class model.
When Claude uses extended thinking, it generates a chain-of-thought reasoning trace before producing its final answer. This reasoning trace can be streamed to the developer through the API. The concern Anthropic is addressing: a sufficiently large collection of high-quality reasoning traces can be used to train a competing model — a technique sometimes called chain-of-thought distillation. By collecting tens of thousands of Claude's reasoning outputs at scale, a third party could potentially train a model that replicates Claude's reasoning style without paying for access.
Anthropic calls its countermeasure "preserved thinking." Claude's internal reasoning cannot be decoupled from the account that created it, and safety classifiers monitor for high-volume structured extraction of reasoning outputs from a single account or coordinated account cluster — a signature distinct from normal application usage. When the classifier fires, extraction attempts are blocked.
This is the first Sonnet-class model to carry this protection; it was previously available only in Opus-class models. The inclusion at the mid-tier level reflects that Sonnet 5.5's reasoning quality now makes it a viable distillation target, which is itself a measure of capability.
What the Sonnet 5.5 Era Means for Deployment Architecture
The Sonnet 5.5 launch does not simplify Claude deployment decisions — it increases the number of variables developers need to optimize across. The model introduces a wider performance spread between effort tiers than previous Sonnet versions, meaning the same model can behave very differently depending on how it is configured.
For most real-world workloads — coding assistance, document synthesis, UI generation, business process automation — Sonnet 5.5 at Medium or High effort will deliver meaningful cost reductions versus the prior generation while matching or exceeding that generation's quality at Max. According to Anthropic's own testing, Sonnet 5.5 at Medium effort exceeds Sonnet 5's best Terminal-Bench score at less than one-tenth the cost per task.
The calculus changes for the highest-complexity agentic tasks: long autonomous coding runs, multi-file refactoring with ambiguous scope, or research synthesis over very large corpora. Here, the choice between Sonnet 5.5 at Max and Opus 5.5 at High is not obvious from list pricing. Given the per-task cost inversion at Max — approximately $7.60 for Sonnet 5.5 versus $5.98 for Opus 5.5 in independent testing — and Opus 5.5's more controlled subagent dispatch behavior, Opus 5.5 may represent the better value for sustained high-complexity work even at its higher list price per token.
What Sonnet 5.5 most clearly demonstrates is that the boundaries between model tiers are compressing. A mid-tier model now leads its own family's large-tier model on the primary agentic coding benchmark. The gap remaining between Sonnet and Opus narrows with each generation. At the same price point as a leading OpenAI mid-tier model, Sonnet 5.5 sits above that competitor on at least one independent intelligence index — for now.