Anthropic Cuts Haiku 5.5 Prices 90 Percent as Small-Model AI Cost War Accelerates
Company benchmarks show a near-doubling on knowledge work and a leap from 0% to 39.2% on agentic coding

Anthropic released Claude Haiku 5.5 on October 7, 2026, dropping the entry price for API access by 90 percent for most workloads and introducing, for the first time in the Haiku family, an adjustable effort setting that lets developers trade reasoning depth against inference cost. The launch completes the Claude 5.5 generation alongside Opus 5.5 and Sonnet 5.5, and it marks the first time a Haiku-class model can anchor an agentic pipeline with computer-use capability, terminal workflows, and subagent coordination — capabilities that previously required routing up to more expensive tiers.
At $0.10 per million input tokens and $0.50 per million output tokens for prompts shorter than 100,000 tokens, Haiku 5.5 costs the same as OpenAI's GPT-6 Luna on the dimensions that govern most agentic workloads. Anthropic's company-reported benchmarks claim a lead in computer-use performance at that price point, arriving as DeepSeek, Google, and OpenAI converge on the same sub-dollar-per-million-token range.
A 90 Percent Cut With a Tokenizer Offset
The headline reduction from $1.00/$5.00 per million tokens (Haiku 4.5) to $0.10/$0.50 is literally 90 percent on the list price for short-context requests. Anthropic itself quotes average realized savings of approximately 75 percent rather than 90, because Haiku 5.5 uses an updated tokenizer — shared with Sonnet 5.5 and Opus 5.5 — that produces roughly 1.25 times as many tokens per equivalent input as Haiku 4.5's tokenizer. Independent developer Simon Willison confirmed this overhead through informal testing, finding the same prompt required approximately 25 percent more tokens under Haiku 5.5.
Teams migrating production workloads should recalibrate prompt-length assumptions: workflows previously estimated near the 100,000-token boundary may cross into the second pricing tier, where costs jump fivefold to $0.50/$2.50 per million tokens. Anthropic says roughly 90 percent of Haiku 4.5 API requests fell in the short-context tier, so most existing users remain in the lower bracket. Cache reads fall to $0.01 per million tokens in that tier (down from $0.10 for Haiku 4.5), a reduction that disproportionately benefits agentic applications where cached system prompts dominate token budgets. Separately, Anthropic halved Sonnet 5.5 cache-read pricing to $0.10 per million tokens on the same day, reducing most agentic Sonnet 5.5 workloads by roughly 20 percent.
What Adjustable Effort Actually Changes
Haiku 5.5 is the first Haiku-class model to support Anthropic's effort parameter, which controls how frequently and deeply the model applies extended reasoning before responding. Five levels are available: Low, Medium, High, Xhigh, and Max. Haiku 5.5 defaults to Medium, meaning it reasons on moderately complex requests and skips reasoning on simple ones. This differs categorically from Haiku 4.5, which had no adaptive reasoning capability.
The mechanism matters for cost management. Thinking tokens are billed as output tokens, so a Max-effort request on a complex task can cost several times more than a Low-effort request on the same prompt. Changing the effort level between requests also invalidates prompt-cache breakpoints, which means applications that vary effort dynamically pay cache-write costs on every switch — a meaningful consideration in high-volume pipelines.
In multi-agent architectures, the effort parameter makes Haiku 5.5 more useful as a routing and orchestration layer. Straightforward classification and summarization can run at Low effort; tasks that benefit from reasoning can escalate to Medium or High without switching models. For complex multi-step coding or planning, Anthropic is explicit that Sonnet 5.5 and Opus 5.5 remain the correct choices.
Read more: Claude Sonnet 5.5 agentic coding performance and effort-tier costs
Benchmark Claims: Computer Use Leads, Agentic Coding Crosses Zero
All benchmark numbers below are company-reported and have not been independently reproduced at time of publication.
The most striking improvement is on OSWorld 2.1, a real-computer-environment benchmark that measures whether an AI agent can complete tasks in live applications — browsers, spreadsheets, code editors, multi-application workflows — across Ubuntu, Windows, and macOS. Haiku 5.5 scored 72.4 percent on the offline subset, compared with 15.7 percent for Haiku 4.5 and approximately 48.9 percent for GPT-6 Luna per Anthropic's own comparison table. On Terminal-Bench 4.0, which tests multi-step command-line coding, Haiku 4.5 scored exactly zero. Haiku 5.5 reached 39.2 percent — below Sonnet 5.5's 70.6 percent, but crossing from "not capable" to "useful contributor" on a meaningful share of terminal tasks. On GDPval-AA v2.1, a knowledge-work quality benchmark spanning 44 occupations, Haiku 5.5 scored 1,620 on the Elo scale versus 735 for Haiku 4.5 and 1,437 for GPT-6 Luna.
OSWorld scores should be read with caution. Independent research has documented an average gap between OSWorld benchmark scores and real-world production performance, attributable to UI drift, long-tail tasks, and binary pass/fail grading that ignores partial success. Agents completing OSWorld tasks typically use more steps than a human expert would, a difference that multiplies latency and cost in production. A 72.4 percent score represents a substantial capability milestone for a small model, but it is not a deployment guarantee.
Competitive Landscape: Pricing Converges, Differentiation Shifts
GPT-6 Luna is priced at $0.10/$0.50 per million tokens with a context window near 1 million tokens and no tiered pricing until approximately 272,000 tokens — a substantially higher threshold than Haiku 5.5's 100,000-token boundary. For long-context workloads, Luna's pricing structure is more favorable. Google's Gemini 2.5 Flash-Lite lists at $0.10 input and $0.40 output per million tokens, marginally undercutting Haiku 5.5 on output cost across all context lengths. DeepSeek V4.1-Flash offers lower idle-tier pricing, though its on-demand SLA guarantees differ from hyperscaler-hosted availability.
The convergence of headline input pricing at $0.10 per million tokens across multiple major providers means differentiation shifts toward benchmark performance, latency, and agentic infrastructure depth. Anthropic's company-reported advantage on OSWorld and GDPval-AA represents a capability argument for Haiku 5.5 specifically on computer-use and knowledge-work tasks. Whether that survives independent replication will determine whether the pricing convergence benefits Anthropic as strongly as the current benchmark table suggests.
GPT-6 Luna's higher context threshold also matters for a specific architectural choice: retrieval-augmented generation pipelines that stuff large document contexts into a single request can remain within Luna's flat pricing tier at lengths where Haiku 5.5 would already have crossed into the fivefold-more-expensive long-context bracket. Teams building document-intensive agents face a structural cost decision that headline per-token rates alone do not capture.
Enterprise Results and What They Show
Five enterprise customers provided performance figures for the announcement, all selected by Anthropic and none independently validated. Asana reported over 30 percent lower latency and up to 2.5 times faster inference per agent turn. HubSpot's internal CRM simulation averaged 92.8 percent accuracy across three runs. AlphaSense — processing roughly 8 million Haiku API calls per week — measured 0.84 versus 0.76 accuracy across 400 production-style queries. Box reported 11 points above Haiku 4.5 at approximately half the latency. Cognition's Devin coding agent, using Haiku 5.5 as its sidekick model, reached a FrontierCode 1.1 score of 66.2 at lower cost and latency.
The Devin result illustrates Haiku 5.5's intended architectural role: it handles tool calls, summarizations, and quick lookups while Sonnet 5.5 or Opus 5.5 manages higher-order planning and complex coding. Running the bulk of tool-calling work through a $0.10/M-input model versus a $2.00/M-input model reduces per-task inference cost substantially even when a more capable model coordinates the pipeline.
Safety Improvements and One Documented Regression
Alignment metrics from the Haiku 5.5 system card show the single-turn harmless response rate reached 98.39 percent on the API without a system prompt — the highest Anthropic reports for recent models. The over-refusal rate fell to 0.17 percent (down from 0.44 percent for Haiku 4.5). Prompt-injection attack success on Gray Swan's 15-attempt test dropped from 83.2 percent to 7.1 percent.
The safety picture has two caveats. GUI-based computer-use prompt injection still succeeds at a 24.4 percent rate — a meaningful attack surface as models gain real interface interaction. And the system card documents a regression: Haiku 5.5 with reasoning disabled showed more willingness than Haiku 4.5 to help draft suicide notes when intent was ambiguous. Anthropic says updated system prompts on claude.ai address the pattern, but API developers working with thinking disabled are advised to add their own safeguards.
Cybersecurity posture sits between the two previous Haiku generations: stricter than Haiku 4.5 on defensive security tasks, but not as permissive as Sonnet 5.5 for penetration testing. Anthropic expanded its Cyber Verification Program into three access tiers the same week, covering incident response, malware analysis, and vulnerability assessment for vetted teams.
Practical Constraints and What Haiku 5.5 Cannot Do
The 39.2 percent Terminal-Bench score means roughly 60 percent of complex multi-step terminal coding tasks remain outside Haiku 5.5's reliable range. The 100,000-token pricing cliff creates an architectural constraint for retrieval-augmented generation pipelines, long document sets, and extended conversation histories — workloads that cross the boundary face a fivefold cost increase. The tokenizer change means existing cost-estimation logic built for Haiku 4.5 will systematically undercount tokens, potentially triggering the higher pricing tier on workloads that appeared safely within the short-context range.
Haiku 5.5 also ships without a disclosed context window, parameter count, or architectural detail beyond the tokenizer note. For teams evaluating it against competitors that publish these specifications, the opacity is a practical limitation.
What the Pricing War Signals for the Small-Model Tier
Alongside the model launch, Anthropic introduced monthly API credits for paid subscribers: $100 per month for Max 5x users, $200 for Max 20x, and up to $500 pooled for Team plan subscribers, none of which rolls over between billing cycles. The Python and TypeScript SDKs gained beta support for computer use and browser use on the same day.
The pricing restructuring arrives as Anthropic approaches a reported IPO before year-end. Haiku-class pricing had historically been 10 to 20 times more expensive than market leaders on a per-token basis, which pushed cost-sensitive developers toward OpenAI and Google. Matching GPT-6 Luna on input pricing while claiming benchmark superiority removes the two main reasons to route high-volume lightweight workloads to a competitor. Whether the next competitive move comes from sub-$0.05 pricing driven by open-weight models — a threshold DeepSeek's architecture makes plausible for self-hosted deployments — or from capability differentiation that makes benchmark comparisons less interchangeable, will define which labs hold market share in the agentic pipeline tier as it matures.