Anthropic’s new small model targets repetitive work at scale. The launch also lowers Sonnet 5.5 cache costs and introduces API credits for Max and Team subscribers.
Anthropic launched Claude Haiku 5.5 on October 7, positioning it for fast, high-volume workloads. Input and output token prices are 90% lower than Haiku 4.5 for prompts up to 100,000 tokens. After accounting for request sizes and token consumption, the company estimates approximately 75% lower average workload costs. [1]
The intended workloads include document summaries, classification, information extraction and small assignments within larger projects. Anthropic still recommends Sonnet 5.5 and Opus 5.5 for complex agentic coding. [1]
What changes
- Lower pricing: $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. [1]
- Adjustable effort: the first Haiku model with a setting to balance reasoning quality, speed and cost. [1]
- More context: a one-million-token context window and up to 128,000 output tokens, according to the documentation. [2]
- Subagent work: a larger model can delegate a defined task, such as retrieving a fact or summarizing a document. [1]
- Availability: Anthropic’s platform, AWS, Google Cloud and Microsoft Azure are included in the announcement. [1]
Tokens are the units used to process content and meter usage. They do not correspond to a fixed number of words.
Why a 90% price cut produces an estimated 75% saving
The two percentages measure different things. The 90% reduction applies to token rates for prompts up to 100,000 tokens. For longer prompts, rates are 50% lower than Haiku 4.5. [1]
Launch pricing in US dollars per million tokens:
| Usage | Haiku 5.5: prompts up to 100K | Haiku 5.5: prompts over 100K | Haiku 4.5 |
|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 |
| Output | $0.50 | $2.50 | $5.00 |
| Cache reads | $0.01 | $0.05 | $0.10 |
| Cache writes, five minutes | $0.125 | $0.625 | $1.25 |
Source: Anthropic. Claude Platform rates; check other providers’ terms separately. [1]
The estimated 75% average saving also accounts for how many tokens are needed to complete a job. The new tokenizer represents the same text with more units. Technical documentation puts the increase in input tokens at approximately 30%, depending on the content. [2]
For example, a workload consuming one million input tokens and 200,000 output tokens, split into requests in the lower price tier, would cost $0.20 on Haiku 5.5 versus $2 on Haiku 4.5 at basic rates. This assumes identical consumption and excludes caching, tools and retries. Actual consumption may change.
Token prices have fallen. Each application’s savings depend on how much work the model needs to deliver a useful result.
Benchmark gains need context
Anthropic’s launch table shows improvements over Haiku 4.5 in professional work, computer use, reasoning and coding. Haiku 5.5 also leads GPT-6 Luna on the comparable evaluations shown. Sonnet 5.5 remains ahead in every row of that table. [1]
| Evaluation | What it tests | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|---|
| GDPval-AA v2.1 | Professional work, Elo | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 | Professional work, Elo | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1, offline subset | Computer use, partial credit | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity’s Last Exam, no tools | Academic knowledge and reasoning | 45.9% | 10.2% | Not provided | 56.9% |
| Humanity’s Last Exam, with tools | Academic knowledge and reasoning | 57.4% | 18.7% | Not provided | 64.5% |
| Terminal-Bench 4.0 | Professional terminal tasks | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1, Main | Agentic coding | 46.4% | Not provided | 42.4% | 52.1%* |
| Chartography, no tools | Visual reasoning | 46.4% | 6.4% | 29.1% | 61.6% |
Results reported in Anthropic’s launch materials. Sonnet’s FrontierCode result uses Xhigh effort. Professional-work scores are not percentages. [1]
These results do not establish superiority on every task. The System Card says GDPval-AA and AA-Briefcase were run independently by Artificial Analysis. Other evaluations require their own provenance and conditions. Unless otherwise noted, Haiku results use adaptive thinking at maximum effort. [11]
OSWorld contains a particularly important distinction: 72.4% is the average partial-credit score, not the proportion of fully completed tasks. Haiku’s strict pass rate was 37.1%. The test used 82 offline tasks without internet access, at maximum effort, averaging five independent attempts per task. [11]
Effort also changes professional-work scores. Haiku scored 1620 at Max and 1277 at Medium on GDPval-AA, and 1578 versus 1372 on AA-Briefcase. Medium is the API default. According to the document, Medium used about one tenth of Max’s output tokens on the first evaluation and less than one quarter on the second. [11]
On Terminal-Bench, Haiku’s 39.2% was measured without internet access, with safeguards enabled and no fallback model. Sonnet’s evaluation used a fallback for flagged requests. Execution systems and conditions therefore matter alongside the headline scores. [11]
The initial market response
Early responses include media coverage, partner distribution and customer evaluations. They show interest without establishing sustained adoption.
Reuters highlighted the expansion of the Claude 5.5 family with its third release in less than a month. [7] VentureBeat emphasized the price reduction and competition with GPT-6 Luna. [8]
AWS published Amazon Bedrock guidance, describing Opus as the planner and Haiku as the worker for defined assignments. [5] GitHub announced Haiku 5.5 in Copilot, rolling out gradually to Pro, Pro+, Max, Business and Enterprise users. [6]
In testimonials selected by Anthropic, Asana reported task-completion latency falling by more than 30% in its tests. Such reports are useful application signals, but are different from independent audits. [1]
An external measurement is already available from Artificial Analysis. The Medium page consulted on October 7 showed 34 points on its Intelligence Index and approximately 137 output tokens per second. It also reported 12.63 seconds to the first token. [9]
Generating text rapidly after a response starts is different from starting immediately. These figures describe the measurement consulted, not a guarantee for live support or every deployment.
Availability and migration
According to the product page, Free, Pro, Max, Team and Enterprise users can select Haiku 5.5 on Claude.ai. It is also available in Claude Code. Its Anthropic API identifier is claude-haiku-5-5. Model availability does not remove plan limits. [3]
For developers, changing the model name alone may be insufficient. The migration guide calls for reviewing token counts, thinking configuration and older parameters. Response blocks should be selected by type instead of assuming the first block contains the textual answer. [4]
The default effort is Medium. The model specification clarifies that the cache-write rates in the table cover five-minute retention; one-hour writes cost $0.20 or $1 per million tokens, depending on prompt length. [12]
Python and TypeScript SDK updates add computer-use and browser-use support in beta. This enables integration; it does not provide a ready-made agent with access to a company’s systems. [1]
Cheaper Sonnet caching and API credits
Sonnet 5.5 cache reads fall from $0.20 to $0.10 per million tokens. Caching reuses previously processed context, such as repeated instructions and documents. Anthropic estimates approximately 20% lower costs on most agentic work, depending on how much consumption comes from cached context. [1]
The monthly API credits announced for rollout this week have separate rules. [10]
| Plan | Monthly credit |
|---|---|
| Max 5x | $100 |
| Max 20x | $200 |
| Team | $20 per Standard seat and $100 per Premium seat, pooled up to $500 |
Subscribers must link a Claude Console organization. Unused credits expire at the end of the billing cycle and do not roll over. Free, Pro and Enterprise are not eligible. Credits do not cover interactive Claude Code, extra app usage or use through AWS, Google Cloud and Microsoft Foundry. [10]
Safety and the practical implication
Anthropic reports improved alignment results, fewer misaligned behaviors and less cooperation with misuse. It also says cybersecurity safeguards are more restrictive than Haiku 4.5’s, including blocks on penetration testing under standard conditions. These are the provider’s statements about its evaluations and policies. [1]
MnzAI Labs’ editorial assessment: the main implication is how work can be divided within agentic workflows. A larger model can interpret a complex request while Haiku summarizes documents, classifies items or retrieves supporting information.
The useful decision metric is cost per correctly completed task, including retries and review. If lower prices come with sufficient quality in real workflows, previously expensive automation may become viable.
Sources
Sources consulted on October 7, 2026. Links in English.
- Anthropic: announcement, pricing, benchmarks and footnotes.
- Claude Platform: new capabilities and technical changes.
- Anthropic: Haiku product page and availability.
- Claude Platform: migration guide.
- AWS: Amazon Bedrock launch and usage guidance.
- GitHub: Haiku 5.5 in Copilot.
- Reuters: expansion of the Claude 5.5 family.
- VentureBeat: launch coverage and price competition.
- Artificial Analysis: Haiku 5.5 Medium evaluation.
- Claude Platform: API credit amounts, eligibility and rules.
- Anthropic: Haiku 5.5 System Card, sections 8.1, 8.4, 8.9.3, 8.10.2 and 8.10.3.
- Claude Platform: official model specification, default effort and cache rates.
