TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Anthropic has released Claude Haiku 5.5, a small model it describes as its fastest to date and designed for high-volume, cost-sensitive tasks. The company says it costs about 75% less to run on average than Haiku 4.5; benchmark results and customer performance reports come from Anthropic and its cited early testers.
Anthropic has released Claude Haiku 5.5, a small model aimed at fast, high-volume tasks, and says it costs about 75% less to run on average than Haiku 4.5. The model is available now through the Claude Platform and major cloud providers, including Amazon Web Services, Google Cloud and Microsoft Azure, according to the company; Claude inference is also available in-country in India via AWS.
Anthropic positions Haiku 5.5 for workloads such as summarization, classification, database queries and prompt compaction, as well as speed-sensitive customer support and browser use, while its cyber-focused Claude work includes Claude access for cybersecurity teams. It says the model can also act as a subagent alongside Sonnet 5.5 and Opus 5.5 on coding work, but cautions that its larger models remain better suited to complex agentic coding tasks.
The company lists Haiku 5.5 input pricing at $0.10 per million tokens for prompts up to 100,000 tokens and $0.50 above that threshold, a topic explored in a study of Claude pricing and value. Output costs $0.50 and $2.50 per million tokens, respectively. Anthropic says around 90% of requests to its previous Haiku model fell within the lower prompt-length tier. The listed rates are below Haiku 4.5’s $1 input and $5 output prices per million tokens.
Anthropic also announced a 50% cut in Sonnet 5.5 cache-read prices, to $0.10 per million tokens, and says that reduces the model’s cost by around 20% on most agentic work. It is introducing a monthly API credit for Claude Max and Team subscribers to support development on its platform; the announcement does not specify the credit amount.
Lower Costs for Routine AI Work
The launch targets a practical barrier to deploying AI agents: the cost and delay of handling many small, repeated steps. If Haiku 5.5 performs adequately for tasks such as summaries, routing and classification, a lower per-token price could make it more viable to run those operations at high volume or as part of a larger agent workflow.
Anthropic’s pricing and speed claims matter most to developers comparing models for particular jobs, rather than choosing one model for every task. The company’s own framing draws that distinction: Haiku 5.5 is for narrowly scoped and latency-sensitive work, while Sonnet 5.5 and Opus 5.5 remain stronger choices for complex coding. Actual savings will depend on prompt size, output length, cache use and how often a task needs retries or escalation to a larger model.
The additional Sonnet cache-read reduction broadens the announcement beyond the Haiku launch. For teams already using Sonnet in agent workflows, Anthropic says the change can lower costs on typical workloads; the size of the benefit for any particular application depends on its cache-read usage.
As an affiliate, we earn on qualifying purchases.
How Haiku Fits Anthropic’s Lineup
Haiku is Anthropic’s smaller model family, presented for work where speed and operating cost are priorities. Haiku 5.5 follows Haiku 4.5, and the company says its new model is both cheaper to run and faster. Anthropic calls it the fastest model it has released, a company claim rather than an independently verified comparison.
The release adds an adjustable effort setting, which Anthropic says lets users choose a balance between cost and intelligence. Its system card describes the company’s evaluation methods and results. Published benchmark figures include 72.4% on the offline subset of OSWorld 2.1 and 39.2% on Terminal-Bench 4.0; those results should be read alongside the stated test conditions and compared with other models only where the evaluations are comparable.
In its announcement, Anthropic says Haiku 5.5 improved on almost all of its alignment evaluations compared with Haiku 4.5. It also describes safeguards that block penetration testing and other cyber techniques it judges more likely to be used by attackers, while allowing a wider range of defensive tasks than its Sonnet 5.5 safeguards. These are vendor-reported evaluations and policy descriptions; independent testing is not included in the supplied announcement.
“Claude Haiku 5.5 is the cheapest, fastest, and most capable small model we’ve ever released.”
— Anthropic
As an affiliate, we earn on qualifying purchases.
Independent Results Still Pending
The available announcement does not provide independent evaluations of Haiku 5.5 or enough detail to establish how its performance compares across real-world applications. Anthropic says customer testing was consistent with its cost and performance findings, but the reported Asana results reflect one customer’s test suite and comparison model; the company did not identify that model in the supplied material.
It is also unclear how much individual customers will save. Anthropic’s estimate of an average 75% lower running cost does not specify the usage mix behind that average, while the actual bill depends on token volume, prompt length and task design. The announcement does not state the monthly API credit amount, eligibility details beyond Claude Max and Team subscribers, or when full program terms will be available.
cost-effective AI summarization software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Availability and Pricing Details
Developers can use Haiku 5.5 now through the Claude Platform with the model identifier claude-haiku-5-5, or access it through AWS, Google Cloud and Microsoft Azure, according to Anthropic. The company points users to a migration guide and a system card for implementation and evaluation details.
Teams weighing a move can test the model against their own workloads, including response quality, latency and total cost, rather than relying only on headline prices or benchmark scores. Further information on the monthly API credit, including its value and conditions, remains to be provided in the supplied announcement.
high-volume AI processing platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Claude Haiku 5.5?
It is Anthropic’s newly released small model, aimed at high-volume, cost-sensitive and speed-sensitive tasks such as summarization, classification and customer support.
How much does Haiku 5.5 cost?
Anthropic lists input at $0.10 per million tokens for prompts up to 100,000 tokens and $0.50 above that length. Output is $0.50 or $2.50 per million tokens, respectively. The company says average running costs are about 75% lower than Haiku 4.5.
Where can developers access it?
Anthropic says Haiku 5.5 is available now on the Claude Platform, AWS, Google Cloud and Microsoft Azure. On the Claude Platform, its model identifier is claude-haiku-5-5.
Is Haiku 5.5 better for coding than Sonnet 5.5?
Anthropic says Haiku 5.5 can support coding workflows as a subagent, but says Sonnet 5.5 and Opus 5.5 are better suited to complex agentic coding. Haiku is positioned for narrower tasks where speed and cost matter more.
What other pricing changes did Anthropic announce?
Anthropic cut Sonnet 5.5 cache-read pricing by 50%, to $0.10 per million tokens, and says this lowers costs by around 20% on most agentic work. It also announced a monthly API credit for Claude Max and Team subscribers, without specifying the amount in the supplied material.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
