A cheaper, faster small model for the high-volume AI work your team already runs
On Oct 7, 2026, Anthropic released Claude Haiku 5.5, calling it the cheapest, fastest, and most capable small model it has shipped. Anthropic positions it for summaries, classification, database queries, subagents, live support, and browser use, and says it costs about 75% less to run than Haiku 4.5 on average.
What shipped
On October 7, 2026, Anthropic released Claude Haiku 5.5, describing it as the cheapest, fastest, and most capable small model it has shipped. Anthropic says Haiku 5.5 is built for high-volume, cost-sensitive work such as summaries, compactions, database queries, and classification, and that it pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding jobs. Because it is also Anthropic's fastest model at standard speed, the company positions it for live customer support and browser use.
Anthropic says Haiku 5.5 costs about 75% less to run than Haiku 4.5 on average. The API id is `claude-haiku-5-5`. It is available now on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure. It is also the first Haiku-class model with an adjustable effort setting, so you can trade cost against intelligence on the same job.
What it means for a business owner
If your team already runs AI on the same jobs every day (ticket triage, document Q&A, recurring summaries, lookups inside a CRM), the bill is mostly volume, not the one hard agent that runs once. Anthropic's own contrast is useful here: Sonnet 5.5 and Opus 5.5 stay better for complex agentic coding, while Haiku 5.5 is for narrowly scoped tasks that used to be too expensive to run at scale.
The owner move is routing, not a wholesale model swap. Keep the large models on multi-step plans and hard coding. Put Haiku 5.5 on the repetitive narrow work, then measure latency and spend on one workflow before you expand. Alongside the Haiku launch, Anthropic also cut Sonnet 5.5 cache reads in half (to $0.10 per million tokens), which it says makes Sonnet about 20% cheaper on most agentic work.
No Consultiply-Anthropic partnership is implied here. This is an owner reading of Anthropic's announcement and system card.
How you could use this yourself this week
- 01
List your high-volume AI jobs.
Write down summaries, classification, ticket triage, document lookups, recurring report pulls, and any subagent that fetches one narrow fact. Note which model runs each job today. One page is enough.
- 02
Pilot Haiku 5.5 on one narrow workflow.
Pick one high-volume, narrowly scoped job. Run it on Haiku 5.5 for a week of real traffic. Compare latency and spend against your current model before you change anything else.
- 03
Keep Sonnet or Opus on complex agent work.
Anthropic says Sonnet 5.5 and Opus 5.5 remain better for complex agentic coding (for example Terminal-Bench-style work). Do not force Haiku onto those jobs just because it is cheaper.
- 04
Name who watches cost and quality.
One named person owns the routing rule and the weekly cost and quality check. Expand Haiku only when both hold up on the pilot workflow.
The vendor claims, labeled as theirs
Pricing, benchmarks, and customer quotes below are Anthropic's. For prompts up to / over 100k tokens, Anthropic lists Haiku 5.5 cache reads at $0.01 / $0.05, cache writes at $0.125 / $0.625, input at $0.10 / $0.50, and output at $0.50 / $2.50 per million tokens. Haiku 4.5 input was $1.00 and output $5.00. Anthropic says about 90% of Haiku 4.5 requests were in the up-to-100k bucket, and that tokenizer changes also affect tokens per task.
On Anthropic's benchmark table, Haiku 5.5 scores 72.4% on OSWorld 2.1 (offline subset) versus 15.7% for Haiku 4.5, 39.2% on Terminal-Bench 4.0 versus 0.0%, and 45.9% on Humanity's Last Exam with no tools versus 10.2%. Sonnet 5.5 remains higher on several agentic coding measures. Customer quotes via Anthropic include Asana (over 30% latency reduction and up to 2.5x faster inference per agent turn), HubSpot (92.8% on its CRM eval suite), AlphaSense (0.84 versus 0.76 on Ask in Document versus Haiku 4.5, on a class of work doing about 8 million calls a week), Box (11 points higher than Haiku 4.5 at about half the latency), and Cognition/Devin (Haiku as a sidekick in Devin Fusion). Details sit in the Claude Haiku 5.5 System Card.
The honest limits
A cheaper small model does not replace judgment about which jobs are narrow and which need a larger lead model. Anthropic is explicit that Sonnet 5.5 and Opus 5.5 stay stronger on complex agentic coding. Haiku 5.5 helps when you already have volume work that was cost-prohibitive, not when you need the deepest multi-step agent.
Treat this as a routing decision. List the high-volume jobs, pilot Haiku on one narrow workflow, keep Sonnet or Opus on hard agent work, and have a named person watch cost and quality every week.
Sources
Want this working in your business?
Tell us what you are trying to fix and we will point you at the smallest useful next step.
Talk to us