Anthropic has released Claude Haiku 5.5, its cheapest and fastest small model to date. It targets high-volume work like summaries, compaction, classification and subagent tasks. It keeps a 1M token context window and up to 128K output tokens. Pricing starts at $0.10 per million input tokens and $0.50 per million output tokens. That is 90% below Claude Haiku 4.5 for prompts up to 100K tokens.
Is it deployable? Yes, as a hosted API. Haiku 5.5 is generally available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS.
What Anthropic Shipped
Haiku 5.5 is the first Haiku-class model with an adjustable effort setting. Adaptive thinking is on by default, and the effort parameter defaults to medium.
It takes text and images, outputs text, and has a June 2026 knowledge cutoff. Batch jobs support up to 300K output tokens in beta.
It is important to note two key things. Non-default temperature, top_p or top_k values return a 400 error. The new tokenizer also counts the same text as roughly 30% more tokens than Haiku 4.5. The migration guide covers both.
Pricing: A 2-Tier Structure
Pricing splits at 100K prompt tokens. Up to 100K, input costs $0.10 and output $0.50 per million. Cache reads cost $0.01 and 5-minute cache writes $0.125. Above 100K, rates rise to $0.50 input and $2.50 output.
Haiku 4.5 charged $1 input and $5 output. Anthropic says about 90% of Haiku 4.5 requests stayed under 100K tokens. After adjusting for the tokenizer, it estimates Haiku 5.5 runs about 75% cheaper on average. Batch processing takes another 50% off.
GPT-6 Luna lists identical short-context rates. Its higher tier starts only above 272K input tokens, at $0.20 and $0.75. For a 150K-token prompt, Luna is cheaper on list price.
Benchmarks
All figures below are Anthropic-reported (see the system card):
- OSWorld 2.1 (offline subset): 72.4%, versus 48.9% for GPT-6 Luna and 15.7% for Haiku 4.5.
- Terminal-Bench 4.0: 39.2%, versus 16.4% for Luna and 0.0% for Haiku 4.5.
- FrontierCode 1.1 (Main): 46.4%, versus 42.4% for Luna.
- Humanity’s Last Exam: 45.9% without tools and 57.4% with tools.
- GDPval-AA v2.1: 1620, versus 1437 for Luna and 735 for Haiku 4.5.
Sonnet 5.5 still leads every row, including 70.6% on Terminal-Bench 4.0. Anthropic itself recommends Sonnet 5.5 and Opus 5.5 for complex agentic coding.
Best Use Cases
Three workloads fit Haiku 5.5 best:
- The first is subagent work under Opus 5.5 or Sonnet 5.5. At Rogo, a Haiku 5.5 subagent pulls a 10-K revenue line while a bigger model builds the deck.
- The second is high-volume document Q&A and summarization. AlphaSense tested it on a feature handling about 8M calls a week.
- The third is speed-sensitive work like live customer support and browser use.
Haiku 5.5 vs Its Closest Competitors
| Feature | Claude Haiku 5.5 | GPT-6 Luna | Gemini 3.5 Flash-Lite |
|---|---|---|---|
| Input / output (per 1M) | $0.10 / $0.50 | $0.10 / $0.50 | $0.30 / $2.50 |
| Long-prompt pricing | $0.50 / $2.50 above 100K | $0.20 / $0.75 above 272K | Flat rate |
| Cache read (per 1M) | $0.01 | $0.01 | $0.03 + storage |
| Context window | 1M tokens | 1,050,000 tokens | 1,048,576 tokens |
| Max output | 128K tokens | 128,000 tokens | 65,536 tokens |
| Inputs | Text, images | Text, images | Text, image, video, audio, PDF |
| Reasoning control | Adaptive thinking + effort (default medium) | reasoning.effort none to max (default medium) | Thinking supported |
| Computer use | SDK support in beta | Supported (Responses API) | Supported (Preview) |
| Batch discount | 50% | 50% (Batch and Flex) | 50% |
| Knowledge cutoff | Jun 2026 | May 18, 2026 | Not listed on model page |
| Where to run | Claude API, Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS | OpenAI API | Gemini API |
Sources: Anthropic docs, Anthropic announcement, OpenAI model page, OpenAI pricing, Google model page, Google pricing. Standard-tier list prices, verified October 7, 2026.


