Evidence at a glance

2026 10 7 。Evidence
1M tokens, 128K tokens。Evidence
API 300K tokens, beta。Evidence
Input $0.10。≤10 tokens
Input $0.50。>10 tokens
$0.50。≤10 tokens

A Million Tokens Is First a Capacity Limit

Anthropic released Claude Haiku 5.5 on October 7, 2026, positioning it as a small model for high-volume, cost-sensitive work. The launch materials name summarization, classification, information extraction, and customer support as intended uses. They also describe it as a possible subagent inside a system built around larger models. Haiku 5.5 accepts text and image inputs, returns text, and is available through the Claude API and multiple cloud platforms.

The announcement places two easily confused figures side by side: a 1M-token context window and a starting input price of $0.10 per million tokens. The first describes how much context the model can handle. The second is not a single rate that applies to every prompt length. For a technical lead, the question is not simply whether a million tokens fit. It is which pricing tier a request enters at the lengths the system will actually send.

The Price Threshold Changes What “Cheap” Means

Haiku 5.5 uses prompt-length pricing tiers. For prompts of up to 100K tokens, input costs $0.10 per million tokens. Above 100K, the rate is $0.50 per million. Output has corresponding tiers: $0.50 per million output tokens for shorter prompts and $2.50 for longer ones. The 1M context window means the model can handle longer inputs, but it does not mean that a million-token prompt remains at the lowest rate. Context capacity and the scope of the low-price tier are separate product conditions.

The comparison with the previous Haiku 4.5 is striking, but the basis of the comparison matters. Haiku 4.5 has standard rates of $1 per million input tokens and $5 per million output tokens. By list price, Haiku 5.5 is 90% cheaper in the prompt tier up to 100K and 50% cheaper in the tier above it. Those differences create the possibility of lower costs, but they do not establish how much a particular business task will save. Prompt length and output volume still determine the total bill.

A Migration Estimate Must Include the Tokenizer

Anthropic's migration guide warns that the same text produces about 30% more tokens on Haiku 5.5 than on Haiku 4.5. That makes a per-token price comparison less straightforward. Even when the rate per million tokens falls, the same business text may generate more billable tokens on the new model. If many requests sit near the 100K tier boundary, a change in token count may also affect which pricing tier applies. Multiplying the old system's token volume by the new model's unit price is therefore not a sound migration estimate.

Anthropic estimates that Haiku 5.5 will have an average running cost about 75% lower than Haiku 4.5. That is the company's estimate of an average, not a discount guaranteed for every workload. A more useful migration calculation starts with actual request samples and measures token counts on both models, the prompt-length tier, and the amount of generated output. The Batch API permits outputs up to 300K tokens, but that capability is marked beta. It should not be treated as interchangeable with a stable production capability when setting operating assumptions.

Adjustable Effort Fits a Routing Strategy

Haiku 5.5 supports adaptive thinking and exposes an effort parameter that developers can use to adjust how deeply it reasons. The default is medium. Public documentation establishes this control at the calling interface, but does not disclose the model's parameter count, network architecture, or training method. “Small model” is therefore a product positioning, not evidence for a particular internal design. The practical interface for an engineering team is the ability to vary reasoning effort by task rather than assuming every request needs the same treatment.

That makes Haiku 5.5 more naturally a layer in a routing strategy than a replacement for every model handling complex work. Summarization, classification, and extraction can often be isolated as relatively bounded steps. Harder steps can receive more effort or be escalated to Sonnet or Opus. Anthropic also says that complex agentic programming remains better suited to larger models. The benefit in practice depends on keeping simple requests on the lower-cost path, and on ensuring that escalation and result-checking costs do not consume the savings.

Benchmarks Show Progress, Not a Substitute for Validation

Anthropic reports a score of 72.4% for Haiku 5.5 on the offline subset of OSWorld 2.1, compared with 15.7% for Haiku 4.5. On Terminal-Bench 4.0, Haiku 5.5 scores 39.2%, Haiku 4.5 scores 0.0%, and Sonnet 5.5 scores 70.6%. These figures support a limited conclusion: on the tests reported by the vendor, the new Haiku improves substantially over its predecessor, while still trailing the larger Sonnet on Terminal-Bench.

Benchmark scores cannot be converted directly into a team's production success rate. The OSWorld result of 72.4% is for an offline subset, not for every computer-use scenario. The Terminal-Bench result likewise does not establish that complex agentic programming can broadly be handed down to Haiku. Before deployment, start with frequent, well-bounded requests and use your own traffic to measure token distribution, output size, escalation rates, and task quality. Then compare costs across the full workflow. If prompts often exceed 100K tokens, or the cost of failure makes escalation and review routine, a low list price is not enough to justify the choice.