Evidence at a glance
The New Price Makes Request Length the Battleground
Anthropic released Claude Haiku 5.5 on October 7, 2026, positioning it as a fast, low-cost model for frequent, low-latency, or well-scoped tasks. It is not meant to replace larger Sonnet or Opus models; it is aimed at work such as summarization, classification, routing, browser operations, and coding subtasks. For technical leaders, the important change is not simply another model release. Haiku’s price is now in a range where it can be compared directly with alternatives in its class.
Haiku 4.5 cost $1 per million input tokens and $5 per million output tokens. Haiku 5.5 drops to $0.10 and $0.50 respectively for prompts up to 100,000 tokens, matching OpenAI’s GPT-6 Luna. Above that threshold, Haiku 5.5 rises to $0.50 and $2.50. Luna does not raise its rates until 272,000 tokens, and then only to $0.20 and $0.75, so the cost curves diverge quickly for long inputs.
The Tokenizer Adds a Cost Beyond the Sticker Price
A per-million-token rate does not directly tell you the cost of a task, because the same text can be split into different numbers of tokens. Simon Willison used his token-counting tool on a long prompt and found that Haiku 5.5 counted about 25% more tokens than Haiku 4.5. Anthropic’s documentation says the same text will typically count as about 30% more. These figures describe different samples and should not be collapsed into a single precise multiplier; public materials do not specify the increase for Chinese prompts.
Long tasks can therefore face two cost increases at once: the tokenizer makes the request larger, and crossing 100,000 tokens moves it into a rate tier five times higher. Anthropic estimates that Haiku 5.5’s average operating cost is about 75% lower than Haiku 4.5’s, based on request-length distribution, and notes that around 90% of Haiku 4.5 requests were below 100,000 tokens. That is an estimate based on its workload mix, not a discount every team should expect. Haiku 5.5 has a one-million-token context window and a 128,000-token maximum output, but capacity is not the same as an economical operating range. Being able to fit text in context does not mean it is wise to send it all.
Reasoning Effort Is Adjustable, but Never Cost-Free
Haiku 5.5 enables adaptive thinking by default, with effort set to medium. Users can adjust the reasoning effort, but cannot turn reasoning off. That exposes a trade-off between capability and cost: this is not simply a cheap model with a fixed speed, since actual runtime also depends on the task and selected effort. Public materials do not disclose its internal architecture or training method, so its product settings are not enough to explain how it achieves faster inference.
Willison used a simple comparison: generating an SVG of a pelican riding a bicycle. Low effort cost 0.0936 cents and took seven seconds; max effort cost 3.3826 cents and took five minutes and nine seconds. Medium and higher settings all produced a good bicycle frame, unlike the low setting. This small test cannot predict general task performance, but it illustrates a more useful engineering point: even when an individual call is inexpensive, waiting time can vary by orders of magnitude. Lower effort may suffice for classification or simple format conversion, while cutting reasoning on a task that needs planning and multiple steps can turn savings into rework.
The Benchmark Gains Are Large; Deployment Still Needs Validation
Anthropic’s published evaluations show substantial gains over Haiku 4.5: GDPval-AA v2.1 rises from 735 to 1,620, the offline subset of OSWorld 2.1 from 15.7% to 72.4%, and Terminal-Bench 4.0 from 0.0% to 39.2%. These figures support a limited conclusion: in the vendor’s evaluations of knowledge work, computer use, and agentic coding, the new model improves markedly over its predecessor.
They do not establish whether the model is more reliable on a particular company’s tasks, nor do they guarantee consistent latency or cost. The public announcement does not provide enough information to independently reproduce all results, and it includes neither Chinese-specific scores nor a real-world latency distribution. Before deployment, teams should test their own request samples and track prompt-token counts, input and output charges, response time, and task success rate, comparing effort settings under the same evaluation. Long-tail inputs deserve particular attention: an attractive average cost does not offset bill spikes from the small share of requests that cross a pricing threshold.
Route Workloads to It Rather Than Betting on One Average Price
Haiku 5.5 is better treated as one cost tier in a task router than as a reason to move every request to one model. Short, frequent, well-scoped tasks are candidates for low or medium effort. For long inputs, teams should first assess whether compression, splitting, or routing to another model is more economical. If production traffic often approaches 100,000 tokens, the router should account for request length; otherwise, a seemingly cheap default may lose its advantage on the most expensive inputs.
Anthropic has also added monthly API credits equal to the subscription fee and allows users to disable automatic reloading. The credits expire at the end of the month rather than rolling over. This lowers the barrier to trying the API and lets teams stop requests when the balance runs out, avoiding surprise charges, but it is less suitable for workloads that fluctuate too much to use the allowance consistently. The actionable conclusion is not that Haiku 5.5 is simply cheap. Calculate costs using your own tokenizer results and request distribution, then set budget and routing rules for long inputs. Its price advantage is real, but so is the boundary written into its pricing table.