Evidence at a glance

$0.10/$0.50 Input/ token, Luna≤100KInput
$0.50/$2.50 Input/ token>100KInput
100KInput 5xEvidence
Haiku 4.5 75%Evidence
Haiku 43,Luna 38Intelligence Index
effort Index Approx. use 162K token,LunaApprox. 50KEvidence

A Cheaper Model Has to Do More Than Answer Faster

Anthropic released Claude Haiku 5.5 in October 2026, nearly a year after Haiku 4.5. The company positions it as a small model for high-throughput, cost-sensitive work. It is available on the Claude Platform and in Claude Code, and Anthropic’s recommended pattern is not to hand it every task on its own, but to pair it with Opus 5.5 or Sonnet 5.5 for subtasks such as summaries, context compaction, and database queries.

That positioning changes the comparison that matters. A same-price contest with GPT-6 Luna is striking, but the more practical question for technical leaders is whether a primary model can safely delegate enough work to a cheaper model to reduce expensive reasoning calls without losing the savings to validation, retries, and recovery. Devin’s launch-day integration offers a concrete example: Cognition reported a 58.4% score on FrontierCode 1.1 and roughly one-eighth the per-task cost of Sonnet 5, recommending Haiku as a sidekick under an Opus 5.5 lead in Fusion. Those are vendor-reported results on a specific evaluation, not a general cost guarantee for coding work.

There Are Two Different Bills Behind “Same Price”

For prompts up to 100K tokens, Haiku 5.5 is priced at $0.10 per million input tokens and $0.50 per million output tokens, the same as GPT-6 Luna. Once a Haiku prompt exceeds 100K tokens, the rates rise to $0.50 and $2.50 per million input and output tokens, respectively, five times the lower tier. Luna’s long-input surcharge begins at 272K tokens. “Same price” therefore describes the standard tier, not every request a production system may send.

The second bill is the number of tokens needed to finish the task. On Artificial Analysis’s Intelligence Index, Haiku 5.5 scored 43 at maximum effort, compared with Luna’s 38, but used about 162K output tokens per task on average, against roughly 50K for Luna. At lower effort, Haiku scored 38 with about 55K tokens, while Luna reached the same score with about 50K. The evaluation had not yet fully modeled Haiku’s fivefold price tier above 100K input tokens, so these scores should not be converted into a firm cost-per-task conclusion.

Effort Turns Delegation into a Scheduling Choice

Haiku 5.5 is the first Haiku model to support Anthropic’s effort settings and adaptive thinking. Callers can adjust how much reasoning the model applies, making the question of how much work to delegate a workflow budget decision rather than a simple model choice. A straightforward, verifiable task can receive less effort; a riskier or more involved task can receive more, or be handed back to a stronger primary model. The published material describes these controls but does not disclose the model architecture or training method, so it does not support attributing the gains to a particular technique.

This design fits the proposed division of labor. A primary model breaks down the goal and judges whether the result is acceptable, while Haiku handles many narrow execution tasks; if an output fails the checks, it can be escalated. Cost control becomes a matter of deciding how much reasoning each task class gets and when to escalate, rather than simply selecting the cheapest model. The materials do not provide success-rate or latency data for this escalation pattern in real businesses, so teams still need to measure end-to-end outcomes rather than rely on the price of an individual call.

Benchmark Scores Do Not Add Up to One Overall Ranking

The available evaluations support the claim that Haiku leads on some tasks, not that it beats Luna across the board. Beyond the Intelligence Index, Haiku 5.5 scored 33% on Terminal-Bench 4.0, compared with 13% for Luna. But on AutomationBench-AA, Haiku scored 35%, while other models in the comparison scored between 53% and 60%. Artificial Analysis said excessive refusals affected Haiku’s result and that it planned to retest after addressing the issue. That finding points both to a deployment risk and to a limitation in the evaluation that still needs review.

Knowledge and hallucination metrics also cannot be reduced to whichever number looks better. On AA-Omniscience, Haiku’s accuracy was 36% and its hallucination rate was 40%; Luna scored 44% and 77%, respectively. Haiku hallucinated less but also answered fewer questions correctly. That looks like a trade-off between caution and answer coverage, not proof that Haiku is more reliable for high-stakes work. The independent Terminal-Bench result also differs from the promotional figure, reinforcing the need to read benchmark scores alongside test conditions, refusal handling, and retest status.

Start by Putting It in a Pipeline That Can Fall Back

Haiku 5.5 also supports a 1M-token context window and up to 128K output tokens, but a large window is not the same as a cheap one: inputs above 100K enter a more expensive tier. Its newer tokenizer also counts roughly 30% more tokens for the same text than Haiku 4.5, so teams should not simply reuse the older model’s token budgets. These differences matter for long-document summaries, retrieval-result consolidation, and agents that run over extended sessions.

A sensible deployment is to start with subtasks that can be checked, retried, and escalated when they fail, while recording input length, output tokens, refusal rate, retry count, and escalation rate for each task class. Then compare Haiku with Luna on your own workload instead of treating an Intelligence Index score as the total bill. If tasks often exceed 100K input tokens, produce large outputs, or fail in ways the primary model cannot detect, matching list prices is not a sufficient reason to switch. For large volumes of small, clearly bounded execution tasks, Haiku is much closer to the role it was designed to fill.