# Haiku 5.5 and the Cost of Useful Work

- Date: 08 Oct 2026 (2026-10-08T16:04:05.000Z)
- Summary: Haiku 5.5 sharply reduces short-context token prices, but context thresholds, tokenization, and reasoning effort complicate the savings. Practitioner subscription comparisons reinforce the case for measuring accepted work rather than nominal inference dollars.
- Tags: `digest`, `ai-discourse`, `claude-haiku-5-5`, `agent-economics`, `llm-pricing`, `coding-agents`, `subscriptions`

## Sources

1. [Anthropic - Introducing Claude Haiku 5.5](https://www.anthropic.com/claude-haiku-5-5) (website)
2. [Simon Willison - Claude Haiku 5.5](https://simonwillison.net/2026/Oct/7/claude-haiku-5-5/) (website)
3. [Theo / t3.gg - Does the $200 Codex plan suck now?](https://www.youtube.com/watch?v=nYA0yASgaZI) (youtube)

## Executive Summary

Cheap inference is making narrowly scoped agent work more affordable, but the headline token price is becoming a less reliable guide to the final bill. Anthropic’s Haiku 5.5 launch combines sharply lower prices with a long-context price jump, a changed tokenizer, and adjustable reasoning effort. Its accompanying cache discounts and subscription API credits make the economic story broader than one small model.

Practitioner discussion points in the same direction: evaluate useful work completed, not nominal dollars of inference consumed. That is a sound framing, although the available workload comparisons do not establish a universal winner between subscriptions or models.

## Haiku’s Price Cut Has Boundaries

[Anthropic’s announcement](https://www.anthropic.com/claude-haiku-5-5) lists Haiku 5.5 at **$0.10 per million input tokens and $0.50 per million output tokens** for prompts up to 100,000 tokens. Those rates are one-tenth of Haiku 4.5’s. Above that prompt threshold, the new rates become **$0.50 input and $2.50 output**—five times the short-context price.

The distinction changes the deployment argument. Anthropic positions Haiku for summaries, compaction, classification, quick lookups, and narrowly scoped coding subagents. It explicitly retains Sonnet and Opus as better choices for complex agentic coding. Cheaper access to a small model is not a claim that it should replace the model leading an entire engineering task.

[Simon Willison’s hands-on account](https://simonwillison.net/2026/Oct/7/claude-haiku-5-5/) adds another qualification: his token-counter comparison found that the same long prompt used roughly 1.25 times as many tokens with Haiku 5.5 as with Haiku 4.5. Anthropic also acknowledges the tokenizer change and estimates an average running-cost reduction of around 75%, rather than equating the 90% short-context rate cut with savings on every workload.

Adjustable reasoning effort adds another variable. Willison’s familiar SVG drawing test took seven seconds at low effort, versus more than five minutes at maximum effort. That is an illustration, not a production benchmark, but it exposes why “fast and cheap” still depends on configuration and the task’s quality requirement.

## The Subscription Argument Is About the Denominator

In [“Does the $200 Codex plan suck now?”](https://www.youtube.com/watch?v=nYA0yASgaZI), Theo discusses reduced perceived subscription value alongside cheaper model pricing. His useful distinction is between API-equivalent dollars, subscription limits, and the cost of completing a task.

He reports a Terminal-Bench comparison in which GPT-6.1 Soul at xhigh effort roughly matched Opus 5.5’s score, while costing about $0.96 per task against $13.11. These are his own results using native coding harnesses, not an independently reproduced comparison or a general productivity estimate. Harness choice, reasoning effort, and workload all matter.

The important interpretation survives that limitation: a subscription can deliver fewer nominal API dollars without necessarily delivering fewer completed tasks. The reverse also holds. A costly speed setting can consume an allowance quickly without adding proportionate value. Theo’s video includes sponsorships and speculative explanations of provider economics; those explanations should not be mistaken for disclosed costs or margins.

Anthropic’s launch offers a more concrete change. Sonnet 5.5 cache reads fall from $0.20 to $0.10 per million tokens; Anthropic estimates around 20% lower costs on most agentic tasks. It also announces monthly API credits rolling out this week: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team subscribers. Willison notes that the credits do not roll over and that disabling auto-reload lets requests stop when the balance runs out.

## The Bigger Story

Yesterday’s argument was that agent authority needs verification. Today’s economics reinforce that principle: the relevant denominator is **accepted, useful work**, not merely generated output. A cheaper attempt that needs repeated correction is not automatically a cheaper result.

Haiku 5.5 makes a stronger case for separating narrow supporting work from complex lead-agent work. It does not prove that such separation improves every workflow. Context length, tokenization, reasoning effort, cache reuse, and failure rates now belong in the comparison alongside the advertised price.

## Further Reading

- **Anthropic’s Haiku 5.5 announcement**, linked above: pricing tiers, intended workloads, evaluation claims, and subscription-credit details.
- **Willison’s account**, linked above: hands-on tokenizer and reasoning-effort observations that qualify the headline savings.
