# Claude Sonnet 5.5 Tests the Economics of Coding Agents

- Date: 29 Sept 2026 (2026-09-29T16:01:29.000Z)
- Summary: Early hands-on accounts suggest Sonnet 5.5 may be useful for codebase investigation and planning, but a failed task and variable token costs complicate simple speed-and-price claims. The meaningful comparison is cost per verified result within a specific workflow.
- Tags: `digest`, `ai-discourse`, `claude-sonnet-5-5`, `coding-agents`, `model-routing`, `evaluation`

## Sources

1. [Simon Willison - Claude Sonnet 5.5 first look](https://simonwillison.net/2026/Sep/28/claude-sonnet-5-5/) (website)
2. [Theo / t3.gg - Sonnet 5.5 hands-on comparison](https://www.youtube.com/watch?v=8WbW_n95wc4) (youtube)

## Executive Summary

The case for a faster, less expensive coding model depends on where it sits in the workflow. [Simon Willison’s first look at Claude Sonnet 5.5](https://simonwillison.net/2026/Sep/28/claude-sonnet-5-5/) relays Anthropic’s speed and cost claims but also records a task that consumed its thinking budget without delivering the requested SVG. In a [hands-on video](https://www.youtube.com/watch?v=8WbW_n95wc4), Theo argues for a more specific role: using Sonnet as a subordinate model to investigate a codebase and prepare a plan, rather than automatically making it the default for every coding task. These are early, task-dependent observations, not a settled ranking of models.

## What Happened

According to Willison’s account of Anthropic’s announcement, Sonnet 5.5 is claimed to be more than 30% faster and up to 30% cheaper for most work while retaining the previous Sonnet 5 price. The distinction between a model’s price and the cost of completing a task matters here: speed or nominal token economics do not guarantee a cheaper successful outcome. Willison describes an attempted SVG task that reached the maximum thinking limit after 128,000 tokens and approximately $1.28 without producing the SVG. That is one reported failure, not an estimate of the model’s overall success rate, but it is an unusually clear counterexample to treating release metrics as an end-to-end productivity measure.

Theo’s creator-run comparison reaches a different, compatible conclusion about usefulness. On his own large-codebase planning benchmark, he says Sonnet 5.5 performed slightly better than Opus 5.5 at about half the cost. He sees value in routing preparatory work—codebase exploration, investigation, and planning—to a faster model beneath a larger coding workflow. Yet he also reports heavy token use, weak design-generation results, and a pricing comparison that becomes less straightforward when cached inputs are taken into account. He cautions against setting reasoning effort to maximum by reflex: extra thought can add expense and may not improve the answer. Neither his benchmark nor Willison’s SVG trial establishes how the model will perform across other repositories or tasks.

## Why It Matters

A model release can improve one part of a workflow while making another part unexpectedly costly. Researching a codebase, planning a change, implementing it, and generating a visual artifact demand different kinds of judgment and have different failure modes. Theo’s proposed subordinate-model role is interesting because it assigns Sonnet a bounded job with a concrete output—the plan—rather than treating a new release as a universal upgrade. Willison’s failure illustrates the other half of that design problem: a run can be fast in principle and still spend its entire budget without producing the artifact.

That makes *cost per verified result* a more useful question than cost per token or the price of a single successful-looking demo. A fair comparison would hold the task and acceptance criteria constant, count retries and cached-input effects, and check whether the plan or artifact actually helped the next step. These are implications of the two accounts, not measurements either account supplies for general use.

## The Bigger Story

This refines yesterday’s view that coding agents make output cheap while leaving problem selection and verification scarce. There is now a second layer of design: deciding which model should perform each step, how much reasoning budget it gets, and when to stop a fruitless run. A cheaper investigator can make a larger agent more effective if its findings are accurate and usable. An unbounded reasoning loop can erase that advantage even on a seemingly small task.

For now, the strongest conclusion is narrow. Sonnet 5.5 deserves evaluation as a planning and investigation worker, but public release claims and a small number of creator trials cannot settle its default place in a coding stack. The workflow, the task, and the stopping rule remain as important as the model name.

## Further Reading

- [Simon Willison’s Claude Sonnet 5.5 first look](https://simonwillison.net/2026/Sep/28/claude-sonnet-5-5/) — release-claim context alongside a concrete unsuccessful task.
- [Theo’s hands-on Sonnet 5.5 comparison](https://www.youtube.com/watch?v=8WbW_n95wc4) — a builder’s account of planning utility, token costs, and design limitations; interpret its benchmark as a personal test.
