AI Digest

Haiku 5.5 and the Limits of Cheap Delegation

Theo’s Haiku 5.5 review shows why cheap parallel workers still need substantive verification and resource limits. His pull-request failures, alongside Simon Willison’s voice-coding account, distinguish affordable implementation from reliable completion.

Haiku 5.5 and the Limits of Cheap Delegation

Executive Summary

Cheaper parallel work becomes expensive when the worker mistakes a status flag for a review or exhausts a shared API quota. Theo’s hands-on Haiku 5.5 review supplies a concrete qualification to the model’s pricing story: it looks promising for bounded exploration and categorization, but his pull-request experiment exposed failures in verification and resource management.

The useful distinction is between delegating attempts and delegating judgment. A small model can make more attempts affordable without being reliable enough to decide that the work is finished. Today’s strongest evidence is this practitioner test, not a comprehensive comparison of model reliability.

The Pull-Request Test Exposes the Boundary

In “finally a good small model”, Theo tries Haiku 5.5 on a backlog of roughly 1,500 open pull requests. One task identifies performance-related changes; another categorizes the backlog into a browsable summary.

The first problem is semantic: Haiku calls a candidate “merge ready” because GitHub reports it as clean. Theo observes that it does not appear to have thoroughly audited the change itself. The distinction matters because a clean GitHub status is not evidence that a proposed implementation is correct or worth landing. He passes a candidate to Opus for substantive inspection instead.

The second problem is operational. Parallel per-PR fetches consume most of his 5,000-request hourly GitHub allowance. The script saves an error response as though it were API data, then fails while processing it. Haiku says it will wait for the quota to reset, but Theo sees no actual wait action: the thread stops until he intervenes.

These are observations from one workflow, not a measured failure rate. Nevertheless, they reveal costs that token-price comparisons miss. A cheap worker can consume a scarce external resource, interrupt other work, and leave recovery to its operator. Theo restarts the task with Opus orchestrating Haiku workers; the video does not establish the completed outcome of that revised run.

Where More Attempts Can Pay Off

Theo’s positive case is narrower than replacing the lead coding model. He wants Haiku for searching, reading documentation, categorizing data, and exploring alternative approaches—especially where results can be checked cheaply.

He illustrates that argument with Anthropic’s egg-drop simulation demo, not an independently reproduced benchmark. As recounted in the video, Opus alone succeeds after 25 attempts, taking about three and a half minutes and costing $0.47. Opus directing ten Haiku subagents makes 86 attempts, finishes in under a minute, and costs $0.14.

The interpretation is plausible: when candidates are independent and success has a clear test, inexpensive parallel attempts can beat slower serial exploration. It does not follow that more workers improve ambiguous code review. The egg-drop example has a success check; “this PR should be merged” demands a more substantive judgment.

This sharpens yesterday’s cost-of-useful-work argument. Context gathering, verification, and retries belong in the bill, but so do shared quotas and recovery. Cheap delegation is attractive where checking the answer remains easier than producing it—not simply wherever the task contains code.

A Human Supervisor Still Finishes the Job

Simon Willison’s voice-built newsletter index offers a smaller, complementary example of the same boundary. He describes spending about half an hour talking to Codex while cooking, producing a Django model, migration, imports, archive pages, and search integration.

The feature was nearly ready, not finished. Reviewing the pull request, he changed a subprocess-based Git import to an API-based approach and refined the presentation. Another half hour of typing-based prompting brought it to deployment.

His conclusion is about multitasking, not abandoning the keyboard. Voice helped him specify and steer a familiar feature; precise corrections and final review still suited typing. Together, the two accounts reinforce a practical view of coding agents: make participation cheaper and delegation broader, while keeping acceptance criteria and the authority to declare completion explicit.

Further Reading

  • Theo’s review, linked above: the live PR-audit failures are more instructive than the headline benchmark charts.
  • Willison’s writeup, linked above: a detailed first-hand account separating voice-driven implementation from review and deployment.
Back to archive