Krosoft

AI Digest

Coding Agents Made Output Cheap, Not Judgment

Simon Willison’s 2026 retrospective and Zhipu’s reported infrastructure-agent results point to the same constraint: agents accelerate implementation when people define useful goals and provide fast, verifiable feedback. More output alone does not resolve product judgment or accountability.

Coding Agents Made Output Cheap, Not Judgment

Executive Summary

The scarce part of software work is shifting from producing code to deciding what should exist and proving that it works. In a newly published retrospective on 2026 in LLMs, Simon Willison describes coding agents crossing a threshold into everyday usefulness, then asks why engineering nevertheless feels harder. A concrete, if company-reported, example appears in Jack Clark’s account of Zhipu’s infrastructure agent: engineers defined the objective and boundaries, while the agent proposed changes and an experimental environment supplied measurable feedback. The two accounts converge on a narrower story than “software builds itself”: agents can accelerate execution when humans can specify a goal and check the result.

What Happened

Willison’s annotated keynote traces the move from late-2025 model improvements to coding agents he considers reliable enough for daily use. He recalls StrongDM’s “software factory” approach, in which people neither write nor review the code directly, as an exploration of how to verify agent-produced software without traditional line-by-line inspection. That is a provocative example from his retrospective, not proof that unreviewed code is generally safe.

He makes the limit tangible with his own projects. Agents helped him build a Python JavaScript interpreter and WebAssembly runtime, but the result prompted a harder question: did anyone need those artifacts? Later he generated playable games from a concept, only to find that something resembling a game is not the same as one with a compelling gameplay loop. More capable models expanded what he could produce; they did not supply the product judgment that makes the output worthwhile. His formulation of the new engineering task is to define goals, specify constraints, and choose tools clearly enough for the model to work effectively.

Clark’s newsletter describes Zhipu’s report on using its GLM-5.3-powered Infra Agent to optimize the launch of GLM-5.3 Flash. According to the company, engineers set objectives and system boundaries, the agent performed analysis and code changes, and tests supplied fast, local, objectively comparable feedback. Zhipu reports moving from initial adaptation to production readiness in under two weeks and tripling end-to-end throughput against its initial baseline. Those are Zhipu’s own results, not independently reproduced measurements or a general estimate of agent productivity. The more transferable detail is the loop: narrow a hypothesis to a particular kernel, parameter, or code path, run an inexpensive test, and let evidence determine whether the change improved anything.

Why It Matters

The phrase “software factory” can obscure where accountability went. In these examples, human involvement does not disappear; it moves upstream into problem selection and boundary-setting and downstream into verification. That complicates the idea that the most important metric is the quantity of code an agent can produce. A team might generate far more code and still fail to produce a useful product—or optimize the wrong workload remarkably efficiently.

Willison says his job feels more intellectually demanding because agents handle easier tasks and his ambitions have grown. His observation is personal rather than a workforce survey, but it explains a real tension in the factory metaphor: automation can remove routine implementation while concentrating the difficult choices. Zhipu’s account offers the constructive counterpart. When success has a measurable definition and feedback is quick, an agent can explore implementation options at a pace that would be awkward for a person. Where success is subjective, or evaluation arrives only after customers experience a failure, that loop is much harder to close.

The Bigger Story

This reinforces a developing view of agents as powerful executors inside a designed system, rather than autonomous substitutes for its designers. The next competitive advantage may be less about prompting an agent to write more and more about building trustworthy feedback: representative tests, clear constraints, and a way to detect whether the apparent win survives outside a demonstration. Willison’s retrospect also warns against mistaking the ease of making an artifact for progress on the original problem.

Further Reading

  • 2026 in LLMs (so far) — Willison’s annotated keynote, especially his reflections on software factories, agent-enabled ambition, and the harder work left to engineers.
  • Import AI 474 — Clark’s wider research roundup includes the Zhipu account and his interpretation of its significance.
Back to archive