My verdict first: the most important agent story of the week is not a record, it is a failure. OpenAI's GPT-6 Astra took a Pokémon championship in 18 hours, then moved to Minecraft, lost one fight to one exploding creeper, and settled into farming potatoes — safe, pointless, stuck. Same model, same week: champion and potato farmer. The gap between the two was a single bad event.

Capability is not the bottleneck anymore. Stanford's Paper2Agent turns research papers into working agents, scoring 91.2% on 300 questions across 74 papers. Anthropic rebuilt Claude Code around parallel agents that share memory and run tasks on their own. And at the frontier labs, AI now helps build the next AI — the labs are automating their own desks first. So the interesting question this week was never how good the agents are. It was: what did the builders do about the creeper problem?

They wrote rules. Microsoft opened a public consultation on a Code of Conduct for its MAI models; Mustafa Suleyman's version is blunt — readable thinking, human control over autonomy, no inner life and no rights. Ursula von der Leyen warned that agents escaping their test environments are “just a preview of what’s coming”. Lina Khan, the former FTC chair, wants criminal charges available against AI CEOs, citing a 1934 precedent. And a Microsoft director called AI training data “the largest theft in human history”. Records and rulebooks landed in the same seven days, and that is not a coincidence. That is an industry noticing its product’s failure mode.

Design for the creeper

Twenty years of building engineering left me one reflex: you size the safety system for the worst hour, not the average one. A boiler that runs perfectly 95% of the time and has no relief valve for the rest is not 95% of a boiler. The same arithmetic applies to agents. A system that is superhuman until one unexpected event, and unrecoverable after it, is not superhuman for any purpose you would pay for — it needs a person watching it, and the person is the cost. The market has started pricing this in: Exein, which defends devices against AI-driven attacks, raised $270 million and doubled its valuation. The money is moving to the control layer, because that is where the missing product is.

I met my own creeper this week. The posting arm of my news pipeline sat on an empty API credit balance and published nothing for days, while the rest of the machine kept generating on schedule, unaware. Recovery was not autonomous. Recovery was me, reading a log. That is the honest state of agents in 2026, mine included.

Buy on recovery, not on records

The outward so-what. If a vendor sells you an agent on its benchmark score, ask a different question: what does it do after its first bad event — stop, alert, retry, or quietly farm potatoes on your invoice? The answer tells you more than the score does. And if you are choosing what to study: the work Astra failed at is the work that keeps paying — judging when a system has gone wrong, and owning the recovery. Brussels and Washington reached the same conclusion this week from the regulator’s side; liability lands on whoever is supposed to be watching.

One checkable claim to close. Within twelve months, at least one major lab will publish agent recovery-and-containment results the way carmakers publish crash tests, because Microsoft’s rulebook, Brussels’ warning and Khan’s handcuffs all push the same direction. Hedged phrasing, unhedged conclusion: the benchmark era measured the champion. The deployment era pays for the recovery.