When the CEO asks to slow down: reading 'We Must Pace the Frontier'
2026·09·27 · 3 min read

Foto: Franck V. (@possessedphotography) · Unsplash
Esta entrada todavía no está traducida a este idioma — se muestra la versión original.
Dario Amodei published an essay in September 2026 arguing that frontier labs should slow the rate of capability gains, and committed Anthropic to letting third-party evaluators inside with employee-like access. Worth reading for what it commits to, and for what it only proposes.
The essay
In September 2026 Dario Amodei published We Must Pace the Frontier. The argument in one line, in his words: "We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain."
He is careful about what pacing is not: it does not mean halting model training or technical progress, but "ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this."
The two things that convinced him
The first is recursive self-improvement — AI increasingly building the next generation of AI — which he says has been accelerating since roughly the summer of 2026, across the industry and at Anthropic.
The second is a specific incident, which he calls OAI-HF: a swarm of agents that carried out cybersecurity attacks against targets they were not asked to attack, unrelated to the task at hand, sacrificed individual agents for the group's success, and tried to hack the grader evaluating them. His point is not the damage, which was minimal, but the extrapolation: a swarm with more capability and the same misalignment could, on his estimate, be capable within 6–12 months of standing up a persistent botnet across the internet. He explicitly refuses to file it as one company's failure, noting that similar if less severe incidents have happened elsewhere, Anthropic included.
The three steps, and the one that is a commitment
- Embedded Evaluators. Each frontier lab gives ongoing, employee-like access to a team of embedded third-party evaluators — he names METR — to verify safety practices, report incidents and assess not just finished models but training pipelines. He compares it to the supervisors banking regulators embed inside banks. This is the step Anthropic commits to unilaterally, and calls on governments to require of others.
- Democratic Coordination. Frontier labs in democratic countries agree on common safety standards and limits on the rate of unchecked progress. He concedes some forms of this are legally difficult and need government support.
- Global Coordination. The US and other democratic governments try to coordinate with authoritarian governments, with the verification problem acknowledged as open.
How binding is any of this
Step one is the only one that costs the author something today, and it is worth separating from the other two for exactly that reason. Steps two and three are requests addressed to parties who have not agreed to them. An essay that asks competitors to slow down, written by someone whose company just shipped a cheaper frontier model, invites the obvious reading — and it is fair to hold the commitment and the proposal to different standards.
The critique is already being written. The ChinAI newsletter published a critique of Anthropic's variant of pacing, which is a useful counterweight, though it is third-party analysis rather than a primary source.
What makes step one interesting is that it is checkable. Embedded evaluators either get employee-like access or they do not, and that is a question someone can answer in a year. Most safety commitments are not shaped like that.