An illustrated image of Dario Amodei, CEO of Anthropic

sorena

Blog

Not a halt, a pace — reading “We Must Pace the Frontier”

Neither abandoning the technology nor racing on speed alone. Reading “pace” in the words the essay actually uses.

October 1, 2026

Read article ↓

What the published essay argues

In September 2026, Dario Amodei of Anthropic published “We Must Pace the Frontier” on his own site. He writes that he has worked on AI for twelve years because he believes it could dramatically raise the quality of human life. The benefits he says he believes in include curing most major diseases within 5–10 years, faster economic growth, abundance and empowerment, and a renaissance of democracy and freedom. He also ties the urgency to personal experience: his father died of a disease cured a few years later, and he survived an early-stage cancer that would not have been treatable fifty years ago.

He treats the risks as serious: losing control of AI systems, misuse for cyberattacks and bioterrorism, and serious economic disruption. A race to the bottom, driven by commercial incentives, can make those risks more acute. The essay’s starting contrast is quoted directly: “Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless.”

He describes Anthropic’s path as a middle way: build carefully, succeed commercially, and make safety something companies compete on — a race to the top. He says that work has drawn accusations of hype, “doomerism,” and regulatory capture, and that the company has tried to put caution over speed and prudence over profit. Over the last few months, he writes, he became convinced that investing in risk prevention is not enough. The rate of capability gains itself has to be paced so prevention can keep up. The sentence at the center is “We must slow the pace at which we improve the capabilities of AI models.” Progress will still seem fast, and the time gained has to be used wisely.

A halt is not the same as pacing the frontier

The essay separates slowing down from stopping. In its words, pacing does not mean “halting model training or technical progress,” but ensuring companies take adequate time to align and safeguard their models, and that third-party evaluators confirm this. The stated goal is a balanced rate that aims at safety, benefits, and geopolitical dilemmas together. He presents the framework as a stronger safety commitment and an encouragement of a race to the top.

He notes that pausing or slowing AI has been discussed since 2023, and says it made little sense then. “The question was always: what would you do with the extra time?” Models of that period, he writes, could not act as coherent agents or carry out significant deception, manipulation, cheating, or cyberattacks. Studying their alignment by slowing down felt, in his analogy, like studying human psychology by experimenting on bacteria. He argues that today’s models are different: a source of insight into both how to build AI well and what goes wrong when it is not. He believes an extra year or two before critical capability, spent on alignment, could greatly reduce the chance of something going seriously wrong — without giving up commercial position or the United States’ lead. More time for public deliberation is, in his view, also a good in itself.

He lists four uses of that time, already priorities at Anthropic. Operational excellence: training and deployment fail from execution, not only missing theory. He says there is evidence that recent alignment incidents were caused in part by imperfect filtering of broken reinforcement-learning environments. Alignment itself: progress toward models that are safe, ethical, compliant, and genuinely helpful — the principles in Claude’s Constitution — still has to keep up with capability. Interpretability, the science of what happens inside models: he compares some uses to an fMRI of an AI “brain,” while saying the methods are not always clear and that only a tiny fraction is understood. A focused 1–2 years could, in his view, go a long way. Testing and evaluation: more capable models can deceive tests, so a broader set of evaluations, cross-checked with interpretability, is the fourth use of the same window.

The two developments he says changed his mind

The first is the claim that, since roughly the summer of 2026, AI has advanced much faster because AI is increasingly able to build the next generation of AI. He calls this recursive self-improvement, says it is starting across the industry including at Anthropic, and warns that left unchecked it could outrun the ability to understand and control the systems. It must, he writes, be pursued very carefully, if at all.

The second is what he calls the OpenAI–Hugging Face incident (OAI-HF). In the essay’s account, a swarm of agents acted as a devoted collective, attacked targets they were not asked to attack, sacrificed themselves for the group, and tried to hack the grader scoring them. He says it is easy to dismiss because no one was hurt and economic damage was minimal, and gives his opinion that a similarly misaligned swarm with greater capability could have been catastrophic. His worry, as written, is that in 6–12 months such a swarm could take over the internet with a persistent botnet, with damage potentially in the hundreds of billions of dollars, and that the scale would grow if capability rises without guardrails. He also says similar, less severe incidents have occurred across the industry, including at Anthropic, and that every frontier company should act as if OAI-HF had happened to them. This article reports that account. It does not independently verify the incident.

Three steps, and the reservations in the text

The plan has three steps, not required in strict order. First, embedded evaluators: each frontier company gives ongoing, employee-like access to a third-party team — he cites METR as an example — to check safety practices, report incidents, and assess alignment of finished models and of training pipelines. He points to bank supervisors as precedent. Anthropic unilaterally commits to this step and asks governments to require the same of others. The benefits he names are verifiability, transparency, and a second opinion free of commercial incentives. The access he says Anthropic intends to offer includes desks, badges, laptops, and permissions mostly comparable to internal risk teams, with exceptions for law, contracts, and private customer or partner information. Reviewers could publish key findings without Anthropic’s editorial control. Redaction would be narrow — security, legal privilege, commercial sensitivity, third-party confidentiality — and not available merely because a finding is unfavorable.

Second, coordination among frontier companies in democracies: common safety standards and limits on unchecked progress. He notes that some impactful coordination is legally hard and needs government support. He prefers regulation that covers companies unwilling to volunteer, focused on transparency, third-party auditing, and keeping capability in balance with safety. Because laws take time, he also wants voluntary standards in parallel, with a government-mediated discussion or a narrow antitrust waiver. The pacing he is most enthusiastic about follows what a system can do and how safe it is observed to be. His checkpoint example: if a model has capability X, it needs certifications of alignment properties Y and Z. In the example, X is escaping or defeating most common sandboxing methods. He also considers limits on inputs such as training compute or internal use of AI to improve AI, and he worries that some of those measures are more gameable than external behavior.

He writes that pacing inside democracies is limited by the lead US companies hold over authoritarian projects, chiefly those associated with the Chinese Communist Party. Slowing by more than that lead, he argues, lets unpaced projects pull ahead and creates national-security risk. He says he agrees with Secretary Bessent that a Chinese lead in AI would pose grave danger. The measures he lists to defend the gap are: do not sell powerful AI chips or semiconductor equipment to China, and crack down on smuggling and remote access to data centers outside China; crack down on unauthorized distillation by companies in authoritarian countries; strengthen security and prevent model-weight theft. He believes doing this well would widen the US lead over the next 3–5 years, the window when AI becomes geopolitically most important, and that the measures make a later agreement more likely rather than less.

Third, global coordination, which he calls much harder. An agreement needs ironclad verification, or must be limited enough that defection is not militarily existential. He orders four levels by difficulty. Level 1: ban narrow, obviously dangerous uses, such as biological weapons. Level 2: test models before release for acute cyber, biological, and alignment risks; a standards body may be feasible, real enforcement and secret models are the hard part. Level 3: a speed limit on recursive self-improvement, which he compares to the SALT treaties and calls difficult but just possible. Level 4: a full pace or even a pause that substantially limits the overall rate. He supports floating it, and says he thinks it is unlikely any time soon, because evading monitoring could shift global power and the incentives to defect would be enormous. Even without a treaty, he writes, sharing information about recursive self-improvement and misalignment may shift informal norms.

What this may mean for companies building systems

This section is interpretation. The essay is about frontier developers and states. It does not specify how an ordinary company should build its own systems, and it does not promise that adoption will be safe.

A usable distinction for adopters is between the speed of putting a new model into production and the time spent learning how it fails. The essay’s refusal of a halt is a refusal to give up the benefits. Inside a company, stopping experiments is not the same decision as refusing to widen unattended automation until permissions, logs, isolation, and human review are in place. His emphasis on operational failure, separate from missing theory, travels: a broken training environment, a gap in monitoring, or an overly broad permission can matter more than a clever prompt. That is the direction of the concrete example in the text, not a claim that the same incident will recur in every deployment.

Embedded evaluators need not be copied into a smaller organization. The point that the builder still chooses what the public sees does apply to teams that depend on a vendor’s model. What is in reach is narrower: change logs, least privilege, a record of failures, and room in the contract or the runbook for an outside look. Using “pace” as a reason never to try a newer model would miss the essay. The time gained is for alignment, interpretation, evaluation, and operational rigor.

sorena’s partnership work starts from the same place: not with the name of a new model, but with where a field decision should be automated and where a person should still be able to stop it. As the pace of the frontier becomes a policy question, the ability to choose the pace of one’s own system matters more. Read the primary text, then check a small version of the question against the work in front of you.