Thursday, May 22, 2026 · London Edition
Anthropic CEO Calls to Slow Down AI Development
Technology7 min read

Anthropic CEO Calls to Slow Down AI Development

Anthropic CEO Dario Amodei urges slowing AI development with a three-step plan and third-party safety evaluators like METR. Here's what it means.

Key takeaways

  1. 1Anthropic CEO Dario Amodei urges slowing AI development with a three-step plan and third-party safety evaluators like METR.
  2. 2That logic now has a prominent dissenter from within its own ranks.
  3. 3Amodei's position carries unusual weight precisely because of where he sits.
  4. 4Anthropic is not a policy think tank or a university research group.
E
Editorial
13 September 2026
ShareXFacebook

Article details

Canonical link
Published
13 September 2026
Last reviewed
13 September 2026
Byline
Editorial
External sources
None attached
Table of contents

Anthropic CEO Calls for Slowing AI Development

For most of the past four years, the dominant logic inside frontier AI labs has been simple: move faster than the competition, ship more capable models, and worry about consequences later. That logic now has a prominent dissenter from within its own ranks. Dario Amodei, the chief executive of Anthropic — one of the three or four companies actually building the systems at the center of this debate — has publicly argued that the industry needs to pull back and that the time for doing so is now.

Amodei's position carries unusual weight precisely because of where he sits. Anthropic is not a policy think tank or a university research group. It is a commercial AI lab competing directly with OpenAI and Google DeepMind for talent, compute, and customers. When a leader of such a company calls for restraint, it is worth examining both what he is actually proposing and why he might be saying it at this particular moment.

His argument, laid out in a substantial essay published in September 2026, centers on a concept he calls "pacing the frontier." Strip away the jargon and the idea is straightforward: the rate at which the most capable AI systems are developed and deployed should be deliberately moderated, with safety evaluations keeping pace rather than trailing months or years behind model releases. Dario Amodei slow down AI is not just a slogan — it is, according to his own framing, a structural proposal with specific mechanisms attached.

The Three-Step Plan to Pace the Frontier

The Three-Step Plan to Pace the Frontier — Religious text about festival of tabernacles with handwritten margin notes
The Three-Step Plan to Pace the Frontier — Religious text about festival of tabernacles with handwritten margin notes

Amodei's essay outlines a three-part framework, though the published summary characterizes the writing as "winding," suggesting the argument is more nuanced — and perhaps more contested internally — than a clean numbered list might imply.

The core of the plan is sequencing: before new frontier models reach broad deployment, they should pass through a defined set of safety checkpoints. This is not a novel idea. What makes Amodei's version distinctive is his willingness to commit Anthropic itself to the discipline, rather than simply calling on governments to impose it on the industry at large. Self-imposed pacing, backed by third-party verification, is the structural bet he is making.

The second element involves transparency. Anthropic has committed to sharing model access with independent evaluators before those models go public. This kind of pre-deployment evaluation is becoming a baseline expectation in policy circles — the UK AI Safety Institute and its US counterpart, the US AI Safety Institute (US AISI), have both pursued frameworks for exactly this kind of third-party assessment since 2023. Amodei's commitment aligns Anthropic with that emerging norm.

The third element is the most philosophically charged: a willingness to accept slower commercial progress as a trade-off for lower catastrophic risk. That trade-off is easy to articulate and politically appealing. Executing it while competitors are not similarly constrained is a different matter entirely.

Third-Party Evaluators: METR's Role in AI Oversight

Third-Party Evaluators: METR's Role in AI Oversight — a close up of a button on a wall
Third-Party Evaluators: METR's Role in AI Oversight — a close up of a button on a wall

One concrete commitment that distinguishes Amodei's announcement from pure rhetoric is the specific naming of METR — Model Evaluation and Threat Research — as a third-party evaluator that will receive access to Anthropic's models.

METR is not a government body. It is an independent research organization that has developed publicly documented methodology for assessing the risk profiles of advanced AI systems, particularly around what researchers call "dangerous capability evaluations." These assessments probe whether a model can, for instance, assist in creating biological agents, conduct sophisticated cyberattacks, or otherwise amplify catastrophic risks at scale. METR has previously worked with multiple labs and has published its evaluation frameworks, which gives its assessments a degree of external credibility that internal reviews cannot replicate.

Granting METR access before a model ships represents a meaningful procedural commitment. It does not guarantee safety — no evaluation framework can fully anticipate emergent behaviors at deployment scale — but it establishes an accountability structure that was largely absent during earlier waves of AI releases. For context, independent pre-deployment evaluations of this kind were essentially non-existent before 2023; today they are becoming a standard expectation for any lab that wants to maintain credibility with policymakers and enterprise customers.

The UK AI Safety Institute, which conducted pre-deployment evaluations of GPT-4o and several other frontier systems under a series of voluntary agreements, has described similar arrangements as "foundational" to its oversight mandate. Anthropic's commitment to METR fits squarely into that emerging ecosystem.

Why This Moment Matters for the AI Industry

The timing of Amodei's essay is not incidental. The AI industry in mid-2026 is at an inflection point that differs qualitatively from earlier moments of rapid capability growth. Systems are beginning to demonstrate autonomous task completion across extended time horizons — what researchers call "long-horizon agency" — and the gap between what models can do in controlled demonstrations and what they do in unmonitored deployment is narrowing.

Government attention has intensified correspondingly. The European Union's AI Act entered its most consequential enforcement phase in 2025, with obligations for so-called General Purpose AI systems that directly affect frontier labs. The US AISI, established under the executive order on AI in late 2023 and subsequently codified by Congress, has expanded its mandate to include mandatory pre-deployment testing for models above defined capability thresholds. The UK's AI Security Institute has published risk frameworks that explicitly reference catastrophic and irreversible harms as the scenarios requiring the highest regulatory attention.

Against this backdrop, a voluntary commitment from a major lab to pace its own frontier development serves a dual purpose: it is a genuine safety measure, and it is also a bid to shape the terms of the regulatory conversation before those terms are imposed externally.

Criticism and Skepticism: Is a Slowdown Realistic?

Not everyone finds the proposal compelling, and the skepticism comes from multiple directions.

From the competitive standpoint, the objection is straightforward. If Anthropic slows down and OpenAI, Google DeepMind, Meta AI, and a growing number of Chinese labs do not, Anthropic does not actually reduce the pace of frontier AI development — it simply cedes ground in a race that continues with or without it. This is sometimes called the "unilateral disarmament" problem, and it is not a bad-faith argument. It reflects a genuine structural dilemma for any single actor in a competitive market.

From the technical community, a different concern surfaces: the adequacy of current evaluation methods. Even METR's published frameworks, rigorous as they are, assess models against known threat categories. Emergent capabilities — behaviors that appear at scale without being predictable from smaller models — are definitionally difficult to evaluate before they exist. Some AI safety researchers argue that pacing without more fundamental advances in interpretability and alignment research is a delay tactic rather than a solution.

There is also a political economy argument. The history of voluntary industry self-regulation in technology is not encouraging. Social media platforms made extensive commitments about content moderation that proved difficult to enforce under competitive pressure. Commitments made by lab leadership are not binding on future leadership, future shareholders, or future competitive environments.

None of this means Amodei's proposal is without value. But it does mean the proposal deserves scrutiny rather than acceptance as sufficient.

What This Means for the Future of AI Safety

Whatever the specific merits of Amodei's three-step plan, his willingness to make the argument publicly from inside a leading lab shifts the terms of the debate in ways that matter.

For years, the dominant narrative from frontier labs has been that safety and capability are complementary — that building more powerful AI is itself the path to safer AI, because only sufficiently capable systems can solve the alignment problem. Amodei is not abandoning that view entirely, but he is acknowledging something the safety research community has argued for longer: that the pace of deployment can outrun the pace of understanding, and that this gap is not merely theoretical.

The practical implication is that safety evaluation infrastructure needs to be treated as a prerequisite for deployment, not an afterthought. METR's involvement with Anthropic's models is one step in building that infrastructure. The US AISI's testing frameworks, the EU's GPAI obligations, and the UK's emerging mandatory reporting requirements are others. Taken together, they represent the early architecture of an accountability system for AI that did not exist five years ago.

Whether that architecture is adequate to the pace of development is still an open question. What Dario Amodei slow down AI advocacy does, at minimum, is put a prominent voice inside the frontier AI ecosystem behind the proposition that slowing down is sometimes the correct choice. That is a more consequential shift in tone than it might first appear.


Source: [The Verge](https://www.theverge.com/ai-artificial-intelligence/994337/anthropic-ceo-slow-down-ai-development)

Comments

No comments yet. Be the first.

Leave a comment