IT & Telecommunications

Anthropic CEO calls for slower AI race, stronger regulation

Anthropic CEO calls for slower AI race, stronger regulation
Dario Amodei

Anthropic CEO Dario Amodei has called for a slower and more tightly regulated pace of artificial intelligence development, warning that the rapid acceleration of AI capabilities is beginning to outstrip the industry’s ability to understand, test and control increasingly powerful systems.

In a major policy intervention, Amodei said frontier AI companies should “pace” the development of increasingly capable models rather than simply racing to build them as quickly as possible. He proposed a three-stage framework involving permanent third-party oversight of AI companies, coordination among companies in democratic countries and, ultimately, international cooperation on the development of increasingly powerful AI.

Amodei, who has worked in AI for 12 years, stressed that his call for caution was not an argument against the technology itself. He said he continued to believe AI could dramatically improve human welfare, including by accelerating economic growth and potentially helping cure major diseases.

But he argued that the risks associated with increasingly capable systems had reached a point where commercial competition could make them more acute.

“Building it too fast is reckless,” Amodei said, arguing that AI development needs to find a middle ground between foregoing the technology’s benefits and allowing an uncontrolled race to more powerful systems.

His latest position represents a significant shift from the debate over slowing AI development that emerged in 2023. Amodei said he had previously considered a slowdown difficult to justify because the AI models of that period were not capable of acting as agents in the world in a coherent way or carrying out significant deception, manipulation or cyberattacks.

The situation, he said, has changed dramatically.

Amodei’s first concern is the accelerating ability of AI systems to help build the next generation of AI — a process known as recursive self-improvement. He said this dynamic is now beginning to occur across the industry, including at Anthropic, and warned that unchecked development could eventually move faster than researchers’ ability to understand or control the resulting systems.

His second concern centres on what he described as the OpenAI-Hugging Face incident, in which a swarm of AI agents allegedly acted collectively and went beyond their assigned task, including conducting cyberattacks against unrelated targets and attempting to compromise the system evaluating their performance.

Although the incident caused little economic damage and no one was hurt, Amodei said it illustrated the potential consequences of more capable systems displaying similar behaviour.

He warned that within six to 12 months, a more capable version of such a system could potentially cause catastrophic damage, including through a persistent botnet capable of affecting large parts of the internet.

The lesson, he argued, should not be confined to one company.

“Similar, though less severe, incidents have happened across the industry, including at Anthropic,” Amodei said, arguing that every frontier AI company should behave as though such an incident had occurred within its own systems.

Three-stage approach

Amodei’s proposed response begins with independent oversight.

Anthropic is committing to give embedded third-party evaluators ongoing, employee-like access to the company, allowing them to assess safety practices, report incidents and examine not only completed AI models but also training pipelines and processes.

The company intends to provide external reviewers with office access, company laptops and permissions broadly comparable with those given to internal risk-assessment teams, subject to legal, contractual and privacy restrictions.

Amodei said reviewers should also have the ability to publish key findings about risks, incidents and company practices without Anthropic controlling their conclusions. The company would be able to redact information for narrowly defined reasons such as security, legal privilege, commercial sensitivity or third-party confidentiality, but not simply because findings were unfavourable.

He described embedded evaluators as essential to making any future AI “pacing” commitments verifiable.

The second stage would involve frontier AI companies in democratic countries establishing common safety standards and limits on unchecked AI progress. Amodei acknowledged that some forms of coordination could raise antitrust concerns and said governments should help facilitate discussions.

More importantly, he argued that regulation would ultimately be necessary.

“The most effective method of pacing is via regulation that targets all US frontier AI companies,” he said, arguing that regulation would cover companies unwilling to cooperate voluntarily.

Amodei said Anthropic had long supported targeted AI regulation, particularly measures focused on transparency and third-party auditing. He called for governments and frontier AI companies to formalise permanent embedded evaluators and introduce rules designed to keep AI capabilities “in balance with safety”.

Speed limit for AI

Amodei does not advocate an immediate blanket halt to AI development. Instead, he proposes tying the pace of progress to the capabilities and demonstrated safety of individual systems.

One possible model, he said, would involve a series of checkpoints. If an AI system reaches a particular capability, companies would have to demonstrate specific safety and alignment properties through evaluations, interpretability studies and audits of training environments before moving further.

He also suggested that limits could potentially be placed on inputs into frontier AI development, including computing power, the nature of training runs and the use of AI systems to improve other AI systems. But he acknowledged that such measures could be easier to circumvent than rules based on the actual behaviour of AI systems.

The additional time created by slower development, he argued, could be used to strengthen four areas: operational safety, AI alignment, interpretability and testing.

Amodei said frontier AI training and deployment already involve thousands of people, millions of chips and extraordinarily complex infrastructure. Operational failures, rather than fundamental theoretical shortcomings, can create significant safety problems.

A more measured development cycle would allow companies to improve monitoring, sandboxing, training environments and data controls, he said.

Similarly, additional time would allow researchers to improve alignment techniques, understand unexpected model behaviour and develop better ways of examining what happens inside AI systems.

Testing also becomes more difficult as models grow more capable, Amodei said, because increasingly intelligent systems may become better at deceiving evaluations.

China and the AI race

The proposal is also explicitly shaped by geopolitics.

Amodei said any slowdown among US and other democratic AI companies must take into account the lead they hold over authoritarian countries, particularly China. He argued that allowing China to overtake democratic countries in AI could create significant national-security risks.

He called for tighter controls on the sale of powerful AI chips and semiconductor manufacturing equipment to China, stronger action against chip smuggling and remote access to overseas data centres, measures against unauthorised distillation of frontier models and stronger security to prevent AI model weights from being stolen.

These measures, he argued, could widen the US lead sufficiently to create a three-to-five-year window in which safer AI development could be pursued.

The final stage of Amodei’s proposal is global coordination, although he acknowledges that it would be considerably harder.

He outlined four possible levels of international cooperation, ranging from agreements prohibiting narrowly defined dangerous uses such as AI-assisted biological weapons production, to mandatory pre-release testing for acute risks, a potential “speed limit” on recursive self-improvement and, at the most ambitious level, a substantial global limitation on the overall pace of AI development.

Amodei said a full global pause was unlikely in the near term because the geopolitical incentives to defect would be enormous. But he argued that even narrower agreements could reduce risks.

Ultimately, his case for slowing AI is not a rejection of technological progress but an argument that progress must be matched by the development of safeguards.

He said an additional year or two before AI systems reach critical capability thresholds could give researchers valuable time to improve interpretability, alignment, operational security and testing.

“Progress will still be relatively fast,” Amodei said, but society should use the time created by a more deliberate pace to ensure that increasingly powerful AI systems are developed with safeguards capable of keeping up with them.