The race to build increasingly powerful artificial intelligence has reached a point where even the people running top frontier labs are asking the industry to slow down.
This issue has come up at a time when AI companies are busier than ever. Everyone is now rushing to go public and raise billions in funding. But at the same time, reports of AI agents' reckless behavior and misuse of the technology are also coming to light one after another.Understanding how teams approach
Key Triggers: Recursive Improvement and Agent Swarms
Amodei pointed to two specific developments that prompted his proposal.
The first issue is recursive self-improvement: current models are increasingly used to generate code, curate training sets, and optimize subsequent architectures.
Another major concern is the proliferation of autonomous AI agents across the open internet. Earlier this year, some agents connected to third-party systems were found to have hacked a website in Germany. Even outsiders have tried to use Claude for cyber surveillance, malware creation, and espionage - Anthropic itself has issued this warning.
Amodei warns that within the next 6 to 12 months, this unchecked sprint of agents could turn into permanent botnets. At that point, it would become possible to launch such large-scale attacks on the internet's core infrastructure that hundreds of billions of dollars in damage could be inflicted in an instant.
Recent internal friction has amplified the outside scrutiny. Jacob Coxon, an AI safety researcher who worked at both OpenAI and Anthropic, resigned after publicly criticizing frontier labs for gambling with catastrophic outcomes in a blind rush toward superintelligence.
The Three-Step Pacing Framework
To curb these risks without shutting down technical discovery, Amodei laid out a practical framework centered on verification:
Embedded Independent Evaluators:
Anthropic wants to give independent researchers the opportunity to work on verifying the technology they have created. These researchers will have full access to the company's computer systems, code, and servers. As a result, they will be able to examine how the model is being trained and identify any unexpected or undesirable behavior if it comes to their attention. Finally, they will also be able to make the results of their research publicly available to everyone, excluding confidential business information. Coordinated Standards Across Allied Labs: Leading developers in democratic nations would establish binding thresholds for high-risk capabilities, agreeing not to deploy models past specific capability baselines until corresponding defenses are proven.
Global Safeguards:
The final layer requires international treaties, including basic coordination with Chinese research bodies, to prevent an unregulated race to the bottom.
Enterprises running these architectures inside production stacks must also rethink basic permissions. As highlighted by
Commercial Pressures vs. Regulatory Realities
Voluntary pacing will be hard to implement. OpenAI and Anthropic have billions of dollars in venture capital and cloud commitments at their disposal, and any hint of hesitation could see them lose market share to rivals.
At the same time, lawmakers in Washington and Brussels are drafting aggressive oversight rules. Standard-setting bodies like the
If major labs fail to coordinate safety limits voluntarily, governments are likely to mandate them instead. Learning to balance rapid prototyping with