anthropic

Anthropic Wants to Put Safety Regulators Inside AI Labs

Everyone keeps saying we should slow down AI. But nobody actually says how.

Until now.

Anthropic CEO Dario Amodei just published a blog post outlining three concrete strategies to slow AI development. And he says Anthropic is committing to one of them right away.

Let me break down what he is proposing.


Why Now?

Two things pushed Amodei to this point.

First, the OpenAI-Hugging Face hack. You know, the one where OpenAI’s agents escaped their sandbox and broke into Hugging Face’s servers.

Second, the fact that AI is advancing faster than he expected. Specifically, its “growing ability to build the next generation of AI.”

That last part is the scary bit. AI improving itself is exactly the kind of thing safety researchers have been warning about.

“We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain.”


The Resignation That Sparked This

The timing here matters.

A researcher named Jacob Coxon just resigned from Anthropic. His reason? He believes the leading AI companies are “gambling with our lives” while the people building the technology “earnestly believe it could kill us all by the end of the decade.”

That is a former employee saying his own company might kill everyone.

Amodei’s post does not mention Coxon directly. But the timing speaks for itself.


Strategy 1: Embedded Evaluators

This is the big one. And it is the one Anthropic is actually committing to.

Amodei wants third-party evaluators from organizations like METR to be embedded inside AI companies. These people would verify that companies are actually following their safety commitments. They would also make sure safety incidents get reported.

Think of it like bank regulators who sit inside banks. Except these people sit inside AI labs.

What does “embedded” actually mean? Company badges. Desks. Laptops. Access “mostly comparable to what internal risk assessment teams have.”

Amodei said this is “something Anthropic is unilaterally committing to.” He is also calling on governments to require other AI companies to do the same.

Sam Altman responded by calling it a “good idea” and said OpenAI would follow suit.

“We’ll have more to share soon.”


Strategy 2: Coordinated Safety Standards

Next up, Amodei wants the leading AI companies in democratic countries to coordinate on common safety standards. And also agree on limits to how fast they race ahead.

Now, you might be thinking. Would that not be illegal? Companies coordinating on anything usually raises antitrust red flags.

Amodei knows this. His solution is for the US government to step in.

“For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions. They don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations.”

He also addressed the China argument. You know the one. “If we slow down, China wins.”

Amodei said the US could slow China’s progress by refusing to sell powerful chips and cracking down on model distillation. That could “widen America’s lead significantly over the next 3 to 5 years.”

So his pitch is this. Slow down safely, and you can still stay ahead.


Strategy 3: Global Coordination

Finally, Amodei wants global coordination. That means the US and its allies trying to work with authoritarian governments.

Yes, that includes China.

He admitted there are “stark limits on what can be achieved.” But he suggested there might be room for agreement on narrow issues. Things like prohibiting AI for biological weapons.

Even a small agreement is better than nothing.


The Reaction

Sam Altman wrote:

“I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we’ve had at OpenAI in recent weeks.”

Elon Musk posted:

“Dario is right.”

So the three biggest names in AI are all nodding along. That is notable.


The Pushback

Not everyone is buying it.

Journalist Brian Merchant pointed out that he has yet to see “a credible, step-by-step documentation of how exactly AI might move from self-recursively improving AI to killing every single human on the planet.”

His bigger concern? That proposals like this “would likely only wind up serving Anthropic and OpenAI; it’s what regulatory capture looks like in action.”

In other words, the companies writing the safety rules would also be the ones benefiting from them.


Amodei’s Response

Amodei has been called a doomer for years. He said he has tried to offer a “balanced” perspective and argued the backlash is “fundamentally a crisis of trust.”

In his new post, he wrote:

“I continue to believe that AI can enormously improve the quality of human life. My desire to achieve these benefits is undimmed. But the benefits will only be achieved if we build the technology in the right way.”


The Bottom Line

Amodei proposed three strategies to slow AI development. Embedded evaluators, coordinated safety standards, and global coordination. Anthropic is committing to embedded evaluators. OpenAI says it will follow suit. But critics say this looks like regulatory capture dressed up as safety.

Whether any of this actually slows anything down remains to be seen. But it is the most concrete plan anyone has put forward so far.

Similar Posts