According to CNBC, Anthropic CEO Dario Amodei proposed embedding long-term independent third-party safety assessors in leading AI companies, granting them access close to internal risk teams and the authority to independently publish their findings under limited restrictions. He explicitly cited the example of embedded banking regulators and promised that Anthropic would provide third-party assessors with system access equivalent to that of internal risk teams.
In a long article over the weekend, Amodei also warned that within the next 6 to 12 months, a misaligned group of agents could take over most of the internet and cause losses of hundreds of billions of dollars. Previously, Jacob Coxon, a former researcher at Anthropic, resigned and warned that leading labs were racing to develop potentially uncontrollable systems.
Bank Examiners Can Shut Down Banks, but AI Assessors Have No Veto Power
Several banking regulatory experts and AI assessment professionals pointed out that this proposal lacks the enforcement power that bank regulators possess. Julie Andersen Hill, Dean of the University of Wyoming Law School, said that government examiners are stationed at large banks and can instruct banks to stop certain activities, limit growth, force management changes, or even shut down the bank in extreme cases. However, Amodei's proposed assessors can only investigate and report, without the power to stop model training or release. "This is a fundamental difference because the power of bank regulators is much greater."
Albert Ziegler, AI lead at cybersecurity company XBOW, said that black-box testing can reveal whether a model has the ability to perform dangerous tasks, but major risks may only appear in condition combinations that the assessors have never triggered. "We really don't have veto power." Assessors can force companies to make informed decisions before releasing, but the final decision-making power remains with the company.
Questioning 'Independence': The Role of METR and 'Audit Washing' Criticism
Amodei named the nonprofit organization METR as a potential embedded assessor, while another former Anthropic researcher, Joe Benton, recently left to join METR. Deborah Raji, a researcher at the University of California, Berkeley, pointed out that this arrangement involves conflicts of financial, ideological, and personal interests, and that METR has mainly focused on Anthropic and OpenAI in recent years, "allowing all sorts of strange things to happen." She believes that merely gaining access does not equal independence, and that an institution independent from the company should decide who is qualified to assess, what can be checked, and where the results are reported. Otherwise, "it's like the company hiring a regular friend to check its homework."
Raji said that the ultimate standard is whether adverse findings result in consequences — "if you do an audit and nothing happens, it's audit washing." Hill said, "If you truly believe AI has the power to destroy society, then you must have an independent supervisor capable of pulling the plug."
Sarah Heck, public policy director at Anthropic, responded on Wednesday that AI companies cannot manage oversight and safety issues through "honorable codes": "We can't check our own homework, and that's very clear."
