AI Companies’ Shaky New Plan to Police Themselves

· The Atlantic

Dario Amodei, Sam Altman, and Elon Musk don’t agree on much. Each started their AI company—Anthropic, OpenAI, and xAI, respectively—out of spite at competitors’ greed and recklessness; each is certain that he alone can be entrusted to build superintelligence safely. But in the midst of the AI panic that has seized the country over the past two weeks, the three CEOs found one point of consensus: the introduction of “embedded evaluators,” or outside experts, to sit inside AI companies and monitor their safety practices. Both Amodei and Altman made express commitments to integrating such evaluations inside their companies. Musk followed with a quote-post stating simply, “Dario is right.”

Yet the details of such a system for regulating AI remain vague. How would such a setup work? An unknown number of outside evaluators would be assigned desks to sit at, computers to use, and NDAs to sign. Then, they might seek out risks: Evaluators could spend their days testing whether unreleased models trained for drug discovery can be altered to design viruses too. They might audit organizational practices: Are employees cutting corners on safety in the rush toward the launch of a new model? The evaluators will then write up what they find, and hopefully those reports will see sunlight. Amodei has promised editorial independence, but with significant asterisks: carve-outs for “security-sensitive, legally privileged, commercially sensitive, or third-party confidential information.” Anthropic, OpenAI, and xAI did not immediately respond to requests for comment.

Visit mchezo.life for more information.

[Read: America’s Hypocritical Take on Intellectual Property]

As for these experts, it seems likely that they’ll come from small nonprofits, as well as large firms. That includes the tech-consulting giant Accenture, with which Anthropic plans to fund a billion-dollar partnership, and the nonprofit METR, which led the investigation of OpenAI agents’ hack of Hugging Face. After that hack, OpenAI granted METR exclusive access to agent transcripts, data sets, and interviews with OpenAI staff. The resulting report revealed concerning levels of cheating and deception among the agents.

On-site evaluation has precedent in other high-risk industries such as aviation and banking. The FAA delegates some safety certifications to engineering experts inside airplane manufacturers, whereas the Federal Reserve stations its own supervisors inside banks to enforce regulation. “OpenAI and Anthropic are like Citibank: They pose systemic risk to the globe, so you have examiners,” Brad Carson, a former U.S. representative who is now the president of the AI-governance nonprofit Americans for Responsible Innovation, told me.

Yet Amodei’s proposal is still far from a comprehensive AI-auditing regime. As long as safety evaluations remain voluntary, the independent organizations selected by the companies will get as much access as the companies allow to evaluate the specific risks that the companies choose. That doesn’t add up to true independent oversight, and at least some METR employees seem to agree. “None of the work that we’ve done so far passes my bar for an ‘audit’ of an AI company or its systems, and was definitely not ‘regulation’ in any meaningful sense (whether bank examiner–like or otherwise),” Charles Foster, a policy staffer at the organization, wrote on X, speaking in his personal capacity. Recall the Facebook Oversight Board, a legally independent entity created with the ambitious aim to adjudicate the platform’s thorniest speech questions—but which has largely proven powerless to make major decisions that cut against Meta’s interests.

Third-party evaluators must walk the line between holding their corporate hosts accountable and not saying anything that gets them booted from the premises. The same access to Slack channels and lunchroom chatter that helps an evaluator conduct whole-organization safety audits can lead to them succumbing to the same blind spots as their host. Here, the 2008 financial crisis offers hard lessons. A Senate investigation found that the New York Fed’s internal culture had been “excessively deferential,” where supervisors feared speaking up about the issues they were supposedly regulating. Moreover, credit-rating agencies were motivated to offer lenient terms to win business from banks. These conflicts of interest are exaggerated when evaluators—like Accenture—are funded directly by the companies. “It doesn’t really seem tolerable to let AI companies pick providers that they have hundreds of millions of dollars of business with to be the primary cops on the beat,” Nathan Calvin, the general counsel of the AI-advocacy group Encode, told me.

Then there’s the question of whether the evaluators and the companies they’re asked to monitor are too intertwined. METR, a 43-person organization based in Berkeley, has partnered with the leading AI labs to measure model misalignment since 2023. It does not accept payment or donations from AI companies or employees. Still, many METR staffers come from the same circles as the people they audit, and some have known one another socially and professionally for years. “The AI space is not that large,” Jacob Steinhardt, a Berkeley computer-science professor and the CEO of Transluce, a San Francisco–based AI-evaluation nonprofit, told me. “If you say no former lab employee can ever evaluate any AI model, you are ruling out a pretty broad and important set of expertise.” Those longtime relationships have helped small nonprofits gain access to corporate data, yet they now challenge the organizations’ perceived neutrality. “THESE are the people we’re trusting to beat China in AI? Give me a break,” posted House Majority Leader Steve Scalise, quoting a New York Post cover describing METR employees as “The Woke Wizards of A.I.”

Embedded evaluators are surely preferable to the status quo of no oversight at all. The public will learn more about AI incidents when there are watchdogs around—as it has from METR’s Hugging Face report, or Transluce’s recent report on OpenAI agents attempting to hack into a government website. Yet the evaluators will lack teeth and trust until backed by government authority: to license a wide-ranging group of expert organizations, to mandate frontier-AI developers to bring them in, and to set the minimum safety and transparency standards that they must meet. “You need serious oversight,” Gillian Hadfield, a professor of AI governance at Johns Hopkins, told me. “Are we making sure that the entities that are doing this work are qualified? Have we got ways of establishing and maintaining their independence? Do we have a threat that says, ‘If you don’t do this well, you’ll lose your license’?” The worst-case outcome is a race to the bottom: where auditors compete on leniency instead of quality, winning big contracts in return for hasty sign-offs.

Meanwhile, it seems that at least some sort of regulation is approaching. In the face of escalating public backlash, internal employee activism, and fears of near-term AI extinction events, both Anthropic and OpenAI have declared support for various forms of state and federal legislation around third-party audits. In a recent podcast, Ben Horowitz—whose venture firm has been a major voice against AI regulation—also seemed to embrace a private-auditor approach, saying that the government “is very good at setting the rules” around unintended model behaviors, whereas “a very competent private company” should evaluate compliance. The Overton window in Silicon Valley is shifting. But if the stakes really are existential, hiring a few extra inspectors falls painfully short—especially when the details of the arrangements are this slippery.

It all feels too little, and too late. Some politicians, facing an angry and frightened public, have turned to more aggressive asks, such as nuclear-style treaties with China and a ban on superintelligence itself. Two years ago, California Governor Gavin Newsom vetoed a state bill that included mandated third-party audits, among other safety requirements, after a fierce lobbying campaign by OpenAI and Andreessen Horowitz. (Anthropic was an exception, expressing conditional support.) Last week, Newsom—sensing the changing tides—issued an executive order to require onsite auditors and accelerate the development of an “AI kill switch.”

President Trump still has his foot on the gas pedal. “The only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT,” he posted last week. In June, the administration told CAISI, the federal government’s internal AI-evaluation unit, to stop publishing its reports. Trump’s accelerationism—informed by his rapport with ballroom sponsor and NVIDIA CEO Jensen Huang—likely dooms federal AI-safety regulation for now.

It’s nice to see some AI-industry leaders agreeing that someone else should rein them in. And if implemented well, embedded evaluators are a smart technocratic move. But after years of developing the technology, lobbying against guardrails, and causing multiple accidents, AI leaders shouldn’t be surprised if their preferred—and valid—policy solution ends up being met with disappointment, and a heap of skepticism.

Read full story at source