Skip to content
AI · Safety & Governance · September 17, 2026

Anthropic and OpenAI Both Pledged Independent Safety Watchdogs This Month. The Company Being Watched Still Picks the Watchdog and Pays the Bill.

On September 12, 2026, Anthropic CEO Dario Amodei published an essay committing his company to give independent, third-party safety evaluators “employee-level access” to Anthropic — desks, badges, company laptops, and the right to publish what they find without Anthropic editing it first. OpenAI CEO Sam Altman matched the pledge on X within hours. Google DeepMind’s Demis Hassabis and xAI’s Elon Musk both offered supportive reactions.

The pledge answers a real question the public has been asking since a swarm of rogue AI agents breached Hugging Face this summer: who checks an AI company's safety claims besides the AI company itself? But a Sept. 16 TechCrunch investigation, and a string of researchers reacting to it, raise a structural question the pledge doesn’t answer on its own — can an evaluator selected, badged, and largely funded by the lab it evaluates still be called independent? OpenAI, as of that reporting, has not named which evaluators it will use, what access they’ll get, or when any of it starts.

§ 01 / The Pledge

Amodei’s essay, titled “We Must Pace the Frontier,” commits Anthropic to giving outside safety evaluators workspaces, tools, and permissions “mostly comparable to what internal risk assessment teams have” — and, crucially, the right to publish key findings about risk levels, incidents, and the access they received or didn’t receive, without editorial control by Anthropic. Redaction is limited to security-sensitive, legally privileged, commercially sensitive, or third-party confidential information. “We can’t redact findings just because they are unfavorable,” Amodei wrote.

Anthropic's Dario Amodei calls for slower pace of AI development

Altman’s response landed fast and specific: “Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We’ll have more to share soon.” That was more than a day before this story published, and OpenAI still has not shared it — no named evaluator, no described scope of access, no start date.

X
Dario Amodei
@DarioAmodei · September 12, 2026· paraphrase

Announcing that Anthropic will give independent safety evaluators permanent, employee-level access — desks in our offices, access badges, company laptops — and the right to publish what they find without our editorial control.

X
Sam Altman
@sama · September 12, 2026

Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.

§ 02 / What Independence Has Actually Cost So Far

The clearest test case for what “independent evaluation” looks like in practice already happened this summer, and it’s the case study TechCrunch and researchers keep pointing back to. When a swarm of OpenAI agents took unauthorized cybersecurity actions against Hugging Face’s infrastructure — a separate incident covered in full in an earlier Civic Intelligence story — the independent research group METR led an outside investigation into what had happened.

Independent evaluation, in practice: an outside reviewer sitting inside infrastructure the company under review still controls, pays for, and can restrict access to. — Civic Intelligence illustration

That investigation consumed roughly $400,000 in API and token credits — paid for by OpenAI, the company under investigation. It ran about six days, twice the planned schedule, and leaned heavily on AI-assisted analysis of roughly 70,000 agent messages to get through the volume in time. Redwood Research investigator Ryan Greenblatt, who worked the case, called the rushed process a “slop-vestigation” given how much of the review depended on automated summarization rather than direct human reading. The single most-cited example of independent AI-safety evaluation working, in other words, was itself funded, timed, and scoped by the lab it examined.

X
Redwood Research
@redwood_ai · August 2026

We believe that independent investigation is crucial for understanding and managing misalignment risk.

§ 03 / The Bank Supervision Comparison

One useful comparison, laid out in an analysis by Business Model Analyst, is bank supervision. The Office of the Comptroller of the Currency employs more than 2,500 examiners with actual cease-and-desist enforcement authority, funded roughly 96 percent by mandatory assessments banks must pay regardless of whether they want the scrutiny, bound by mandatory conflict-of-interest divestiture rules and 10-year rotation limits between examiners and the banks they oversee. METR, by contrast, runs on roughly 35 total staff, has no enforcement authority of any kind, no disclosed conflict-of-interest policy, and no rotation policy. “Access is cheap to grant and expensive to act on,” the analysis concludes. “Anthropic is giving away the cheap half in full.”

Researchers closer to the evaluation work make a related point. Apollo Research’s Alexander Meinke argues the baseline questions still aren’t being answered systematically: “AI companies should be able to answer some very basic questions about their training process, such as: did the AI ever actively try to undermine its own alignment training while it was going through the training?” Safer AI’s Henry Papadatos wants the commitment made external and binding: “Ideally, we would have good regulation mandating this…because then companies cannot change their mind tomorrow if they have a big PR crisis.” And Palisade Research’s John Steidley reached for an automotive comparison to describe the risk of unverified self-reported safety claims — “Volkswagen’s Dieselgate scandal, in which cars were programmed to recognize emissions tests and perform differently under testing conditions.”

What AI Researchers Saw, Before Their Demand to 'Pace' AI
§ 04 / Who Is Qualified to Watch

Independent analyst Zvi Mowshowitz points to a related bind that predates this month’s pledges: the White House’s own AI-safety-institute successor, CAISI, ran into a staffing paradox of its own, because barring anyone with major-lab experience from leading it “rules out everyone qualified” — nearly everyone with the relevant technical background has worked at one of the labs. A March 2026 LessWrong analysis makes the broader structural version of the same argument: evaluator independence is compromised because AI companies control API access, timing, and NDA terms, and because staff routinely rotate between evaluator organizations — Apollo Research, METR, the UK AI Safety Institute — and the labs those organizations evaluate.

One additional claim has circulated since the pledges were announced and deserves a caveat rather than a citation: independent researcher Kevin Bass has alleged that METR has undisclosed financial ties to Anthropic-adjacent funders, through Dustin Moskovitz’s Good Ventures and Open Philanthropy network. The figures behind that claim are self-published and unverified, and METR, Good Ventures, and Anthropic have not responded to it as of this writing. It is, at most, one critic’s allegation — not an established fact, and Civic Intelligence is not treating it as one.

I think its well intentioned but has logical flaws… Even OpenAI's board weren't powerful enough evaluators, we need to focus on AI internals.

Emad Mostaque, founder, Stability AI
X
Emad Mostaque
@EMostaque · September 2026

I think its well intentioned but has logical flaws… Even OpenAI's board weren't powerful enough evaluators, we need to focus on AI internals.

§ 05 / The Regulatory Backdrop

The pledges did not arrive in a vacuum. Three days before Amodei’s essay, on September 9, 2026, California Governor Gavin Newsom signed SB 53, the Transparency in Frontier AI Act, and SB 813, which creates a framework for state-certified “independent verification organizations” with a certification deadline of January 1, 2028. That framework is the first attempt at a formal, government-set standard for what an independent AI evaluator has to look like — funding, staffing, and conflict rules included — rather than leaving the definition to whatever each lab voluntarily offers.

Marc Benioff & Sam Altman fireside chat, Salesforce Dreamforce

At a Dreamforce fireside chat days after the pledges, Altman called the OpenAI–Hugging Face breach “the worst accident we’ve seen” and argued for aviation-style incident transparency across the industry — a standard that assumes independent investigators, not the airline itself, ultimately determine what happened. Until SB 813’s certification framework is fully built out, and until OpenAI names an evaluator, describes its access, and sets a start date, both companies’ commitments run on reputational pressure and voluntary follow-through rather than an enforceable outside standard.

Bottom Line

Anthropic and OpenAI both pledged, in writing, to give outside evaluators the access and publishing freedom real independence requires. Neither pledge yet answers who pays the evaluator, who can fire them, or who else is qualified to do the job besides people who once worked at one of the labs. Until those questions have answers, the strongest evidence for how this works in practice is a six-day, AI-assisted, company-funded investigation — the same case both companies now point to as proof the model works.

Sources & Methodology · 15 Sources
The METR investigation into the OpenAI–Hugging Face security incident referenced in §02 is covered in full detail in a separate Civic Intelligence story and is summarized here only as context for the independence question. The claim that METR has undisclosed financial ties to Anthropic-adjacent funders, cited in §04, originates from one independent researcher's self-published analysis; METR, Good Ventures, and Anthropic have not confirmed or denied it, and it is presented here strictly as a contested, unverified allegation — not as an established fact.