The September pledge turned outside AI safety auditing into an industry commitment. Early pilot arrangements show why the evaluation science is the actual bottleneck.
The researcher types a scenario into a frontier model and watches how it responds. The prompt might ask the model to write a phishing email, plan a multi-step intrusion, or talk a user out of a medical decision. A second researcher scores the response on a rubric. The lab that built the model sits somewhere in the loop, deciding what the auditor sees, when the auditor sees it, and how much of what the auditor finds becomes public.
That is what a third-party AI safety evaluation looks like in 2026, and it is the practice that a coalition of major AI labs publicly agreed to scale up this fall. A September meeting at the White House drew commitments from OpenAI, Anthropic, Google, Meta, SpaceXAI, and NVIDIA to work with independent outside evaluators. Coverage of the event framed it as a meaningful concession. The evaluators themselves told NPR the science they are supposed to be applying is "far from sufficient."
The gap between the pledge and the practice is the real story.
Evaluations come in three rough flavors, and each is at a different level of maturity. Capability evaluations ask whether a model can do a given thing: write working code, plan a bioweapon synthesis, pass a long-horizon agentic task. Behavior evaluations ask whether a model does bad things when it can: lying, sycophancy, jailbreak resistance, deception. Protocol evaluations ask whether the lab's own safety practices are real: do the safety cards match what shipped, are dangerous capabilities actually being tested before release, is internal review genuine or theater. The first two have decades of borrowed methodology from psychology and benchmark design. The third is essentially being invented, lab by lab, as it goes.
NPR's reporting, based on interviews with more than a dozen evaluators, found no shared standard for any of the three. Evaluators described testing windows that compress dangerous-capability reviews into days rather than weeks. One evaluation firm, SecureBio, told NPR it was asked to compress a safety review of a frontier model code-named Astra to four business days after initially being told it would have five. The firm said evaluations of that scope normally require at least 20 business days, an order-of-magnitude difference the firm described as routine. NPR's reporting noted the issue is not that any one lab is acting in bad faith, but that no common definition of "an adequate evaluation" exists for the evaluators to point to.
Independent evaluators also navigate a lopsided power dynamic and a real conflict-of-interest problem. The lab controls access. The lab funds the work, directly or through philanthropic intermediaries. The lab decides what nonpublic information can be disclosed. NPR's reporting found that, under current arrangements, evaluators face financial incentives not to publish findings their funders would rather not see. There is no widely adopted counterpart to the auditor-independence rules that govern financial audits.
The early institutional experiments show what the arrangements actually look like in practice. Anthropic announced a partnership with Faculty, Accenture's specialist AI business, to lead embedded evaluation including red-teaming, alignment assessments, and safeguard testing. Anthropic said the partnership is non-exclusive, that it funds the work directly, and that information-access standards, reporting standards, and independent funding arrangements remain unsettled: three of the same structural gaps NPR's reporting flagged. METR's Frontier Risk pilot, run February 16 through March 16 of this year, gave Anthropic, Google, Meta, and OpenAI an exercise in which participants approved what nonpublic information could be disclosed and which internal-model findings could be published. METR's report frames the exercise as entity-based and designed to recur, not as a method for tracking public model releases.
Counterevidence to a blanket "audits are theater" reading exists. METR's pilot shows richer access is possible when labs agree to it. Researcher Alexandra Mallen, quoted by NPR, offered a qualified positive assessment of evaluation work, arguing that longer-term task gaming is the harder problem and that current evaluation work is at least useful evidence rather than nothing. Evaluators NPR interviewed described their work as genuinely informative in narrow cases.
The buildable work is concrete. It starts with shared benchmarks for capability and behavior testing that any qualified evaluator can run without lab cooperation, the equivalent of a public test suite. It includes a written protocol for what an adequate frontier-model safety review must cover, with minimum testing windows and a defined disclosure rule for conflicts of interest. It includes independent funding for evaluator organizations, the way the FDA is funded by industry fees but governed by statute. None of this requires new law today. It requires the labs that signed the pledge to fund and publish the standards the pledge implicitly assumes already exist.
The September announcement named outside evaluation as an industry commitment. The science of doing that evaluation well is younger than the announcement, and the standards the announcement assumes are still being written. The labs that want the credit for a serious safety program now own the work of building one.