The push to establish independent safety oversight for the world’s leading artificial intelligence models is facing mounting skepticism from the very researchers tasked with conducting it. As companies like OpenAI and Anthropic promote the concept of embedding outside safety evaluators into their operations, third-party testers are warning that such arrangements often lack the necessary access to provide meaningful scrutiny. The debate over internal transparency arrives just as local lawmakers escalate their own oversight efforts.

Executives from OpenAI, Anthropic, Google, and Meta are scheduled to testify at a New York City Council hearing regarding AI risks, following a recent letter from OpenAI to the council recommending specific safeguards. The convergence of local legislative pressure and growing disillusionment among independent researchers highlights a critical bottleneck in AI governance: the reliance on voluntary corporate cooperation to assess systemic risks.

The structural limits of embedded oversight

The concept of embedded evaluation—where external researchers are granted internal access to probe models for dangerous capabilities—has gained traction as a compromise between corporate secrecy and public accountability. However, the mechanics of these agreements heavily favor the developers. AI companies retain the ability to withhold underlying training data, restrict access to specific model weights, or redact public findings before they are published. This dynamic effectively reduces independent audits to guided tours, where evaluators can only assess what they are permitted to see.

Marius Hobbhahn, CEO of Apollo Research, an organization that evaluates the safety of models from major AI developers, noted that he has "become very cynical over the last three years" regarding corporate commitments. Anthropic, an AI lab originally founded with an explicit focus on safety, and OpenAI, the creator of ChatGPT, have both engaged with outside evaluators. Yet, the fundamental asymmetry remains unresolved. If an evaluator requests comprehensive access and is only granted a curated subset of data, the resulting safety certification risks serving as a public relations shield rather than a rigorous technical audit.

Commercial velocity outpaces governance frameworks

While the debate over safety testing access stalls, the commercial and infrastructural expansion of the AI sector continues to accelerate. The dual reality of the industry is evident in the concurrent activities of its leading figures. Even as OpenAI navigates safety hearings and sends policy recommendations to the New York City Council, its broader ecosystem continues to drive market movements. Recently, shares of Cerebras, an AI chipmaker competing in the hardware space, climbed 6 percent after OpenAI chief executive Sam Altman publicly described the company as a "close partner."

This juxtaposition illustrates the core tension facing regulators and safety researchers alike. The capital and hardware requirements of the AI race are forging deep, market-moving alliances that outpace the development of standardized safety protocols. The upcoming testimonies before the New York City Council by executives from Meta, Google, Anthropic, and OpenAI represent an attempt by local governments to fill the regulatory vacuum left by stalled federal legislation. However, without a mechanism to mandate unvarnished access for third-party evaluators, legislative hearings risk focusing on theoretical safeguards while the actual technical oversight remains constrained by corporate discretion.

The friction between independent evaluators and AI developers suggests that voluntary safety commitments are reaching their practical limits. As local municipalities begin demanding answers and commercial partnerships continue to drive rapid expansion, the industry faces a structural reckoning. The effectiveness of future AI governance will likely depend not on corporate promises, but on whether external auditors can secure the unmediated access required to actually verify them.

With reporting from The Information, CNBC Technology.

Source · The Information