Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?
Anthropic and OpenAI plan to integrate independent safety evaluators inside their labs, but researchers warn that true oversight requires more than just access.

Leading artificial intelligence labs Anthropic and OpenAI have announced plans to embed independent safety evaluators directly within their organizations. This move aims to address growing concerns about the rapid advancement of powerful AI models and the potential risks they pose to society.
While researchers have welcomed this unprecedented level of access to the labs' internal processes, many remain cautious. They warn that meaningful oversight is impossible without absolute transparency, structural independence, and eventually formal regulation.
Industry experts point out that for embedded evaluators to function effectively, they must operate without corporate interference or conflicts of interest. Without these safeguards, internal monitoring risks becoming little more than a superficial compliance exercise.
The debate highlights the broader challenge of governing rapidly evolving technologies where voluntary corporate commitments often outpace established legal frameworks. Comprehensive regulation will likely be needed to ensure accountability across the entire industry.
For the global tech community, the outcome of this experiment will set an important precedent for how artificial intelligence safety is managed, monitored, and regulated in the years to come.



