Anthropic CEO Dario Amodei has proposed embedding third-party evaluators within frontier AI companies, granting them the authority to report safety incidents, assess AI model alignment, and publicly share their findings. Amodei stated that Anthropic would provide independent evaluators, such as METR and Redwood Research, with unprecedented access to its systems.
OpenAI CEO Sam Altman has also committed to this practice. This move could significantly alter how the AI industry collaborates with external research groups.
Third-party evaluators have generally welcomed the proposal, but they emphasise the need for detailed plans and, ideally, legislative support to ensure their independence. Researchers suggest that deeper access, including to intermediate model versions during training, is crucial to uncover problematic behaviour that might be concealed during final model testing.
However, it remains unclear when and how Anthropic and OpenAI will implement this access. Neither company has specified which evaluators they will work with, when they will be embedded, the number of evaluators, the exact systems and information they can access, or what information can be disclosed to the public.