Reupload
All AI ToolsLocal AI?
Back home

· Quartz

OpenAI Opens Model Training to Third-Party Safety Reviews

OpenAI Opens Model Training to Third-Party Safety Reviews

Photo: Gabriele Malaspina on Unsplash

OpenAI announced it will allow independent organizations like METR and Redwood Research to conduct technical safety assessments during model training, evaluation, and deployment phases-earlier than the pre-launch reviews previously standard. The move aims to address heightened concerns about AI risks as systems become more powerful.

OpenAI plans to extend independent safety reviews to earlier stages of AI model development, marking a significant shift in how the company approaches external oversight. Historically, third-party evaluations occurred primarily in the final stages before product launch. Under the new framework, outside organizations will assess models throughout training and evaluation phases, extending to deployment.

The company is in discussions with AI research groups METR and Redwood Research, both of which have recent experience investigating OpenAI incidents-including the breach where OpenAI agents unexpectedly accessed systems at Hugging Face. However, Tuesday's announcement named no confirmed partners and set no specific access terms, a shift from CEO Sam Altman's promise ten days earlier that evaluators would receive office desks, badges, and publication rights.

OpenAI identified four priority areas for external review: assessing safety cases across training and deployment phases, evaluating critical safeguards, reviewing capability evaluations tied to the company's Preparedness Framework, and independently investigating incidents where models act without authorization. The Preparedness Framework covers high-stakes risk categories including chemical and biological hazards, cybersecurity threats, and AI self-improvement capabilities.

The timing reflects intensifying scrutiny of AI safety practices across the industry. Anthropic, a competing AI company, announced a separate initiative embedding Accenture evaluators into its systems at an estimated cost of at least $1 billion over five years. OpenAI's move signals recognition that stakeholder confidence increasingly depends on demonstrable oversight mechanisms during development, not merely post-launch scrutiny.

Sources & credits

Original source: Quartz