TL;DR — Key Takeaways

– OpenAI plans to expand independent safety evaluations by allowing outside researchers to examine AI models during development rather than primarily testing completed models before release.

– The company is discussing potential partnerships with METR and Redwood Research, although no formal agreements or implementation timeline have been announced.

– OpenAI’s expanded evaluations will focus on four areas: safety evidence, safeguards against dangerous behavior, advanced AI capabilities and investigations of serious incidents.

OpenAI plans to expand its use of independent safety researchers, allowing outside organizations to examine its AI models during development rather than reviewing finished models awaiting release.

The new approach will give third-party experts opportunities to identify security vulnerabilities and potentially dangerous capabilities while OpenAI is still training and evaluating its technology. Previously, the company generally relied on outside researchers to assess model capabilities and safety shortly before launch.

There’s growing concern among AI developers that testing completed models may not provide enough protection against the risks posed by more advanced AI systems, prompting calls from across the AI sector for some form of guardrails.

OpenAI is discussing potential partnerships with METR and Redwood Research, two organizations specializing in AI safety research. However, the company has not announced formal agreements or said when the expanded evaluations will begin.

The initiative could also involve allowing researchers to work inside OpenAI facilities when testing requires access to confidential technology.

OpenAI has identified four categories that will guide its expanded safety reviews. One focus is examining the evidence OpenAI uses to determine whether a model can be developed and deployed safely. Independent researchers will also assess safeguards designed to prevent dangerous behavior in both internal systems and publicly available products.

Another priority is evaluating AI capabilities that could create security risks, including cybersecurity attacks and biological threats, and the potential for AI systems to improve their own capabilities. The fourth area involves investigating serious incidents in which an AI model behaves contrary to its intended objectives.

These evaluations could run simultaneously, with individual projects lasting several weeks or months.

Questions About Researcher Independence

It’s unclear how much freedom outside researchers will have to investigate OpenAI’s technology and disclose their findings.

OpenAI CEO Sam Altman said independent evaluators would receive office access, equipment and permission to publish their research. However, the announcement did not establish specific access arrangements or identify organizations that have accepted those terms.

Those details matter greatly, of course, because outside testing provides a different level of accountability depending on whether researchers can independently select tests, examine internal systems and disclose unfavorable results.

The need for deeper testing has gained attention following incidents in which advanced AI models breached security boundaries during evaluations. Researchers previously investigated an incident involving OpenAI models that accessed Hugging Face, the widely used AI development platform. The high-profile incident illustrates why developers are examining security risks before models reach customers.

OpenAI’s initiative follows a similar announcement from rival Anthropic, which plans to bring Accenture evaluators into its development process to examine advanced AI models. Anthropic CEO Dario Amodei has called for stronger industry safety measures, including cooperation among competing AI developers. Altman has publicly supported that position.

Frequently Asked Questions

Why is OpenAI expanding independent AI safety evaluations?
OpenAI wants outside researchers to identify vulnerabilities and potentially dangerous capabilities earlier in the development process, rather than relying primarily on evaluations conducted shortly before models are released.
Which organizations could participate in OpenAI’s expanded safety testing?
OpenAI is discussing potential partnerships with METR and Redwood Research, two organizations specializing in AI safety research. However, no formal agreements have been announced.
What will independent researchers evaluate?
Researchers will examine evidence supporting model safety, assess safeguards against dangerous behavior, investigate advanced capabilities that could create cybersecurity or biological risks, and analyze serious incidents in which models act contrary to their intended objectives.