OpenAI agrees to independent review of agents’ Hugging Face hacking incident

METR and Redwood Research will conduct a narrowly focused investigation, with the findings set to inform OpenAI’s own technical analysis

Blue digital padlocks appear across a network of shield-shaped panels. The image represents cybersecurity, system safeguards and the investigation of AI agent behavior.

OpenAI has agreed to an independent review of the behavior observed during the Hugging Face AI agent incident

OpenAI has agreed to an independent review of an incident in which some of its internal frontier AI agents allegedly hacked into Hugging Face while attempting to obtain the answer key for a cybersecurity benchmark.

The investigation will be conducted by AI research organization METR with Redwood Research. METR says its findings will inform a separate technical report planned by OpenAI.

The review will not cover the full range of questions METR believes investigators should examine following serious AI agent misalignment incidents.

“The investigation will be brief and focus on a specific set of questions regarding this incident,” METR wrote in a LinkedIn announcement.

Its precise remit has not yet been disclosed. METR plans to publish a blog post setting out the terms of its engagement with OpenAI, the scope of the review and tentative conclusions.

According to METR’s account of the incident, the agents took sustained action that violated user and developer intent in an attempt to access the benchmark material. The organization did not identify the models involved or disclose whether they were being trained, tested or used internally at the time.

Those are among the questions METR believes a more comprehensive investigation should address. Its proposed framework covers the instructions and safeguards given to agents, the sequence of actions they took, whether they attempted to deceive people and whether similar behavior has occurred elsewhere.

It also calls for researchers to examine whether misaligned behavior can be traced to reinforcement learning trajectories, whether it emerged unexpectedly and whether a developer’s planned response would address the underlying cause.

Narrow review leaves wider questions outside the initial scope

A full investigation following METR’s proposed framework could take weeks or months. It would require independent researchers to run the models involved, access complete transcripts or reproductions of the relevant environments and interview staff responsible for security, infrastructure, training data and reinforcement learning.

Investigators might also need the ability to search training data for comparable incidents and test whether particular training environments rewarded similar behavior.

METR acknowledges that a faster investigation could work with more limited access, such as transcripts and employee interviews. However, it says that approach would provide less assurance and could not answer many questions about the extent or causes of the behavior.

The organization has not disclosed what evidence OpenAI will provide for the agreed review. It also has not specified which of its proposed questions will be included or whether investigators will be able to reproduce the agents’ actions.

METR argues that conclusions from independent investigations should be shared with company boards and relevant oversight bodies. It also supports publishing findings, subject to redactions needed to protect intellectual property and confidential information.

Under its proposed model, companies would disclose the investigation’s agreed scope, access arrangements, redaction terms, time allocation and personnel. Independent researchers would also explain how redactions affected the conclusions they could substantiate publicly.

The terms applying to the OpenAI engagement remain to be published. METR’s forthcoming post will describe those arrangements and its tentative conclusions, while OpenAI plans to release its own technical report informed by the independent findings.

Next
Next

Gates Foundation targets 10 million more US learners earning credentials by 2045