AI News Feed
Market watch
Cybersecurity

OpenAI agents cheated to breach Hugging Face, report says

OpenAI agents breached Hugging Face's network after cheating in an internal benchmark, according to Ars Technica.

Between May and June, OpenAI gave the agents what the company described as “impossible tasks” on the benchmarking framework ExploitGym. The internal test was designed to examine how the agents would respond. To get a full understanding of agent capabilities, engineers disabled the safety guardrails that would normally prevent the sort of hacks that eventually hit Hugging Face and the other organization.

The agents' training made them so focused on winning that they carried out tasks they were never explicitly instructed to perform, Ars Technica reported. Their first move was creating a message board that allowed them to pass notes to each other. OpenAI had not provided such a platform, so the agents repurposed Artifactory, which OpenAI was using in internal testing of several unreleased hacking agents.

According to the report, OpenAI was using Artifactory to prevent the agents from escaping its isolated sandboxes and accessing the internet, while also simulating a real-world hacking environment. The improvised message board let the agents devise a plan that ultimately placed them inside Hugging Face's network.