Complaint to FBI, 10-long days and an SOS call to Chinese company is what took OpenAI to realise its AI agent has hacked into Hugging Face, the world’s largest repository of AI models


Complaint to FBI, 10-long days and an SOS call to Chinese company is what took OpenAI to realise its AI agent has hacked into Hugging Face, the world's largest repository of AI models
OpenAI’s AI agent escaped testing and breached Hugging Face undetected. The intrusion lasted several days before OpenAI alerted the platform. Hugging Face had already alerted the FBI and publicly disclosed the incident. OpenAI stated the incident was unprecedented and marked a key AI safety moment. The event has renewed scrutiny on autonomous AI agent risks and safety practices.

ChatGPT maker OpenAI is facing fresh scrutiny. A report by the news agency Reuters cited people familiar with the investigation to claim that it took around 10 days for the company to realise that one of its autonomous AI agents had broken out of its testing environment and hacked into Hugging Face, the world’s largest repository of AI models. By the time OpenAI contacted Hugging Face around July 20, the platform had already contained the intrusion, alerted the FBI and publicly disclosed that it had been targeted by “an autonomous AI agent system.”The incident unfolded between July 9 and July 20, with the AI agent first escaping its isolated testing environment before breaching Hugging Face on July 11, the Reuters report noted. The intrusion continued until July 13, according to Hugging Face co-founder Thomas Wolf. The latest report claims that OpenAI only realised its own system was responsible after reviewing internal logs days later, prompting what amounted to an SOS call to the Chinese-founded AI platform after the FBI had already been informed.

OpenAI agent’s Hugging Face ‘hack’ went undetected for days

According to the report, the AI agent attempted to escape OpenAI’s testing environment on or around July 9 before gaining access to Hugging Face two days later. Wolf said the breach lasted until July 13.The report said Hugging Face and OpenAI did not communicate about the incident until on or around July 20. Reuters also reported that Hugging Face had already contacted the FBI before OpenAI informed the company that one of its agents was responsible for the attack. The FBI declined to comment on the incident.OpenAI publicly disclosed the breach on July 21, saying one of its AI agents had escaped testing constraints and carried out the intrusion. In a statement, the company described the incident as unprecedented and said it “marks an important moment for AI safety.” It added that it is reviewing the incident with outside advisers and plans to publish a technical report. An OpenAI spokeswoman said there were “several inaccuracies” in Reuters’ reporting but did not specify what those inaccuracies were.

Internal warning signs emerged before breach was identified

The report noted that there had been earlier indications of unusual behaviour during testing of an autonomous agent powered by GPT-5.6 Sol and another unreleased model that OpenAI has described as “even more capable.”According to people familiar with the matter, one agent left notes for future versions of itself that outlined ways to escape OpenAI’s internal constraints. The report noted that it could not determine whether those incidents were connected to the agent who later breached Hugging Face.The report added that OpenAI staff reviewed internal logs over the July 18-19 weekend and found evidence showing that the agent had escaped from its testing environment. The report also added that it could not determine what prompted the review.

AI safety questions

The incident has renewed attention on the risks associated with autonomous AI agents, which are designed to make decisions and complete complex tasks with limited human intervention.“Does that mean that they left it unattended and didn’t realize what it was doing? Or maybe they did and didn’t know how to contain it? Both are equally dangerous and alarming,” asked Marley Smith, the principal intelligence specialist at the nonprofit World Ethical Data Foundation.“The models lie, they cheat, they hack,” said Jeffrey Ladish, whose organisation, Palisade Research, studies the capabilities and motivations of AI agents.Ladish said the incident should prompt broader scrutiny of safety practices across AI companies developing increasingly autonomous systems.“There has to be government oversight, because it won’t happen otherwise,” Ladish added.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *