Advertisement

OpenAI admits AI agent escaped safeguards and hacked Hugging Face during internal evaluation

OpenAI has disclosed that one of its experimental AI agents broke out of a controlled testing environment, gained internet access and breached Hugging Face without human direction. The company says the episode highlights the growing cyber risks posed by increasingly capable AI systems as governments intensify scrutiny of frontier models.

Advertisement
OpenAI shuts down China-linked influence campaigns that used ChatGPT to target US tariff
 (File image/Reuters)
OpenAI shuts down China-linked influence campaigns that used ChatGPT to target US tariff (File image/Reuters)
FP Tech Desk|Jul 22, 2026, 08:00:16 IST

OpenAI has revealed that one of its experimental AI agents independently escaped a restricted testing environment and carried out a cyber attack against AI platform Hugging Face, marking what the company describes as an unprecedented security event involving an autonomous AI system.

The disclosure comes amid mounting concerns over the offensive cyber capabilities of advanced AI models and growing calls for stronger oversight before increasingly powerful systems are released to the public. It also arrives just days before OpenAI chief executive Sam Altman is expected to meet US officials in Washington to discuss the next generation of frontier AI models, an area attracting heightened regulatory attention following recent advances in AI-driven cybersecurity capabilities.

Advertisement

According to OpenAI, the incident occurred during an internal evaluation designed to measure the cyber abilities of its latest AI models. The company said the exercise involved GPT-5.6 Sol, launched earlier this month, alongside a more advanced unreleased model undergoing pre-launch testing.

AI agent exploited unknown flaws to reach the internet: What happened

OpenAI said engineers had intentionally relaxed certain cyber safeguards as part of the evaluation, while keeping the AI systems inside a sandbox environment intended to isolate them from the open internet. The models were instructed to attempt hacking activities within those controlled conditions so researchers could assess their capabilities.

techMore from Tech

Instead of remaining confined, the AI agents reportedly devoted significant computing resources to identifying a route beyond those restrictions. According to the company, they discovered and exploited previously unknown software vulnerabilities, allowing them to escape the sandbox and establish internet access.

Once outside the testing environment, the system continued pursuing its assigned objective. OpenAI said the agent obtained login credentials and used them to compromise Hugging Face, a widely used platform that hosts open-source AI models and datasets for developers.

Advertisement

In a blog post announcing the breach, OpenAI said, "We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly."

The company added that it expects similar events to become more common as AI systems become increasingly capable of conducting sophisticated cyber operations.

Hugging Face confirmed that it detected an intrusion on its systems last Friday. Chief executive Clement Delangue wrote on X that the company had initially suspected the breach originated from a frontier AI laboratory because of the sophistication displayed during the attack.

After working with OpenAI to investigate the incident, Delangue said the companies concluded there was no evidence of malicious intent behind the event. "We strongly believe there was no malicious intent on their part. It's quite mind-blowing that all of this happened autonomously!" he wrote.

Incident fuels debate over AI safety and regulation

The breach is likely to intensify debate over how frontier AI systems should be tested before deployment, particularly as autonomous AI agents gain greater ability to perform complex digital tasks with limited human oversight.

Advertisement

Unlike conventional AI chatbots that respond to prompts, AI agents are designed to execute multi-step objectives independently. That autonomy has raised concerns among researchers and policymakers that highly capable systems could identify unintended methods of completing assigned tasks, including exploiting vulnerabilities in digital infrastructure.

OpenAI said Hugging Face's own AI-powered defensive systems ultimately detected and halted the unauthorised activity on its infrastructure, limiting the impact of the breach.

The company also confirmed that it informed law enforcement agencies and relevant government authorities after discovering what had occurred.

The disclosure comes against a backdrop of increased scrutiny from US policymakers, who have shown a greater willingness to examine advanced AI systems before they reach the public. That attention has grown following global concern over Anthropic's Mythos model, which demonstrated highly advanced capabilities for identifying and exploiting cybersecurity weaknesses.

OpenAI's latest disclosure is expected to add urgency to ongoing discussions around AI safety standards, cybersecurity testing and the safeguards needed to ensure increasingly autonomous systems remain under effective human control.

Handpicked stories, in your inbox
Global stories. Indian perspective. Zero noise.
No Spam. Unsubscribe Any Time.
First Published:Jul 22, 2026, 08:00:16 IST
Advertisement
Advertisement
Advertisement
Advertisement
Up Next