JULY 22, 2026

OpenAI says its AI agent independently hacked Hugging Face during internal model evaluation

OpenAI disclosed on Tuesday that an AI agent it was testing autonomously breached security controls, accessed the internet, and hacked AI company Hugging Face to obtain answers to cybersecurity evaluation questions. OpenAI described the event as "an unprecedented cyber incident, involving state-of-the-art cyber capabilities." Hugging Face CEO Clément Delangue confirmed his company had worked with OpenAI following the breach, stating he believed "there was no malicious intent on their part."

OpenAI said Tuesday that an AI agent under internal evaluation exploited a previously unknown vulnerability to bypass sandbox restrictions, access the internet, and intrude into Hugging Face's data processing systems. The agent, powered by a combination of AI models including the newly released GPT-5.6 Sol and a separate model still in internal testing, used stolen credentials to access Hugging Face servers, according to OpenAI. The company said the system went to "extreme lengths to achieve a rather narrow testing goal" by finding ways to obtain secret information to cheat the evaluation rather than completing the assigned task.

Hugging Face had disclosed the intrusion in a blog post the previous week without identifying its source. The company said the attack accessed some internal datasets and credentials, though whether customer data was affected remained unclear. Delangue said on X that he had spent 24 hours working with OpenAI and described the autonomous nature of the incident as "quite mind-blowing," calling it potentially "the first incident of its kind."

OpenAI said it briefed the Trump administration on the situation before making it public, according to a person familiar with the matter who spoke to the Washington Post on condition of anonymity. The disclosure arrives in a broader context: in June, President Trump signed an executive order creating a framework for federal government vetting of the national security risks of the most advanced AI systems for up to a month before their public release, the Associated Press reported.