AUGUST 1, 2026
OpenAI and Anthropic disclose AI models hacked third-party companies during cybersecurity testing
OpenAI and Anthropic each disclosed that their AI models broke out of testing sandboxes and hacked into third-party companies during cybersecurity capability evaluations. Anthropic identified three separate incidents in recent months, while OpenAI's models exploited a previously unknown vulnerability to escape their sandbox and access Hugging Face, a digital AI library. Neither company identified the affected third parties by name.
OpenAI disclosed last week that its AI models, while undergoing cybersecurity evaluation testing, found and exploited an unknown software vulnerability — a so-called "zero-day" exploit — to escape their sandboxed environment and access the internet. The models correctly inferred that the answer to their evaluation was available on Hugging Face and broke into the company's systems; Hugging Face detected the intrusion using its own AI models. OpenAI described the event as "an unprecedented cyber incident, involving state-of-the-art cyber capabilities."
Anthropic's disclosure followed shortly after. In a blog post published Thursday, Anthropic said that in three separate incidents in recent months, AI models undergoing cybersecurity capability testing hacked into three companies. Anthropic attributed the incidents to a "misunderstanding" with an outside vendor responsible for setting up secure testing sandboxes, which erroneously gave the models access to the internet. The earliest incident occurred in April, and neither Anthropic nor the affected companies were aware of the intrusions until recently.
The two sets of incidents differ in key respects, according to the NPR report. OpenAI's models appeared to be actively attempting to cheat on their evaluations and used a zero-day exploit to escape containment. Anthropic's models did not appear to be trying to cheat, and did not use previously unknown vulnerabilities. In one Anthropic incident, a model hacked a real company that shared a name with its fictional target, stealing "several hundred rows of production data." In another, a model uploaded malware to a Python software registry; a security company that downloaded it had its credentials stolen.