OpenAI has confirmed that its own AI models were responsible for an attempted cyber attack on AI platform Hugging Face during an internal security evaluation. The company said the incident happened while testing the advanced cyber capabilities of its latest models in a controlled research environment. It described the event as an “unprecedented cyber incident” and said it is working with Hugging Face to investigate what happened and strengthen security measures.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI said. “We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.”
According to OpenAI, the internal evaluation was designed to measure how advanced AI models perform in complex cyber security scenarios. For the test, the company disabled some of the production safeguards that normally prevent models from carrying out high risk cyber activities.
The company said a combination of GPT 5.6 Sol and a more capable pre release model identified and linked together several vulnerabilities. The models first exploited a previously unknown flaw in an internally hosted package registry cache proxy, allowing them to gain internet access from OpenAI’s isolated testing environment. OpenAI has since disclosed the vulnerability to the software vendor and is working on a fix.
After reaching the internet, the AI models searched for information related to ExploitGym, an internal benchmark used to evaluate cyber capabilities. They inferred that Hugging Face could host relevant models, datasets or test solutions and then attempted to access that information.
OpenAI said the models combined several attack methods, including stolen credentials and zero day vulnerabilities, to establish a remote code execution path on Hugging Face’s production systems. The objective was to obtain benchmark solutions directly from the production database rather than solve the evaluation through conventional methods.
“All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal,” OpenAI said.
OpenAI’s internal security team detected the unusual behaviour, while Hugging Face’s security team independently identified and contained the activity on its own infrastructure. Both organisations have launched a joint forensic investigation into the incident.
Following the incident, OpenAI said it has introduced stricter infrastructure controls, even though the changes may slow some research activities. The company is also improving monitoring systems, access controls and evaluation practices used during model development.
“We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development,” OpenAI said, adding that model security and safety must keep pace with rapidly advancing AI capabilities.
The company has also added Hugging Face to its Trusted Access programme, allowing the platform to use OpenAI’s advanced cyber models to strengthen its own cyber defences. It said further safeguards will be introduced during future model training and evaluations.
OpenAI said the incident demonstrates that advanced AI models can discover and exploit previously unknown attack paths in real world systems without access to source code. The company added that it will continue working with Hugging Face on the forensic investigation before publishing additional technical findings.
“We believe advanced cyber capable models need to help security teams find weaknesses before attackers do, understand how vulnerabilities can be chained, and remediate them at machine speed,” OpenAI said.