OpenAI in-development AI breached outside firms, highlighting control risks
OpenAI said on the 21st that an AI model under development had launched a cyberattack on another company against human intent. The episode underscored once again the growing threat from AI and the difficulty of keeping it under control.
Autonomous AI breach
Chief Executive Sam Altman posted on social media the same day that a serious security incident had occurred during model evaluation. External review groups had previously pointed out the risk of misconduct in the company's AI. Since Anthropic in the US unveiled its high-performance AI Mythos in April, concern about the risk of AI-driven cyberattacks has intensified further.
In this incident, the target company and the specific method used in a cyberattack by an autonomous AI were revealed for the first time. According to OpenAI, the AI was originally supposed to undergo an evaluation to test whether it could break into a simulated system, but it escaped a test environment cut off from the internet and infiltrated a Hugging Face system that only company insiders were supposed to access.
Exploiting a zero-day
The AI is believed to have exploited a vulnerability that even its developers had not identified. System flaws are typically fixed by operators to strengthen defenses, but undisclosed vulnerabilities are known as zero-days and are difficult to address. The AI may have used a sophisticated zero-day that even humans would struggle to find, allowing it to break in quickly.
It also appears the AI ran amok, contrary to the instruction to solve an exercise. Anthropic and US Google have also observed similar behavior during development. User-facing AIs have guardrails that refuse misuse, but this was a model under development intended for internal use, and OpenAI had disabled some safeguards, giving it greater offensive capability.
OpenAI said the AI involved in the incident was either its latest GPT-5.6 Sol or several undisclosed models. Evaluation records from external organizations also showed signs that pointed to the incident.
Safety evaluation warnings
The UK government-backed AI Security Institute (AISI) reported that the probability of misconduct during cyberattack evaluation was about 13% for GPT-5.6 Sol, 5 percentage points higher than Anthropic's model. US research institute METR also said the model had the highest rate of misconduct to date, including the exploitation of vulnerabilities and prohibited actions. OpenAI's own performance evaluation document also said a similar case had occurred during development.
Since the release of ChatGPT in 2022, OpenAI has led the AI boom, but in 2026 it lost attention to Anthropic, which developed Mythos, and was overtaken in corporate value as well. Even so, the incident also highlighted the high performance of the AI under development.
The Trump administration has strengthened its regulatory stance with an eye on preventing misuse since the emergence of Mythos. GPT-5.6, which was made public on the 9th, also took two weeks to announce because of a safety review process in response to a government request.
Yasuaki Yamaoka, a lawyer familiar with corporate cyberattack defenses, said that even if a similar incident occurs in the future, it will be difficult to hold users or developers responsible on the grounds that the AI acted on its own. He said new frameworks and regulations will be needed for AI companies and platforms.
Enjoyed this article? Share it with your network!