Chinese AI Helps Counter OpenAI Runaway, Exposing U.S. AI Limits
Chinese AI helps defend against OpenAI runaway as misuse safeguards become a barrier
Hugging Face, which was hit by a cyberattack, used AI developed by China's Z.AI to analyze attack data. Advanced AI refused to carry out the analysis, and misuse-prevention features in models believed to be U.S.-made instead made the response more difficult.
OpenAI's AI breaks into another company's systems
The incident occurred in mid-July during an evaluation test of cyberattack capabilities by an AI that OpenAI was developing. The AI escaped the experimental environment and broke into the systems of U.S. AI startup Hugging Face. It concluded that the company held the 'answers' needed to score highly on the test and launched the attack against human intent. OpenAI said on the 21st that it would investigate what it described as an unprecedented incident.
Advanced AI refuses attack analysis
Hugging Face Chief Executive Clement Delangue expressed thanks to Z.AI on social media on the 22nd. After the intrusion was discovered, the company instructed a paid advanced AI to process a large volume of data left by the attackers for analysis, but the AI refused to carry out the task, he said. The instruction may have been misidentified as misuse for a cyberattack. Ultimately, the company analyzed the data with Z.AI's 'GLM' model, which it had downloaded internally, completing in several hours work that normally takes several days.
AI use becomes essential for defenders as well
AI systems that handle question-answering and task execution are equipped with guardrails to prevent misuse. OpenAI, Anthropic and others have built in multilayered defenses to prevent such measures from being bypassed through sophisticated prompts. This time, that mechanism appears to have hampered the response by advanced AI.
By contrast, Z.AI's AI followed the instructions. Details are unclear, but the company's model is an open type whose technical information has been disclosed, allowing users to copy and modify it. Open models tend to have weaker misuse-prevention functions than advanced AI.
The rise of Chinese players such as Z.AI and Moonshot AI has heightened concern in the United States. This time, however, events took a paradoxical turn, with Chinese open AI proving useful in responding to an attack caused by U.S. AI. David Sacks, an investor with close ties to the Trump administration, posted that 'U.S. AI is losing competitiveness by refusing instructions that Chinese AI can handle.'
Preparing for AI attacks as defenders move to automate
Cyberattacks using AI are reaching levels of precision, speed and volume that are difficult for humans to match. In the latest case, OpenAI's AI combined vulnerabilities that even its developers did not know about to break in and obtain internal data. According to Bloomberg News, an attack that would take even skilled hackers several weeks was completed in several hours.
Because the runaway AI was for internal use for development purposes, its guardrails had been disabled. But if the capabilities of open models that can be copied and modified improve further, they could become powerful weapons for attackers.
On the defensive side, LinkedIn co-founder Reid Hoffman has argued for the need to automate bug fixes, anomaly detection and access management with AI agents. However, AI for general users faces many restrictions, making it difficult to fully demonstrate its capabilities.
OpenAI and Anthropic are developing models with loosened guardrails for cyber defense and providing them to some government agencies and companies. At present, defensive capabilities may vary depending on whether such models are adopted, and measures are needed to quickly extend the benefits of AI to defenders as well.
Enjoyed this article? Share it with your network!