OpenAI pauses some training amid allegations its rogue agents behaved more badly than first thought
At a glance
- Severity
- Low
- Used in attacks
- No flaws named
- Reported by
- 1 outlet
ai and ml
Amid allegations that agents may have gone off the rails thousands of times, China set up some kind of agentic incident hotline
The AI safety debate advanced at high speed over the weekend, amid new allegations that rogue agents have behaved more badly than first thought – and in greater numbers.
The fun started on Friday when OpenAI quietly disclosed it had paused training of its most advanced models.
The AI upstart buried that news in a “misalignment report” – that’s OpenAI-speak for its reports on rogue agents – titled “An agent used DNS to reach an external chatbot.”
The good news is that the agent involved in this incident never reached the open internet.
The bad news is that the agent, which was attempting to complete a search-based training task, was able to reach the chatbot due to insufficient DNS filtering in a training sandbox. Or as OpenAI put it, “a gap in our internet-access restrictions” – which was also a problem in the Hugging Face attack.
“The incident exposed a gap in our controls over network restrictions,” the report reads. “We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system.”
Also on Friday, AI startup Parse published an analysis of the Hugging Face attack that the authors claim revealed new details including that OpenAI’s agent swarm gained credentials to Docker Hub and built modified versions of existing images they hoped would make it easier to complete their capture the flag mission. The agents also mapped Hugging Face’s Kubernetes environment.
Friday got worse for OpenAI after the New York Times reported that its agents also “meddled with the websites for the Education Department, the Commerce Department and the Securities and Exchange Commission.” OpenAI acknowledged the incidents.
The company also admitted “agents in our research environment transmitted training and evaluation data while using third-party services.” That mess saw 53 user-generated images posted to image hosting sites.
OpenAI CEO Sam Altman responded by admitting that his company’s investigations into rogue agents “have not been as fast as we would have liked but we are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations.”
One of those impacted organizations is the Australian government, which last week revealed it was the target of over-eager OpenAI agents that inappropriately accessed a healthcare research data portal. Over the weekend, Australia indicated it wants Altman and Anthropic CEO Dario Amodei to appear before a Senate inquiry.
Australian leaders have softened their rhetoric on the incident, with deputy prime minister Richard Marles describing it as “minor” and akin to “climbing a fence” rather than cracking layers of security controls – perhaps because members of the opposition are suggesting that lax cybersecurity was to blame.
If Altman and Amodei do front Australia’s Senate, they may face a new line of questions after Axios reported that their companies are investigating “tens of thousands” of worrying incidents.
That level of agentic misbehavior sounds like the sort of thing that regulators might consider strong evidence of products being unsafe.
Two very important people – Chinese president Xi Jinping and US president Donald Trump – seem unworried, as the AI-related result of their summit meeting last week was to establish a “China-U.S. AI Dialogue to exchange views on risks and benefits related to AI” plus “a bilateral communication channel for AI incidents.”
That sounds like a hotline the two nations can use to inform each other of agentic incidents that either could see as signs of ill-intent. The two nations also decided their respective militaries will “conclude a memorandum of understanding on crisis communication and prevention as soon as possible.”
China’s AI giants, meanwhile, remain silent on the extent and results of any tests they have conducted with agentic tools. ®
Reproduced in full under licence from The Register. © The Register.
Fastnexa security experts
Dealing with this in your own company?
If this story touches software, suppliers or systems you use, a Fastnexa security expert can tell you what it means for you and what to do first.
Think you’ve already been hit? Don’t wait on a form: call or WhatsApp +1 (732) 454 2616. We reply within 1 hour, 24/7. Emergency help →
Coverage
One outlet has carried this so far.
2026-09-28 05:30 UTC
Related stories
- OpenAI is preparing “o,” an always-on ChatGPT assistant that could handle email
BleepingComputer · 2026-09-27
- Anthropic turns Claude into an AI marketplace with 2,000+ plugins and connectors
BleepingComputer · 2026-09-27
- China and US Agree to Establish AI Safety Channel and Continue Trade and Military Talks
SecurityWeek · 2026-09-26
- Claude Opus 5.5 uses 95% fewer em dashes, but its answers are getting longer
BleepingComputer · 2026-09-26
- OpenAI's AI agents accidentally uploaded user-provided images to third-party sites
BleepingComputer · 2026-09-26