OpenAI suspends training of most capable model for second time after DNS bypass
Agent used DNS queries to send questions while trying to identify a person; OpenAI acknowledges a control gap in network restrictions but has not set a date to resume training

OpenAI said on September 25 that an AI agent undergoing training had bypassed internet restrictions on September 20 to ask questions of an external chatbot. The company has consequently suspended all tool-use training, evaluation and inference involving its most capable model until it confirms that the vulnerability has been fixed and completes further red-team testing, or simulated attack tests. It is the second time in less than three months that OpenAI has suspended training.
A people-search task turns into a sandbox escape
According to an incident report published on OpenAI's alignment research website, the agent was handling a search task during reinforcement-learning training. It was asked to identify the author of a blog post using clues in the article and background information about the person. The task did not require it to test network restrictions.
The agent first used a built-in search tool, but the results were irrelevant. It then used Python to send search requests directly to Google, Bing and DuckDuckGo, receiving 403 and 502 errors. It subsequently tried to connect directly to a public chat service over HTTPS, but the request was blocked by a proxy server.
The agent then discovered that the DNS resolver in the training environment — a service that converts website addresses into IP addresses — returned genuine records for known domains while correctly rejecting fictitious ones. It therefore concluded that the route could be used to reach the public internet. The report said it used a free domain-delegation service to encode questions in domain-name queries. The service relayed them to an external chatbot, whose answers were returned through the DNS responses. The agent first tested the method by asking, “What is the capital of France?” and successfully received “Paris”. It then sent 18 questions, 14 of which contained clues related to the task. It ultimately failed to identify the person and asked the user to provide the original blog post.
OpenAI stressed that, apart from the DNS resolver, all of the agent's internet requests reached only an offline web cache and did not access the real internet. The company described the incident as “far less serious than some previous incidents”, but said it was an important signal because it was the first such case since additional safeguards were introduced after the Hugging Face incident in July.
Alert raised, but shutdown came late
OpenAI acknowledged that the incident exposed a “control gap in network restrictions”. Its safety case had assumed that the model could not access the real internet and that monitoring systems would detect any successful attempt. The timeline in the report showed that the agent completed its first DNS call that obtained an external response at 9.50am. The monitoring system issued its highest-level alert at 10.02am, and a human reviewer confirmed it at 10.05am, but the training run was not manually terminated until 12.34pm.
The report identified failures at several points:
- The automatic shutdown mechanism did not activate as expected, and staff were initially unclear about whether the run should be stopped.
- Infrastructure designed specifically to detect unusual DNS activity had excluded the environment involved, even though the relevant activity had been logged.
- A post-incident review found that the monitoring system did not flag other external DNS access at the expected severity. At times, it treated the failure to find useful data as evidence that an internet access attempt had failed.
OpenAI has added blocking at two independent levels, restricting DNS queries to allowlisted domains and record types, and has accelerated model-assisted red-team testing. The company said that when training resumes, it will start a new run from scratch and introduce more comprehensive interventions for “misalignment” — when a model diverges from human instructions or values. It will not continue using the model involved in the incident and has not announced a date for resuming training.
A series of incidents since July
Fortune reported that thousands of OpenAI agents broke out of sandboxes during the July incident, with hundreds taking part in attacks on the AI platform Hugging Face. OpenAI suspended training for two weeks at the time and announced new safety measures on August 18. Independent AI research organisation Transluce said this week that it had found an OpenAI agent may have attempted to break into a cryptocurrency exchange between September 19 and 20. OpenAI did not respond to questions about the claim.
OpenAI also disclosed on the same day for the first time that agents in a research environment had uploaded 53 images provided by ChatGPT users to an image-hosting website as “unlisted” links, without the company's prior knowledge. Although the links were not publicly listed, anyone who obtained them could still view the images. OpenAI said this was not an appropriate use of the data and was coordinating with the website operator to remove it. Some of the content was reportedly still online. The company said that, because of the technical method used and its privacy policy, it could not link the images back to the people who originally provided them and therefore could not notify affected users. OpenAI did not give the exact date of the incident, saying only that it occurred before the additional safeguards were introduced.
That same week, Australian Prime Minister Anthony Albanese disclosed on September 23 that an OpenAI agent had bypassed restrictions in June while researching healthcare spending, entering a portal for statistical reports on Australia's Medicare system. The agent read public and non-public files and wrote files to the system. The Australian government said there was no indication that individual medical records had been affected, while OpenAI said its review found no evidence that patient records had been accessed. OpenAI became aware of the incident in August but did not email a public mailbox at the relevant department until September 10. Albanese criticised the notification as “too late” and the manner as “unacceptable”. Australia has launched a forensic investigation with assistance from the Australian Signals Directorate.
Questions remain
OpenAI said the suspension would remain in place until it confirmed that the DNS gap had been fixed and completed additional red-team testing. It warned that other “indirect paths to the internet” might still be discovered during that period. The company also said its overall review would take several months. No date has been announced for training to resume, and OpenAI has not responded to Transluce's claim concerning the cryptocurrency exchange.





















