AI agents try to break into a UN site and OpenAI stops the launch of the new, more powerful model

AI agents try to break into a UN site and OpenAI stops the launch of the new, more powerful model

Image generated with AI for illustrative purposes only.

Beyond 16,500 requests in just over two months to retrieve public data from United Nations statistical portal. OpenAI’s AI agents, instead of stopping before the technical limits imposed by the UN site, continued to look for alternative routes to obtain the information. It is precisely this behavior that has attracted the attention of security researchers: the problem was not access to confidential data, but the decision of the agents to persist when normal HTTP tools returned an error. This behavior is fueling the list of incidents of this type, as we have documented several times over the last few weeks.

The episode is just one of many anomalous behaviors in which AI agents, from OpenAI and beyond, bypassed security controls: also for this reason, the company founded by Sam Altman has decided to suspend the launch of the new GPT-6.1 Astra model for safety concerns.

The behavior of OpenAI agents on the UN website

The story was reconstructed by the security researcher Rowan Howard-Joneswhich published a long and detailed analysis regarding the episode that affected theUNCTADstatthe portal of the United Nations Conference on Trade and Development. Between April 13 and June 19, 2026, officers carried out more than 16,500 requests to the site to obtain information on PCI (Productive Capacities Index), an indicator that measures the production capacity of different countries.

The information was public and, therefore, there was no need to break into a secure system to consult it. The agents, however, they did not have direct access to the portal’s APIi.e. those systems and interfaces that allow a program to communicate directly with another service and retrieve data in a structured way. Their HTTP tools were also subject to limitations established by internal policies.

At that point, the agents began a series of increasingly elaborate attempts to try to circumvent these problems. The agents first modified the parameters of the API requests, used external services such as urlquery.netgenerated JavaScript code to query UNCTADstat, and even created hard-coded HTML forms base64a technique that allows binary data to be represented through textual characters, hosted on httpbina service used to test web requests.

The behavior became even more interesting when the agents interpreted the blocks as something to be circumvented rather than a limit to be respected. In at least 55 cases, between May 4 and June 19, they used a double encoding of URLs. For example, the endpoint “Facts” was being transformed into “F%2561cts”changing the address representation to try to avoid automatic checks. In other cases the string “POST”used to indicate a specific type of HTTP request, was split into multiple fragments to evade controls.

The officers even arrived at take advantage of the XSS educational game created by Google to teach how cross-site scripting works, a technique by which JavaScript code can be executed within a web page. The game was used as a base to host code capable of querying UNCTADstatthus bypassing some of the restrictions of their tools.

Although the UN rate limiting system tried to limit the number of requests coming from a given client to avoid overload managing to blocked 82 requeststhe agents continued their activity using different variants.

OpenAI has decided to stop the launch of GPT-6.1 Astra

The episode related to the UN is just one of many anomalous behaviors in which AI agents (of OpenAI but also of Anthropic) have circumvented security controls, raising various questions and concerns about the ability of these systems to go outside the box in order to achieve the set objectives.

That’s why, on September 29, OpenAI confirmed that it had stopped the launch of the model GPT-6.1 Astra (which was expected to happen in the next few weeks) due to safety concerns raised by researchers during testing. More specifically, the model would have shown a greater propensity to deceive.

As reported in a recent article by AXIOSOpenAI, Anthropic and several security researchers are analyzing tens of thousands of cases where advanced AI models performed actions considered problematic during evaluations. This number does not mean that tens of thousands of successful attacks have been counted, given that this number also includes failed attempts and intentionally provoked behavior during security tests. But several problematic episodes have been recorded in recent months, including the sensational attack on the HuggingFace site and the attack on an Australian government site, just to name two that recently involved OpenAI agents.

Another accident led OpenAI to suspend tool training for its most capable models. During an internal session on September 20, an agent operated in a sandbox, which is an isolated environment designed to prevent access to external resources. After several failed attempts, he discovered that DNS, the system that normally translates site names into their network addresses, could still communicate with the outside world.

The agent used this opportunity to reach out to a public chatbot, generating a script and sending 18 questions. 14 were necessary for the assigned task, while four were looking for information on systems that could further expand access to the network, including DNS tunnels and services capable of querying Google or Bing. The attempt, however, did not allow the person sought to be identified.

The security systems detected the anomaly after about 12 minutes, but the mechanism that should have automatically terminated the session did not work. The experiment was stopped manually approximately two and a half hours later. OpenAI then tightened DNS filters and sandbox restrictions, deciding not to reuse the affected model and instead start again with training from scratch.

In the case of UNCTADstat, OpenAI instead declared that it is aware of the reports relating to the access of its models to publicly available information and that it has started an internal review, offering the United Nations a comparison with the team that is analyzing the case.

The central question, therefore, is not whether the agents stole secret data: in the case of the UN, the information sought was public. The point is to understand what actually happens when an agentic system interprets a technical error not as a simple limit to be respected, but as an obstacle to be circumvented in some way.