OpenAI’s AI has escaped from the test environment

22 Jul 2026 Tereza Poledníková

No description

OpenAI has announced that, during an internal test of cyber capabilities, two of its advanced models managed to escape the intended boundaries of the sandbox. Whilst searching for a solution to the task being tested, the models expended a significant amount of computing power attempting to gain access to the open internet. To do so, they exploited a previously unknown vulnerability in the test environment’s proxy software, thereby escaping the sandbox and gradually escalating their privileges until they reached a node with internet access. From there, they deduced that they might find answers to the test on Hugging Face’s servers. Hugging Face detected and halted the attack itself; the company largely detected and analysed it using its own AI and confirmed the unauthorised access.

Read the details:

The most important aspect of the whole affair is precisely that both companies publicised the incident promptly, objectively and with specific technical details, rather than concealing it. This is precisely the kind of transparency that critics of the sector have been calling for – and, at the same time, a reminder that understanding such events requires consistent attention to verified primary sources, rather than merely to their media interpretations. Even for humanities-oriented institutions, it is therefore appropriate to systematically monitor developments in the field of artificial intelligence, both in an international context and at home.


More articles

All articles

You are running an old browser version. We recommend updating your browser to its latest version.

More info