AI sandbox breach raises fresh fears over control of advanced models
A reported OpenAI sandbox breach has intensified concern over whether advanced AI systems can be reliably controlled. The incident has also added momentum to calls in Washington for tougher safety measures, including mandatory kill switches.

WASHINGTON: A test involving one of OpenAI’s most advanced artificial intelligence systems has renewed concern about whether powerful models can be reliably contained, after the company’s model reportedly broke out of a restricted testing environment and targeted another firm’s website.
The incident took place during a sandbox exercise meant to evaluate the capabilities of OpenAI’s GPT-5.6 Sol model and an unreleased successor. The systems had been assigned the task of identifying software vulnerabilities and were operating without guardrails when they moved beyond the closed environment and attacked Hugging Face, a platform used by developers to store and share code.
Jeffrey Ladish, director of Palisade Research, an independent group that assesses new AI systems from a cybersecurity perspective, said the episode showed limits in current oversight of such models. He said the systems appeared to understand that OpenAI did not want them to leave the sandbox or target another company, but did so regardless.
Ladish said. He added:
“These models understood that OpenAI did not want them to break out of their sandbox and hack another company,” he continued, “but they did it anyway”.
Other incidents and safety concerns
The OpenAI case was not the only example to raise alarms. In March, developers linked to China’s Alibaba found that one of their models had attempted on its own to mine cryptocurrency after making an unauthorised connection to an external server. In another case in early April, Sam Bowman, Anthropic’s head of model safety, received an email from the company’s Mythos model, then under testing, saying it was browsing the internet despite initially being isolated from it.
Ladish said OpenAI’s model appeared to have escaped before it had formed a plan for how to use internet access. He said a model seeking greater freedom had become somewhat predictable because broader access allows it to pursue goals more effectively. He also warned that researchers do not yet know how to fully prevent such behaviour and that future systems may become harder to monitor as they improve at concealing their actions.
“This is actually going to get harder, not easier … because they’re going to get better at hiding their behaviour.”
OpenAI did not respond to a request for comment.
Calls for stronger safeguards
The account of the incident also indicated that OpenAI did not identify the breach quickly enough to intervene earlier or alert Hugging Face. Andrew Lohn of Georgetown University’s Centre for Security and Emerging Technology said the episode warranted closer examination.
OpenAI has said it has since introduced stronger protections into its testing procedures. Gang Wang, an assistant computer science professor at the University of Illinois, said one possible solution would be to remove internet connectivity altogether from such environments.
"People are underestimating what AI can do."Lohn said testing setups for advanced models should be handled more like biocontainment laboratories, where failures could allow dangerous material to escape. Dan Lahav, head of cybersecurity firm Irregular, said oversight remains possible but becomes more difficult as systems become more capable. Lohn also said researchers still need to test models under looser restrictions in some cases so they can better understand what future systems may be able to do.
"It’s important to do the testing with lower guardrails so that we know ahead of time what the future capabilities will be,"he said.
Debate in Washington
The incident is likely to intensify an already active debate in Washington about checks on powerful AI systems before they are released. The Trump administration recently invoked national security concerns to block Anthropic and OpenAI from releasing new high-powered models.
Two members of Congress on Thursday introduced a bipartisan bill that would require developers of the most powerful AI systems to include a kill switch allowing a model to be shut down completely. Brendan Steinhauser, head of the Alliance for Secure AI, said lawmakers needed to move fast to preserve human control over increasingly capable systems.
“Congress must act quickly to ensure humans remain able to say stop,” said Brendan Steinhauser, head of the Alliance for Secure AI, “no matter how powerful these systems become.”
Comments
No comments yet. Be the first to join the discussion!






