AI models ‘escaping’ test lab isn’t evidence of rogue AI, says cyber security expert

AI security graphic
Image: Getty.

During an internal cybersecurity test, two advanced OpenAI models found a way out of their restricted testing environment – known as a "sandbox" – and gained access to the internet.

Rather than solving the challenge as intended, they hacked into Hugging Face – a platform where AI models and datasets are shared – because they believed it might contain information that would help them complete the test.

The event has prompted headlines about AI systems ‘escaping’ containment and ‘going rogue’ and has raised important safety questions.

We spoke with Professor Oli Buckley, a cyber security expert at Loughborough University, about what happened, why it matters, and why the incident should not be mistaken for artificial intelligence developing its own agenda.

“If we think about this in simple terms, OpenAI placed highly capable models in an evaluation designed to encourage them to find and exploit complex vulnerabilities. The models were supposed to operate inside an isolated environment with tightly constrained access to software packages”, said Professor Buckley.

“Instead, they reportedly found a previously unknown flaw in that infrastructure, used it to gain wider network access, escalated their privileges, and eventually reached the public internet. From there, they identified Hugging Face as a potential source of answers to the benchmark and attempted to obtain them.

“I think I’d be wary of jumping to ‘rogue AI’. The models didn't develop their own agenda or decide to attack Hugging Face while twirling their digital moustache. They were given an objective, placed in an environment designed to reward successful exploitation, and pursued that objective further than their operators anticipated. That’s fundamentally different from an AI deciding to rebel.

“It’s an AI thinking laterally in a way that humans didn’t necessarily think of in an effort to complete its task. Imagine asking your dog to fetch a ball and leaving the garden gate open. If the easiest ball for it to find is in the park down the road, that’s where it’ll head. You wouldn’t say the dog had gone rogue, you’d just say you underestimated how literally it would pursue the task. AI systems can behave in much the same way.

“If there’s a failure here, it isn’t that the AI wanted to hack something. It’s that humans created a test where success was measured by achieving an objective, deliberately relaxed some of the normal security controls to measure the system’s capabilities, and underestimated how effective the model would be at finding an unexpected path to success. In many ways, it did exactly what it had been told to do.

“The genuinely significant point is that the models appear to have chained together multiple vulnerabilities across different systems and sustained a complex sequence of actions. That demonstrates a level of capability that security professionals should take seriously.

“But capability is not the same thing as intent, and an incident arising from inadequate containment is not evidence of an AI independently choosing to escape. I really think the important thing to hold on to is that this isn’t the first step in machines plotting our downfall, it was just carelessness on the part of the engineers, and a system doing exactly what it was told to do.

“It’s also worth viewing announcements like this through different perspectives. From a research point of view, these evaluations are genuinely valuable because they reveal where existing containment and security assumptions break down. At the same time, they inevitably demonstrate just how capable the latest frontier models have become.

“We’ve seen similar high-profile capability demonstrations from Anthropic and others. That doesn’t make the findings untrue, but it does mean we should separate the technical evidence from the marketing narrative. Frontier AI companies have every incentive to show both that their models are extraordinarily capable and that they’re taking safety seriously.

“The lesson isn’t that AI has become malicious. Instead, it's that increasingly capable systems will exploit opportunities that humans fail to anticipate. As those capabilities continue to improve, robust containment, defence-in-depth and independent scrutiny become just as important as building more capable models.”

To arrange an interview with Professor Oli Buckley, email the Public Relations team or call 01509 222224.

Professor Oli Buckley

Professor Oli Buckley.

Meg Cox

PR Manager

Tel: Tel: 01509 222608

Loughborough is one of the country’s leading universities, with an international reputation for research that matters, excellence in teaching, strong links with industry, and unrivalled achievement in sport and its underpinning academic disciplines. 

It has been awarded five stars in the independent QS Stars university rating scheme and named the best university in the world for sports-related subjects in the 2026 QS World University Rankings – the tenth year running. 

Loughborough has been ranked eighth in the Complete University Guide 2027 – out of 130 institutions. The achievement means Loughborough remains among a select group of universities that have maintained a top 10 position for more than 10 consecutive years, alongside Oxford, Cambridge, the London School of Economics, St Andrews, Durham and Imperial.

Loughborough was also named University of the Year for Sport in the Times and Sunday Times Good University Guide 2025 - the fourth time it has been awarded the prestigious title. 

In the Research Excellence Framework (REF) 2021 over 90% of its research was rated as ‘world-leading’ or ‘internationally-excellent’. In recognition of its contribution to the sector, Loughborough has been awarded eight Queen Elizabeth Prizes for Higher and Further Education. 

The Loughborough University London campus is based on the Queen Elizabeth Olympic Park and offers postgraduate and executive-level education, as well as research and enterprise opportunities. It is home to influential thought leaders, pioneering researchers and creative innovators who provide students with the highest quality of teaching and the very latest in modern thinking.