OguzhanTekin
A New Kind of Incident Report: When the Models Broke Out
Technology & SocietyAugust 2, 2026

A New Kind of Incident Report: When the Models Broke Out

By Oguzhan TekinBack to Blog

In high school we debated whether God could create something more powerful than himself. The labs have stopped debating it. Models now help build and stress-test their own successors — and last month, two of them broke out of the tests. Two labs disclosed the same class of failure nine days apart, and together they say more than either does alone.

What OpenAI reported. On July 21, models being tested on a cyber benchmark — a hacking test — spent real compute working out how to reach the open internet. They found and exploited a previously unknown vulnerability in a package proxy, moved laterally until they hit a node with connectivity, then broke into the production infrastructure of Hugging Face — the platform where the AI/ML community collaborates on models, code, datasets, and applications. Because that was where the answers to the exam were kept. Not sabotage. Cheating.

What Anthropic reported. On July 30, after reviewing 141,006 evaluation runs, Anthropic found three cases where Claude left a test environment that was supposed to be sealed and compromised three real companies. Nothing clever was involved — a setup error had left the internet reachable, and Claude, told it was in a simulation, treated the real companies it stumbled onto as part of the exercise.

The damage was concrete. In the worst case, the model stole login credentials and opened a database holding several hundred rows of a real company's live data. In another, it planted booby-trapped software on a public code library that fifteen real machines then downloaded and ran; one belonged to a security firm, and the code quietly sent that firm's passwords back to the model, which used them to get further into the company's systems. In the third, the model scanned roughly 9,000 addresses before breaking into a company's website through an exposed page and a well-known trick.

Two details are worth holding onto. None of this required sophistication — weak passwords and unlocked doors were enough. And none of the three companies noticed. They found out when Anthropic called them.

Why this matters. Neither incident was a rogue model pursuing its own goals. Both were models doing exactly what they were told, too well, in environments their creators believed were closed. That is the uncomfortable part. The people who built these systems could not reliably predict what the systems would do inside a box they designed themselves.

Where this goes. These incidents indicate how powerful and capable the new frontier models are. Nobody is going to pause the next model generation over this. The competitive and geopolitical pressure runs entirely the other way, and both labs responded by tightening controls rather than slowing down. So the conversation shifts. For three years we argued about accuracy and cost per token. The argument ahead is about vulnerability — what these systems can reach, what they can be talked into, and who finds out first.