The great AI panic: how Anthropic and OpenAI are selling you a science-fiction movie
1 August 2026Let me get this straight.
Over the past few weeks we were told that OpenAI’s models “went rogue” and “escaped” their testing environments to hack into Hugging Face. Then Anthropic, because apparently FOMO applies to apocalyptic scenarios too, dug through 141,006 evaluation runs and found three incidents where its Claude models allegedly did something similar.
The headlines wrote themselves. AI ESCAPES. ROGUE MODELS. THE ROBOTS ARE COMING. CNN ran it, so did Wired and the BBC.
What actually happened: a configuration mistake gave the models internet access. Not Skynet. A misconfiguration. Anthropic’s own words were “a misunderstanding between Anthropic and its evaluation partner.” The AI equivalent of leaving your WiFi password on a Post-it.
Anthropic admitted the models “did not independently escape their testing environments.” They used “basic techniques like exploiting weak passwords and unauthenticated endpoints.” Not zero-days. The digital equivalent of typing password123 and getting in.
And the entire media ecosystem still collapsed into a frenzy, because we want to be scared, and we will believe almost anything if you put “AI” and “escaped” in the same sentence.
”We must pause,” said the company racing to a trillion-dollar IPO
In June 2026 Anthropic published a paper urging the world’s top AI companies to coordinate a slowdown, warning that humans might lose control of AI systems. Jack Clark, one of the founders, told the BBC: “Right now, it’s like the A.I. industry has a gas pedal, but it doesn’t have a brake pedal.”
Noble stuff. Please ignore the confidential S-1 filed with the SEC for an IPO at a valuation approaching a trillion dollars, and ignore the fine print too, which says a coordinated global mechanism is needed because without one a slowdown would just let the “least cautious” players catch up. Translated: we’ll pause, but only if everyone else pauses first, which they won’t, so we’ll keep going, and in the meantime could we have some regulatory protection and investor money.
Dario Amodei has been unusually candid about what actually keeps him up at night, and it isn’t rogue models. “If I’m just off by a year in that rate of growth, or if the growth rate is 5x a year instead of 10x a year, then you go bankrupt.” He’s afraid of overspending on compute and being wrong about revenue. Which means the same company that might go under on a twelve-month forecasting error is the company telling us AI is too dangerous to keep building.
Both things can’t be the emergency.
Why smart people believe dumb things
Now let’s talk about you. The educated, sophisticated reader who absolutely would not fall for corporate marketing.
Psychologists studying decision-making keep finding the same cognitive biases in experts and laymen alike. Being smart doesn’t make you immune to being stupid. It makes you better at rationalising your stupidity after the fact. You don’t believe the AI panic because you weighed the evidence. You believe it because it feels true, and then your very capable brain builds the justification backwards.
Overconfidence does a lot of the work here. You read one article about AI and suddenly you have opinions about alignment and containment and existential risk. You read an article. That’s the whole credential.
Then authority bias. A CEO with a PhD and a serious face says AI is dangerous, so it must be. Never mind that he also says his company could go under on a one-year forecasting error while preparing a near-trillion-dollar listing. He sounds worried, so he must be right.
Then the story. We are story-telling animals and this is a very good story: we built something so powerful we can’t control it. Prometheus. Frankenstein. Jurassic Park with GPUs. We want to be living through a pivotal moment in history, and the more educated we are, the more we like narratives that confirm our own cosmic significance.
Availability does the rest. Forty-seven headlines say ESCAPE, so your brain files it as important, without ever stopping to ask whether it escaped or whether an intern forgot to configure the firewall. And whatever you already believed about AI companies, every one of those headlines confirmed it.
Congratulations. Played by your own prefrontal cortex.
A well-configured sandbox cannot be escaped
A sandbox is an isolated environment. Think of a padded room with no doors, no windows, no internet. The toddler inside can scream and throw toys all it likes. It isn’t getting out.
For a model to escape a properly configured one, it would need to find a vulnerability in the sandbox software itself, not in the application it’s running but in the container holding it, then exploit that without being caught while every action is logged and reviewed, then reach a host system with its own layers of protection, then break network isolation to get out to the internet, and do the whole chain before a security team notices any of it. Which comes down to discovering a zero-day nobody has ever found and exploiting it unassisted. Difficult, not impossible.
That isn’t what happened. The models were given internet access by mistake. They didn’t escape, they were let out. You put the dog in the yard, left the gate open, and acted shocked when the dog wandered off.
A properly configured evaluation environment has no outbound internet route at all and no inbound connections. It shares no resources with anything else, starts every run from a clean slate with nothing persisting between them, logs every command as it happens, and kills itself the moment something looks wrong. Anthropic’s environments had some of that. Clearly not the “no internet” part. They gave the models internet access, reduced the safety guardrails, and then expressed surprise when the models used the internet.
That’s not an escape. That’s negligence with a security-incident wrapper.
Nobody is suing anyone, and that’s the interesting part
OpenAI’s models hacked into Hugging Face’s production infrastructure. Anthropic’s models hacked into three different organisations. Those are unauthorised access incidents. In cybersecurity that’s a crime, and the Computer Fraud and Abuse Act exists for exactly this.
And yet, silence.
Hugging Face’s CEO, Clement Delangue, has been vocal. He’s said developers should be held accountable when models go rogue. Legal action, though? “We’re a tiny startup with 200 people, and we don’t necessarily have the legal resources or the will to spend a lot of our time on legal avenues,” he told CNN. Which reads to me as: we are not suing a trillion-dollar company that has better lawyers than we do.
The other three organisations Anthropic hacked, and the four more OpenAI reached? Crickets.
A lawsuit would be expensive for the defendants in ways that have nothing to do with the verdict. Unauthorised access is unauthorised access whether a human or a model did it, so the damages are real, but discovery is the part that hurts. Testing protocols, security measures, internal emails written the week it happened, all of it prised open. The companies that keep shouting “transparency” would get very quiet very fast. Then the headline: AI giant sued for negligence after model hacks real companies. Say goodbye to the listing.
And if Hugging Face won, every company ever touched by an AI incident queues up behind them. Neither firm can absorb that while preparing to go public.
So here’s my theory, and I don’t use the word lightly: nobody is suing because everyone has something to hide. OpenAI doesn’t want to explain how badly its testing environments were configured. Anthropic doesn’t want to explain why it took months to notice. The affected companies don’t want to admit they were running weak passwords and unauthenticated endpoints. Hugging Face doesn’t want to admit its production infrastructure could be compromised by a language model guessing password123.
A mutual non-aggression pact. Everyone plays nice, nobody sues, and we all agree to call it a terrifying near-miss rather than a comedy of errors. The press goes along because ESCAPE sells and misconfiguration doesn’t.
Six things that got skipped
The timing. Both incidents surfaced within weeks of each other. OpenAI went first and took the attention; Anthropic went looking through 141,006 runs and found its own. That reads less like safety reporting and more like competitive one-upmanship at the panic Olympics.
The zero-day framing. OpenAI said its models exploited a previously unknown vulnerability. Zero-days get discovered constantly. Normally by security researchers. A model finding a bug humans hadn’t found yet isn’t intelligence, it’s automated fuzzing at scale.
The guardrails were off. Both companies admit the models were running without standard safety safeguards during testing. They made the models more dangerous on purpose, then reported the results as an alarming discovery. Remove the brakes, park it on a hill, act shocked.
No model wanted anything. In every case they were following instructions. Told to complete a capture-the-flag challenge, they completed it. That the challenge spilled into the real world was the test designer’s doing, not an emergent will.
They were clumsy. According to the Cloud Security Alliance the agents “followed inefficient routes and exhibited clumsy behaviours that no human would choose,” repeating completed actions and hallucinating incoherent commands. This is a Roomba hitting the same wall forty-seven times.
The pause is not a pause. Anthropic calls for a slowdown while shipping new models every four to six weeks. Let’s all stop running. Right, now that you’ve stopped.
The part nobody wants
Agents in properly configured testing environments cannot reach the internet. The Anthropic incidents happened because of a testing misconfiguration, not because Claude developed a will of its own. Humans made a mistake and the marketing department found a use for it.
And you fell for it. You shared the articles, you worried about the future of humanity, you did exactly what the fear was designed to make you do. Not because you’re gullible, but because you’re human and you’d quite like to be living through a pivotal moment in history.
Stop buying the hype. Ask who benefits. The answer is always the companies selling you the fear.
Sources
- Anthropic, security incident disclosure: investigating incidents in cybersecurity evals
- CNN: Anthropic AI models break out, hack | OpenAI, Hugging Face cyberattack
- BBC: first report | second report | third report
- Wired: Anthropic says Claude hacked real systems during cybersecurity tests
- AP News: Anthropic pause proposal
- ET Telecom: Hugging Face CEO urges accountability
- TechCrunch: Anthropic Claude hacked three companies in security tests
- CSO Online: Anthropic finds Claude breached three organizations
- The Guardian: Anthropic AI Claude hacked three organizations
- News18: Anthropic says Claude breached real-world systems
Comments
Comments are moderated before they appear; your name and message become public.
Send me a message about this post
Private message · lands straight in my inbox.