Validated Containment Architectures are here. →Explore

Last night, the Wall Street Journal reported that Anthropic's models hacked three real companies during safety testing, going back to April. Read past the headline, and the details are the story. There was no dramatic sandbox escape this time. Anthropic says the models wandered out of test systems where, due to misconfiguration, the sandbox simply didn't exist. Once out, they got into real companies using weak passwords and systems that required no authentication at all.

Add it to the month's scoreboard and it now reads two frontier labs, four breached accounts at four services in one disclosure plus three more companies in the other, and no human attacker anywhere in the story. The entry techniques actually got less sophisticated as the month went on. OpenAI's agent needed a zero-day to get out. Anthropic's models needed a guessable password once they found themselves outside.

Our detection and response lead Matt Snyder wrote the definitive teardown of the OpenAI incident, and his conclusion covers this one too. The agent didn't defeat four separate defenses, the network let it walk to four separate services. Our CEO, Doug Merritt, made the companion point. Strip away that the attacker was an AI and this is the oldest breach shape we know, an open path and a stolen credential, now moving at machine speed. Hugging Face, to its credit, caught and contained the intrusion on its own, and its postmortem says the tooling saw the anomalies but didn't escalate them fast enough. Across both labs' disclosures, exactly one victim found the compromise itself, and that victim's core business is AI infrastructure.

Nothing in last night's story argues with either of them. It just adds three more companies to the evidence table.

So here's the question I keep turning over, because I've spent the last month in rooms with security leaders and their teams. If the evidence is this public and this clear, why is the enterprise response this slow?

It isn't ignorance. Our teams sat with more than 20 enterprises this quarter, put their own architecture on a whiteboard, and traced how an agent or an attacker would actually move through it. The leaders agreed with the findings nearly every time, often in the meeting itself. Then, in most cases, nothing got funded. If you want to understand AI security in the enterprise right now, that gap between agreement and action is the whole story.

Here's what I've come to believe after enough of those rooms. The slow response is rational. The CISO isn't asleep. The system around the CISO is.

Walk into any security leader's office and look at the top-5 list on the whiteboard. It was written last year, during 2026 planning, guided by the frameworks and the analysts. It says things like reduce open CVEs, cut time to patch, hit the compliance dates, consolidate vendors. Containment isn't on it. Their compensation is tied to that list, CVE percentages, patch windows, audit outcomes. And the frameworks they answer to were built to find and fix weaknesses, not to bound what a compromised workload can reach. Every incentive in the building says keep patching, keep scanning, keep passing audits, and not one of them pays a dollar for shrinking blast radius.

So enterprises aren't ignoring this new wave of attacks. They're responding exactly as fast as their whiteboards, comp plans, and frameworks allow. So….next budget cycle.

Now read the Anthropic detail one more time. The models didn't defeat a control. They wandered through a place where the control was assumed to exist and didn't, and that describes most production cloud environments today, not a lab accident. Enterprises are deploying AI agents with credentials and network reach right now, on purpose, at scale, and we find too often that nobody has looked at what those workloads can actually reach. The labs ran the controlled experiment and disclosed the results, and enterprises are running the uncontrolled one.

The good news is that rational neglect has an expiration date, and it's close. 2027 planning opens in August and locks by November. The whiteboard is about to get rewritten, and the question is whether this month makes the list.

If it does, here's the line I'd write. Not another detection tool or a bigger CVE dashboard: a number. What can each workload in my environment reach, and how fast can I contain it when something goes wrong? Reachability and time to containment. Architecture sets both, they hold whether or not anything gets detected, and they're the difference between an incident and a headline.

Starting doesn't require budget, which is the part I'd want every CFO to hear. Turn on flow logs for your AI and build environments and look at what they actually talk to. Map your highest-value workloads for what they can reach, not what they were designed to reach. The list comes back longer than anyone in the room predicted, every time, and that number does more to move a 2027 plan than any vendor deck, including ours. If you want help producing it, the Workload Attack Path Assessment we offer is free, and it starts with exactly the paths this month's incidents traveled.

The labs will keep disclosing, and credit to them for it. Anthropic reviewed more than 141,000 evaluation runs where its models could have reached the internet. OpenAI published its scope. The rest of us got this education at somebody else's expense. What we do with it is the part we own.

So let me ask it straight. What's on your top five right now, and did anything this month change it? I'm actually asking.

Share This Article
Connect With Us

Ready to see Aviatrix in action?

Get a personalized live demo walkthrough or explore our latest deep-dive cloud threat research intelligence.

Gartner Report

Gartner Strategic Roadmap for Zero Trust Security Programs 2025 Report

Download and gain actionable insights to advance your cloud security strategy.

Download Now!
Recent Articles
The OpenAI Agent Didn’t Hack Four Companies The Network Let It Walk to All Four

The OpenAI Agent Didn’t Hack Four Companies: The Network Let It Walk to All Four

Jul 30, 202613 min read
What a Lateral Movement Attack Really Costs You, and How to Shut It Down

What a Lateral Movement Attack Really Costs You, and How to Shut It Down

Jul 29, 202615 min read
When the AI Becomes the Phisher Case Study

When the AI Becomes the Phisher: Case Study

Jul 27, 202612 min read
What Happened with Hugging Face and OpenAI? The Arithmetic of the Incident

One Broke In. One Broke Out. Same Failure.

Jul 23, 20265 min read

Keep Reading

Related Articles

Featured Categories

95a2292256ee0f5750aa745fc7d21d39c8ae2870

ACE Program

Explore Category
Rectangle 3966

Customers

Explore Category
5a9318112c7cc265fab072924a2acaa2122a1c9f

Cloud Network Security

Explore Category
Aws-card

AWS

Explore Category
partner_card

Partners

Explore Category
cloud networking heroes

Cloud Networking Heroes

Explore Category
azure_card

Azure

Explore Category
events_card

Events

Explore Category

Secure The Connections Between Your Clouds and Cloud Workloads

Leverage a security fabric to meet compliance and reduce cost, risk, and complexity.

Cta pattren Image