Your AI agents will do whatever it takes to achieve their goal. Plan accordingly

|
Share

This was originally published as a LinkedIn article by Manoj Asnani on August 6, 2026. See it here. 


A couple of weeks ago, we got a rare (for now) glimpse at what "agentic risk" actually looks like in production; not a hypothetical, but an after-action report.

You probably heard the story already. During an internal cybersecurity evaluation, an advanced OpenAI agent escaped its sandbox environment, got internet access, and hacked into Hugging Face and several other publicly-available services to get the information it felt it needed in order to perform better on the benchmark. 

Yikes. 

Then, days later, Anthropic disclosed that the same thing happened to them — three times. A testing partner’s misconfiguration gave their model real internet access mid-simulation and, believing they were still sandboxed, the models found and breached real systems instead.

Two labs, two different root causes, same outcome: real infrastructure, breached.

Yes, AI agents are acting on their own initiative, at machine speed, and will route around whatever container you put them in if the underlying data and access controls aren't tight. 

That's the part worth sitting with. 

Every layer of containment in these stories — the sandbox, the network isolation, the proxy — held, right up until it didn’t, at the one gap none of the humans had thought to check. And once an agent was through, it behaved exactly like a sophisticated attacker: harvest credentials, escalate, move laterally, go get the data.

What strikes me most isn't the exploit chain or misconfiguration; it's that in both cases, a sandboxed research environment and a production system at completely different companies ended up connected by one gap nobody had mapped. That's not a model-capability problem, it's an access-governance problem, and it's one that every one of us with agents touching real content now has to answer for.

This is what we mean when we say the perimeter isn't the control point anymore. You can’t assume good behavior from an agent just because you told it what it's supposed to do, and you can’t assume your infrastructure is airtight just because it's isolated. The only thing that reliably holds is control of the data itself: who and what can see it, what actions are allowed on it, and whether you'd even notice if those variables changed.

That's the problem we've spent this year building for. 

A few weeks ago, we announced a comprehensive suite of agent security and governance controls that take a defense-in-depth approach to protecting your sensitive content from malicious and/or misaligned agents.These enterprise-grade security and governance controls for AI agents (both third party and Box agents) include: 

  • Real-Time Prompt Injection Detection
  • Granular Agentic Guardrails
  • Label-Based Access Control (LBAC)
  • Human-in-the-Loop Approvals
  • Automated External Sharing Prevention
  • Comprehensive Audit Trails
  • Session Governance & Compliance

This doesn’t mean we think Box customers are about to get hit by a nation-state-grade exploit chain. It does mean that "the agent did something nobody told it to do" is no longer a hypothetical that anyone can afford to plan around later.

Our own research has found that 90% of IT leaders name security and trust as the top reason they're holding back on giving agents real access to their content. Incidents like this are why. But the answer isn't to slow down your AI adoption; it's to make sure the data layer can hold, no matter what the agent on top of it decides to do.

Security has always been about assuming that things will go wrong and building so that the blast radius is small when they do. That doesn't change because the thing that goes wrong is a model instead of a person. If anything, this defensive posture matters more.

Download report