Your AI Agent Runs Code. Here's What Actually Contains It.

AI agent sandboxing decides whether generated code can escape to your host, your secrets, or other tenants. Here's how containers, gVisor, and microVMs actually work-and where each breaks down.

Cover art for Your AI Agent Runs Code. Here's What Actually Contains It.

A developer on your team asks a coding agent to analyze a CSV and plot the results. The agent writes Python, runs it, and hands back a chart. That whole interaction takes about four seconds. What most people do not think about is the gap between "the agent runs Python" and "nothing bad happened." That gap is filled-or not-by a sandbox, and the engineering choices inside it are not obvious.

This post explains how AI agent sandboxing actually works, from the simplest container boundary to the hardware-enforced microVM layer that serious production systems use. No ops experience required.

Why a running agent is a problem even when it behaves

The root cause of agent execution risk is what security people call ambient authority: an agent process inherits all the permissions of its execution environment. It has database credentials, network access, and filesystem permissions-not because it needs all of them for every operation, but because that is how processes work in a traditional OS model.

This is what makes the "confused deputy" framing useful. An attacker does not need to hack anything. They just need to persuade the confused deputy to use the authority it already has. A prompt-injected instruction inside a retrieved document can do that without touching any traditional vulnerability.

Agent sandboxing is the practice of running an agent, its tools, or generated code inside a restricted execution environment. The sandbox limits which files, networks, credentials, processes, and computing resources the agent can access. If the agent makes an unsafe decision or processes malicious instructions, those boundaries reduce the potential impact on production systems and sensitive data.

The three layers that actually matter

Most discussions of sandboxing flatten what is really a layered stack into a single word: "container." Here is what the layers are.

Layer 1: Linux namespaces and cgroups (the container)

Linux namespaces and cgroups provide process-level isolation. A container shares the host kernel but has a restricted view of the filesystem, network, and process table. This is what Docker gives you. It is fast and cheap, but the key phrase is shares the host kernel. All containers on a host share the same kernel. Linux namespaces, cgroups, and seccomp filters create the illusion of isolation, but they are just software boundaries that can be broken.

This is not theoretical. CVE-2024-21626, "Leaky Vessels," is a patched runc vulnerability in which a leaked file descriptor referencing the host's working directory allowed a crafted container to escape to the host filesystem. It predates most agent tooling, but the containers those agents run in are built on the same primitives.

Layer 2: Syscall filtering (seccomp or gVisor)

The next layer is a syscall filter. Seccomp lets you define exactly which system calls a process is allowed to make and block everything else by default. An agent sandbox does not need to call ptrace, does not need to mount filesystems, does not need most of the roughly 400 syscalls Linux exposes.

The problem with seccomp is that it filters syscalls but runs in the same kernel the workload exploits. Seccomp also cannot inspect pointer arguments, limiting its visibility into memory contents of permitted syscalls.

gVisor takes a different approach. Instead of trusting the host kernel, it runs a user-space kernel called Sentry that intercepts every system call the application makes.

OpenAI's Code Interpreter uses gVisor for sandboxing. gVisor serves as a user-space kernel that implements a substantial portion of the Linux system call interface, providing strong isolation by intercepting and handling system calls made by applications, thus preventing direct interaction with the host kernel.

Layer 3: The microVM (its own kernel)

Teams that take this seriously usually add a third layer: a microVM. The key difference is that a microVM boots its own kernel rather than sharing the host's. A kernel exploit inside the VM can only damage that VM's kernel-the host kernel is unreachable by design.

Either approach gives each sandbox a dedicated kernel, so a crash or kernel exploit inside one sandbox cannot affect the host or another sandbox.

~150 msE2B Firecracker microVM cold startfrom snapshot restore
< 5 MiBmemory overhead per microVMversus ~131 MB for a traditional QEMU VM
50,000+concurrent sandbox sessionssupported by Modal's infrastructure

E2B provides a managed cloud sandbox API: the agent sends code over HTTP, E2B executes it in an isolated Firecracker microVM that boots in under 200 ms, and returns stdout/stderr. The speed comes from pre-warming: the key to fast startup is pre-warmed snapshots-a pool of VMs is started ahead of time, brought to a ready state, and a memory snapshot is taken. Incoming requests restore directly from the snapshot rather than booting a kernel from scratch, reducing cold-start time to around 150 ms.

Beagle in action#data-team, 11:02am
The ask
analyst asks Beagle to run a Python snippet against last month's revenue export
Beagle drafts
routes the execution request to the team's configured sandbox, drafts the output with a note that the code ran in an isolated environment, no host access
You approve
analyst approves; the chart posts inline, the sandbox tears down automatically
Do this in your workspace →

How the isolation levels actually compare

The choice between these layers is a trade-off between security guarantees, latency, and operational cost. Here is how they line up:

Isolation approach Kernel sharing Container escape risk Cold-start overhead Example
Plain Docker Shared host kernel Known CVE class Milliseconds Most DIY setups
Docker + seccomp/AppArmor Shared host kernel Reduced, not eliminated Milliseconds Default hardened containers
gVisor (user-space kernel) Host kernel mostly unreachable Structurally harder to land Low OpenAI Code Interpreter
Firecracker microVM Dedicated guest kernel Contained to guest ~150 ms E2B, AWS Lambda
Type-1 hypervisor (Edera) Dedicated guest kernel per agent Contained per zone Low High-security deployments

While gVisor provides improved isolation over containers through user-space kernel implementation, microVMs offer hardware-enforced boundaries that provide stronger security guarantees without the latency overhead added by gVisor, nor the compatibility issues that can come with it.

The part the container cannot see: what lives outside it

Here is the non-obvious part most sandboxing discussions skip. The LLM and the agent itself sit outside the sandbox. Only the code being executed runs inside it.

That means an agent that is legitimately credentialed-say, with a read/write token to your document store-can still cause damage without ever triggering a sandbox boundary. An agent with legitimate database access can still exfiltrate data through allowed API calls if compromised. Behavioral enforcement at the kernel level-controlling API access, network destinations, and process execution-addresses the threats isolation alone misses.

Claude Code takes a complementary approach by exposing graduated permission modes-from fully autonomous execution to mandatory user approval for every tool call-so that the same agent can operate at different trust levels depending on the task and the operator's risk tolerance.

This is a useful framing: the sandbox handles what generated code can do; the approval layer handles what the agent is allowed to decide on its own. They solve different problems. Conflating them is the design mistake that causes incidents.

Running agent-generated code in a shared service
Without Beagle
the agent's Python runs directly on the host, inheriting its database credentials and outbound network access; a poisoned document in the retrieval set is enough to exfiltrate data
With Beagle
code executes inside a Firecracker microVM with egress denied by default; even a successful prompt injection cannot reach the host filesystem or other tenants' sandboxes

The coding agent landscape reflects this spectrum clearly. Isolation choices range from no sandboxing at all (Gemini CLI, Aider, OpenCode run in a local shell) to Docker containers (SWE-agent, OpenHands), to platform sandboxing using Bubblewrap and Landlock on Linux (Codex CLI), to shadow git checkpoints that track state without writing to disk (Cline). The right choice depends on what the agent can access, not just what it writes.

AI agent sandboxing: common questions

What is AI agent sandboxing?

AI agent sandboxing is the practice of running agent-generated code in a restricted environment that limits access to the host filesystem, network, credentials, and other tenants. A proper sandbox has at least a container boundary, usually a syscall filter, and ideally a separate kernel via a microVM. The goal is to contain the blast radius when an agent makes a bad decision or gets manipulated.

Is Docker enough to sandbox an AI agent?

No. A Docker container shares the host kernel, and container escape vulnerabilities-like CVE-2024-21626-are a real, documented attack class. For untrusted code execution, especially in multi-tenant environments, you need at least a syscall filter (seccomp or gVisor) on top of the container, and ideally a microVM with its own dedicated kernel.

How fast do microVM sandboxes start?

With pre-warmed snapshots, Firecracker microVMs used by platforms like E2B start in roughly 150 ms end-to-end. That is slow compared to a plain container (milliseconds) but fast enough for most interactive agent workflows, and the isolation guarantees are meaningfully stronger.

What is the "confused deputy" problem in AI agents?

The confused deputy problem means an agent can cause harm not by being hacked, but by being persuaded to use the legitimate credentials it already holds. A prompt-injected instruction inside a retrieved document can direct an agent to call an API, delete a file, or exfiltrate data-using access the agent was legitimately granted. Sandboxing limits what generated code can do, but does not prevent the agent from making bad API calls with real credentials.

Do sandboxes protect against prompt injection?

Partially. A sandbox limits what happens after the agent is manipulated-it cannot prevent the manipulation itself. If an agent processes a malicious document that contains hidden instructions, the sandbox will not catch that. What it prevents is the injected instruction from escaping the execution environment, reaching host infrastructure, or touching other tenants' data.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle