How AI Agent Code Execution Sandboxes Actually Work

AI agents that run code need more than a container. Here's how Firecracker microVMs and gVisor actually isolate agent-generated code - the mechanics, the tradeoffs, and the real costs.

Cover art for How AI Agent Code Execution Sandboxes Actually Work

E2B ran 40,000 sandboxes a month in March 2024. By March 2025 that number was 15 million. That 375x jump in a year is the clearest signal of something real: the moment teams started shipping agents that actually execute code, they needed a fundamentally different isolation model - and fast.

This post is about that isolation layer. Not the argument for why you need it, but the mechanics of how it works: what Firecracker and gVisor actually do when an agent submits a Python file, where the security boundary sits in each case, and what the numbers look like when you're operating at scale.

Why a container is not a sandbox for agent code

Containers work well when all code inside the container is in the same trust domain as the host - as is often true when you're running your own services. The moment you accept code from outside your trust boundary (users, agents, plugins), treat "shared kernel" as a conscious risk decision, not a default.

A Docker container using the default runc runtime shares the host Linux kernel. That means a kernel exploit inside the container is a host exploit. Hardening tightens policy. It doesn't change the boundary. You can harden a container until it's nearly unusable, but the few syscalls you allow still run privileged code in the shared host kernel.

This is not a theoretical concern. The threat model has two layers: execution isolation - preventing malicious or buggy agent-generated code from escaping to the host system - and agent-layer threats - prompt injection and tool poisoning attacks that subvert what the agent does before any code ever runs. Effective sandboxing requires addressing both.

How Firecracker microVMs work

Firecracker is the strongest common isolation primitive in use today. MicroVMs provide hardware-virtualization isolation in a lightweight form factor. Each execution gets its own kernel, memory space, and virtual hardware, optimized for fast boot times rather than the full weight of a traditional VM.

The mechanics: Firecracker uses KVM to create lightweight VMs with minimal virtual devices. Each microVM boots a stripped-down Linux kernel with a minimal init process. The guest has full kernel isolation from the host; an exploit inside the microVM would need to break through KVM and hardware virtualization boundaries to reach the host.

The security guarantee is qualitatively different from gVisor: even if an attacker escapes the guest VM and exploits a bug in Firecracker itself, they land in a severely restricted environment with no host filesystem access, no network beyond the configured tap device, and no ability to invoke privileged syscalls.

Firecracker's published figures are a boot under 125 ms, under 5 MiB of memory overhead per microVM, and up to 150 microVM launches per second per host. The cold-start cost is real, but production deployments sidestep most of it through snapshot-restore: Firecracker's snapshot-restore mechanism lets you pause a sandbox, preserve its memory and filesystem state, and resume it in 5-30 ms. That's the technique that makes agent tool calls feel instant even with hardware-level isolation underneath.

Beagle in action#eng-agents, 11:02am
The ask
agent submits a 40-line Python script to reproduce a reported bug
Beagle drafts
notes the execution request, checks the sandbox policy, surfaces the result log in-thread with a source link to the run
You approve
the engineer approves; the reproduction output posts back without the agent ever touching the host filesystem
Do this in your workspace →

How gVisor works - and where the tradeoff bites

gVisor takes a different route. gVisor implements a user-space kernel that intercepts system calls from containerized applications. Instead of allowing direct kernel access, gVisor's "Sentry" component provides a compatibility layer that implements the Linux system call interface in user space.

When a process inside a gVisor sandbox calls open() or socket(), the call is intercepted by Sentry rather than the host kernel. Sentry internally simulates the syscall and manages its own virtual file system and network stack.

That's the appeal: no VM boot, no KVM dependency, and it slots into existing container tooling with minimal changes. The cost is runtime overhead. Firecracker provides near-native performance with minimal VM boundary overhead. gVisor's syscall interception adds latency, sometimes 10-30% slower on I/O-heavy workloads. For CPU-bound workloads, both perform reasonably well. For I/O-intensive applications, Firecracker maintains better throughput.

The compatibility gap matters too: gVisor has high compatibility for ordinary code, but some syscalls are unimplemented and syscall-heavy or exotic workloads can break. Firecracker has very high compatibility - it's a real Linux kernel, so if it runs on Linux it runs in the microVM.

Here's who uses each in production:

Isolation Boundary Cold start Runtime overhead Who uses it
Firecracker microVM Hardware (KVM) ~125-150 ms Near-native AWS Lambda, E2B, Vercel
gVisor Userspace kernel ~50-100 ms 10-30% on I/O Modal, Google Cloud Run
Plain container Shared host kernel ~10 ms None Internal, trusted code only

Cloudflare's Dynamic Workers, which entered open beta in April 2026, takes a different approach using V8 isolates rather than microVMs - trading Firecracker's hardware boundary for sub-millisecond startup (roughly 100x faster boot) and 10-100x lower memory per execution context. This makes it viable for high-frequency, short-lived agent tool calls where VM overhead is prohibitive.

What the costs look like at scale

A default 2 vCPU sandbox with 4 GiB of RAM costs about $0.166 per hour, or roughly $0.0028 per minute - meaning a thirty-second agent task is a fraction of a cent. That sounds cheap, and for a single agent it is. The math shifts at scale.

Every "run this test," "try this patch," and "reproduce this bug" from an agent burns time inside an isolated sandbox, and sandbox-hours pile up faster than LLM tokens do. A single always-on coding agent working an 8-hour day consumes roughly 175 sandbox-hours a month. A fleet of five crosses 850. At that point the managed-vs-self-hosted question gets serious: self-hosting the Firecracker layer cuts per-execution cost 60-80% once you clear roughly 500 sandbox-hours a month.

375xE2B sandbox volume growthMarch 2024 → March 2025
~150 msFirecracker cold startbefore snapshot-restore optimization
5-30 msFirecracker snapshot-restoreresume time in production
10-30%gVisor I/O overheadvs. near-native for Firecracker

One number that surprised me when digging into this: an agent working on a Python project across 10 turns has installed packages, written files, and accumulated intermediate outputs. Full sandbox re-initialization on every turn wastes 200-500 ms on environment setup. That's the real case for snapshot-restore - not just boot speed, but preserving stateful agent sessions across turns without paying the full initialization cost each time.

There is also a GPU wrinkle that most comparisons skip. The GPU requirement makes the choice of isolation layer non-trivial. gVisor's user-space kernel intercepts GPU calls at a point that blocks direct PCIe passthrough. Firecracker's hardware virtualization path supports VFIO device passthrough to the microVM, giving the sandbox real GPU access with near-native performance. If your agents need GPU for inference or data work inside the sandbox, Firecracker is currently the only realistic path.

Running an agent that executes code
Without Beagle
agent generates Python, it runs in a container sharing the host kernel - a kernel exploit in agent output is a host exploit
With Beagle
same code lands in a Firecracker microVM with its own kernel; a guest escape lands in a restricted VMM layer with no host filesystem access
Beagle in action#platform, 3:47pm
The ask
'did the agent's test-runner step actually finish? can't see the output'
Beagle drafts
pulls the sandbox execution log, formats the stdout/stderr, and drafts a thread reply with pass/fail counts and a link to the full run
You approve
you approve; the result posts in the channel where the question was asked, with the run ID attached for audit
Do this in your workspace →

AI agent code execution sandbox: common questions

What is an AI agent code execution sandbox?

An AI agent code execution sandbox is an isolated compute environment where agent-generated code runs without access to the host system or other tenants' data. The agent submits code via API, the sandbox executes it in isolation, and returns results. The isolation boundary - container, gVisor, or microVM - determines what an exploit inside the sandbox can reach.

Why can't you just use a Docker container to sandbox agent code?

Docker containers share the host Linux kernel. Any syscall an agent makes reaches real kernel code on the host. Hardening (seccomp, AppArmor, capability drops) tightens the surface but does not change the fundamental boundary. For code you wrote and trust, that is often fine. For code an LLM generated and nobody reviewed, a shared kernel is a risk you are consciously accepting.

What is the difference between Firecracker and gVisor for agent sandboxing?

Firecracker gives each execution a full guest Linux kernel, enforced by hardware virtualization (KVM). gVisor intercepts syscalls in userspace without a real VM. Firecracker is slower to cold-start (125 ms) but has near-native runtime performance. gVisor starts faster (50-100 ms) but adds 10-30% overhead on I/O-heavy workloads and has lower syscall compatibility.

How much does a Firecracker sandbox cost per agent run?

At E2B's published rates, a 2 vCPU / 4 GiB sandbox costs about $0.166 per hour, or $0.0028 per minute. A typical 30-second agent task costs well under a cent. The cost scales with concurrency: a fleet of five always-on coding agents clears roughly 850 sandbox-hours a month, at which point self-hosting the Firecracker layer can cut per-execution cost 60-80%.

Do AI agent sandboxes support GPU workloads?

gVisor does not support direct GPU passthrough because its userspace kernel intercepts calls in a way that blocks PCIe passthrough. Firecracker's hardware virtualization supports VFIO GPU passthrough, giving the sandbox near-native GPU access. Managed providers like E2B do not currently offer GPU sandboxes; teams needing GPU-enabled agent sandboxes run Firecracker on their own GPU clusters.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle