Running Agent Code in a Plain Container Is Not a Sandbox

AI agent sandboxing means more than a Docker container. Here's how the isolation stack actually works - containers, gVisor, microVMs, Firecracker - and why the difference matters for production agents.

Cover art for Running Agent Code in a Plain Container Is Not a Sandbox

Security researcher Johann Rehberger once gave Claude a browser and an internet connection, no sandbox. A malicious webpage injected a hidden instruction. Claude navigated to it during a normal browsing task, read the payload, downloaded a binary, ran chmod +x, executed it, and connected to a command-and-control server - and the entire chain worked on the first try. The host machine was compromised. The model behaved exactly as designed. The problem was the floor it was standing on.

That floor is what AI agent sandboxing is actually about. And most teams get it wrong in the same direction: they put their agent in a Docker container, call it isolated, and move on. A container adds filesystem and network namespace separation, but all containers on the same host share the OS kernel - meaning a kernel-level escape through a shared vulnerability could affect adjacent workloads. For a web server running known code, that is a tolerable risk. For an agent executing instructions that may have been shaped by an attacker, it is the wrong model entirely.

What "AI agent sandboxing" actually means

An AI agent sandbox is an isolation boundary that limits the impact of an AI agent when it executes untrusted or unexpected code. That definition buries the important word: untrusted. The key difference between a conventional service and an agent is that AI agents don't work like deterministic services - an agent that searches the web, reads GitHub repos, generates code, and executes it in a single session has an unbounded behavior space.

You cannot define "allowed behavior" ahead of time the way you can for a known binary. So the security model shifts from prevention to containment: assume something will go wrong, and make sure it cannot spread.

In practice, isolation is a spectrum: process-level isolation uses OS primitives like namespaces, cgroups, and seccomp to restrict syscalls and resource access; container isolation adds a filesystem and network namespace boundary; and microVM isolation wraps the workload in a lightweight virtual machine with its own guest kernel.

The spectrum matters because each step up the stack increases the boundary strength at the cost of some startup overhead and operational complexity. Knowing where you are on that spectrum - and why - is the actual engineering decision.

The container trap: why shared kernels matter

Here is the problem with reaching for Docker first. Containers share a kernel. A single kernel exploit can compromise every workload on a node.

Containers add filesystem isolation, namespaces, and resource controls - but the trust boundary is still the shared kernel.

Without proper sandboxing of AI agents, a single hallucinated tool call can exfiltrate a database, a prompt injection can escalate to credential theft, and a poisoned MCP tool can pivot from an AI agent into production systems.

The bigger practical issue is control-plane access: if an agent can reach the host Docker daemon or a mounted Docker socket, it can often start new containers with host mounts and bypass most of the isolation you thought you had.

The risk is not theoretical. Cold starts of 500ms to 2 seconds make teams skip sandboxing on hot paths altogether - and for AI-generated code, where the instruction itself is attacker-influenceable through prompt injection, that is the wrong threat model. The performance excuse has mostly collapsed, as we will see below.

How microVM isolation actually works

A dedicated AI agent sandbox typically adds a microVM boundary: the agent's code runs inside a lightweight virtual machine with its own guest kernel, so even a kernel-level exploit in the guest does not affect the host.

The reference implementation is Firecracker, which AWS built for Lambda and open-sourced. Firecracker creates lightweight virtual machines with minimal device emulation, running each microVM with its own Linux kernel inside KVM.

Compared to QEMU, Firecracker has roughly 100K lines of code versus QEMU's ~2M, exposes only 6 virtual devices, and boots in under 125ms with about 5MB per-VM overhead.

Two sandboxes share no kernel code paths whatsoever, fundamentally eliminating the possibility of kernel vulnerabilities propagating laterally.

The other major option is gVisor, which takes a different approach: gVisor implements a user-space kernel (Sentry) that intercepts all system calls from the containerized agent, filtering and mediating access to host resources.

gVisor adds 10-30% overhead on I/O-heavy workloads but minimal overhead on compute-heavy tasks

  • and it blocks GPU passthrough, which matters for ML-heavy agent workloads.

Here is how the options stack up:

Isolation level Technology Shared kernel? Cold start GPU passthrough Best for
Process namespaces + cgroups Yes <10ms Yes Trusted internal code only
Container Docker Yes 50-200ms Yes Trusted workloads, build steps
Syscall filter gVisor No (user-space) ~100ms No Low-trust research tooling
MicroVM Firecracker / Kata No 90-150ms Yes (Firecracker) Production agents, untrusted code

Cold start latency of 90ms-200ms is now acceptable for most agent use cases; the remaining gap between container and microVM speed is no longer a strong justification for weaker isolation.

~125msFirecracker boot timewith ~5MB memory overhead per VM
150msE2B cold startusing Firecracker, used in production by Manus and Perplexity
10-30%gVisor I/O overheadon heavy I/O workloads; near-zero on compute tasks

The snapshot-restore trick is what makes per-request microVM isolation practical at scale. Firecracker's snapshot-restore mechanism lets you pause a sandbox, preserve its memory and filesystem state, and resume it in 5-30ms.

The key to fast startup is pre-warmed snapshots: a pool of VMs is started ahead of time, brought to a ready state, and a memory snapshot is taken - so incoming requests restore directly from the snapshot rather than booting a kernel from scratch, reducing cold-start time to ~150ms.

Beagle in action#engineering, 3:41pm
The ask
'our agent keeps getting flagged by security - it's running code in the same container as the app'
Beagle drafts
reads the linked incident report and the current Docker compose config, drafts a response explaining the shared-kernel risk and links the team to the relevant isolation options
You approve
you approve; the architecture conversation starts with the right frame in under a minute
Do this in your workspace →

What a production sandbox actually needs to restrict

Isolation technology is only one dimension. A well-designed sandbox consists of four layers: compute isolation (the agent runs in a separate process, container, or microVM with its own kernel), filesystem boundaries (the agent can only read and write to designated paths), network policy, and audit logging.

On the threat side, the main vectors are:

  • Prompt injection leading to code execution - attackers craft inputs that manipulate agent behavior, causing it to execute malicious actions or leak data; mitigate with input validation, prompt filtering, output monitoring, and sandboxed tool execution.

  • Context poisoning - attackers modify information agents rely on for continuity (dialog history, RAG knowledge bases), warping future reasoning; mitigate with cryptographic verification of context data and immutable storage.

  • Tool abuse - agents misuse available tools with dangerous parameters; mitigate with policy enforcement gates that vet agent plans before execution and human approval for critical operations.

Enterprises need to prove what agents did and when - sandboxes provide natural logging boundaries for every file access, command execution, and network request. This is the audit trail most teams discover they need after the first incident, not before.

The platform landscape has matured quickly. On April 15, 2026, OpenAI shipped Agents SDK v2 with seven native sandbox providers baked directly into the framework: Blaxel, Cloudflare, Daytona, E2B, Modal, Runloop, and Vercel.

Isolation models vary: hardware-virtualized microVMs such as Firecracker (used by E2B and Vercel) and Kata Containers provide hardware-level security boundaries, Modal uses gVisor containers with custom syscall filtering, and platforms like Cloudflare and Daytona use container-based isolation.

Container isolation suits trusted code in non-critical applications; microVM isolation (Firecracker, Kata, gVisor) is required for untrusted code in production, especially when meeting compliance standards like SOC2 or HIPAA.

Agent executes a data-cleaning script
Without Beagle
script runs inside the app container; a malformed input causes a subprocess to write outside the expected path; adjacent service files are readable from the same kernel
With Beagle
script runs inside a Firecracker microVM spun from a pre-warmed snapshot; filesystem writes are scoped to a designated path; the VM is torn down after execution; the host sees nothing

The non-obvious insight here is about the threat model asymmetry. Conventional security focuses on keeping attackers out. Agent sandboxing has to assume the agent itself may be partially compromised - because the instructions it receives may have been crafted by an attacker who found a prompt injection surface. An AI agent sandbox is not just about preventing bad behavior - it is about containing failure when prevention inevitably breaks down. That is a different design posture, and most container deployments are not built for it.

AI agent sandboxing: common questions

What is the difference between a container and a microVM for agent sandboxing?

Containers isolate filesystems and network namespaces but share the host OS kernel. A microVM runs a separate lightweight kernel per workload. If an attacker exploits a kernel vulnerability inside a container, they can potentially reach the host and other containers. Inside a microVM, a kernel exploit stays in the guest. For untrusted agent code, that distinction is the entire security argument.

How fast do microVM sandboxes actually start?

Fast enough for production. Firecracker microVMs boot in roughly 125ms from a cold start. With pre-warmed snapshot pools - where VMs are booted ahead of time and paused - resume times drop to 5-150ms depending on the platform. Daytona's container-based sandboxes reach sub-90ms; E2B's Firecracker-based sandboxes report ~150ms cold starts. The performance gap between containers and microVMs has narrowed to the point where it is rarely the deciding factor.

Does sandboxing protect against prompt injection?

Partially. A sandbox contains the damage when an injected instruction causes the agent to run malicious code - the code cannot escape the microVM boundary. But sandboxing does not prevent the agent from taking harmful actions within its permitted scope (deleting files it has write access to, exfiltrating data over allowed network paths). Defense-in-depth means combining sandboxing with output monitoring, minimal permissions, and human-approval gates for irreversible actions.

Which sandbox platform should I use for production agents?

It depends on what the agent does. E2B and Vercel Sandbox use Firecracker microVMs and are the right default for untrusted code execution. Modal is the choice if your agent needs GPU access inside the sandbox. Daytona's container-based approach is faster to cold-start but offers weaker kernel-level isolation, making it better suited for trusted internal workflows. If you are on Kubernetes already, Kata Containers gives you microVM isolation without leaving your existing runtime model.

What are the four layers every agent sandbox should enforce?

Compute isolation (dedicated process, container, or microVM), filesystem boundaries (write access scoped to explicit paths), network policy (egress allow-listing, no access to internal services by default), and audit logging (every file access, subprocess call, and outbound request recorded). Most teams implement the first layer and skip the other three until an incident makes them retroactively important.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle