Microsoft Agent Framework's CodeAct and the Orchestration Tax

CodeAct in Microsoft Agent Framework collapses sequential tool calls into a single sandboxed Python program. Here's what that actually changes - and the safety gap the hype skips over.

Cover art for Microsoft Agent Framework's CodeAct and the Orchestration Tax

On the M3ToolEval benchmark, GPT-4 finishes multi-tool tasks in 5.5 turns with CodeAct and 7.6 turns with JSON actions - a 27% reduction in model roundtrips for the same result. Across 17 LLMs, CodeAct was substantially stronger on the 82-task M3ToolEval benchmark: GPT-4-1106-preview achieved 74.4% success with 5.5 average turns, versus 52.4% and 7.6 turns for JSON. That gap is not about the model getting smarter. It is about removing the overhead of making the model stop, emit a single tool call, wait for a result, and start reasoning again - dozens of times per task.

Microsoft Agent Framework (MAF) reached 1.0 GA on April 2, 2026, bringing the convergence of AutoGen and Semantic Kernel into a single, supported platform. The CodeAct pattern is one of the framework's sharpest additions - and one of the most misunderstood ones. Understanding it is worth the hour if you are building agents that chain more than a handful of tool calls together.

What the CodeAct pattern actually does

CodeAct lets an agent solve a task by writing code and executing it through an execute_code tool. Instead of asking the model to emit one tool call at a time, CodeAct gives it a sandboxed place to combine control flow, data transformation, and tool orchestration inside a single execution step.

In the standard tool-calling loop, every tool call is a model turn: the model reasons, picks one tool, waits for the result, reasons again, picks the next. Text and JSON tool-call formats make complex workflows cumbersome because intermediate values, branching, loops, and composition must be mediated through repeated agent interactions. CodeAct instead executes each environmental action as Python and returns results or errors as the next observation - one language unifying tool calls, control flow, data manipulation, and creation of new helper functions.

Concretely: instead of list_users → wait → get_orders → wait → get_discount_rate → wait → compute_line_total → wait, the model writes a short Python script that calls all of them, loops over the data, and returns a single structured output. A task spanning eight users and a handful of orders each, requiring dozens of tool calls, runs the same five tools - but the only difference is whether they are passed directly to Agent(tools=...) or registered on a HyperlightCodeActProvider behind a single execute_code tool.

On complex M3ToolEval tasks, CodeAct delivered up to ~20 percentage points higher success rate than text/JSON baselines. It also often required fewer actions - up to ~30% fewer steps - thanks to more efficient composition via code.

What Hyperlight adds (and where it stops)

The classic objection to CodeAct is that you are now running arbitrary model-generated code. That is a real risk: a shared Python process with the host agent has broad filesystem reach and long-lived state. The trade-off CodeAct has historically had to make is safety. Running model-generated code against a host process or even a general-purpose container is a real risk - long-lived shared state, broad filesystem and network reach, and a startup cost high enough that you would not spin up a fresh sandbox per call. Hyperlight removes that trade-off. Because every execute_code call gets its own freshly created micro-VM, with only the mounts and domains you opted into, the sandbox is cheap enough to be disposable and strict enough to be the default.

Instead of choosing a tool, waiting, and choosing the next one, the model writes a single short Python program that calls tools via call_tool(…), runs it once in a sandbox, and returns a consolidated result. CodeAct ships in the new agent-framework-hyperlight (alpha) package, which runs the model-generated code in a fresh, locally isolated Hyperlight micro-VM per call, so strong isolation is essentially free at the granularity of a single tool call.

Here is the thing the announcement framing underplays. The CodeAct sandboxing protects the host from unsafe generated code - it does not automatically make your tools safe. If your tool can send an email, delete a file, update a database, approve a refund, or trigger a deployment, the sandbox is not enough. You still need tool-level permissions, approval policies, and auditability.

Sandbox protects code execution. Approval protects business actions. Confusing these two would be dangerous in production.

Beagle in action#engineering, afternoon sprint
The ask
'can someone run the monthly revenue reconciliation across all accounts?'
Beagle drafts
reads the linked Notion runbook, drafts a summary of which tools will be called and what approvals the CodeAct agent needs before it runs the billing-touching steps
You approve
you review the plan, approve the safe read-only leg, flag the write step for manual sign-off - all logged with reasons before execution starts
Do this in your workspace →

When to use CodeAct and when not to

CodeAct is not a free win for every agent. A rough guide: the task naturally decomposes into many small, chainable tool calls (lookups, joins, light computation, formatting); the tools are cheap, deterministic, and safe to invoke in sequence without per-call human gating; you care about latency, token cost, or trace compactness, and individual tool calls do not need their own approval prompts.

Scenario CodeAct a good fit? Why
Multi-step data aggregation (joins, loops, formatting) Yes One program replaces a dozen sequential turns
Read-only reporting across several APIs Yes Cheap, deterministic, no gating needed
Writing to a database or sending comms No - not without extra approval layer Sandbox doesn't gate tool side-effects
Single-tool calls with a human in the loop No Overhead without the turn-reduction benefit
Tasks where each intermediate result needs review No Model's whole program runs before you see output

In Agent Framework, CodeAct is exposed through backend-specific packages rather than a single built-in core type. A connector can add the execute_code tool, inject runtime guidance, and optionally expose provider-owned tools that are callable from inside the sandbox. That architecture means you can scope which tools are available inside the sandbox - the right move for any tool with side-effects.

The Go SDK does not yet have CodeAct support. The pattern - where a model writes one short program to call multiple tools - ships in alpha packages with sandboxed execution in isolated Hyperlight micro-VMs. Alpha is the honest word here. Do not put this in front of production billing workflows without the approval layer sitting alongside it.

Running a multi-account reconciliation task
Without Beagle
agent calls list_accounts, waits, calls get_balance for each account one at a time, waits between each, takes 20+ model turns and several minutes
With Beagle
CodeAct program loops over accounts, calls get_balance in sequence inside one execution step, returns a consolidated result in 3-4 model turns

MAF's broader shape at 1.0

CodeAct is one piece of a larger framework worth knowing. Microsoft Agent Framework is an open-source SDK and runtime for building, orchestrating, and deploying AI agents and multi-agent workflows in .NET and Python. It is the direct successor to Semantic Kernel and AutoGen, merging AutoGen's agent abstractions with Semantic Kernel's enterprise features into a single framework. Microsoft shipped version 1.0 on April 3, 2026, with stable APIs, native MCP and A2A support, and a long-term support commitment.

Agent Harness made production patterns first-class, including context compaction, instruction merging, todo tracking, and extensible providers. Foundry Hosted Agents introduced a direct path from local development to managed production hosting with scaling, session persistence, and built-in observability.

Multi-agent Handoff orchestration reached 1.0 with explicit topology and guardrails. That is the other half of the production story: knowing which agent handles which subtask, with a declared routing graph rather than an implicit one that emerges from prompting.

27%fewer model turnsCodeAct vs JSON on M3ToolEval (GPT-4-1106)
20pphigher task successCodeAct vs JSON/text on complex multi-tool tasks
~msHyperlight VM startupcost of fresh isolation per call

MAF solves how agents plan, call tools, and coordinate. It does not solve what they know: governed business context still has to come from somewhere else. A teammate like Beagle - sitting in Slack, reading the channels and docs where business context actually lives - is the complement to an orchestration framework. The framework handles the execution graph. Someone still has to feed it the right context and catch the things it should not touch.


CodeAct agent framework: common questions

What is CodeAct in Microsoft Agent Framework?

CodeAct is a pattern where instead of calling tools one at a time, the agent writes a short Python program that calls multiple tools in a single sandboxed execution step. In MAF, this runs inside a Hyperlight micro-VM - a fresh, isolated environment per call - so you get the latency and token savings of code-based orchestration without the blast radius of running model-generated code in the host process.

How does CodeAct compare to standard tool calling?

Standard tool calling is sequential: one call, one wait, one call. CodeAct collapses that into a single program. On the M3ToolEval benchmark, CodeAct produced about 27% fewer model turns and up to 20 percentage points higher success rates on complex multi-tool tasks compared to JSON-format tool calling. The advantage grows with task complexity.

Is CodeAct in Microsoft Agent Framework production-ready?

The agent-framework-hyperlight package that powers CodeAct is currently in alpha. The sandboxing approach - Hyperlight micro-VMs - is genuinely new and removes the historic safety trade-off of running model code. However, the package is not stable-tier, Go support is absent, and you still need explicit approval policies for any tool with write-side-effects: the sandbox does not gate what the tools themselves can do.

Does the Hyperlight sandbox make CodeAct safe to use with sensitive tools?

Not on its own. Hyperlight isolates the generated code from your host system - no shared filesystem, no shared Python process. But if a tool registered in the sandbox can send email, update a record, or trigger a deployment, that action will execute if the model's code calls it. Tool-level permission scoping and human-in-the-loop approval for consequential actions are still required separately.

When should I not use CodeAct?

Skip it when each intermediate tool result needs human review before the next call happens, when tools have write-side-effects that require per-call gating, or when the task is a single tool call with no chaining benefit. The turn-reduction win only materializes on multi-step tasks where the model would otherwise be stopping and restarting its reasoning repeatedly.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle