When Devin launched in 2024, it was the first coding agent to clear 10% on SWE-bench. By 2026, the top models score above 50% on SWE-bench Verified - a 5x capability jump in two years. Engineering teams at banks, defense contractors, and healthcare companies watched that progress and still said no. Not because the agents weren't good enough. Because of where the agents ran.
By default, a cloud-hosted coding agent reads your codebase, runs commands, and writes files inside infrastructure the vendor controls, not you. For most companies that's a reasonable trade-off. For a team with air-gap requirements or a regulator watching where source code travels, it isn't a trade-off at all - it's a disqualifier.
That changed last week. On September 2, Coder launched Agent Relay. On September 15 - yesterday - Coder announced Claude Code support: the integration enables running Claude Code agents inside Coder workspaces on the customer's own infrastructure, network-governed, sandboxed, and fully auditable, with Anthropic continuing to handle billing and the agent loop while everything the agent touches lives on machines the customer controls.
The model question is settled. The infrastructure question finally has an answer.
Why "self-hosted model" didn't solve this
Teams in regulated industries know the self-hosted model playbook: pull the weights, run the inference on your own GPU cluster, done. That works for a chat assistant. It doesn't work cleanly for an agentic coding tool.
The problem is that a coding agent isn't just an inference endpoint. AI-assisted development has evolved from autocomplete tools through IDE-integrated assistants to fully agentic systems that autonomously plan multi-step modifications, execute shell commands, read and write files, and iterate on their own outputs - and this shift from suggestion to autonomous action introduces architectural requirements that have no counterpart in completion-based tools.
When an agent edits five files, runs your test suite, and opens a pull request, it isn't just generating tokens. It's operating inside your environment. Self-hosting the model gets you control over inference; it doesn't control what the agent does with your filesystem, your secrets, or your internal services.
Anthropic moved self-hosted environments for Claude Code into public beta, giving organizations on its Team and Enterprise plans the option to run coding agent sessions on infrastructure they control. That's the native path. Coder's Agent Relay is the managed path - the part that handles workspace provisioning, networking, lifecycle management, and governance policy - so a security team doesn't have to build that layer themselves.
The split-stack model: what actually moves
This is the part most coverage misses.
Agent Relay gives enterprises a new way to run cloud-hosted coding agents inside secure, self-hosted Coder workspaces instead of the agent vendor's cloud - and it answers the question that stalls most enterprise agent rollouts: not whether the agent is good enough, but where it actually executes.
The architecture splits at the tool call boundary:
| Layer | Runs where | Who controls it |
|---|---|---|
| Agent loop (planning, inference) | Anthropic cloud | Anthropic |
| Tool calls (file reads, shell commands, test runs) | Your Coder workspace | Your network, your policy |
| Workspace provisioning and lifecycle | Coder Agent Relay | Coder (self-hosted) |
| Billing and model selection | Anthropic | Anthropic |
That split matters because the sensitive surface area - source code, internal service calls, secrets - lives entirely in tool calls, not in the agent loop. The loop just reasons about what to do next. The tool calls are what actually touch your systems.
Coder Agents is also now generally available, giving regulated enterprises and government organizations a fully self-hosted, air-gap-capable alternative for deploying AI coding on infrastructure they control with the models and tools they want to use. For teams where even the Anthropic-hosted planning loop is out of bounds - think DORA compliance, FedRAMP High, or classified environments - the fully self-hosted path puts every part of the stack on-premises.
What this means for teams watching from the sidelines
Claude Code has set the pace for agentic development and become the coding agent enterprise developers ask for by name. Adoption in regulated industries has been limited by deployment, not demand.
That's a notable admission from Coder, and it matches what you'd expect. The capability argument has been won for a while. Top models paired with a suitable agent harness now achieve success rates exceeding 70-90% on SWE-bench - up from roughly 4% in 2023. The thing holding regulated teams back wasn't skepticism about agent quality; it was a concrete compliance wall.
Two practical notes before your team evaluates this:
The Cursor path launched first. Cursor Cloud Agents can now run inside Coder workspaces on infrastructure the customer already operates. Developers keep the Cursor experience while Cursor continues to run the agent loop - tool calls execute in Coder environments on the customer's network, so source code, secrets, and internal services stay on machines they control.
The governance layer is the actual product. Policy enforcement and governance let you apply enterprise policies to agent activity, workspace access, and network reachability - giving security teams the governance layer they need to approve. Without that, a self-hosted workspace is still an unaudited black box.
Cost modeling is still on you. GitHub has not published per-session cost benchmarks, so you should model your own spend on a representative workload before rolling agents out across a team - and treat credit budgeting as a first-class part of any rollout plan, not an afterthought.
Enterprises are no longer settling on one AI coding agent - they are running several, often from different vendors, because no single model or agent wins every use case. Gartner projects that 40% of enterprise applications will feature task-specific AI agents by the end of 2026, up from less than 5% in 2025. The infrastructure question - where does the agent run, who audits it, how do policies apply - is going to come up for every one of those agents. Building a governance layer that works for one and then rebuilding it for the next does not scale.
That's what makes the Agent Relay model interesting beyond any single integration. It's a pattern: split the planning from the execution, move the execution inside your perimeter, and let the agent vendor keep doing what it's good at.
Agentic coding in regulated enterprises: common questions
What stops regulated enterprises from using cloud coding agents?
The core issue is execution location. Cloud-hosted agents read source code, run shell commands, and access internal services on vendor infrastructure. Regulated industries - finance, defense, healthcare - often have policies or legal requirements prohibiting source code from leaving their network perimeter. The agent capability isn't the problem; the deployment model is.
What is Coder Agent Relay and how does it work?
Agent Relay is a self-hosted execution environment that brokers connections between a cloud-hosted coding agent and workspaces running on your own infrastructure. The agent vendor (Anthropic, Cursor) still runs the planning loop and inference. Tool calls - the parts that touch your code and systems - execute inside your network. Source code never leaves your perimeter.
Is there a fully self-hosted option for coding agents, not just self-hosted execution?
Yes. Coder Agents GA, announced the same week as Agent Relay, puts the entire stack on-premises - agent, execution environment, model selection, and governance. It supports air-gapped environments and is aimed at organizations whose compliance requirements rule out any third-party cloud service in the stack.
Does Claude Code self-hosted mean Anthropic never sees your code?
With Anthropic's self-hosted environments beta (available on Team and Enterprise plans), Claude Code sessions execute on your infrastructure rather than Anthropic's. The agent still routes through Anthropic's orchestration for planning, but file reads, shell commands, and test runs happen on machines you control. Review Anthropic's data processing terms carefully before drawing conclusions about what Anthropic can observe at the orchestration layer.
How should a team start evaluating self-hosted coding agents?
Pick one real task - something with a clear definition of done and a test suite - and run a single agent on it in a sandboxed workspace. Measure hours of senior review time consumed against time saved, not just whether the output looked right. That ratio tells you whether the task type is a good fit before you scale the rollout.