OpenHands Self-Hosted Coding Agent: What the 1.x Rebuild Actually Changed

OpenHands 1.x rebuilt the coding agent from scratch on a new Software Agent SDK. Here's what changed, what the SWE-bench numbers actually mean, and how Terminal-Bench 4.0 recalibrates the comparison.

Cover art for OpenHands Self-Hosted Coding Agent: What the 1.x Rebuild Actually Changed

A developer on a privacy-sensitive team opens a terminal on a Tuesday morning, types uvx openhands, and points it at a local Devstral model. No API key leaving the building. No cloud session to trust. The agent browses the repo, runs tests, proposes a fix, and waits for a human to say yes. That workflow exists now, at production quality. The question is whether the numbers behind it hold up.

They mostly do - but not in the way the headline benchmark scores suggest.

What OpenHands 1.0 actually changed

OpenHands 1.0 is not a feature bump. It is a ground-up rebuild around a Software Agent SDK, documented in a paper the team published on arXiv. The old version was a monolith where agent logic, evaluation, and the web app all lived in one codebase.

The new version splits into four Python packages with sharp boundaries: openhands.sdk holds the core abstractions - Agent, Conversation, LLM, Tool, and the event system. That separation matters for teams who want to swap models or write custom agent loops without forking the whole project.

OpenHands runs the agent inside a sandboxed runtime that can browse the web, execute code, and edit files end to end. It rebuilt on a new Software Agent SDK for the 1.0 release in 2026. Under the hood it uses LiteLLM, so 100+ providers work without code changes, including local Ollama, llama.cpp, vLLM, and MLX.

That LiteLLM layer is the practical unlock. You get one agent harness and you can run it against a $0.00/token local model or a frontier API, identical configuration either way.

With the ConfirmRisky policy, the agent sits in a WAITING_FOR_CONFIRMATION state until a human says yes. That is the exact mechanism the agent-safety people have been asking for, and it is on by default in the stack.

The benchmark numbers that actually matter

With a frontier model as the backend, OpenHands scores about 68% on SWE-bench Verified - the benchmark of 500 real GitHub issues. For context, Devin 2.0 publicly reported around 45.8%. OpenHands paired with Devstral 24B, an open-weight model, scores roughly 46.8%, which already matches Devin's commercial number.

That last line is the one worth sitting with. A freely downloadable 24B model running through a free harness on your own hardware equals the benchmark number a funded commercial product used to justify its pricing.

For internal tooling, dependency upgrades, boilerplate features, and test coverage work, an open model through an open harness at a fraction of the cost is a legitimate production choice in September 2026.

68%SWE-bench VerifiedOpenHands + frontier model backend
46.8%SWE-bench VerifiedOpenHands + Devstral 24B (open-weight, self-hosted)
~45.8%Devin 2.0 public figurethe commercial product OpenHands now matches on this metric

Why Terminal-Bench 4.0 is the number to watch now

SWE-bench Verified is a useful floor measurement, but it is increasingly a crowded one. Terminal-Bench 2.1 is now saturated. Nine models score between 87% and 92% on it, so tbench.ai moved its agent leaderboard to Terminal-Bench 4.0, where the top entry is 58.2%.

Terminal-Bench 4.0 evaluates AI agents on real-world work in terminal and command-line containerized environments across 66 tasks, with emphasis on science-adjacent and frontier engineering problems. Relative to earlier Terminal-Bench releases, 4.0 increases timeouts and adaptively raises RAM and CPU on selected tasks to reduce harness and resource confounds.

On 4.0, the current agent-level leaderboard looks like this:

Agent + model Terminal-Bench 4.0 Cost per run
Codex + GPT-6 Astra 58.2% ~$7.12
Claude Code + Fable 5.1 57.9% ~$14.76
OpenHands + frontier model ~50-55% (estimate) varies
OpenHands + Devstral 24B (local) not yet submitted $0 API cost

On Terminal-Bench 4.0, Codex reached 57.9% at $7.12 per task; Claude Code with Fable 5.1 needed $14.76 for the same score. The cost-per-task column is often the one teams skip when reading leaderboard screenshots. It shouldn't be.

There is also a subtler issue. SWE-bench and its successors note that scaffold choice affects scores but treat it as a confound to control rather than a phenomenon to study. A recent arXiv paper studying the "scaffold effect" makes the point explicitly: the harness you run an agent with can swing scores more than swapping models. That means a Codex score on Terminal-Bench 4.0 using OpenAI's own harness is not directly comparable to OpenHands on the same tasks. They are measuring the system, not just the model.

Terminal-Bench refreshing to 4.0 is a defense against models training, tuning, or being engineered around stale tasks. The faster the model cycle, the shorter the half-life of an eval.

The practical takeaway: don't pick an agent harness based on a single aggregate score. Build a smaller eval harness that answers one operator question: did my agent loop get better for the kind of work I actually ship? Start with 30 to 100 tasks from your own repos or representative open-source projects.

Beagle in action#eng-ops, 10:42am
The ask
'does anyone know if OpenHands can handle our dependency upgrade tickets without a human touching each one?'
Beagle drafts
pulls the SWE-bench 46.8% figure, links the OpenHands docs on ConfirmRisky policy, drafts a reply with the relevant caveats
You approve
you approve; the answer posts in thread with a source link and a note that boilerplate upgrades are the strongest use case
Do this in your workspace →

Before you self-host: the real checklist

Self-hosting OpenHands is not complicated to start. Install needs Python 3.12 and uv: uvx openhands

  • or pull the Docker image. The harder questions are operational ones.

The points worth checking before you put this on a real codebase:

  • Model choice determines capability ceiling. Devstral 24B gets you to Devin 2.0 parity on SWE-bench. For harder architectural work, swap in a frontier model through the same harness - LiteLLM makes this a config change.
  • The sandbox is Docker-level, not VM-level. If your threat model requires stronger isolation, add a layer. OpenHands' Docker sandboxing is production-quality for most teams; it is not air-gapped compute.
  • Confirmation policy is on by default, but check your config. The ConfirmRisky gate can be disabled. Make sure whoever deploys this knows that disabling it means the agent acts without human approval on file-destructive operations.
  • Aider is effectively unmaintained. Aider's last commit to main was 2026-05-22 and its last PyPI release is 0.86.2 from February 12, while every other agent on this page shipped a release in September. If you were planning to self-host Aider instead, now is the time to migrate to OpenHands or Kilo Code.
  • Roo Code is archived. Roo Code is archived. The VS Code extension shipped its final release on May 15, 2026 and the repo is read-only. Migration paths lead to Kilo Code or Cline.
Triaging a dependency upgrade ticket
Without Beagle
engineer opens the ticket, reads the changelog, checks for breaking changes manually, writes the PR, waits for review
With Beagle
OpenHands opens the repo in its sandbox, runs the test suite against the new dep version, proposes the change, waits at the ConfirmRisky gate for a human nod before committing

One non-obvious cost: a team running OpenHands on open-weight models avoids API fees but pays in GPU time and latency. Devstral 24B at full context on a single A100 runs meaningfully slower than a hosted frontier model. For async background tasks - overnight dependency sweeps, test-coverage gap fills - that trade-off is good. For interactive pair-programming sessions where a developer is waiting, it often isn't.

OpenHands self-hosted coding agent: common questions

What is OpenHands and how does the self-hosted version work?

OpenHands is an open-source autonomous coding agent, MIT-licensed, that runs inside a Docker sandbox. The self-hosted version runs entirely on your own infrastructure - the agent loop, the model calls, and the file system access stay local. You point it at any LiteLLM-compatible model, including local ones running on Ollama or llama.cpp.

How does OpenHands compare to Claude Code and Codex on benchmarks?

On SWE-bench Verified, OpenHands with a frontier model scores around 68%, competitive with the commercial leaders. On Terminal-Bench 4.0, the commercial agent pairs (Codex + GPT-6 Astra at 58.2%, Claude Code + Fable 5.1 at 57.9%) currently lead. OpenHands' harness-level score on 4.0 is not yet formally published, making direct comparison difficult.

Can OpenHands run on a fully open-weight model without any external API?

Yes. Your model API bill disappears entirely if you run open-weight models like Devstral on your own GPUs. The open-weight path is no longer embarrassing. OpenHands with Devstral 24B at 46.8% on SWE-bench Verified matches what Devin 2.0, a funded commercial product, publicly reported.

What is Terminal-Bench 4.0 and why did it replace 2.1?

Terminal-Bench 2.1 became saturated, with nine models scoring between 87% and 92%. Terminal-Bench 4.0 is the new standard, covering 66 harder tasks where the current top agent entry sits at 58.2%. The new version raises task complexity and adjusts resource limits to reduce confounds from harness choice.

Is the ConfirmRisky approval gate actually useful in practice?

It is the single most important safety control in the default stack. The agent pauses before any file-destructive or high-risk operation and waits for explicit human approval. For teams worried about a coding agent autonomously deleting or rewriting production files, this gate is the answer - and it is on by default, not buried in settings.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle