A developer on a privacy-sensitive team opens a terminal on a Tuesday morning, types uvx openhands, and points it at a local Devstral model. No API key leaving the building. No cloud session to trust. The agent browses the repo, runs tests, proposes a fix, and waits for a human to say yes. That workflow exists now, at production quality. The question is whether the numbers behind it hold up.
They mostly do - but not in the way the headline benchmark scores suggest.
What OpenHands 1.0 actually changed
OpenHands 1.0 is not a feature bump. It is a ground-up rebuild around a Software Agent SDK, documented in a paper the team published on arXiv. The old version was a monolith where agent logic, evaluation, and the web app all lived in one codebase.
The new version splits into four Python packages with sharp boundaries: openhands.sdk holds the core abstractions - Agent, Conversation, LLM, Tool, and the event system.
That separation matters for teams who want to swap models or write custom agent loops without forking the whole project.
OpenHands runs the agent inside a sandboxed runtime that can browse the web, execute code, and edit files end to end. It rebuilt on a new Software Agent SDK for the 1.0 release in 2026. Under the hood it uses LiteLLM, so 100+ providers work without code changes, including local Ollama, llama.cpp, vLLM, and MLX.
That LiteLLM layer is the practical unlock. You get one agent harness and you can run it against a $0.00/token local model or a frontier API, identical configuration either way.
With the ConfirmRisky policy, the agent sits in a WAITING_FOR_CONFIRMATION state until a human says yes. That is the exact mechanism the agent-safety people have been asking for, and it is on by default in the stack.
The benchmark numbers that actually matter
With a frontier model as the backend, OpenHands scores about 68% on SWE-bench Verified - the benchmark of 500 real GitHub issues. For context, Devin 2.0 publicly reported around 45.8%. OpenHands paired with Devstral 24B, an open-weight model, scores roughly 46.8%, which already matches Devin's commercial number.
That last line is the one worth sitting with. A freely downloadable 24B model running through a free harness on your own hardware equals the benchmark number a funded commercial product used to justify its pricing.
For internal tooling, dependency upgrades, boilerplate features, and test coverage work, an open model through an open harness at a fraction of the cost is a legitimate production choice in September 2026.
Why Terminal-Bench 4.0 is the number to watch now
SWE-bench Verified is a useful floor measurement, but it is increasingly a crowded one. Terminal-Bench 2.1 is now saturated. Nine models score between 87% and 92% on it, so tbench.ai moved its agent leaderboard to Terminal-Bench 4.0, where the top entry is 58.2%.
Terminal-Bench 4.0 evaluates AI agents on real-world work in terminal and command-line containerized environments across 66 tasks, with emphasis on science-adjacent and frontier engineering problems. Relative to earlier Terminal-Bench releases, 4.0 increases timeouts and adaptively raises RAM and CPU on selected tasks to reduce harness and resource confounds.
On 4.0, the current agent-level leaderboard looks like this:
| Agent + model | Terminal-Bench 4.0 | Cost per run |
|---|---|---|
| Codex + GPT-6 Astra | 58.2% | ~$7.12 |
| Claude Code + Fable 5.1 | 57.9% | ~$14.76 |
| OpenHands + frontier model | ~50-55% (estimate) | varies |
| OpenHands + Devstral 24B (local) | not yet submitted | $0 API cost |
On Terminal-Bench 4.0, Codex reached 57.9% at $7.12 per task; Claude Code with Fable 5.1 needed $14.76 for the same score. The cost-per-task column is often the one teams skip when reading leaderboard screenshots. It shouldn't be.
There is also a subtler issue. SWE-bench and its successors note that scaffold choice affects scores but treat it as a confound to control rather than a phenomenon to study. A recent arXiv paper studying the "scaffold effect" makes the point explicitly: the harness you run an agent with can swing scores more than swapping models. That means a Codex score on Terminal-Bench 4.0 using OpenAI's own harness is not directly comparable to OpenHands on the same tasks. They are measuring the system, not just the model.
Terminal-Bench refreshing to 4.0 is a defense against models training, tuning, or being engineered around stale tasks. The faster the model cycle, the shorter the half-life of an eval.
The practical takeaway: don't pick an agent harness based on a single aggregate score. Build a smaller eval harness that answers one operator question: did my agent loop get better for the kind of work I actually ship? Start with 30 to 100 tasks from your own repos or representative open-source projects.
Before you self-host: the real checklist
Self-hosting OpenHands is not complicated to start.
Install needs Python 3.12 and uv: uvx openhands
- or pull the Docker image. The harder questions are operational ones.
The points worth checking before you put this on a real codebase:
- Model choice determines capability ceiling. Devstral 24B gets you to Devin 2.0 parity on SWE-bench. For harder architectural work, swap in a frontier model through the same harness - LiteLLM makes this a config change.
- The sandbox is Docker-level, not VM-level. If your threat model requires stronger isolation, add a layer. OpenHands' Docker sandboxing is production-quality for most teams; it is not air-gapped compute.
- Confirmation policy is on by default, but check your config. The
ConfirmRiskygate can be disabled. Make sure whoever deploys this knows that disabling it means the agent acts without human approval on file-destructive operations. - Aider is effectively unmaintained. Aider's last commit to main was 2026-05-22 and its last PyPI release is 0.86.2 from February 12, while every other agent on this page shipped a release in September. If you were planning to self-host Aider instead, now is the time to migrate to OpenHands or Kilo Code.
- Roo Code is archived. Roo Code is archived. The VS Code extension shipped its final release on May 15, 2026 and the repo is read-only. Migration paths lead to Kilo Code or Cline.
One non-obvious cost: a team running OpenHands on open-weight models avoids API fees but pays in GPU time and latency. Devstral 24B at full context on a single A100 runs meaningfully slower than a hosted frontier model. For async background tasks - overnight dependency sweeps, test-coverage gap fills - that trade-off is good. For interactive pair-programming sessions where a developer is waiting, it often isn't.
OpenHands self-hosted coding agent: common questions
What is OpenHands and how does the self-hosted version work?
OpenHands is an open-source autonomous coding agent, MIT-licensed, that runs inside a Docker sandbox. The self-hosted version runs entirely on your own infrastructure - the agent loop, the model calls, and the file system access stay local. You point it at any LiteLLM-compatible model, including local ones running on Ollama or llama.cpp.
How does OpenHands compare to Claude Code and Codex on benchmarks?
On SWE-bench Verified, OpenHands with a frontier model scores around 68%, competitive with the commercial leaders. On Terminal-Bench 4.0, the commercial agent pairs (Codex + GPT-6 Astra at 58.2%, Claude Code + Fable 5.1 at 57.9%) currently lead. OpenHands' harness-level score on 4.0 is not yet formally published, making direct comparison difficult.
Can OpenHands run on a fully open-weight model without any external API?
Yes. Your model API bill disappears entirely if you run open-weight models like Devstral on your own GPUs. The open-weight path is no longer embarrassing. OpenHands with Devstral 24B at 46.8% on SWE-bench Verified matches what Devin 2.0, a funded commercial product, publicly reported.
What is Terminal-Bench 4.0 and why did it replace 2.1?
Terminal-Bench 2.1 became saturated, with nine models scoring between 87% and 92%. Terminal-Bench 4.0 is the new standard, covering 66 harder tasks where the current top agent entry sits at 58.2%. The new version raises task complexity and adjusts resource limits to reduce confounds from harness choice.
Is the ConfirmRisky approval gate actually useful in practice?
It is the single most important safety control in the default stack. The agent pauses before any file-destructive or high-risk operation and waits for explicit human approval. For teams worried about a coding agent autonomously deleting or rewriting production files, this gate is the answer - and it is on by default, not buried in settings.