SWE-bench Pro's top models cluster within a 29.7-point band. Run the same models through DeepSWE and the spread blows out to 69.8 points. That gap is not noise-it is the benchmark telling you something important about what it is actually measuring.
DeepSWE is a coding agent benchmark released by Datacurve in May 2026, with the full paper landing on arXiv on July 8. Its premise is simple and its implications are not: most public coding benchmarks are built from merged GitHub pull requests, which means the fixes almost certainly leaked into model pretraining data, and the grading tests were written for one specific solution rather than for the problem. DeepSWE was designed to close both gaps at once.
What DeepSWE actually tests
DeepSWE measures how well frontier coding agents handle real engineering work-not toy functions or LeetCode puzzles, but multi-step tasks inside active open-source repositories. It was created by Datacurve and released in May 2026.
The benchmark includes 113 tasks spread across 91 repositories in five languages-TypeScript, Go, Python, JavaScript, and Rust-and each task asks the model to implement a feature or fix a bug inside a real codebase, then verifies the result with hand-written tests that check software behavior rather than implementation details.
The contamination problem it targets is real and growing. Most public agentic coding benchmarks follow SWE-bench in mining merged fixes from public GitHub repositories, which creates two problems: the fixes and their discussion were likely seen during pretraining, so a high score can reflect recall rather than problem-solving; and each task is graded by the tests that shipped with its merged fix, which were written to confirm one specific fix rather than grade an arbitrary solution, so they can fail a correct alternative or pass an incomplete one.
Every DeepSWE task is written from scratch by Datacurve's engineers, not adapted from existing GitHub commits or pull requests-which means no model has seen the solution during pretraining, a growing problem as benchmarks leak into training data and scores inflate without real improvement.
The grading gap is striking. When an independent LLM judge re-reviews graded runs, it disagrees with DeepSWE's verifier about an order of magnitude less often than with SWE-Bench Pro's inherited tests: 1.4% versus 32.4%. That 32.4% disagreement rate on SWE-bench Pro is not a minor calibration issue-it means roughly one in three graded results on that benchmark is genuinely arguable.
DeepSWE evaluates each task with the mini-swe-agent harness, where each task is solved in an isolated container with no internet access.
Because graders only score the committed patch in a separate container, some easier shortcuts no longer work: an agent cannot monkey-patch the test framework, and because the CTRF report records each task-defining test by name, dropping tests or forcing an early exit shows up as missing or failed results rather than a pass.
What the leaderboard actually shows
The spread difference is the headline. DeepSWE separates frontier coding agents across a wider band-a 69.8-point range versus 29.7 points on SWE-Bench Pro for the eight models with public reports. A wider spread aids resolution by making differences between configurations easier to resolve.
GPT-5.6 Sol from OpenAI currently leads the DeepSWE leaderboard with a score of 0.727 across 13 evaluated AI models.
GPT-5.5's 67% at $7.23 per task makes OpenAI the clear value leader on DeepSWE-it delivers near-top accuracy with the lowest token usage and fewest steps of any frontier model.
On the open-weight side, Hy4 preview by Tencent is the top-ranked open-source model on DeepSWE, with a score of 0.643.
There is also a notable absence. Notably absent from the leaderboard are DeepSeek and Mistral. Neither has been evaluated on DeepSWE v1.1 as of June 2026. Given that DeepSeek V4 scores well on SWE-bench Pro, that gap is worth watching-it is either a methodological blind spot or a vendor choice to avoid a harder test.
Pass rate alone hides what an agent spends to get there. DeepSWE tracks three cost-shaped measures alongside pass rate: median output tokens, median wall-clock duration, and median dollar cost per trial. This matters more than it looks. A model that completes 70% of tasks but burns $14 per task changes the economics of agentic pipelines entirely compared to one hitting 65% at $3 a task.
Why the task scope changes what you learn
Despite being about half the length of SWE-bench Pro's prompts, DeepSWE's prompts describe tasks whose reference solutions touch 5.5x more code. That asymmetry is the point. A short, ambiguous task spec that requires touching a lot of real code is much closer to what a developer actually hands an agent than a long, exhaustive prompt that all but scaffolds the solution.
The benchmark also lives alongside a confusing naming collision worth knowing: DeepSWE's name is shared, confusingly, with an unrelated 2025 open-source coding agent from Agentica and Together AI-a model, not a benchmark-a collision worth noting for anyone cross-referencing results under the same name.
How DeepSWE compares to other coding evals
There is now a small cluster of benchmarks trying to solve the same contamination problem from slightly different angles:
| Benchmark | Tasks | Languages | Grading method | Contamination control |
|---|---|---|---|---|
| DeepSWE | 113 original | TS, Go, Py, JS, Rust | Hand-written verifiers | All tasks written from scratch |
| SWE-bench Pro | ~300 | Python-heavy | Inherited fix tests | None (mined from merged PRs) |
| Terminal-Bench 2.1 | varies | Broad CLI tasks | Per-task tests | Tasks authored with solutions |
| FrontierSWE | varies | Python | Functional tests | Partial-some authored tasks |
The integration split of APEX-SWE has expert engineers author end-to-end systems graded by functional tests against live services. Terminal-Bench also authors each task with its own solution and tests, but by design measures command-line "terminal mastery" broadly rather than software engineering specifically-by its own account, software engineering is the largest but not the majority category.
The practical upshot: if you are selecting a coding agent for agentic pipeline work-not one-shot completions-DeepSWE is currently the highest-signal public eval for that use case. That said, a wider spread aids resolution but is not by itself a capability claim: it discriminates between agents only insofar as the resulting rank order tracks an external signal of quality. Run it against your actual repositories before committing.
For the open-weight picture, it is worth noting that GLM-5.2- an open-weight MoE model from Zhipu AI released June 13, 2026 under the MIT license -has published scores on several of these benchmarks. But a teammate like Beagle, triaging model questions in a Slack channel, would flag the important caveat: GLM-5.2 is verbose, generating about 43K output tokens per task versus 16K for GPT-5.5, and its cheap per-token price is real, but effective cost-per-task is closer to the frontier than the sticker suggests. That is exactly the kind of second-order cost that DeepSWE's per-task tracking is designed to surface.
DeepSWE benchmark: common questions
What is DeepSWE, and how does it differ from SWE-bench?
DeepSWE is a long-horizon software engineering benchmark created by Datacurve and released in May 2026. It measures how well frontier coding agents handle real engineering work-multi-step tasks inside active open-source repositories. Unlike SWE-bench, its tasks are written from scratch rather than mined from merged pull requests, eliminating the pretraining contamination that makes high SWE-bench scores hard to interpret.
Why does SWE-bench Pro show such a narrow score spread between top models?
SWE-bench grades tasks using tests that shipped with merged fixes, which were written to confirm one specific fix rather than grade an arbitrary solution. This makes grading unreliable and inflates scores for models that have seen similar patterns in training data. The result is an artificially compressed leaderboard where models that differ meaningfully in practice appear nearly identical on paper.
Which model leads the DeepSWE leaderboard right now?
GPT-5.6 Sol from OpenAI currently leads the DeepSWE leaderboard with a score of 0.727, and the current average across 13 evaluated models is 0.5. The top open-weight model is Tencent's Hy4 preview at 0.643-but several major open-weight models, including DeepSeek V4, have not yet published DeepSWE results.
Does DeepSWE measure cost efficiency, not just accuracy?
Yes-this is one of its design differentiators. Pass rate alone hides what an agent spends to get there. DeepSWE tracks three cost-shaped measures alongside pass rate: median output tokens, median wall-clock duration, and median dollar cost per trial. A model that scores 70% but costs $14 per task changes the ROI of an agentic pipeline significantly compared to a model hitting 65% at $3.
Can I run DeepSWE myself on a model not yet on the leaderboard?
Datacurve releases the benchmark, its verifiers, and the full record of evaluation trajectories publicly. The harness is mini-swe-agent, and each task runs in an isolated container. If you want to evaluate a model that hasn't published results-DeepSeek V4 and Mistral are obvious gaps-you can run it yourself using the public tooling.