DeepSeek Harness (dsh) - released as a developer preview on August 13, 2026 - reached roughly 197,000 GitHub stars within eight days of launch. That is a faster adoption rate than any tool DeepSeek has previously shipped. The star velocity is real, but it is the architectural bet underneath that is worth understanding.
Most coding agent frameworks treat the agent loop as infrastructure: you configure it, you do not replace it. DeepSeek Harness takes the opposite position. It is an open-source agent runtime released under MIT on August 13, 2026, in which the model adapter, the tool registry, the session log, and the agent loop itself are all replaceable plugins.
That is the claim. What it looks like in practice is worth unpacking.
What "everything is a plugin" actually means
The core is Cordis - a plugin framework where plugins contribute services, typed events, and reversible effects to a shared context. Every part of the product is a plugin: the model adapter, the tool registry, the session log, the agent loop itself. There is no privileged core to patch - you extend dsh by mounting a plugin beside the others.
That means you can swap out the model provider, replace the sandbox implementation, or change the loop strategy without touching the core. The plugins communicate through Cordis services and events.
The practical upshot: teams that have been forking OpenCode or patching Cline internals to change how sessions persist or how tools are scoped can instead write a plugin that mounts alongside the defaults. No fork, no upstream merge conflicts.
It is a program you install and run - the npm package is @deepseek-ai/dsh - and it connects whatever model you give it (DeepSeek V4 Flash by default) to your filesystem, your terminal, the web, and other agents.
One important disambiguation before you install:
a completely unrelated Python package called deepseek-harness has been on PyPI since May and also exposes a dsh command.
They are different products from different maintainers. @deepseek-ai/dsh is the coding-agent harness on npm, released August 13.
The four presets - and which one actually matters for evaluation
dsh ships with four agent presets, each loading a different plugin set. Standard is a full coding agent: file editing, shell, file and web search, skills, planning, goals, subagents, and workflows. It is the right default for most work.
The mode that matters for anyone running evals is Minimal. Minimal mode strips it to two tools: persistent bash and str_replace_editor. This is explicitly designed for benchmarking. If you are running SWE-bench or similar evals, you want minimal mode because it reduces the harness's influence on results.
This is a genuine advance over most open-source harnesses, which expose no clean separation between "the agent loop in production" and "the agent loop for measurement." When you benchmark OpenHands or Cline in their default configurations, the scaffolding around the model is doing a lot of work - context injections, skill invocations, planning prompts - and that work is hard to isolate. With dsh, you lock the plugin set, swap only the model plugin, and compare trajectories under identical conditions.
Every run produces an append-only session log: system prompts, reasoning traces, tool calls and their results, subagent scheduling, context injections. The Trajectory view lets you inspect these by source. You can resume, fork, search, and replay from the same event stream.
Reviewers are calling Trajectory "DevTools for agents," which is about right. It is the kind of observability most agent frameworks add as an afterthought.
How it compares to the other open-source harnesses right now
OpenCode (~202k stars, MIT) is still the de facto open-source alternative to Claude Code: provider-agnostic, local models, polished TUI. dsh is positioned as a direct competitor, but the architectural differences are real.
| OpenCode | DeepSeek Harness (dsh) | OpenHands 1.0 | |
|---|---|---|---|
| License | MIT | MIT | MIT |
| Interface | TUI | Web UI (port 3080) | Browser (Agent Canvas) |
| Agent loop | Fixed | Plugin (swappable) | SDK-based, composable |
| Benchmark mode | No dedicated preset | Minimal mode | Configurable |
| Session replay | No | Yes (Trajectory) | Partial |
| Status | Stable | Developer preview | 1.0 GA |
| Default model | Any | DeepSeek V4 Flash | Any |
OpenHands shipped its 1.0 release with production-ready Docker sandboxing, built-in security policies, resource limits, and a plugin system.
It autonomously completes roughly 68% of SWE-bench Verified tasks, which puts it in the same conversation as commercial agents. OpenHands 1.0 has the production edge; dsh has the extensibility edge and the faster plugin ecosystem growth.
What the hype is hiding
dsh is billed as v0.1; the package actually reads 0.1.0-rc.5, with no GitHub release or tag behind it. The developer preview label is honest. Plugin contracts will break before a stable release, which means anything you build on the current API today will need to be revisited.
dsh runs as a local web app on port 3080 rather than in your terminal, and it is explicitly a developer preview that promises to break.
The lack of a native CLI mattered enough to the community that
a community project called oh-dsh - from a university open-source club - shipped Desktop, Web, and TUI builds from a single ohdsh command within a fortnight of launch.
That gap between the official release and immediate community patches is telling. The architecture is sound enough that builders are investing quickly, but the first-party surface is still rough. This is not a tool to drop into production today; it is a tool to build alongside now, so you are not catching up in six months when it stabilizes.
One more honest note on the star numbers. Press coverage reports roughly 50,000 GitHub stars in its first 12 hours. GitHub stars are attention, not adoption. The plugin count on launch day was inflated by the README asking authors to add a topic tag for discoverability. The real signal is whether the plugin contracts stabilize quickly and whether third-party loop implementations start appearing with benchmark results attached. Watch the Trajectory logs, not the star count.
A teammate like Beagle is worth thinking about here not for running the agent itself, but for the surrounding workflow: surfacing the right dsh changelog or plugin to a team who just asked in Slack which harness to use for their Friday eval run.
DeepSeek Harness: common questions
What is DeepSeek Harness (dsh)?
DeepSeek Harness is DeepSeek AI's open-source agent runtime, released under MIT on August 13, 2026, in which the model adapter, the tool registry, the session log, and the agent loop are all replaceable plugins.
You install it via npm (@deepseek-ai/dsh), connect it to any OpenAI-compatible model endpoint, and compose its behavior through configuration patches rather than code forks.
Is dsh model-locked to DeepSeek models?
No. dsh has first-class pairing with DeepSeek V4 Flash - self-hosted or via API - while staying model-open. Researchers have already run it with Claude Sonnet 5 through a local OpenAI-compatible shim. The model adapter is a plugin like everything else, so swapping providers is a configuration change.
What is dsh Minimal mode for?
Minimal mode strips the agent to two tools - persistent bash and str_replace_editor - and is explicitly designed for benchmarking. It exists so you can hold the harness constant and vary only the model. This is the mode to use for SWE-bench or any other eval where you want the scaffold's influence on results to be as small and measurable as possible.
How does dsh compare to OpenHands 1.0?
The two are complementary in practice. OpenHands 1.0 is an architectural redesign aimed at production-level requirements, moving away from monolithic designs toward a modular system based on event-sourcing and optional isolation. dsh trades production maturity for deeper extensibility: its agent loop is itself a plugin, which OpenHands 1.0 does not offer. Teams that need SOC 2 audit trails today should stay on OpenHands; teams experimenting with novel loop strategies should look at dsh.
Should my team use dsh now?
DeepSeek Harness is available under the MIT license as a developer preview; its APIs and plugin contracts may still change before a stable release. Use it now if you are building agent infrastructure or running controlled evals and can absorb breaking changes. Hold off if you need a stable surface for a production workflow - the rc versioning means contracts will shift before 1.0.