DeepSeek Harness: The Open-Source Coding Agent Where the Loop Is a Plugin

DeepSeek Harness (dsh) shipped August 13, 2026 and hit 197k GitHub stars in eight days. Here is what its plugin-first architecture actually means for teams building coding agents.

Cover art for DeepSeek Harness: The Open-Source Coding Agent Where the Loop Is a Plugin

DeepSeek Harness (dsh) - released as a developer preview on August 13, 2026 - reached roughly 197,000 GitHub stars within eight days of launch. That is a faster adoption rate than any tool DeepSeek has previously shipped. The star velocity is real, but it is the architectural bet underneath that is worth understanding.

197kGitHub stars in 8 daysfastest DeepSeek project ever
4runtime presetsStandard, Code, Minimal, Creator
0.1.0-rc.6current package versionstill a developer preview

Most coding agent frameworks treat the agent loop as infrastructure: you configure it, you do not replace it. DeepSeek Harness takes the opposite position. It is an open-source agent runtime released under MIT on August 13, 2026, in which the model adapter, the tool registry, the session log, and the agent loop itself are all replaceable plugins.

That is the claim. What it looks like in practice is worth unpacking.

What "everything is a plugin" actually means

The core is Cordis - a plugin framework where plugins contribute services, typed events, and reversible effects to a shared context. Every part of the product is a plugin: the model adapter, the tool registry, the session log, the agent loop itself. There is no privileged core to patch - you extend dsh by mounting a plugin beside the others.

That means you can swap out the model provider, replace the sandbox implementation, or change the loop strategy without touching the core. The plugins communicate through Cordis services and events.

The practical upshot: teams that have been forking OpenCode or patching Cline internals to change how sessions persist or how tools are scoped can instead write a plugin that mounts alongside the defaults. No fork, no upstream merge conflicts.

It is a program you install and run - the npm package is @deepseek-ai/dsh - and it connects whatever model you give it (DeepSeek V4 Flash by default) to your filesystem, your terminal, the web, and other agents.

One important disambiguation before you install: a completely unrelated Python package called deepseek-harness has been on PyPI since May and also exposes a dsh command.

They are different products from different maintainers. @deepseek-ai/dsh is the coding-agent harness on npm, released August 13.

The four presets - and which one actually matters for evaluation

dsh ships with four agent presets, each loading a different plugin set. Standard is a full coding agent: file editing, shell, file and web search, skills, planning, goals, subagents, and workflows. It is the right default for most work.

The mode that matters for anyone running evals is Minimal. Minimal mode strips it to two tools: persistent bash and str_replace_editor. This is explicitly designed for benchmarking. If you are running SWE-bench or similar evals, you want minimal mode because it reduces the harness's influence on results.

This is a genuine advance over most open-source harnesses, which expose no clean separation between "the agent loop in production" and "the agent loop for measurement." When you benchmark OpenHands or Cline in their default configurations, the scaffolding around the model is doing a lot of work - context injections, skill invocations, planning prompts - and that work is hard to isolate. With dsh, you lock the plugin set, swap only the model plugin, and compare trajectories under identical conditions.

Every run produces an append-only session log: system prompts, reasoning traces, tool calls and their results, subagent scheduling, context injections. The Trajectory view lets you inspect these by source. You can resume, fork, search, and replay from the same event stream.

Reviewers are calling Trajectory "DevTools for agents," which is about right. It is the kind of observability most agent frameworks add as an afterthought.

Beagle in action#eng-platform, 10:22am
The ask
'which harness mode should we use for our internal eval run?'
Beagle drafts
reads the dsh docs, drafts a summary of Standard vs Minimal trade-offs with a note about Trajectory logging
You approve
you hit approve; the recommendation posts with a direct link to the preset comparison - no context-switching to docs
Do this in your workspace →

How it compares to the other open-source harnesses right now

OpenCode (~202k stars, MIT) is still the de facto open-source alternative to Claude Code: provider-agnostic, local models, polished TUI. dsh is positioned as a direct competitor, but the architectural differences are real.

OpenCode DeepSeek Harness (dsh) OpenHands 1.0
License MIT MIT MIT
Interface TUI Web UI (port 3080) Browser (Agent Canvas)
Agent loop Fixed Plugin (swappable) SDK-based, composable
Benchmark mode No dedicated preset Minimal mode Configurable
Session replay No Yes (Trajectory) Partial
Status Stable Developer preview 1.0 GA
Default model Any DeepSeek V4 Flash Any

OpenHands shipped its 1.0 release with production-ready Docker sandboxing, built-in security policies, resource limits, and a plugin system.

It autonomously completes roughly 68% of SWE-bench Verified tasks, which puts it in the same conversation as commercial agents. OpenHands 1.0 has the production edge; dsh has the extensibility edge and the faster plugin ecosystem growth.

What the hype is hiding

dsh is billed as v0.1; the package actually reads 0.1.0-rc.5, with no GitHub release or tag behind it. The developer preview label is honest. Plugin contracts will break before a stable release, which means anything you build on the current API today will need to be revisited.

dsh runs as a local web app on port 3080 rather than in your terminal, and it is explicitly a developer preview that promises to break. The lack of a native CLI mattered enough to the community that a community project called oh-dsh - from a university open-source club - shipped Desktop, Web, and TUI builds from a single ohdsh command within a fortnight of launch.

That gap between the official release and immediate community patches is telling. The architecture is sound enough that builders are investing quickly, but the first-party surface is still rough. This is not a tool to drop into production today; it is a tool to build alongside now, so you are not catching up in six months when it stabilizes.

Adding a custom tool to a coding agent
Without Beagle
fork the upstream repo, find where tools are registered in the monolith, patch around the fixed loop, manage merge conflicts every release
With Beagle
write a dsh plugin that mounts a tool via ctx.tools, configure it in a patch file, run npx @deepseek-ai/dsh web - no fork

One more honest note on the star numbers. Press coverage reports roughly 50,000 GitHub stars in its first 12 hours. GitHub stars are attention, not adoption. The plugin count on launch day was inflated by the README asking authors to add a topic tag for discoverability. The real signal is whether the plugin contracts stabilize quickly and whether third-party loop implementations start appearing with benchmark results attached. Watch the Trajectory logs, not the star count.

A teammate like Beagle is worth thinking about here not for running the agent itself, but for the surrounding workflow: surfacing the right dsh changelog or plugin to a team who just asked in Slack which harness to use for their Friday eval run.

DeepSeek Harness: common questions

What is DeepSeek Harness (dsh)?

DeepSeek Harness is DeepSeek AI's open-source agent runtime, released under MIT on August 13, 2026, in which the model adapter, the tool registry, the session log, and the agent loop are all replaceable plugins. You install it via npm (@deepseek-ai/dsh), connect it to any OpenAI-compatible model endpoint, and compose its behavior through configuration patches rather than code forks.

Is dsh model-locked to DeepSeek models?

No. dsh has first-class pairing with DeepSeek V4 Flash - self-hosted or via API - while staying model-open. Researchers have already run it with Claude Sonnet 5 through a local OpenAI-compatible shim. The model adapter is a plugin like everything else, so swapping providers is a configuration change.

What is dsh Minimal mode for?

Minimal mode strips the agent to two tools - persistent bash and str_replace_editor - and is explicitly designed for benchmarking. It exists so you can hold the harness constant and vary only the model. This is the mode to use for SWE-bench or any other eval where you want the scaffold's influence on results to be as small and measurable as possible.

How does dsh compare to OpenHands 1.0?

The two are complementary in practice. OpenHands 1.0 is an architectural redesign aimed at production-level requirements, moving away from monolithic designs toward a modular system based on event-sourcing and optional isolation. dsh trades production maturity for deeper extensibility: its agent loop is itself a plugin, which OpenHands 1.0 does not offer. Teams that need SOC 2 audit trails today should stay on OpenHands; teams experimenting with novel loop strategies should look at dsh.

Should my team use dsh now?

DeepSeek Harness is available under the MIT license as a developer preview; its APIs and plugin contracts may still change before a stable release. Use it now if you are building agent infrastructure or running controlled evals and can absorb breaking changes. Hold off if you need a stable surface for a production workflow - the rc versioning means contracts will shift before 1.0.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle