Qwen3-Coder-Next Runs a Coding Agent on One GPU

Qwen3-Coder-Next is an open-weight coding agent model with 80B total parameters but only 3B active per token. Here's what its SWE-Bench score actually means for teams who want to self-host.

Cover art for Qwen3-Coder-Next Runs a Coding Agent on One GPU

A developer opens their cloud API bill on a Monday morning and sees $340 charged over the weekend - a long-running coding agent had been filing issues, generating patches, and running tool loops while nobody was watching. The work was legitimate. The bill was not surprising, exactly, but it was the kind of number that makes a team ask whether they should be running this workload themselves.

That question now has a more credible answer than it did six months ago. Alibaba's Qwen team released Qwen3-Coder-Next on February 3, 2026, under the Apache 2.0 license. It is an ultra-sparse Mixture of Experts model with roughly 80 billion total parameters but only about 3 billion activated per token, built on the Qwen3-Next architecture and tuned for agentic, repository-level coding.

That gap - 80 billion stored, 3 billion active - is the entire story. Get it wrong and you either dismiss the model for having "only" 3B active parameters, or you worry it needs a data center to run. Neither is right.

What "3B active parameters" means for a self-hosted coding agent

Active parameters set your inference compute, not total parameters. Compute cost at inference scales roughly linearly with active parameters. A 32B-active model costs approximately 10× more per token to run than a 3B-active model on equivalent hardware. Qwen3-Coder-Next exploits this hard: the feed-forward layers hold 512 routed experts plus one shared expert, and the router activates only 10 of the routed experts for any given token. Because the bulk of the network sits idle on each forward pass, the model touches roughly 3 billion of its 79 billion non-embedding parameters at inference time.

Total parameters still matter for storage and VRAM allocation, but they do not burn compute per token. At Q5_K_M quantization the model occupies around 64 GB; at Q4_K_M, roughly 52 GB. Because only 3B parameters are active per token, inference is fast - typically 60-120 tokens per second on a single H100.

The minimum requirement of ~52 GB RAM at Q4_K_M excludes all devices with 32 GB , which is an honest constraint. A MacBook Pro M3 Max with 64 GB runs it, an M3 Pro with 36 GB does not. That is the real hardware floor, and it is worth checking before you get excited about the benchmarks.

The SWE-Bench numbers in context

SWE-Bench Verified is a curated set of real GitHub issues from popular Python repositories. The score measures how often an agent, given a repo and a bug report, can produce a patch that passes the existing test suite. It is the closest public benchmark to actual software engineering work - though it still favors Python and doesn't test anything that requires a human to assess.

On SWE-Bench Verified using the SWE-Agent scaffold, Qwen3-Coder-Next scores 70.6. DeepSeek-V3.2 at 671B parameters scores 70.2, and GLM-4.7 at 358B parameters scores 74.2.

On the more challenging SWE-Bench Pro, Qwen3-Coder-Next scores 44.3, above DeepSeek-V3.2 at 40.9 and GLM-4.7 at 40.6.

The comparison to DeepSeek-V3.2 is the one worth sitting with. DeepSeek-V3.2 has 671B total parameters and is not practically self-hostable by most engineering teams. Qwen3-Coder-Next matches its SWE-Bench Verified score with a model that fits on consumer hardware. These results support the claim from the Qwen team that Qwen3-Coder-Next achieves performance comparable to models with 10-20× more active parameters, especially in coding and agentic settings.

What these numbers do not show: closed frontier models still lead. Claude Sonnet 4.6 scores 79.6% on SWE-Bench Verified and Claude Opus 4.6 scores 80.8%

  • both roughly 10 percentage points above Qwen3-Coder-Next. Most production coding tasks fall below the threshold where the gap to Claude Sonnet 4.6 is visible - meaning for routine work, Qwen3-Coder-Next is now "good enough." That "good enough" claim is defensible for high-volume, repetitive agentic loops. It is not defensible for one-off tasks where accuracy on the first attempt matters more than cost.
70.6%SWE-Bench Verifiedvs 70.2% for DeepSeek-V3.2 at 671B total params
44.3%SWE-Bench Probeats DeepSeek-V3.2 (40.9%) and GLM-4.7 (40.6%)
$0.12 / $0.80per million tokens (in/out)on OpenRouter as of February 2026
~52 GBQ4_K_M storageminimum viable VRAM/RAM for local deployment

Where it falls short on real agentic loops

This model supports only non-thinking mode and does not generate thinking blocks in its output. That matters for agent scaffolds that rely on chain-of-thought reasoning to plan multi-step tasks. Qwen3-Coder models with larger active parameter counts have hybrid thinking modes; Coder-Next trades that for inference efficiency.

One hands-on review tested it on a Next.js webhook task and found that it invented a package import that was not installed. When pointed to the error it corrected itself, but only after the reviewer named the file. Claude Code on the same task would have read package.json first. That pattern - strong raw code generation, weaker autonomous repo exploration - shows up consistently. For standalone code generation it is the first open model one reviewer would actually keep in a toolchain. For full agentic loops (file editing, multi-step plans), it is still one generation behind - tool-use reliability is the gap, not raw coding ability.

The model card notes it is designed specifically for coding tasks and may not be optimal for general conversational use. That is unusual candor from a model card. It is also accurate: this is not a general-purpose assistant with coding bolted on, it is a coding agent with general capability treated as a secondary concern.

Beagle in action#engineering, 10:47am
The ask
'can someone check if the new webhook handler PR is safe to merge?'
Beagle drafts
reads the linked diff and the repo's existing test output, drafts a reply summarizing coverage gaps and one dependency it cannot verify locally
You approve
you approve the summary; it posts with a source link to the test run - no API cost, no context lost in thread noise
Do this in your workspace →

Training method: why RL on executable environments matters

The technical report explores how far strong training recipes can push the capability limits of models with small parameter footprints. To achieve this, the team performs agentic training through large-scale synthesis of verifiable coding tasks paired with executable environments, allowing learning directly from environment feedback via mid-training and reinforcement learning.

This is the non-obvious part of the release. The MoE efficiency is well-covered. What gets less attention is that the model was not just trained on code text - it was trained on agent trajectories with real execution feedback. The difference between "generate code that looks right" and "generate code that passes tests in a live environment" is exactly what SWE-Bench measures. Training on executable environments is why a 3B-active model can close most of the gap to much larger models on that specific benchmark.

Through its training recipe, Qwen3-Coder-Next excels at long-horizon reasoning, complex tool usage, and recovery from execution failures, ensuring robust performance in dynamic coding tasks. The recovery-from-failure piece is what makes or breaks a coding agent in practice - a model that gives up or hallucinates a solution after a single failed execution is not useful for unattended agent loops.

Running a coding agent over a weekend
Without Beagle
every token billed at cloud API rates; $300+ for a 48-hour bug-triage loop; all code transits Anthropic or OpenAI servers
With Beagle
model loaded once on a 64GB workstation; per-token cost is electricity; repo never leaves your network; a teammate like Beagle drafts the Slack summary when the agent posts results

Open-weight coding agent: common questions

What is Qwen3-Coder-Next and how does it differ from Qwen3-Coder?

Qwen3-Coder-Next is Alibaba's efficiency-focused coding agent model, released February 4, 2026. Where the earlier Qwen3-Coder 480B model activated 35 billion parameters per token (requiring server-scale hardware), Coder-Next cuts that to 3 billion active parameters across an 80B total MoE, making it deployable on a single H100 or two consumer GPUs at Q4 quantization.

What hardware does Qwen3-Coder-Next actually need to run?

At Q4_K_M quantization, the model occupies roughly 52 GB, which means a 64 GB MacBook M3 Max, an RTX 5090 (32 GB VRAM, plus system RAM offloading), or a single H100 80 GB runs it comfortably. A 32 GB device will not. At Q5_K_M you need around 64 GB, and BF16 requires approximately 160 GB.

How does Qwen3-Coder-Next score on SWE-Bench compared to closed models?

It scores 70.6% on SWE-Bench Verified - roughly 9 percentage points behind Claude Sonnet 4.6 (79.6%) and Claude Opus 4.6 (80.8%). It matches or beats DeepSeek-V3.2 (70.2%) and beats GLM-4.7 (74.2% total, but with 358B total parameters). For most routine coding agent workloads, the gap to frontier closed models is not the deciding factor; privacy, cost, and rate limits are.

Does Qwen3-Coder-Next support thinking mode?

No. Unlike larger models in the Qwen3 family, Coder-Next does not generate chain-of-thought reasoning blocks. This simplifies prompt engineering and reduces output token overhead on long agent loops, but limits multi-step planning on tasks that genuinely benefit from extended reasoning before committing to an action.

What API pricing is available if I don't want to self-host?

On OpenRouter, Qwen3-Coder-Next is priced at $0.12 per million input tokens and $0.80 per million output tokens, with cached reads at $0.07 per million. At that rate, a 500-turn coding agent loop averaging 2,000 tokens per turn costs roughly $0.84 in output tokens - well under a dollar per substantial agentic session.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle