Cohere North Mini Code Is Fast on One GPU, Not the Smartest

Cohere's North Mini Code ships 30B parameters but activates only 3B per token, runs on a single H100, and costs nothing in licensing. Here's what the benchmarks actually say before you deploy it privately.

Cover art for Cohere North Mini Code Is Fast on One GPU, Not the Smartest

Cohere released North Mini Code on June 9, 2026

  • and the marketing copy calls it an "agentic coding model." The benchmarks tell a more specific story. It scores 27.6 on the Artificial Analysis Intelligence Index, slightly below Qwen3.6 on coding, but hits up to 2.8x the output throughput of Devstral Small 2 on identical hardware. That gap between "smartest" and "fastest on one GPU you own" is the whole decision.

What North Mini Code actually is

North Mini Code is a 30B total / 3B active parameter Mixture of Experts model trained specifically for agentic coding. It is the first model in Cohere's North family of code agent models, released under the Apache 2.0 license. The MoE design matters in practice: it uses 128 experts with 8 activated per token. You get the representational capacity of a large model at the inference cost of a much smaller one.

It offers a 256K-token context window with up to 64K tokens of output, supports interleaved reasoning and tool use via JSON schema, and is released open-weight under the Apache 2.0 license.

Cohere has spent most of its life shipping closed, enterprise-focused models. North Mini Code 1.0 is a notable turn: a compact, fully open-weight coding model built specifically for agentic workflows - the kind that plan, edit files, run tools, and self-correct across a long task rather than answering a single prompt.

North Mini Code is available for free on Hugging Face (bf16, fp8, w4a16), OpenRouter, and Model Vault - Cohere's fully managed inference platform. It has been specifically trained for compatibility with OpenCode, but works with most coding agents.

The hardware math for private deployment

Cohere lists a minimum requirement of a single H100-class GPU when the model is served in the FP8 numeric format. That is a concrete deployment target, not a vague "runs locally" claim you have to stress-test yourself.

30B total / 3B active means it fits on 1× H100 at FP8 and serves fast - around 199 output tokens per second on Cohere's API, with a 256K-token context window. At 199 tok/s, a 500-token code edit takes under three seconds. For an agent running dozens of sequential tool calls, throughput compounds.

30B / 3Btotal / active parametersMoE: knowledge of large, cost of small
256K tokenscontext window64K max output tokens
199 tok/soutput throughput (Cohere API)2.8× faster than Devstral Small 2 on same hardware
1× H100minimum deployment targetFP8 precision, no multi-GPU rack required

Cohere released North Mini Code under the Apache 2.0 license, which allows commercial use, modification, and private deployment with no per-seat or per-token fees. You are responsible for the hardware you run it on, but the weights themselves carry no licensing cost.

Compare that to frontier API pricing: GPT-5.5 sits at $30 per million output tokens, while Claude Opus 4.7 and Gemini 3.1 Pro come in around $12 to $25 per million output tokens. A self-hosted H100 on a cloud instance runs roughly $3-4/hr in spot pricing. At 199 tok/s, you can generate around 700,000 tokens an hour - putting your effective per-million-token cost well under $10 once the box is busy. The break-even is not exotic.

What the benchmarks honestly say

Here is where you need to read carefully. Cohere leads with agentic software engineering. The numbers support that positioning on one axis and undercut it on another.

North Mini Code scores 33.4 on the Artificial Analysis Coding Index but only 14% on GDPval-AA and 37% on τ²-Bench Telecom - strong at code, weak at broader agentic reasoning.

North Mini Code ranks #238 of 438 AI models for overall intelligence, #107 of 210 for coding, and #170 of 192 for agentic tasks in one independent tracker. That last number is the one to sit with. Cohere frames the model around agent harnesses like SWE-Agent and OpenCode. Cohere optimized it for terminal work, repository-level software engineering, and tool use across multiple agent harnesses. That focus matters when reading the benchmark table because a coding agent is evaluated together with its prompt format, tools, parser, and execution environment. In the right harness, with the right prompt format, the numbers look better. Out of that harness, the ranking tells you to be cautious.

Dimension Score / Rank Note
Artificial Analysis Coding Index 33.4 Competitive for its size class
Artificial Analysis Intelligence Index 27.6 Slightly below Qwen3.6 on coding
Agentic task ranking #170 / 192 Narrow, structured tasks only
GDPval-AA 14% Weak on broader autonomous reasoning
τ²-Bench Telecom 37% Weak on multi-domain tool tasks
Output throughput ~199 tok/s 2.8× Devstral Small 2 on same iron

Its coding capability is stronger than its general reasoning profile, making it better suited to implementation assistance than complex autonomous problem solving.

Treat vendor-reported benchmarks as a starting point and validate on your own repository before committing. That applies here more than most: the agentic scores are sensitive to which harness you plug it into.

Beagle in action#eng-tools, 2:47pm
The ask
'is North Mini Code good enough to replace our Codex API calls for the PR review bot?'
Beagle drafts
pulls the Cohere docs and the Artificial Analysis index, drafts a summary of the coding vs agentic ranking gap with the relevant benchmark rows
You approve
team sees the #170/192 agentic rank alongside the 199 tok/s throughput figure - makes a harness-specific test plan rather than a blanket swap
Do this in your workspace →

Where it fits in a real private deployment

Its small active footprint makes it suitable for local deployment while remaining competitive with larger open-source models on software engineering and terminal-based agentic benchmarks. Designed tasks include: agentic software engineering for repo-level code changes inside harnesses like SWE-Agent and OpenCode; terminal-based agents driving shell tools end-to-end across multi-turn tasks; and local and on-device coding where the 3B active parameter footprint enables low-latency inference.

The pattern it slots into is what practitioners call SLM-first routing. The agent runs the local model on every step. A lightweight router watches for trouble - a low-confidence response, a schema violation, a tool call that doesn't parse - and only then escalates that one turn to a frontier cloud model. The result lands back in the local loop and execution continues.

Multiple 2025-2026 practitioner reports put the share of agentic tasks retained in the efficient local lane at 80-90% - a rule of thumb, not a peer-reviewed measurement, and one that shifts with the workload and the confidence threshold. For coding tasks that are well-structured and narrow - the kind North Mini Code was trained on - that retention rate is plausible. For open-ended reasoning tasks, it isn't.

Use North Mini Code for cost-sensitive coding support and experimentation. Choose a stronger model for production agents, difficult debugging, or client-facing technical analysis.

A teammate like Beagle, working inside Slack or Teams, would surface that tradeoff early - flagging when an incoming question requires multi-step reasoning rather than structured code generation, before you've committed to a model swap.

Replacing a cloud coding API with North Mini Code privately
Without Beagle
every PR review bot call hits a $25/M output API; costs spike on long agentic sessions; data leaves the network
With Beagle
North Mini Code on one H100 at ~$3.50/hr spot; 199 tok/s throughput; zero per-token meter; data stays inside your VPC

The decision boundary is sharper than it looks: if your agent task is structured, repository-local, and tool-call-heavy, North Mini Code's throughput and $0 licensing are hard to match. If the task requires open-ended judgment across domains, you'll hit the agentic ranking ceiling fast.


North Mini Code for private AI: common questions

What hardware does North Mini Code need to run privately?

Cohere lists a minimum requirement of a single H100-class GPU when the model is served in FP8 precision.

A hosted endpoint is available via Cohere's Chat V2 API using the model identifier north-mini-code-1-0 if you'd rather not run infrastructure. It's also available through Ollama as north-mini-code-1.0 for local experimentation.

Is North Mini Code actually free to use commercially?

Cohere announced it in June 2026 and released the weights under the Apache 2.0 license, which permits commercial use, modification, and private deployment without per-seat fees. For an organization that wants to run AI on its own terms, the license is as important as the benchmark scores.

How does North Mini Code compare to Qwen3 Coder for self-hosted coding agents?

North Mini Code is not the smartest small model on paper - it scores slightly below Qwen3.6 on coding - but is one of the fastest, hitting up to 2.8x the output throughput of Devstral Small 2 on identical hardware.

Qwen 3 Coder handles the bulk of routine coding tasks at a fraction of frontier cost, and the dense 32B and MoE 235B variants are self-hostable on consumer or prosumer hardware. If throughput on a single GPU matters more than raw coding benchmark rank, North Mini Code has an edge. If benchmark scores on general reasoning are the deciding factor, Qwen 3 Coder is the better default.

What is North Mini Code bad at?

Its very low agentic performance is a clear limitation for multi-step automation, tool orchestration, and unattended engineering tasks. The 14% score on GDPval-AA and 37% on τ²-Bench Telecom confirm it struggles outside structured, code-first tasks. Do not deploy it as a general-purpose agent without harness-specific validation.

Where can I download the weights?

The model card is CohereLabs/North-Mini-Code-1.0 on Hugging Face. In FP8 it fits on a single H100, so a self-hosted deployment is realistic for one high-memory GPU. Quantized variants (bf16, fp8, w4a16) are all available.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle