Cohere released North Mini Code on June 9, 2026
- and the marketing copy calls it an "agentic coding model." The benchmarks tell a more specific story. It scores 27.6 on the Artificial Analysis Intelligence Index, slightly below Qwen3.6 on coding, but hits up to 2.8x the output throughput of Devstral Small 2 on identical hardware. That gap between "smartest" and "fastest on one GPU you own" is the whole decision.
What North Mini Code actually is
North Mini Code is a 30B total / 3B active parameter Mixture of Experts model trained specifically for agentic coding. It is the first model in Cohere's North family of code agent models, released under the Apache 2.0 license. The MoE design matters in practice: it uses 128 experts with 8 activated per token. You get the representational capacity of a large model at the inference cost of a much smaller one.
It offers a 256K-token context window with up to 64K tokens of output, supports interleaved reasoning and tool use via JSON schema, and is released open-weight under the Apache 2.0 license.
Cohere has spent most of its life shipping closed, enterprise-focused models. North Mini Code 1.0 is a notable turn: a compact, fully open-weight coding model built specifically for agentic workflows - the kind that plan, edit files, run tools, and self-correct across a long task rather than answering a single prompt.
North Mini Code is available for free on Hugging Face (bf16, fp8, w4a16), OpenRouter, and Model Vault - Cohere's fully managed inference platform. It has been specifically trained for compatibility with OpenCode, but works with most coding agents.
The hardware math for private deployment
Cohere lists a minimum requirement of a single H100-class GPU when the model is served in the FP8 numeric format. That is a concrete deployment target, not a vague "runs locally" claim you have to stress-test yourself.
30B total / 3B active means it fits on 1× H100 at FP8 and serves fast - around 199 output tokens per second on Cohere's API, with a 256K-token context window. At 199 tok/s, a 500-token code edit takes under three seconds. For an agent running dozens of sequential tool calls, throughput compounds.
Cohere released North Mini Code under the Apache 2.0 license, which allows commercial use, modification, and private deployment with no per-seat or per-token fees. You are responsible for the hardware you run it on, but the weights themselves carry no licensing cost.
Compare that to frontier API pricing: GPT-5.5 sits at $30 per million output tokens, while Claude Opus 4.7 and Gemini 3.1 Pro come in around $12 to $25 per million output tokens. A self-hosted H100 on a cloud instance runs roughly $3-4/hr in spot pricing. At 199 tok/s, you can generate around 700,000 tokens an hour - putting your effective per-million-token cost well under $10 once the box is busy. The break-even is not exotic.
What the benchmarks honestly say
Here is where you need to read carefully. Cohere leads with agentic software engineering. The numbers support that positioning on one axis and undercut it on another.
North Mini Code scores 33.4 on the Artificial Analysis Coding Index but only 14% on GDPval-AA and 37% on τ²-Bench Telecom - strong at code, weak at broader agentic reasoning.
North Mini Code ranks #238 of 438 AI models for overall intelligence, #107 of 210 for coding, and #170 of 192 for agentic tasks in one independent tracker. That last number is the one to sit with. Cohere frames the model around agent harnesses like SWE-Agent and OpenCode. Cohere optimized it for terminal work, repository-level software engineering, and tool use across multiple agent harnesses. That focus matters when reading the benchmark table because a coding agent is evaluated together with its prompt format, tools, parser, and execution environment. In the right harness, with the right prompt format, the numbers look better. Out of that harness, the ranking tells you to be cautious.
| Dimension | Score / Rank | Note |
|---|---|---|
| Artificial Analysis Coding Index | 33.4 | Competitive for its size class |
| Artificial Analysis Intelligence Index | 27.6 | Slightly below Qwen3.6 on coding |
| Agentic task ranking | #170 / 192 | Narrow, structured tasks only |
| GDPval-AA | 14% | Weak on broader autonomous reasoning |
| τ²-Bench Telecom | 37% | Weak on multi-domain tool tasks |
| Output throughput | ~199 tok/s | 2.8× Devstral Small 2 on same iron |
Its coding capability is stronger than its general reasoning profile, making it better suited to implementation assistance than complex autonomous problem solving.
Treat vendor-reported benchmarks as a starting point and validate on your own repository before committing. That applies here more than most: the agentic scores are sensitive to which harness you plug it into.
Where it fits in a real private deployment
Its small active footprint makes it suitable for local deployment while remaining competitive with larger open-source models on software engineering and terminal-based agentic benchmarks. Designed tasks include: agentic software engineering for repo-level code changes inside harnesses like SWE-Agent and OpenCode; terminal-based agents driving shell tools end-to-end across multi-turn tasks; and local and on-device coding where the 3B active parameter footprint enables low-latency inference.
The pattern it slots into is what practitioners call SLM-first routing. The agent runs the local model on every step. A lightweight router watches for trouble - a low-confidence response, a schema violation, a tool call that doesn't parse - and only then escalates that one turn to a frontier cloud model. The result lands back in the local loop and execution continues.
Multiple 2025-2026 practitioner reports put the share of agentic tasks retained in the efficient local lane at 80-90% - a rule of thumb, not a peer-reviewed measurement, and one that shifts with the workload and the confidence threshold. For coding tasks that are well-structured and narrow - the kind North Mini Code was trained on - that retention rate is plausible. For open-ended reasoning tasks, it isn't.
Use North Mini Code for cost-sensitive coding support and experimentation. Choose a stronger model for production agents, difficult debugging, or client-facing technical analysis.
A teammate like Beagle, working inside Slack or Teams, would surface that tradeoff early - flagging when an incoming question requires multi-step reasoning rather than structured code generation, before you've committed to a model swap.
The decision boundary is sharper than it looks: if your agent task is structured, repository-local, and tool-call-heavy, North Mini Code's throughput and $0 licensing are hard to match. If the task requires open-ended judgment across domains, you'll hit the agentic ranking ceiling fast.
North Mini Code for private AI: common questions
What hardware does North Mini Code need to run privately?
Cohere lists a minimum requirement of a single H100-class GPU when the model is served in FP8 precision.
A hosted endpoint is available via Cohere's Chat V2 API using the model identifier north-mini-code-1-0 if you'd rather not run infrastructure. It's also available through Ollama as north-mini-code-1.0 for local experimentation.
Is North Mini Code actually free to use commercially?
Cohere announced it in June 2026 and released the weights under the Apache 2.0 license, which permits commercial use, modification, and private deployment without per-seat fees. For an organization that wants to run AI on its own terms, the license is as important as the benchmark scores.
How does North Mini Code compare to Qwen3 Coder for self-hosted coding agents?
North Mini Code is not the smartest small model on paper - it scores slightly below Qwen3.6 on coding - but is one of the fastest, hitting up to 2.8x the output throughput of Devstral Small 2 on identical hardware.
Qwen 3 Coder handles the bulk of routine coding tasks at a fraction of frontier cost, and the dense 32B and MoE 235B variants are self-hostable on consumer or prosumer hardware. If throughput on a single GPU matters more than raw coding benchmark rank, North Mini Code has an edge. If benchmark scores on general reasoning are the deciding factor, Qwen 3 Coder is the better default.
What is North Mini Code bad at?
Its very low agentic performance is a clear limitation for multi-step automation, tool orchestration, and unattended engineering tasks. The 14% score on GDPval-AA and 37% on τ²-Bench Telecom confirm it struggles outside structured, code-first tasks. Do not deploy it as a general-purpose agent without harness-specific validation.
Where can I download the weights?
The model card is CohereLabs/North-Mini-Code-1.0 on Hugging Face. In FP8 it fits on a single H100, so a self-hosted deployment is realistic for one high-memory GPU. Quantized variants (bf16, fp8, w4a16) are all available.