Under concurrent load on a single H100, Mellum 2 runs 79% faster than Qwen3-8B and 21% faster than Qwen2.5-7B - not because it is a smarter model, but because its MoE architecture only touches 2.5 billion of its 12 billion parameters per forward pass. That gap between headline parameter count and active compute is the whole story.
JetBrains open-sourced Mellum 2 on June 1, 2026, under Apache 2.0, explicitly targeting the use cases where Claude Code and OpenAI Codex cannot go: air-gapped environments, compliance-sensitive teams, and shops that won't route proprietary code through an external API. This is a narrower claim than "we built a frontier model." It is worth taking seriously on its own terms.
What "focal model" actually means
JetBrains describes Mellum 2 as a "focal model" - fast and well-scoped for high-frequency tasks inside larger AI systems, not a replacement for every model in the stack, but a way to make the stack faster, cheaper, and easier to control. In practice, that means: a sub-agent that classifies an incoming prompt, a step that compresses retrieval context before it goes to a larger reasoner, a tool-calling pass that does not need GPT-4o-level reasoning to succeed.
Modern AI systems increasingly rely on multiple model calls - routing, retrieval, summarization, planning, validation, tool use - and many of these operations are latency-sensitive and do not require the largest available model. Mellum 2 targets those workloads.
This design philosophy has a concrete implication for how you read the benchmarks. The LiveCodeBench v6 score (37.2 for the Instruct variant) is lower than some 9-14B dense models, and the model is behind top comparators on MMLU-Redux and GPQA Diamond. The interpretation for engineers: Mellum 2 is optimized for throughput and latency in component roles - routing, RAG summarization, repeated agent steps - not for being the highest absolute performer across every academic benchmark.
That is not spin. It is an honest framing, and it is the right one for the target use case. What trips people up is comparing it to GLM-5.3 or Devstral on SWE-bench and concluding it is weak. It is not competing there.
The throughput numbers, translated
JetBrains benchmarked Mellum 2 against Qwen2.5-7B and Qwen3-8B on a single H100 GPU, using input and output sizes representative of real production code completion workloads. In single-request mode, it matches Qwen2.5-7B almost exactly - 192 tokens per second versus 193. Under concurrent load, which is where production deployments actually operate, it pulls 21% ahead of Qwen2.5-7B and 79% ahead of Qwen3-8B.
That 79% gap only appears under concurrency. At one request at a time, it vanishes. With only 2.5B parameters active per token, the architecture is designed to behave more like a 2.5B model than a conventional 12B dense model from an inference perspective - relevant for teams routing high volumes of requests.
To make that concrete: a team running 50 concurrent code-summarization calls through a self-hosted agent pipeline would see Mellum 2 drain its request queue nearly twice as fast as a Qwen3-8B setup with equivalent GPU memory. That matters when your agent loop has a tight SLA and your GPU is already shared across multiple services.
Where Mellum 2 fits (and where it does not)
Unlike its predecessor, which operated as a "focal" model that only focused on a single task like code completion inside an editor, Mellum 2 functions as a full-fledged coding assistant that can generate and edit code, call external tools, execute multi-step agentic workflows, hold long conversations, and use explicit reasoning.
JetBrains publishes six checkpoints - Base-Pretrain, Base, Instruct-SFT, Thinking-SFT, Instruct, Thinking - exposing each pipeline stage for research transparency, which is rare among IDE vendors. If you want to fine-tune Mellum 2 on your internal codebase, you have the right artifacts to start from base weights rather than an opinionated instruction-tuned version.
Apache 2.0 and native vLLM support - including tool-calling via --tool-call-parser hermes - simplifies production adoption and private hosting.
Early community reports flag Ollama compatibility issues with the custom MoE architecture , so if your team's local setup runs on Ollama today, test before you commit. vLLM is the safer path at launch.
One thing JetBrains does not address in the release: whether Mellum 2 will be integrated into IDE products such as AI Assistant or Junie, and on what timeline, is not addressed in the model card. The open-source release and any IDE integration are currently separate tracks.
Here is the comparison laid out:
| Model | Total params | Active params/token | SWE-bench Verified | License | Context |
|---|---|---|---|---|---|
| Mellum 2 (Instruct) | 12B | 2.5B | Not targeted | Apache 2.0 | 131K |
| Devstral Small 1.1 | 24B | 24B (dense) | 53.6% | Apache 2.0 | 128K |
| Qwen3-8B | 8B | 8B (dense) | - | Apache 2.0 | 128K |
| Qwen3.5-9B | 9B | 9B (dense) | - | Apache 2.0 | 128K |
Devstral is the right pick when the bottleneck is solved SWE-bench-style multi-file engineering tasks. Mellum 2 is the right pick when the bottleneck is throughput on a private, shared inference server running dozens of short, repeated steps inside a larger agent pipeline.
--tool-call-parser hermes flag and a link to the relevant checkpointWhat is genuinely new versus incremental
Mellum 2 is not a frontier model and does not claim to be. JetBrains' stated aim is the infrastructure layer of agentic AI systems - routing, retrieval pipelines, and sub-agent tasks - as well as private on-premises deployment. It is the follow-on to a 4B-parameter model debuted in late 2024 as a proprietary code completion tool before being open-sourced in April 2025. Unlike its predecessor, Mellum 2 is open from day one.
What is genuinely new: the active-compute framing applied to a purpose-built coding MoE, with full checkpoint transparency and a clean Apache 2.0 license. What is incremental: the underlying idea of a "small fast component model inside a multi-model stack" has been around since mixture-of-experts routing papers, and Mellum 2 is not the first small coding model with these numbers. What is missing: independent reproduction of the benchmark scores, and any clarity on IDE integration timelines.
The honest verdict is that Mellum 2 earns attention for teams with two specific problems: proprietary code they cannot send to an API, and a production agent pipeline running high-frequency short tasks. For everyone else, Devstral Small or a well-quantized Qwen3 variant on Ollama is still the simpler path.
Mellum 2 as a self-hosted coding agent model: common questions
What is Mellum 2 and who made it?
Mellum 2 is a 12B model engineered for production AI latency, throughput, and cost, built from scratch by JetBrains and released under the Apache 2.0 license. It is a Mixture-of-Experts architecture designed for agent sub-tasks, RAG pipelines, and private code deployments - not a frontier coding benchmark contender.
How does Mellum 2's MoE architecture affect inference cost?
Mellum 2 uses a Mixture-of-Experts architecture with 64 total experts and 8 activated per token. The full model has 12B parameters, but inference only touches 2.5B per forward pass - a roughly 5x reduction in compute compared to running all 12B. Under concurrent load, this translates to measurably higher throughput than same-class dense models.
Can Mellum 2 run locally or in an air-gapped environment?
Yes. It runs entirely on hardware you control, and is explicitly designed for deployment scenarios where Claude Code and OpenAI Codex cannot go: air-gapped environments, compliance-sensitive organisations, and teams that do not want to route every inference call through an external API. vLLM is the recommended serving layer; community reports flag some Ollama compatibility issues with the custom MoE format.
How does Mellum 2 compare to Devstral Small on coding benchmarks?
They target different jobs. Devstral Small 1.1 achieves 53.6% on SWE-Bench Verified, surpassing all other open models on that benchmark, while remaining lightweight enough to run on a single 4090 GPU or Apple Silicon machine. Mellum 2 does not publish a SWE-bench number; its Instruct variant scores 37.2% on LiveCodeBench v6. The meaningful difference is active compute: Mellum 2's 2.5B active parameters give it a throughput edge at high concurrency that a 24B dense Devstral cannot match.
What are Mellum 2's known weaknesses?
The LiveCodeBench v6 (37.2) trails Qwen3.5-9B (63.7) and Ministral-3-14B (42.4); GPQA Diamond (40.9) and MMLU-Redux (78.1) are below most models in the comparison set. It is not designed for frontier-level tasks - JetBrains explicitly positions Mellum 2 as a component model. It is also text and code only - no image or multimodal input - and carries an 8,192-token context for some task types, shorter than the 128K+ windows now common on full coding agents.