Separate Compute Cost from Storage Cost Before Running Open-Weight Models

Two frontier open-weight models landed this week - Mistral Large 4 and Reflection AI's Beam. Before your team plans around either, understand why the active-parameter count and the total-parameter count are completely different bills.

Cover art for Separate Compute Cost from Storage Cost Before Running Open-Weight Models

Two frontier open-weight models landed within 24 hours of each other this week. Reflection AI unveiled Beam, its first open-weight model, on October 5.

Mistral released Mistral Large 4 into public API preview on October 6. The headlines on both led with a single enormous number. That number is almost entirely the wrong thing to look at if you are deciding whether to run either model yourself.

The right number is two numbers, and they live in completely different budget lines.

Why the active/total split is the number that actually matters

A mixture-of-experts model splits its parameters into many small specialist subnetworks. For any given token, only a subset of those specialists - the "active" parameters - actually runs. The rest sit in memory, doing nothing, still taking up VRAM.

This creates two completely separate cost surfaces. An MoE is cheap to compute and expensive to hold: you pay VRAM for every parameter even though only a fraction fires per token. Getting those confused is how teams end up staring at an out-of-memory error on hardware they thought was sufficient.

Beam is a 501-billion-parameter sparse mixture-of-experts model, with only 23 billion of those parameters active on any given token.

ML4 holds 1.05 trillion parameters in total, with about 49 billion active per token - under 4.7 percent of the total.

Do the math on what that means for storage. A trillion-parameter model at FP8 precision weighs roughly 1.05 terabytes of weights alone. At mixed FP8/FP4, you need roughly 760 GB of GPU memory just to hold ML4; at full FP8, the estimate rises to about 1.26 TB.

At the minimum viable config - eight H100 80 GB cards giving 640 GB - ML4 sits at the absolute floor with no margin at 4-bit.

Storage still charges by the total: roughly 501 GB for Beam at FP8 plus a heavy KV cache at its 1-million-token context - open weights do not mean cheap to run.

What each model's efficiency claim actually says - and what it leaves out

Reflection's selling point for Beam is intelligence per token, not peak capability: it claims to match GLM-5.2 on reasoning while using 3 to 4 times less inference compute. That is a meaningful claim if it holds up. But read the method.

Reflection estimates compute as roughly twice the active parameter count multiplied by mean generated tokens, excluding prompt prefill, attention operations, and serving overhead. It calls this an approximate comparison rather than measured inference cost. Prompt prefill and attention at a 1-million-token context window are not small omissions.

On Reflection's own benchmark table, Chinese rivals still lead on most rows: DeepSeek V4.1 Flash beats Beam on Terminal Bench v2.1 (90.6 vs 80.1) and DeepSWE v1.1 (74.2 vs 44.4), and Kimi K3 leads on SWE Bench Pro v2-Hard (88.2 vs 77.2). Beam's argument is not that it tops the leaderboard but that it approaches that territory at a lower per-token compute budget.

Because only 49 billion parameters work on each token, Mistral Large 4 does roughly the per-token work of a 49-billion-parameter dense model, which one early measurement put at 116 tokens per second. That is a usable throughput for many agentic workflows. But every expert must still sit in memory, so self-hosting is a multi-GPU server project, not a workstation one.

501B / 23BBeam's total vs active params4.6% of parameters do the work per token
1.05T / 49BML4's total vs active params<4.7% active; the rest sits in VRAM
3-4xBeam's claimed compute advantagevs GLM-5.2, by Reflection's own estimate method
~760 GBML4 minimum GPU memoryat mixed 4-bit/FP8 quantization

Two things you cannot evaluate yet - and why they matter before you plan hardware

Both models are announced, not delivered. The weight files for Beam and ML4 are due later in October. That lag matters for two reasons beyond the obvious "you can't run it yet."

First, the license. Mistral Large 3 shipped under Apache 2.0. Mistral's documentation lists Large 4 as "Open" without naming a license, and reporting suggests the weights will arrive under a custom Mistral license whose terms were not detailed. Until those terms are published, nobody can confirm whether commercial self-hosting will be unrestricted.

Open weights do not mean free commercial use. Beam's Apache 2.0 license is clearer - Reflection confirmed Apache 2.0 weights are coming this month

  • but the weights themselves are not yet verifiable.

Second, the hardware requirements. Weights for ML4 are promised around October 27, and the license and hardware requirements are unpublished. The model is a 1T-parameter MoE far beyond a standard GPU host, so it needs multi-GPU H100-class nodes. Nothing is plannable or procurable yet.

A team that starts sizing a self-hosted cluster right now is pricing against specs that haven't shipped. That is not due diligence - it is speculation with a purchase order attached.

Beagle in action#infra-planning, 10:22am
The ask
'can someone summarize what we actually know vs what's still TBD on ML4 self-hosting?'
Beagle drafts
reads the Mistral announcement, the Yottalabs hardware analysis, and the license thread; drafts a two-column summary: confirmed specs vs open questions
You approve
you approve it; the channel has a sourced status doc instead of a thread of half-remembered tweets
Do this in your workspace →

How to actually compare these two models before the weights land

You can evaluate both through their APIs right now. That is the right move. ML4 launched October 6 in public preview, available only through Mistral Studio and the Mistral API.

Beam is available in selective preview, with weights and an Apache 2.0 release planned later this month.

Run your actual workload through both APIs. Not a benchmark someone else ran - your tasks, your context lengths, your tool-call patterns. Then compare:

  • Latency on your p95 prompt length, not a synthetic short prompt
  • Output quality on your hardest five tasks, scored by the people who will use the output
  • Cost per completed task, not cost per token (MoE token counts can diverge from dense-model intuitions at long contexts)
  • Refusal rate on edge cases your team actually hits

The comparison table below shows what is known as of this week - use it as a checklist of what to verify when weights ship.

Dimension Beam (Reflection AI) Mistral Large 4
Total / active params 501B / 23B 1.05T / 49B
Context window 1M tokens (256K beta API) 1M tokens
License Apache 2.0 (at weights release) Unannounced; "custom Mistral license" reported
Weights available Later October 2026 ~October 27, 2026
Minimum GPU config (FP8) ~500 GB VRAM ~1.26 TB VRAM
Modalities Text only Text + images
Benchmark lead areas SWE-Bench Verified (80.9) Legal, finance, cyber, visual grounding
Preview API pricing Not published $1.36 / $4.18 per M tokens (in/out)

The efficiency architecture is real and it matters. Active-parameter compute costs are genuinely lower than a dense model of the same nominal size. But the storage cost is set by the total, and for a trillion-parameter model, that storage bill is a multi-H100-node conversation. Know which bill you are talking about before the meeting with your infrastructure team.


Open weight model self-hosting: common questions

What does "active parameters" mean for MoE models?

In a mixture-of-experts model, only a subset of the total parameters - the "active" ones - processes any given token. The rest sit in GPU memory but do not compute. Active parameters set your inference cost and speed. Total parameters set your VRAM requirement. For Beam (23B active / 501B total) and ML4 (49B active / 1.05T total), these are very different numbers.

How much GPU memory does Mistral Large 4 require to self-host?

At FP8 precision, ML4 requires roughly 1.26 TB of GPU memory - about sixteen H100 80 GB cards - to hold the full weights. At a mixed 4-bit/FP8 quantization, the floor drops to around 760 GB, which fits on a cluster of ten to twelve H100s. The license terms are not yet published, so confirm commercial self-hosting rights before sizing hardware.

Is Reflection AI's Beam better than Mistral Large 4?

They are optimized for different things. Beam leads on several coding benchmarks (SWE-Bench Verified at 80.9) and is text-only with Apache 2.0 licensing. ML4 is natively multimodal, covers 160+ languages, and posts stronger results on legal, financial, and cybersecurity tasks. Neither model's weights are independently verifiable yet - run both APIs against your specific workload before deciding.

When will Mistral Large 4 open weights be available?

Mistral has committed to releasing ML4 weights by the end of October 2026. Reporting from multiple outlets points to October 27 as the target date, though Mistral's official announcement only says "by end of month." The license terms will be published alongside the weights - check those before planning commercial self-hosted deployment.

What is the real inference cost difference between MoE and dense models?

For a given token, an MoE model activates only its active-parameter slice - so Beam runs like a 23B dense model per token, not a 501B one. Reflection claims this gives Beam a 3-4x compute advantage over GLM-5.2 at comparable benchmark scores. But that estimate excludes prompt prefill and attention costs, which grow significantly at long contexts. Test on your actual context lengths.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle