Qwen3.8-Max Open Weights Ship, But Read the Fine Print

Alibaba released Qwen3.8-2.4T-A95B weights on August 12 - the first Qwen Max-class model you can download. Here's what the open weights include, what they don't, and what independent testing found.

Cover art for Qwen3.8-Max Open Weights Ship, But Read the Fine Print

The three previous Qwen Max models - 3.5-Max, 3.6-Max, 3.7-Max - never shipped weights at all. All three stayed API-only. On August 12, 2026, that changed: Alibaba published the checkpoint for Qwen3.8-2.4T-A95B through its ModelScope community account, making this the first Qwen-Max-tier model to receive an open-weight release. That is genuinely new. What those weights contain - and what they quietly leave out - is the part worth reading carefully before you plan a deployment around them.

What Qwen3.8-Max actually is

Qwen3.8-Max is a 2.4-trillion-parameter mixture-of-experts (MoE) model with 95B active parameters, a 1M-token context window, and native text, image, and video input. That is the hosted product. Qwen3.8-2.4T-A95B is the open-weight base model - 2.4T parameters, 95B active, MoE architecture, 256K native context - while Qwen3.8-Max is Alibaba's cloud version built on this model, augmented with production post-training. The open weights give you the base; the API gives you the enhanced version.

That distinction matters immediately for three things:

  • Vision input. The downloadable checkpoint is text-only. The hosted API stays multimodal with 1M context and built-in tools at $2/$6 per 1M; weights are text-only with thinking required-on.

  • Thinking mode. Qwen3.8-2.4T is thinking-only, while Qwen3.8-Max is hybrid. If you need the model to answer without a chain-of-thought trace - for latency-sensitive pipelines, or just to save output tokens - that option isn't in the weights.

  • Context window. The model's native context is 262,144 tokens, while the 1 million figure is a ceiling reached through extension.

None of this makes the release bad. It makes it different from what the launch headlines implied.

The "95B active" figure doesn't shrink the download

Here is the number that trips teams up. The "95B active" number reduces work per token; it does not turn a 2.4-trillion-parameter checkpoint into a 95-billion-parameter download.

The official BF16 weights total 4.892TB. Official FP8 is 2.496TB. The current four-bit files are roughly 1.31-1.45TB. Even Unsloth's smallest posted one-bit GGUF is 397GB before runtime overhead.

Even at aggressive quantization, the full weight set doesn't fit on a single consumer GPU or even a single high-end workstation. It's built for serving infrastructure using vLLM or SGLang, not for a laptop or a single RTX card.

For reference, INT4 at 1M context breaks down like this:

Component Size
Weights (Q4, all experts resident) ~1,344 GB
KV cache at 1M tokens ~328 GB
Runtime overhead ~167 GB
Total resident ~1,839 GB
Active-expert offload (weights only ~53 GB) ~419 GB total

Active-only cuts weights to 53.2 GB and the total to 419 GB, which fits four H200s - but the per-token PCIe penalty of cold-expert fetch hurts throughput.

NVIDIA's NIM guide for Qwen3.8-2.4T-A95B is unusually clear: only GB300-NVL72 hardware is supported in that release, Kubernetes is required, and deployment needs at least four nodes.

4.89 TBBF16 weightsofficial checkpoint size on Hugging Face
~1.3 TB4-bit quantizedsmallest practical serving footprint (experts resident)
$2 / $6input / output per 1M tokenshosted API, Alibaba Model Studio
262Knative context in the weightsvs 1M in the hosted API

What the benchmarks actually showed

The headline claim is that Qwen3.8-Max scores 86.1 on OSWorld-Verified, ahead of GPT-5.6 Sol Max at 83.2, Claude Fable 5 at 85.0, and Gemini 3.1 Pro at 76.2. Those numbers look strong.

Qwen's static benchmarks replicated; its agentic benchmarks did not. That is now a pattern across three vendors in two weeks. When a launch chart leads with a long-horizon agentic number, assume the harness is doing some of the work until someone neutral reproduces it.

Qwen3.8-Max is a specialist wearing a generalist's launch chart. It is genuinely excellent at algorithmic coding and tool use, priced well, and licensed generously. It is mediocre at agentic software engineering, unreliable on facts, and expensive per completed job.

That last point deserves its own sentence. Despite $2/$6 tokens, it is the second-most-expensive of 20 models per completed task, because it burns far more tokens getting to an answer in forced-thinking mode. Thinking-always-on is fine when reasoning depth is the point; it is expensive overhead when you just want to classify a support ticket.

Beagle in action#eng-ops, 10:41am
The ask
'can someone pull the SWE-bench score for Qwen3.8-Max from the vendor page and compare it to the neutral harness result?'
Beagle drafts
finds the Alibaba launch post (82.3 on SWE-bench Verified) and the independent re-run result, drafts a side-by-side with a note on harness differences
You approve
you approve; the comparison posts in-thread with both source links so the eval discussion has something concrete to argue about
Do this in your workspace →

The 27B is the actual self-host story

Unlike the 3.7 generation, which never shipped open weights, Qwen 3.8 now offers both deployment models: Qwen3.8-Max as a frontier API at aggressive per-token pricing, and Qwen3.8-27B as Apache 2.0 weights that serve from a single GPU.

The 27B version went live on Hugging Face on August 14, 2026 as Qwen/Qwen3.8-27B and Qwen/Qwen3.8-27B-FP8. Apache 2.0 is confirmed in the actual LICENSE file, not only in the model card front matter. Commercial use is permitted and there is no revenue share clause.

Qwen3.8-27B is the self-host tier: Apache 2.0, 28B dense with a vision encoder, roughly 56GB of VRAM at BF16, 28GB at FP8, or 14 to 17GB quantized, serving through vLLM or SGLang on a single rented GPU.

The 27B version leads on all four benchmarks where it is compared against its predecessor. The jumps on Terminal Bench and SWE-bench Pro stand out, because neither is one-shot question answering - both measure multi-step agent behavior, and going from 51.7 to 73.0 is a serious difference on that axis.

The practical split most teams land on: the big model through an API for the hardest reasoning, the small model self-hosted for volume work. That is a reasonable default. A teammate like Beagle routing triage-level questions to the 27B and escalating complex reasoning to the Max API would spend the token budget where it earns it.

Evaluating Qwen3.8 for a production agent workflow
Without Beagle
read the launch post, see "open weights," plan a self-hosted 2.4T deployment, discover the hardware requirement is 15+ H100s at INT4 three weeks later
With Beagle
check the checkpoint size, confirm the text-only limitation, route hard jobs to the API and volume jobs to the 27B on a single GPU

The license for the Max-tier weights also needs a read before production use. Keep copyright and permission notice; above 100M MAU or $20M monthly revenue, prominently display the model name; MaaS or AI Work Assistant businesses above $50M TTM aggregate revenue need a separate Qwen license, with an internal-use carve-out if you do not expose the model, outputs, or capabilities to third parties. For most teams that is not a problem. For teams building external-facing AI products, it is worth a lawyer's eye.

Qwen3.8-Max open weights: common questions

What is Qwen3.8-2.4T-A95B and how does it differ from Qwen3.8-Max?

Qwen3.8-2.4T-A95B is the downloadable open-weight checkpoint: 2.4 trillion total parameters, 95 billion active per token, text-only, thinking mode always on, native context 262K. Qwen3.8-Max is the hosted API built on this checkpoint, adding vision input, non-thinking mode, 1M-token context, and built-in tools. They share weights but are different products.

Can you self-host Qwen3.8-Max on a single server?

No. At 4-bit quantization the weights alone occupy roughly 1.3TB, requiring 15 or more A100 80GB GPUs just for weight storage before KV cache and runtime overhead. The practical self-host option in this family is Qwen3.8-27B, which runs on a single GPU at 14-17GB quantized under Apache 2.0.

How accurate are the Qwen3.8-Max benchmark numbers?

Static coding benchmarks replicated in independent testing; agentic and long-horizon numbers did not. The OSWorld-Verified and coding scores are credible. Multi-step software engineering scores were lower under neutral harnesses than in Alibaba's launch comparisons. Validate against your own tasks before committing.

Does the Qwen3.8-Max license allow commercial use?

The custom Qwen3.8-Max License permits commercial use with conditions. Attribution is required at scale; MaaS or AI Work Assistant businesses above $50M TTM revenue need a separate license. Internal use that doesn't expose the model to third parties has a carve-out. The 27B version ships under Apache 2.0 with no such restrictions.

What is the API price for Qwen3.8-Max?

The Model Studio API charges $2 per million input tokens and $6 per million output tokens, with cached input at $0.25 per million. The hosted API is the economically sensible path for most teams: running the 2.4T checkpoint yourself costs more per token in GPU hours than the API price once you factor in idle time.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle