GLM-5.3: What the Open-Weight Coding Gains Actually Cost

GLM-5.3 jumps from 4.6 to 28.3 on Terminal-Bench 3.0 using post-training alone-no new base model. Here's what the benchmark gains hide about token costs and the license change.

Cover art for GLM-5.3: What the Open-Weight Coding Gains Actually Cost

GLM-5.3 is an open-weight model from Z.ai that upgrades GLM-5.2 without touching the underlying base model. That sentence is doing a lot of work. On the harder Terminal-Bench 3.0 eval, the jump from GLM-5.2 to GLM-5.3 is 4.6 to 28.3 - nearly sixfold. Every point of that gain came from a new post-training pass. The weights landed on Hugging Face on August 25, and the hardware and licensing math changed in ways the launch headlines mostly skipped.

What the benchmark table actually shows

The gains are concentrated in long-horizon, agentic coding tasks - exactly the framing Z.ai uses, calling GLM-5.3 stronger specifically at "complex coding and long-horizon tasks" with open-source state-of-the-art results on Terminal-Bench 3.0 and Agents' Last Exam.

The full picture across the model card:

Benchmark GLM-5.2 GLM-5.3 Change
Terminal-Bench 3.0 4.6 28.3 +515%
DeepSWE 46.2 66.9 +45%
SWE-Marathon 19.4 42.5 +119%
FrontierSWE 67.5 78.1 +16%
AutomationBench 26.2 48.2 +84%
Terminal-Bench 2.1 81.0 88.2 +9%

Terminal benchmarks measure whether a model can chain shell commands, read output, and recover from errors across many steps

  • the kind of multi-step recovery that makes the difference between a coding agent that finishes a task and one that stalls waiting for a human. GLM-5.2 was weak there. Going from near-bottom to first among open models in one point release changes which model you reach for in agent pipelines.

The honest caveat: the two most quotable claims - the 50% coding gain and the doubled SWE-Marathon score - are Zhipu internal numbers with no public harness to reproduce them. That doesn't make them false, but they stay unverified until independent evaluators rerun the suites.

Most named-suite GLM-5.3 scores were run or assembled by Z.ai. Public tasks do not make a vendor run independent.

The verbosity problem and what it does to your bill

Z.ai list pricing is $1.40/M input and $4.40/M output. That sounds cheap. It is not cheap per task.

GLM-5.2 averaged about 43K output tokens per Index task, with 37K of them pure reasoning - versus 16K for GPT-5.5. At $4.40/M output, 43K tokens costs roughly $0.19 per task in output alone. At volume - say 2,000 agent tasks per day - that's $380/day, or about $139,000/year, before input costs. The cheap per-token price is real, but effective cost-per-task is closer to the frontier than the sticker suggests.

Under the "Max" effort level, GLM-5.2 pushes to peak intelligence but utilizes nearly 85K output tokens per task. Switching to "High" effort sacrifices only a few points in performance while effectively halving the required token output

  • a real lever if you're running it in high-volume loops.
28.3Terminal-Bench 3.0 scoreup from 4.6 on GLM-5.2
43Kavg output tokens per taskat High reasoning effort (GLM-5.2)
$1.40 / $4.40input / output per M tokensunchanged from GLM-5.2 to GLM-5.3
8× B200sneeded for full 1M-context servingper Z.ai's own recipe
Beagle in action#eng-infra, 11:02am
The ask
'should we switch the agent pipeline to GLM-5.3 or stick with GLM-5.2?'
Beagle drafts
pulls the model card, pricing page, and Terminal-Bench table; drafts a reply with the per-task cost delta at the team's current task volume
You approve
you approve the number-checked summary; the team has a decision, not a research task
Do this in your workspace →

GLM-5.3 vs GLM-5.2: the license change matters more than the benchmarks

GLM-5.2 shipped MIT - full stop, no strings. GLM-5.3 ships under a custom GLM-5.3 License.

The flagship weights are permitted for commercial use in most cases, but any Model-as-a-Service business clearing $10 billion in revenue over a trailing 12-month period has to pass a Z.ai security review before continuing commercial use.

That's a narrower restriction than the user-count clauses in licenses like Llama's, and it won't affect the overwhelming majority of teams self-hosting GLM-5.3 - but read the actual license text on the Hugging Face repo before you build a product around it if you're anywhere near that revenue tier.

GLM-5.3-Flash is 320B and MIT licensed, which makes it the easier self-host on both hardware and terms. GLM-5.2 remains MIT, so it stays a valid choice for a team that needs permissive terms.

GLM-5.2 GLM-5.3 (flagship) GLM-5.3-Flash
Parameters 753B ~744B 320B
Active per token ~40B ~40B ~18B
License MIT Custom (GLM-5.3) MIT
Context 1M tokens 1M tokens 1M tokens
Multimodal No No Yes
Min GPUs (FP8) 4-5 × 80 GB 4-5 × 80 GB smaller
Full-context recipe - 8 × B200 -
API input/output price $1.40 / $4.40 $1.40 / $4.40 lower

The self-hosting math is straightforward: GLM-5.3 shipped at ~753B parameters, against GLM-5.2's 744B. That difference changes almost nothing - the same 4-bit build still lands in the 370-400 GB range and needs four to five 80 GB datacenter cards at minimum.

The full-context recipe uses eight B200s.

The API stayed at $1.40/M input, $4.40/M output, $0.26/M cached. At those prices, $64,000 of GPUs buys you roughly 14 billion output tokens - decades of personal usage. The hardware case for GLM-5.3 is privacy and always-on agents, not economics.

The emergent cyber capability nobody planned for

As Z.ai scaled post-training, cyber capability developed faster than they expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks.

That's Z.ai's own model card talking. It picked up a large, apparently unplanned boost in cyber capability, particularly on exploitation-chain tasks. A team evaluating GLM-5.3 for internal code review or dependency auditing should read that as a feature. A team deploying it in a less controlled setting should read it as a risk surface that wasn't part of the original design.

The pattern here is worth noting beyond this specific model: post-training at scale produces emergent behavior that pre-training alone does not predict. Z.ai did a more extensive safety review before releasing the weights than for any prior version - hence the two-week gap between the August 14 API launch and the August 25 weight release. Whether that review was sufficient is not knowable from the outside. It is knowable that the lab flagged the issue rather than ignoring it.

Choosing between GLM-5.2 and GLM-5.3 for an agentic coding pipeline
Without Beagle
you default to GLM-5.2 because MIT is clean, ignore the Terminal-Bench 3.0 gap, and miss that Flash closes most of it at MIT and half the parameter count
With Beagle
you benchmark your actual task distribution against Terminal-Bench 3.0 task types, run GLM-5.3-Flash on MIT, and escalate to the flagship only when the task complexity warrants the license review

A teammate like Beagle can surface the relevant model card sections and benchmark tables in-thread when your team is mid-decision - the kind of lookup that otherwise eats a Slack thread for thirty minutes.

GLM-5.3 open-weight coding model: common questions

What is genuinely new in GLM-5.3 versus GLM-5.2?

GLM-5.3 reuses the GLM-5.2 base model entirely; every improvement comes from a new post-training pass. The largest reported gains are on long-horizon, agentic coding evals: Terminal-Bench 3.0 jumps from 4.6 to 28.3, SWE-Marathon more than doubles. Single-turn code completion scores improve modestly. If your workload is short completions, GLM-5.2 is nearly identical.

Is GLM-5.3 actually open source?

GLM-5.3 is open-weight under a custom license, not OSI open source. For individuals and any business under $10B in trailing-12-month revenue, the permissions match MIT in practice: use, modify, fine-tune, distribute, sell. GLM-5.2 and GLM-5.3-Flash carry full MIT licenses with no such trigger.

How much does self-hosting GLM-5.3 actually cost?

At 4-bit quantization, the model lands in the 370-400 GB range and needs four to five 80 GB datacenter cards at minimum.

The full-context recipe uses eight B200s.

The practical answer for most teams: start with the API, measure task success, and self-host only when control or sustained utilization justifies the eight-GPU complexity.

What is Terminal-Bench and why does the 3.0 score matter?

Terminal-Bench measures multi-step terminal agent performance - chaining shell commands, reading output, and recovering from errors without human intervention. GLM-5.2 scored 4.6 on version 3.0, which is near the bottom of the open-weight field. On Terminal-Bench 3.0, the gain to GLM-5.3 is dramatic: 28.3 versus 4.6. That gap suggests GLM-5.2 struggled badly with whatever harder task distribution 3.0 introduces, and GLM-5.3's post-training specifically closed it.

Should teams worried about output token costs use GLM-5.3?

Yes - with the "High" effort mode, not "Max." Switching from Max to High effort sacrifices only a few benchmark points while effectively halving the required token output. At $4.40/M output tokens, that halving is significant at any real agent task volume. Run High as your default; reserve Max for the tasks where accuracy loss is measurable and costly.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle