GLM-5.3 is an open-weight model from Z.ai that upgrades GLM-5.2 without touching the underlying base model. That sentence is doing a lot of work. On the harder Terminal-Bench 3.0 eval, the jump from GLM-5.2 to GLM-5.3 is 4.6 to 28.3 - nearly sixfold. Every point of that gain came from a new post-training pass. The weights landed on Hugging Face on August 25, and the hardware and licensing math changed in ways the launch headlines mostly skipped.
What the benchmark table actually shows
The gains are concentrated in long-horizon, agentic coding tasks - exactly the framing Z.ai uses, calling GLM-5.3 stronger specifically at "complex coding and long-horizon tasks" with open-source state-of-the-art results on Terminal-Bench 3.0 and Agents' Last Exam.
The full picture across the model card:
| Benchmark | GLM-5.2 | GLM-5.3 | Change |
|---|---|---|---|
| Terminal-Bench 3.0 | 4.6 | 28.3 | +515% |
| DeepSWE | 46.2 | 66.9 | +45% |
| SWE-Marathon | 19.4 | 42.5 | +119% |
| FrontierSWE | 67.5 | 78.1 | +16% |
| AutomationBench | 26.2 | 48.2 | +84% |
| Terminal-Bench 2.1 | 81.0 | 88.2 | +9% |
Terminal benchmarks measure whether a model can chain shell commands, read output, and recover from errors across many steps
- the kind of multi-step recovery that makes the difference between a coding agent that finishes a task and one that stalls waiting for a human. GLM-5.2 was weak there. Going from near-bottom to first among open models in one point release changes which model you reach for in agent pipelines.
The honest caveat: the two most quotable claims - the 50% coding gain and the doubled SWE-Marathon score - are Zhipu internal numbers with no public harness to reproduce them. That doesn't make them false, but they stay unverified until independent evaluators rerun the suites.
Most named-suite GLM-5.3 scores were run or assembled by Z.ai. Public tasks do not make a vendor run independent.
The verbosity problem and what it does to your bill
Z.ai list pricing is $1.40/M input and $4.40/M output. That sounds cheap. It is not cheap per task.
GLM-5.2 averaged about 43K output tokens per Index task, with 37K of them pure reasoning - versus 16K for GPT-5.5. At $4.40/M output, 43K tokens costs roughly $0.19 per task in output alone. At volume - say 2,000 agent tasks per day - that's $380/day, or about $139,000/year, before input costs. The cheap per-token price is real, but effective cost-per-task is closer to the frontier than the sticker suggests.
Under the "Max" effort level, GLM-5.2 pushes to peak intelligence but utilizes nearly 85K output tokens per task. Switching to "High" effort sacrifices only a few points in performance while effectively halving the required token output
- a real lever if you're running it in high-volume loops.
GLM-5.3 vs GLM-5.2: the license change matters more than the benchmarks
GLM-5.2 shipped MIT - full stop, no strings. GLM-5.3 ships under a custom GLM-5.3 License.
The flagship weights are permitted for commercial use in most cases, but any Model-as-a-Service business clearing $10 billion in revenue over a trailing 12-month period has to pass a Z.ai security review before continuing commercial use.
That's a narrower restriction than the user-count clauses in licenses like Llama's, and it won't affect the overwhelming majority of teams self-hosting GLM-5.3 - but read the actual license text on the Hugging Face repo before you build a product around it if you're anywhere near that revenue tier.
GLM-5.3-Flash is 320B and MIT licensed, which makes it the easier self-host on both hardware and terms. GLM-5.2 remains MIT, so it stays a valid choice for a team that needs permissive terms.
| GLM-5.2 | GLM-5.3 (flagship) | GLM-5.3-Flash | |
|---|---|---|---|
| Parameters | 753B | ~744B | 320B |
| Active per token | ~40B | ~40B | ~18B |
| License | MIT | Custom (GLM-5.3) | MIT |
| Context | 1M tokens | 1M tokens | 1M tokens |
| Multimodal | No | No | Yes |
| Min GPUs (FP8) | 4-5 × 80 GB | 4-5 × 80 GB | smaller |
| Full-context recipe | - | 8 × B200 | - |
| API input/output price | $1.40 / $4.40 | $1.40 / $4.40 | lower |
The self-hosting math is straightforward: GLM-5.3 shipped at ~753B parameters, against GLM-5.2's 744B. That difference changes almost nothing - the same 4-bit build still lands in the 370-400 GB range and needs four to five 80 GB datacenter cards at minimum.
The full-context recipe uses eight B200s.
The API stayed at $1.40/M input, $4.40/M output, $0.26/M cached. At those prices, $64,000 of GPUs buys you roughly 14 billion output tokens - decades of personal usage. The hardware case for GLM-5.3 is privacy and always-on agents, not economics.
The emergent cyber capability nobody planned for
As Z.ai scaled post-training, cyber capability developed faster than they expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks.
That's Z.ai's own model card talking. It picked up a large, apparently unplanned boost in cyber capability, particularly on exploitation-chain tasks. A team evaluating GLM-5.3 for internal code review or dependency auditing should read that as a feature. A team deploying it in a less controlled setting should read it as a risk surface that wasn't part of the original design.
The pattern here is worth noting beyond this specific model: post-training at scale produces emergent behavior that pre-training alone does not predict. Z.ai did a more extensive safety review before releasing the weights than for any prior version - hence the two-week gap between the August 14 API launch and the August 25 weight release. Whether that review was sufficient is not knowable from the outside. It is knowable that the lab flagged the issue rather than ignoring it.
A teammate like Beagle can surface the relevant model card sections and benchmark tables in-thread when your team is mid-decision - the kind of lookup that otherwise eats a Slack thread for thirty minutes.
GLM-5.3 open-weight coding model: common questions
What is genuinely new in GLM-5.3 versus GLM-5.2?
GLM-5.3 reuses the GLM-5.2 base model entirely; every improvement comes from a new post-training pass. The largest reported gains are on long-horizon, agentic coding evals: Terminal-Bench 3.0 jumps from 4.6 to 28.3, SWE-Marathon more than doubles. Single-turn code completion scores improve modestly. If your workload is short completions, GLM-5.2 is nearly identical.
Is GLM-5.3 actually open source?
GLM-5.3 is open-weight under a custom license, not OSI open source. For individuals and any business under $10B in trailing-12-month revenue, the permissions match MIT in practice: use, modify, fine-tune, distribute, sell. GLM-5.2 and GLM-5.3-Flash carry full MIT licenses with no such trigger.
How much does self-hosting GLM-5.3 actually cost?
At 4-bit quantization, the model lands in the 370-400 GB range and needs four to five 80 GB datacenter cards at minimum.
The full-context recipe uses eight B200s.
The practical answer for most teams: start with the API, measure task success, and self-host only when control or sustained utilization justifies the eight-GPU complexity.
What is Terminal-Bench and why does the 3.0 score matter?
Terminal-Bench measures multi-step terminal agent performance - chaining shell commands, reading output, and recovering from errors without human intervention. GLM-5.2 scored 4.6 on version 3.0, which is near the bottom of the open-weight field. On Terminal-Bench 3.0, the gain to GLM-5.3 is dramatic: 28.3 versus 4.6. That gap suggests GLM-5.2 struggled badly with whatever harder task distribution 3.0 introduces, and GLM-5.3's post-training specifically closed it.
Should teams worried about output token costs use GLM-5.3?
Yes - with the "High" effort mode, not "Max." Switching from Max to High effort sacrifices only a few benchmark points while effectively halving the required token output. At $4.40/M output tokens, that halving is significant at any real agent task volume. Run High as your default; reserve Max for the tasks where accuracy loss is measurable and costly.