Read the Hermes 4.3 Model Card Before You Deploy It

Hermes 4.3 is the first open-weight model trained on a decentralized GPU network secured by a blockchain - and it outperformed the centrally-trained version. Here's what the numbers actually say.

Cover art for Read the Hermes 4.3 Model Card Before You Deploy It

Hermes 4.3 scores 93.8% on MATH-500 and 65.5% on GPQA Diamond

  • from a 36B model trained not on a private GPU cluster but across nodes scattered over the open internet. That is the actual news in this release, and it is more interesting than the benchmark scores alone.

Hermes 4.3 is the first Hermes model post-trained entirely on the Psyche network, a distributed training system that uses the DisTrO optimizer to communicate between training nodes spread across data centers over the open internet, secured by Solana blockchain consensus. That sentence has a lot in it. Unpack it before you decide whether it matters.

What Psyche actually changed about training

Decentralized training is not new as a research idea, but making it work at production quality is. The standard objection is bandwidth: gradient synchronization across data centers is slow, which means either you accept worse compute efficiency or you need a smarter optimizer.

Nous Research trained Hermes 4.3 both on Psyche and via the traditional centralized approach specifically to verify effectiveness for production workloads, and the Psyche-trained version outperformed the centralized version on a suite of downstream tasks. That is a meaningful result. It rules out the easy criticism that decentralized training is a clever demonstration that quietly underperforms.

By enabling nodes throughout the world to collaborate on a single training run, Psyche can dramatically reduce the cost of training frontier-level models, leveling the playing field for open-source AI model developers. "Dramatically" is a claim without a number attached. The blog post does not publish a cost comparison between the two runs, which is the figure that would actually settle the question. Take the efficiency framing as directionally plausible, not verified.

The Psyche decentralized training network is clearly an active investment. Hermes 4.3 being explicitly the first model trained this way suggests Nous Research is treating Psyche as a production training infrastructure path, not a one-off experiment. Subsequent model releases will likely also use Psyche, and the decentralized training approach may enable larger parameter counts or faster iteration cycles as the network grows.

What the corpus expansion actually buys you

The post-training corpus grew from 1M samples and 1.2B tokens to approximately 5M samples and ~60B tokens, blended across reasoning and non-reasoning data. That is a 50× expansion in token count. In the Hermes line, corpus scale has historically correlated with structured output quality - the model card emphasizes schema adherence and JSON repair as specific gains.

Reasoning mode can be activated via the chat template with the flag thinking=True or by a system prompt, and system instructions before or after the reasoning system message adjust the model's policies, style, and effort of thinking.

That toggle is the part most teams will actually use. Turning reasoning on costs tokens - the <think>...</think> block runs before the final answer. For a 36B model at self-hosted inference rates, those extra tokens are cheap. On managed API endpoints they add up. The practical move is to route simple classification or retrieval calls with thinking off, and enable it for multi-step planning or code review. A teammate like Beagle, sitting between the user and the model, can apply that routing rule automatically based on task type.

Beagle in action#engineering, 2:47pm
The ask
'can someone summarize the PR diff and flag anything that touches auth?'
Beagle drafts
routes to Hermes 4.3 with thinking=True for the auth-flag step, thinking=False for the summary
You approve
draft posts the plain summary plus a flagged auth block, with source lines cited - you approve before it hits the channel
Do this in your workspace →

The RefusalBench number and why it needs a footnote

Hermes 4.3 achieves state of the art on RefusalBench across all popular closed and open models in being helpful and aligned to the user's values. This is the claim Nous Research leads with in the announcement. The footnote: RefusalBench is a benchmark created by Nous Research itself, designed to test a model's willingness to be helpful in scenarios commonly disallowed by other models.

Hermes 4.3 scored 74.6% on RefusalBench - meaning it answered 74.6% of questions that other aligned models refuse - compared to 59.5% for Hermes 4 70B. The improvement is real. What it means depends entirely on what questions are in the set. Nous Research publishes the full eval responses for inspection, which is the right move, but "SOTA on RefusalBench" in a press headline is a lab grading its own homework.

For teams where over-refusal has caused real workflow problems - code generation that declines to touch security-adjacent logic, agents that refuse to summarize internal HR documents - this is a concrete and checkable improvement. For teams where content guardrails are a compliance requirement, a high RefusalBench score is a feature that needs a second look, not a selling point.

93.8%MATH-500Hermes 4.3 36B, from model card
74.6%RefusalBenchvs 59.5% on Hermes 4 70B
~60B tokenspost-training corpus50× larger than Hermes 4 baseline
$0.13 / $0.40per 1M in/outHermes 4 70B on OpenRouter

Pick the right variant for your workload

Three active Hermes variants are worth comparing directly:

Variant Params Base API (input/output per 1M) When to use
Hermes 4 70B 70B Llama 3.1-70B $0.13 / $0.40 Structured outputs, agent tool calls, cost-sensitive pipelines
Hermes 4 405B 405B Llama 3.1-405B $1.00 / $3.00 Hard reasoning, high-stakes generation, max context
Hermes 4.3 36B 36B ByteDance Seed 36B self-host or Nous Portal On-prem, privacy-first, hybrid reasoning on your hardware

Running a 405B model requires either a multi-GPU cluster or cloud inference; most production deployments use the 70B or smaller variants for cost and latency reasons.

Hermes 4.3 36B outperforms the larger Hermes 4 70B on several benchmarks according to the model card

  • which matters if you are choosing between them for a self-hosted setup. Half the parameters, better scores on some tasks: that is a real pareto improvement, driven by the larger post-training corpus rather than raw scale.

The one honest caveat: open-weight models are maintaining a consistent 3-6 month capability gap behind closed frontier labs, and that gap has held for over 18 months. Hermes 4.3 is competitive at the 36B tier, not at the absolute frontier. If your workload needs GPT-5.5-class reasoning, this is not that. If it needs a steerable, structured-output-capable model that runs on hardware you control, it is worth testing this week.

Running structured-output agent tasks
Without Beagle
a frontier API at $3/M output tokens, every agent step billed, refusal on edge-case prompts slows iteration
With Beagle
Hermes 4.3 36B self-hosted, thinking toggle per task type, no per-token bill above compute cost

Hermes 4.3 open-weight model: common questions

What is Hermes 4.3 and how is it different from Hermes 4?

Hermes 4.3 is a 36B hybrid-reasoning model from Nous Research built on ByteDance's Seed 36B base, released August 2025. The key difference from Hermes 4 is that it was trained on the Psyche decentralized network rather than a centralized GPU cluster, and its post-training corpus grew 50× - from 1.2B to ~60B tokens.

Can Hermes 4.3 replace a frontier model like Claude or GPT-4 in an agent pipeline?

For structured outputs, tool calling, and reasoning tasks in the 36B weight class, yes - with caveats. It trails frontier closed models on absolute capability by roughly 3-6 months according to open-weight tracking data. Where it wins is cost, steerability, and on-premise deployability for teams with data-residency constraints.

What is Psyche and does decentralized training actually work?

Psyche is Nous Research's distributed training network that uses the DisTrO optimizer to sync gradients across nodes over the open internet, with Solana blockchain handling consensus. The Hermes 4.3 model card confirms the Psyche-trained version outperformed the centrally-trained control on downstream tasks - the first published evidence that this approach works at production quality.

Is the RefusalBench score a reliable benchmark?

Treat it cautiously. RefusalBench was created by Nous Research and measures whether the model answers questions other aligned models decline. The 74.6% score is a real improvement over prior Hermes versions, and the full eval responses are published. But a lab-built benchmark measuring the lab's own model's willingness to answer sensitive questions is not independent verification.

What hardware do you need to self-host Hermes 4.3 36B?

A single high-VRAM GPU (an A100 80GB or two 40GB A100s in NVLink) handles 36B at full precision. In 4-bit GGUF quantization - available on HuggingFace as NousResearch/Hermes-4.3-36B-GGUF - it fits on consumer-grade hardware with 24GB VRAM, which changes the cost math substantially for small teams.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle