Hermes 4 Hybrid Reasoning: What Actually Changed

Nous Research's Hermes 4 and 4.3 introduced a toggleable reasoning mode, decentralized training via Psyche, and a deliberately low refusal rate. Here's what that means in practice.

Cover art for Hermes 4 Hybrid Reasoning: What Actually Changed

Hermes 4 405B hit 96.3% on MATH-500 when it launched on August 26, 2025. That number got copied into every coverage headline within a day. Most of those headlines missed the actual story - which is not the leaderboard score, but the three structural decisions underneath it that make Hermes 4 different from anything the open-weight field had shipped before.

This is a post about those three decisions, and what they mean for teams thinking about running open-weight models in production.

What Hermes 4 actually is (not just Llama with better math)

Hermes 4 405B is not "Llama 3.1 but better at math." It is a single checkpoint that switches between standard inference and chain-of-thought reasoning via a <think> toggle - and the 96.3% MATH-500 score is the reasoning-on number, which most hosted deployments do not enable by default.

That distinction matters a lot for cost and latency planning. Nous Research claims hybrid mode raises MATH-500 accuracy from 93.1% to 96.3% on the 405B tier. In direct mode, the model skips the chain-of-thought entirely. Direct mode improves latency up to 28% according to early Ollama telemetry. So you are not choosing between a reasoning model and a non-reasoning model - you are choosing per call, on the same checkpoint. That is architecturally different from how OpenAI structures its o-series, where reasoning is baked into a separate model family.

The trademark of the Hermes series is tool calling, and Hermes 4 keeps it intact. The format is XML: the signatures of available functions are described inside <tools> tags in the system message, and when the model decides to use one, it emits a <tool_call> object with the name and arguments. This scheme is identical to Hermes 3, so agent code written for the previous version works unchanged. That backward compatibility is quietly valuable - it means an existing Hermes 3 agent pipeline can upgrade without touching a prompt template.

The post-training corpus was massively increased from 1M samples and 1.2B tokens to ~5M samples / ~60B tokens blended across reasoning and non-reasoning data.

Hermes 4 training leveraged a modified TorchTitan across 192 NVIDIA B200 GPUs.

96.3%MATH-500 (reasoning mode)Hermes 4 405B, August 2025
70.5%GPQA Diamond405B, open-weight SOTA at launch
60B tokenspost-training corpusup from 1.2B in prior generation
$0.13/Minput tokens, 70B via OpenRouterversus $2-15/M for frontier closed APIs

The refusal design and why it is a production concern

Particularly notable is Hermes 4's performance on RefusalBench, achieving 57.1% in reasoning mode - the highest score among evaluated models, significantly outperforming GPT-4o (17.67%) and Claude Sonnet 4 (17%).

Nous frames this as a feature. According to Nous Research, Hermes 4 aims to be a "neutrally aligned" model that obeys the user rather than rejecting legitimate requests. And for many use cases - creative writing, internal tooling, research assistants where the operator trusts the users - that is genuinely useful.

But there is a specific mechanism worth understanding before you deploy. The 57.1% is also a reasoning-mode number. The model literally reasons itself into compliance with requests it would otherwise refuse. The <think> block processes the request, finds a frame where engagement is defensible, and then responds. That is a qualitatively different failure mode than a model that is simply undertrained on refusal.

The practical implication: deploying Hermes 4 without your own safety layer and assuming the model's native alignment is sufficient is a mistake - not because the model is broken, but because it was not designed for that assumption. The neutral stance is a design choice, not a gap in training. You own the risk layer here.

This is not a knock on the model. It is an honest description of the tradeoff Nous made. Teams building internal agents where operators control the prompts can lean into the steerability. Teams building consumer-facing products need a separate guardrail layer regardless.

Beagle in action#engineering, 11:02am
The ask
'should we use Hermes 4 or keep using the OpenAI API for our internal code review bot?'
Beagle drafts
pulls the Hermes 4 pricing from OpenRouter ($0.13/M input, $0.40/M output on 70B), compares to the team's current GPT-4 spend, and drafts a cost comparison with a note on the refusal-rate tradeoff
You approve
you review the draft, add the guardrail requirement to the decision doc, post to the channel in one edit
Do this in your workspace →

Hermes 4.3: the decentralized training story

Hermes 4.3 was released August 25, 2025, and is built on ByteDance's Seed 36B base model. The notable departure from prior Hermes releases: it was trained using Nous Research's Psyche decentralized training network rather than a traditional centralized GPU cluster. The model card explicitly calls this out as the first Hermes model trained this way.

Hermes 4.3 is the first production model post-trained entirely on the Psyche network, a distributed training network that uses the DisTrO optimizer to efficiently communicate between training nodes spread out through data centers over the open internet and secured by the consensus of the Solana blockchain.

The key question about any decentralized training claim is whether it actually works, or whether it is a research curiosity. Nous addressed this directly. They trained Hermes 4.3 both on Psyche and via the traditional centralized approach, using the same recipe (FSDP+AdamW) for an apples-to-apples comparison, then trained the model a second time on Psyche.

The training run proved stable throughout, averaging 144k tokens/second spread across 24 Psyche nodes. Using DisTrO's overlapped collective strategy, the entirety of the P2P communications were hidden by the training time, effectively achieving equivalent throughput to traditional, centralized training.

The Psyche-trained version of Hermes 4.3 outperformed the traditional centralized version.

That last line deserves to sit alone for a moment. A model trained across internet-connected nodes outperformed one trained on a private cluster using identical data and recipe. That is not a theoretical outcome - it is a concrete benchmark result from a primary source.

Hermes 4.3 was trained with an extended context length of up to 512K tokens and nearly matches, and in some cases exceeds, the performance of Hermes 4.

For teams evaluating the model on pure capability: Hermes 4.3 36B is now SOTA across non-abliterated models on the RefusalBench Leaderboard, surpassing the previous best of 59.5% on Hermes 4 70B. Better steerability, smaller footprint - at 36B it runs on hardware that cannot touch the 70B or 405B variants.

Model Size Context Notable capability Training infrastructure
Hermes 4 14B / 70B / 405B 131K Hybrid reasoning, tool calling Centralized, 192× B200
Hermes 4.3 36B 512K Hybrid reasoning, top RefusalBench Decentralized, 24-node Psyche
Running a complex agentic task with Hermes 4
Without Beagle
engineering team routes every uncertain request to a closed API, paying $5-15/M input tokens with no control over the model's refusal behavior or tool-calling format
With Beagle
self-hosted or via OpenRouter at $0.13/M input, reasoning toggled on per call, tool-call format unchanged from Hermes 3 so the existing agent scaffold needs no edits

What is genuinely new versus incremental

The hype around Hermes 4 focused on the MATH-500 number. That score is real, but it describes a capability the model delivers only when reasoning mode is on - a flag that most API wrappers do not set by default.

The genuinely new things are:

  • A single-checkpoint reasoning toggle. Reasoning-on and reasoning-off are not separate models. One set of weights, controlled by a flag. The cost difference per call is meaningful: developers can disable traces and restore faster, cheaper inference during production calls, because reasoning mode expands outputs substantially.

  • Neutral alignment as a first-class design goal, not an afterthought. The model exhibits reduced "policy rigidity" compared to many proprietary systems. In adversarial roleplay and meta-contexts, Hermes 4 adopts instructed personas and generates immersive in-character output, minimizing default to generic safety warnings. This is deliberate, documented, and changes how you need to think about the safety layer you build around it.

  • Decentralized production training that actually worked. Psyche running Hermes 4.3 at equivalent throughput to centralized training is not incremental. By enabling nodes throughout the world to collaborate on a single training run, Psyche can dramatically reduce the cost of training frontier-level models, leveling the playing field for open-source AI model developers. Whether that scales beyond 36B post-training to full pre-training at frontier scale is the open question - but the infrastructure exists and produced a released model.

What is incremental: the benchmark scores at AIME and GPQA. Nous positioned the 405B as "competitive with DeepSeek R1 and Qwen3 235B," but the community consensus was that DeepSeek R1 at 671B outperforms Hermes 4 on raw reasoning tasks. Hermes 4 is competitive at its size class and price point. It is not a ceiling-raiser on pure math reasoning.

A useful comparison: the 70B variant runs via OpenRouter at $0.13 per million input tokens and $0.40 per million output tokens.

GPT-4 equivalent performance now costs $0.40/million tokens versus $20 in late 2022

  • the market has moved. But even at today's rates, frontier closed models still run $2-15 per million input tokens , a 10-100x gap against the Hermes 70B price. For teams running high-volume internal agents where neutral alignment is a feature and not a liability, that spread is hard to ignore.
Beagle in action#product, 3:45pm
The ask
'we need to compare Hermes 4 70B against our current GPT-4o setup for the code-review agent - can someone pull the cost math?'
Beagle drafts
computes monthly token cost at current usage volume against the $0.13/$0.40 per million input/output rate, flags that reasoning mode should be toggled per-call to avoid output bloat
You approve
the comparison posts in-thread with a recommendation; you approve with an edit noting the guardrail requirement
Do this in your workspace →

Hermes 4 open-weight model: common questions

What is the Hermes 4 hybrid reasoning mode?

Hybrid reasoning lets a single Hermes 4 checkpoint switch between direct response and explicit chain-of-thought reasoning using <think> tags, controlled per call. Reasoning-on improves math and STEM accuracy - MATH-500 goes from 93.1% to 96.3% on the 405B - but expands output tokens significantly. Most hosted deployments do not enable it by default.

How does Hermes 4 compare to DeepSeek R1 and Qwen3?

Hermes 4 405B scores 96.3% on MATH-500 and 70.5% on GPQA Diamond, competitive with Qwen3 235B at a much smaller parameter count. DeepSeek R1 at 671B still outperforms it on raw reasoning by community consensus. Hermes 4 is the better choice when steerability, tool-calling compatibility, and cost matter more than absolute reasoning ceiling.

What is Hermes 4.3 and how is it different from Hermes 4?

Hermes 4.3 is a 36B model built on ByteDance's Seed base, trained on Nous Research's Psyche decentralized network - the first production model to be. It extends context to 512K tokens and outperforms Hermes 4 on RefusalBench. It fits on hardware that cannot run the 70B or 405B variants, making it the practical choice for local or constrained deployment.

Is Hermes 4's low refusal rate safe for production?

That depends entirely on your use case and guardrail stack. Hermes 4's 57.1% RefusalBench score reflects deliberate neutral alignment - the model reasons through requests rather than refusing them by default. For internal tooling with trusted users, that is a feature. For consumer-facing products, you need a separate safety layer. Nous does not provide one; that is by design.

Where can I run Hermes 4 today?

The 70B and 405B are available via OpenRouter at $0.13/M and roughly $1/M input respectively, with FP8 quantized weights on HuggingFace. The 14B runs locally on hardware with 16GB VRAM via LM Studio GGUF builds. Hermes 4.3 36B weights are also available on HuggingFace, with GGUF builds for local inference.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle