Hermes 4's Hybrid Toggle Sounds Simple. The Build Decision Isn't.

Nous Research's Hermes 4 ships a single checkpoint that can reason or not-on demand. Here's what that hybrid mode actually means for teams choosing an open-weight reasoning model.

Cover art for Hermes 4's Hybrid Toggle Sounds Simple. The Build Decision Isn't.

A developer on your team opens the Hermes 4 technical report, sees 96.2% on MATH-500, and drafts a Slack message: "think we can replace our reasoning pipeline with this." What they haven't read yet is the OpenRouter endpoint note, buried below the model description: function calling is unavailable on the 405B hosted endpoint. The benchmark was run with reasoning on. Most hosted deployments don't enable it by default.

That gap - between what a model is trained to do and what a given deployment actually exposes - is the thing worth understanding before you build on Hermes 4.

What hybrid reasoning actually means in a single model

The headline change in Hermes 4 is hybrid reasoning: the model decides whether to answer directly or open a <think>…</think> block to deliberate - something Hermes 3 did not offer. That sounds straightforward, but the implementation detail matters for how you price calls.

A specialized fine-tuning stage introduces an explicit </think> token at a preset reasoning length of 20k-30k tokens, trained via masked loss so the model learns controlled truncation of reasoning traces without degrading answer fidelity. In other words, Nous trained the model not just to think, but to stop thinking at a predictable length - which is the part that makes cost estimation possible rather than a guessing game.

Hybrid reasoning means a single model does two things that used to require separate models: when the question is simple, it answers directly without spending extra tokens; when the task is complex, it opens a deliberation block delimited by <think> and </think> tags, reasons step by step, and only then writes the final answer.

The token cost implication is direct. On OpenRouter, Hermes 4 70B costs $0.13 per million input tokens and $0.40 per million output tokens, with a 131,072-token context window. Reasoning-off calls are cheap. Flip the boolean and you're generating up to 30k tokens of thinking trace before a single output word - those are billed as output tokens. For a latency-sensitive pipeline calling the model hundreds of times a day, that distinction is a line item, not a footnote.

The numbers Nous published, and what they actually describe

In reasoning mode, the 405B posts high scores: 96.2% on MATH-500, 70.6% on GPQA Diamond, and 81.9% on AIME 2024, per Nous's technical report. Those are strong numbers. They are also reasoning-on numbers.

The 96.3% MATH-500 score is the reasoning-on number, which most hosted deployments do not enable by default. So if you're evaluating Hermes 4 against a deployment where reasoning is off, you're not comparing like-for-like with the published figures.

The second number worth isolating is RefusalBench. Hermes 4 scored 57.1% on RefusalBench; GPT-4o scored 17.67%; Claude scored 17%.

RefusalBench measures how often a model will engage with requests that mainstream frontier models refuse; a higher score means the model refuses less. Nous frames this as "neutral alignment" - the model does not have a political or moral stance built into its post-training, and it treats the user as an adult who can decide what to do with a response.

That framing is coherent, and it's also a governance question your team needs to answer before you ship. A model that refuses 57% of what others refuse is genuinely more useful for some applications. It requires deliberate policy at the application layer for others.

96.2%MATH-500 (reasoning on)405B, from Nous technical report
57.1%RefusalBenchversus ~17% for GPT-4o and Claude
~60Bpost-training tokensup from ~1.2B in Hermes 3

The hosted-API gap that the model card doesn't explain

Here is the practical wrinkle. The Hermes 4 405B endpoint on OpenRouter does not accept tools, so function calling is unavailable there. It supports response_format for JSON output, without JSON-schema enforcement.

The model itself is trained for tool use. It supports structured outputs including JSON mode, schema adherence, function calling, and tool use. The gap is between the model's capability and what a single hosted provider has wired up. Hermes 4 uses the Hermes tool format (<tool_call>) and ships automatic parsers built into vLLM and SGLang

  • so if you're self-hosting, function calling works. If you're reaching for the easiest managed endpoint first, on the 405B, it doesn't.

The 70B avoids this problem. It is the variant your team will likely use in production anyway - running a 405B model requires either a multi-GPU cluster or cloud inference; most production deployments use the 70B or smaller variants for cost and latency reasons.

Evaluating Hermes 4 for a tool-calling pipeline
Without Beagle
spin up the 405B on OpenRouter, send a tool call, watch it silently fail - trace the issue across provider docs and the model card before finding the endpoint note
With Beagle
read the OpenRouter endpoint page before writing code; route tool-calling workloads to the 70B, reserve the 405B for pure reasoning tasks where you control hosting

What's genuinely new versus incremental

Nous greatly expanded the fine-tuning data. The post-training corpus went from Hermes 3's roughly one million samples and 1.2 billion tokens to about 5 million samples and 60 billion tokens, blending data with and without reasoning. That jump in scale is what underpins the model's main new feature.

That is not a small increment. The scale of the post-training corpus is roughly 50× larger than Hermes 3. For a team running structured-output pipelines - JSON extraction, schema-validated responses, multi-step tool sequences - the practical consequence is that Hermes 4's instruction-following is measurably more reliable, not because the base model changed, but because the fine-tuning data volume did.

Hermes 4.3, released August 25, 2025, is built on ByteDance's Seed 36B base model and was trained using Nous Research's Psyche decentralized training network rather than a traditional centralized GPU cluster - the first Hermes model trained this way. Whether decentralized training becomes a repeatable production method is a longer conversation, but it is genuinely novel at this scale.

What is incremental: the benchmark comparisons to commercial models are, as always, cherry-picked task categories. The model does not have multimodal input. The 70B does not support image input and can only process text. For teams where vision is a requirement, Hermes 4 is not a drop-in for GPT-4o.

The honest summary: Hermes 4 is the most capable open-weight instruction model Nous has shipped, with a real architectural addition in the hybrid reasoning mode. The 70B, self-hosted or via a provider that surfaces tool calls properly, is the practical deployment target. The 405B is a benchmark vehicle and a research option for teams with GPU infrastructure. That distinction matters more than the MATH-500 number.

Beagle in action#eng-ai, 2:47pm
The ask
'should we move our code-review summarizer to Hermes 4 70B to cut costs?'
Beagle drafts
checks the current Claude Haiku pricing and the Hermes 4 70B rate, estimates monthly token volume from the last 30 days, drafts a comparison with a break-even number
You approve
engineer approves the draft; the answer posts with a linked source - no spreadsheet required
Do this in your workspace →

Hermes 4 open-weight reasoning model: common questions

What is Hermes 4's hybrid reasoning mode?

Hermes 4 is a single model checkpoint that can either answer directly or generate an explicit chain-of-thought before responding. You control it with a reasoning_enabled boolean per request. Reasoning-on calls use up to 30k tokens of thinking trace before the answer, billed as output tokens. Reasoning-off calls behave like a standard instruction model.

How does Hermes 4 pricing compare to GPT-4o?

The Hermes 4 70B costs $0.13 per million input tokens and $0.40 per million output tokens via OpenRouter. GPT-4o runs at $2.50 input and $10 output per million tokens - roughly 19× higher on input and 25× higher on output. The gap narrows when reasoning is enabled on Hermes 4, because thinking tokens add to output volume.

Is Hermes 4 safe to deploy without added guardrails?

That depends on your use case. Hermes 4 scores ~57% on RefusalBench compared to ~17% for GPT-4o, meaning it refuses far fewer requests by design. Nous calls this neutral alignment. For internal tools or research pipelines, this is often desirable. For public-facing products, you need application-layer content policy - the model will not supply it.

Does Hermes 4 support function calling and tool use?

The model is trained for tool use and ships parsers for vLLM and SGLang. However, the Hermes 4 405B endpoint on OpenRouter currently does not accept tool calls. The 70B endpoint also has this limitation on the managed OpenRouter endpoint as of publication, but works natively when self-hosted with vLLM or SGLang.

Should my team pick Hermes 4 70B or 405B?

Almost certainly the 70B for production. The 405B requires multi-GPU infrastructure to self-host and has limited managed-endpoint availability. The 70B fits on a single H200, supports a 131k context window, and is where most of the practical post-training improvements land. Reserve the 405B for benchmarking or high-stakes reasoning tasks where you control the serving stack.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle