What Does Hermes 4 Actually Give You as an Agent Brain?

Nous Research's Hermes 4 is a free-to-download, hybrid-reasoning model built specifically for agent tool use. Here's what's genuinely new, what the numbers miss, and how to pick the right size.

Cover art for What Does Hermes 4 Actually Give You as an Agent Brain?

The Hermes 4 training dataset totals approximately 5 million samples and 19 billion tokens

  • roughly 50 times more tokens than Hermes 3 . That is not a version bump. It is a ground-up rebuild, and it changes what you can actually do with the model in a production agent.

What Hermes 4 actually is (and what it is not)

Hermes is a family of open-source large language models fine-tuned by Nous Research. The models start from open base weights - Meta's Llama series, Mistral, ByteDance Seed, or Alibaba Qwen - and add specialized training data for instruction following, structured output, function calling, and multi-turn conversation. Hermes 4 is not a new foundation model. It is a post-training stack applied on top of other labs' bases, and that distinction matters for how you think about its capabilities.

Hermes 4 is a fine-tuned model family, not an agent. It powers agents. The companion Hermes Agent framework (MIT-licensed) is a separate layer you run on top.

The current family covers four sizes: a 14B based on Qwen 3, a 36B (Hermes 4.3) based on ByteDance's Seed-OSS-36B, a 70B based on Llama 3.1, and a 405B based on Llama 3.1.

The 14B and 36B ship under Apache 2.0; the 70B and 405B inherit the Llama 3 community license.

Variant Base Context License VRAM (Q4_K_M)
Hermes 4 14B Qwen 3 128K Apache 2.0 ~9 GB
Hermes 4.3 36B Seed-OSS 36B 512K Apache 2.0 ~22 GB
Hermes 4 70B Llama 3.1 131K Llama 3 ~42 GB
Hermes 4 405B Llama 3.1 131K Llama 3 multi-GPU

Hybrid reasoning: what the toggle actually does

Hermes 4 introduces what Nous Research calls "hybrid reasoning," allowing users to toggle between fast responses and deeper, step-by-step thinking processes. When activated, the models generate their internal reasoning within special <think> tags before providing a final answer - similar to OpenAI's o1 reasoning models but with full transparency into the AI's thought process.

The benchmark case Nous leads with: the 405B model scored 96.3% on MATH-500 in reasoning mode and 81.9% on AIME'24. Those are legitimate numbers for an open-weight model, but read them carefully. AIME 2025 contains only 30 problems, which makes evaluation particularly sensitive to sampling variance - for all models, average-of-5 results from two independent runs can differ significantly by up to 5-10 percentage points, making side-by-side comparison unreliable. Nous ran evaluations in reasoning mode; the smaller 70B scores are more modest. The 70B variant hit 11.3% on AIME 2025 via Artificial Analysis's independent measurement

  • competitive for the size tier, but not frontier.

The more operationally useful property is the toggle itself. An agent doing a quick Slack lookup does not need to burn reasoning tokens. A multi-step code-generation task probably does. The model can deliberate internally using <think> traces or respond directly, and developers can toggle reasoning off for faster, cheaper inference in production. That is a real architecture win for teams routing different task types through the same model endpoint.

The tool-call format that became a standard

Tools can be specified and invoked via the Hermes Function Calling standard, which places tool definitions (as JSON schemas) in <tools> and invocations and responses in <tool_call> and <tool_response> respectively. What started as Nous's internal format has spread considerably. vLLM now supports parsing tool calls from model output into structured messages via the --tool-call-parser hermes flag

  • and Qwen's own vLLM docs point teams to that exact flag when serving Qwen3 models. The Hermes format is now the de facto open-weight tool-call standard, which means skills and agent harnesses written for Hermes run on other models without changes.

The Hermes Function-Calling dataset is a synthetic instruction-following corpus combining single-function and multi-function tool-calling conversations together with structured extraction, JSON-mode, and agentic interaction examples. The training examples include user requests, tool definitions as JSON schemas, assistant-generated tool calls, tool responses, and final assistant replies - allowing models to learn both API selection and argument generation in multi-turn settings.

One non-obvious consequence: the format's ubiquity means the ecosystem around it is now larger than Nous itself. As of June 2026, 90,881 community-contributed skills are built on the Hermes tool-calling standard. If you're evaluating whether to adopt the format, you're not betting on one lab - you're betting on a standard.

Beagle in action#product-ops, 2:47pm
The ask
'can you pull what the refund policy says for orders over $500?'
Beagle drafts
calls the Notion MCP tool, drafts a reply with the specific clause and a doc link
You approve
you approve in one click; the reply posts with its source logged - no copy-paste, no tab-switching
Do this in your workspace →

Hermes 4.3: the variant nobody is talking about enough

Hermes 4.3 was trained with an extended context length of up to 512K tokens and nearly matches - and in some cases exceeds - the performance of Hermes 4 70B at half the parameter cost. Based on Seed-OSS-36B-Base, Hermes 4.3 is an excellent shape for consumer local inference or enterprise self-deployment.

The context window alone changes what's possible. The 70B and 405B top out at 131K tokens. Nous reports the 36B nearly matches the 70B at half the parameters, thanks to the Seed-OSS base - and it has a 512K context window, which the 70B does not. For an agent that needs to ingest a large codebase, a long Slack thread history, or multiple documents before acting, 512K versus 131K is a meaningful difference, not a footnote.

There is also a training provenance story worth knowing. Hermes 4.3 is the first production model post-trained entirely on Nous's Psyche network - a distributed training system that uses the DisTrO optimizer to coordinate nodes across the open internet, secured by the Solana blockchain consensus. The design can dramatically reduce the cost of training frontier-level models.

Nous verified Psyche's effectiveness for production workloads by training Hermes 4.3 both on Psyche and via the traditional centralized approach. The two runs produced comparable models, which is the result they needed to validate the decentralized training path.

What to be skeptical about

RefusalBench is the headline number Nous leans on hardest. Hermes 4 achieved the highest score among all tested models on RefusalBench, scoring 57.1% in reasoning mode and significantly outperforming GPT-4o (17.67%) and Claude Sonnet 4 (17%). But Nous developed RefusalBench internally, classifying 32 categories of requests that typically result in refusals from frontier models and hand-crafting 166 prompts that cover these categories. A lab benchmarking itself on a benchmark it designed is worth noting. The low-refusal property is real and useful for agent deployments where the model needs to follow instructions precisely - but "minimal refusals" also means you own the safety layer entirely. There is no gated filter to catch a misdirected tool call.

The Hermes Agent framework's latency profile is also worth understanding before committing to it. Independent performance audits of the agent show Hermes turns take 5-13 seconds even when the work is trivial, and a full request commonly runs several minutes. Most of that overhead is not model latency - it's serial round-trips and context retrieval happening before the first token. If your use case needs sub-second responses, Hermes Agent's current architecture needs custom tuning.

Running an agent on a closed API vs. Hermes 4 self-hosted
Without Beagle
every tool call goes through a $5-15/M token rate card, data leaves your infra, and the model's refusal behavior is outside your control
With Beagle
Hermes 4.3 36B on a single H100 runs at roughly $1.49/hr GPU cost; tool-call format and safety thresholds are yours to configure, and a 512K context window handles long agent sessions without chunking

Across eleven open-weight and eight closed models on the Artificial Analysis Intelligence Index, the cheapest qualifying open-weight model - standardized by intelligence - completes a task at roughly one fifth of the cost of a comparable closed model. That ratio does not apply uniformly across all task types, and closed models still lead at the very top of the capability curve. But for the bulk of agent work - structured lookups, tool orchestration, formatted summaries - Hermes 4 sits well within the range where cost routing away from closed APIs makes quantifiable sense.

5M samplesHermes 4 training set19B tokens, ~50× larger than Hermes 3
512K tokensHermes 4.3 context4× the 131K ceiling on the 70B and 405B variants
~1/5 the costopen-weight vs. closedper task, standardized by capability tier (Artificial Analysis)

Hermes 4 open-weight model: common questions

What is Hermes 4?

Hermes 4 is a family of open-weight, hybrid-reasoning language models from Nous Research, released August 2025. It ships at 14B, 70B, and 405B parameter sizes on Llama 3.1 and Qwen 3 bases, with a follow-up 36B variant (Hermes 4.3) on ByteDance's Seed-OSS architecture. All sizes support toggleable chain-of-thought reasoning and native function calling.

How is Hermes 4 different from Hermes 3?

Hermes 3 introduced structured function calling. Hermes 4 added hybrid reasoning - the ability to produce <think> traces before answering - and expanded the training dataset by roughly 50×. Hermes 4.3 added a 512K context window and reduced refusal rates further, while switching the base model from Llama to Seed-OSS for the 36B size tier.

Can I use Hermes 4 commercially without paying Nous Research?

The 14B and 36B variants (Hermes 4 and 4.3) are Apache 2.0 licensed - free for commercial use with no royalties. The 70B and 405B inherit Meta's Llama 3 community license, which allows commercial use under specific terms (primarily a restriction on services that exceed 700 million monthly active users).

Is Hermes 4 good enough to replace a closed frontier model for agents?

For structured tool-calling, long-context summarization, and JSON schema adherence, yes - especially the 36B and 70B variants, which handle the majority of production agent workflows. For frontier-level reasoning on genuinely hard problems (competition math, complex multi-hop research), closed models still hold a measurable edge. The practical approach most teams land on is routing: open-weight models for the high-volume, well-defined steps; closed APIs for the hard judgment calls.

What hardware does Hermes 4.3 36B require to self-host?

At Q4_K_M quantization, Hermes 4.3 36B runs at approximately 22 GB VRAM - fitting on a single 24 GB GPU (RTX 3090/4090 or A10G). The full 36B at higher precision needs a single H100 or two A100s. A Mac with 32 GB unified memory can run it via Ollama without a discrete GPU.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle