Hermes 4 70B costs $0.13 per million input tokens and $0.40 per million output tokens on OpenRouter. Claude Sonnet 4.6 output runs at $15 per million. That is a 37x spread on the token that actually costs you money in an agent loop - and Hermes 4 is not a toy.
What Nous Research actually built
Hermes is a family of open-source large language models fine-tuned by Nous Research, starting from open base weights - Meta's Llama, Mistral, ByteDance Seed, or Alibaba Qwen - with specialized training for instruction following, structured output, function calling, and multi-turn conversation. The fine-tune is the product; the base model is borrowed. When someone says "Hermes 3 405B," they mean Nous Research's fine-tuned version of Meta's Llama 3.1 405B.
Hermes 4 is a family of open-weight large language models released by Nous Research in late August 2025. It is the fourth generation of the Hermes line of fine-tuned, instruction-following models, and its defining feature is hybrid reasoning: each model can answer directly in a fast "non-reasoning" mode or first deliberate inside explicit <think>...</think> tags in a "reasoning" mode, with the behavior selectable at inference time.
The current Hugging Face collection lists four sizes: a 14B based on Qwen 3, a 36B (Hermes 4.3) based on ByteDance's Seed-OSS-36B, a 70B based on Llama 3.1, and a 405B based on Llama 3.1.
The 14B and 36B ship under Apache 2.0; the 70B and 405B inherit the Llama 3 community license. That split matters for commercial use - Apache 2.0 means redistribute freely; Llama 3's community license restricts products with over 700 million monthly active users. Most teams will never hit that ceiling, but compliance teams may notice.
The hybrid reasoning feature and why it changes agent economics
Hermes 3 was a good instruction and agentic model, but it always answered in one go, with no explicit reasoning phase. Hermes 4 adds exactly that: thinking mode with <think> tags, which lifts performance on math, code, and logic tasks without giving up fast answers when they are not needed.
For agent builders, the practical consequence is call-by-call cost control. A routing step or intent classifier does not need chain-of-thought; a multi-step planning step does.
The model facilitates agentic behavior by training it to generate interleaved reasoning and tool use. Multiple tool interactions can occur within a single <think> block. The environment intercepts a special <tool_call> token, executes the specified tool (e.g., a Python interpreter), and appends the result to the context with <tool_response>. The model is then prompted to resume generation from this new state, allowing it to process the tool's output and potentially invoke further tools before closing the <think> block.
That is a more honest tool-use design than many models deliver - the reasoning trace and the tool calls are woven together rather than bolted on separately.
The Hermes 4 dataset consists primarily of newly synthesized reasoning and non-reasoning data, totaling approximately 5 million samples and 19 billion tokens. The data strategy combined a substantial volume of reasoning-focused data with a diverse set of non-reasoning instructions - 3.5 million reasoning samples and 1.6 million non-reasoning samples.
The volume of fine-tuning data grew fivefold compared to Hermes 3.
What Hermes costs versus renting from a frontier lab
The cost gap between self-hosted open-weight and closed frontier inference has become the main commercial argument for models like Hermes 4.
| Deployment option | Input cost ($/M tokens) | Output cost ($/M tokens) | Control |
|---|---|---|---|
| Hermes 4 70B (OpenRouter) | $0.13 | $0.40 | Hosted third party |
| Hermes 4 70B (self-hosted H100) | ~$0.05-0.18 | ~$0.18 | Full |
| Claude Sonnet 4.6 | $3.00 | $15.00 | None |
| GPT-4.1 nano | $0.10 | $0.40 | None |
Self-hosted Llama 4 70B on an H100 at batch=8 costs approximately $0.18 per million output tokens, versus $15 for Claude Sonnet 4.6 output at identical quality ceiling.
The self-hosting break-even versus managed APIs sits at roughly 2 to 5 million tokens per day on reserved GPU capacity over a 12-month window.
With published weights, the API price is a ceiling, not a floor. You can quantize, batch, distill, and right-size hardware in ways that are simply not available when you rent tokens from a closed vendor.
A less-covered number: closed models still accounted for close to 80% of AI token usage over a five-month study period on OpenRouter, as well as nearly 96% of revenue - but closed models cost 87% more to run, at $1.86 per million tokens on average, compared with 23 cents for open models. Teams are still paying the premium. The question is whether quality justifies it for your specific workload.
The Psyche angle is the genuinely new thing - and still speculative
Hermes 4.3 was trained using Nous Research's Psyche decentralized training network rather than a traditional centralized GPU cluster. The model card explicitly calls this out as the first Hermes model trained this way.
Psyche is a decentralized training network built on the Solana blockchain that coordinates heterogeneous GPUs - from consumer RTX 4090s to datacenter H100s - into fault-tolerant training runs.
In April 2025, the organization secured its Series A funding round, with Paradigm leading a $50 million investment that brought total funding to approximately $65 million.
In July 2026, it was reported to be finalizing a $75M+ round at a ~$1.5B valuation.
Nous Research is best known for its Hermes family of fine-tuned language models, which have been downloaded over 33 million times from Hugging Face, and for its Psyche platform, a decentralized training network built on the Solana blockchain.
The honest assessment: whether decentralized training at this scale becomes a repeatable production method or remains an experiment is worth watching. The training results are competitive, which at minimum demonstrates that the approach is viable. Psyche's architecture is not a vanity project - it is a direct attack on the GPU cluster bottleneck that makes frontier model training inaccessible to anyone outside a handful of labs. But the blockchain incentive layer introduces risks (token volatility, participation dynamics, governance) that a normal ML engineering team has no experience modeling.
A teammate like Beagle routing low-stakes queries to a Hermes 4 14B endpoint and high-stakes queries to a 70B endpoint could cut per-call costs by 60-70% relative to sending everything to a single closed-model API - without any change in the user-facing experience.
Hermes 4: common questions
What is Hermes 4 and how does it differ from Hermes 3?
Hermes 4 is Nous Research's August 2025 open-weight model family - 14B, 70B, and 405B parameter sizes - built on Llama 3.1. Its key upgrade from Hermes 3 is hybrid reasoning: a toggleable <think> mode that lets the model deliberate before answering. Hermes 3 had no explicit reasoning phase.
Can Hermes 4 handle tool use and function calling for agents?
Yes. Tool use is a first-class capability.
Hermes 4 uses the Hermes tool format (<tool_call>) and ships automatic parsers built into vLLM and SGLang.
The model is trained to interleave reasoning and tool calls inside a single <think> block, which means it can plan, call a tool, receive a result, and decide whether to call another tool before producing a final answer.
What license do Hermes 4 models ship under?
License depends on size. The 14B (Qwen3 base) and 36B Hermes 4.3 (ByteDance Seed base) ship under Apache 2.0 - fully permissive for commercial use. The 70B and 405B variants are built on Llama 3.1 and inherit Meta's Llama 3 community license, which restricts products serving more than 700 million monthly active users.
Is self-hosting Hermes 4 worth it compared to using a closed model API?
The self-hosting break-even versus managed APIs sits at roughly 2 to 5 million tokens per day on reserved GPU capacity over a 12-month window. Below that, managed inference providers running open-weight models (Together AI, Fireworks, OpenRouter) deliver most of the cost benefit without the operational overhead.
What is Psyche and does it affect the model's quality?
In January 2025, Nous Research announced the launch of Psyche, its decentralized training platform. Hermes 4.3 36B was the first model trained on it. The results are benchmark-competitive with centrally trained equivalents, which proves the method works - but it is still an early-stage system, and the Solana-based incentive layer adds risk factors unfamiliar to most ML teams.