Hermes 4.3 Was Trained on the Internet, Not a Cluster

Nous Research's Hermes 4.3 is the first production model trained on a decentralized GPU network. Here's what that actually changes-and what it doesn't-for teams evaluating open-weight models for agentic work.

Cover art for Hermes 4.3 Was Trained on the Internet, Not a Cluster

An agent workflow breaks mid-run. The model refuses a data-formatting step it decides is "sensitive," and the automation falls apart. That is not a hypothetical. It is the over-refusal failure mode that practitioners reported against Hermes 3 in multi-tool pipelines, and it is the specific problem Hermes 4.3 was built to fix.

Hermes 4.3 36B is a frontier, hybrid-mode reasoning model based on ByteDance Seed 36B, made by Nous Research. It shipped December 2025 - roughly four months after the main Hermes 4 family - and the most interesting thing about it is not the benchmark numbers. It is how it was trained.

What Psyche actually is, and why the Hermes 4.3 result matters

Hermes 4.3 is the first production model post-trained entirely on the Psyche network, a distributed training network that uses the DisTrO optimizer to efficiently communicate between training nodes spread through data centers over the open internet, secured by the consensus of the Solana blockchain.

That sounds exotic. The practical question is whether it works. The Psyche-trained version of Hermes 4.3 outperformed the traditionally centralized version on a suite of downstream tasks - a confirming signal that Psyche is up to the task of training production models. Nous published the centrally trained version as a research artifact alongside the Psyche version, so anyone can diff the evals directly.

The implication is structural, not just technical. By enabling nodes throughout the world to collaborate on a single training run, Psyche can dramatically reduce the cost of training frontier-level models, potentially leveling the playing field for open-source AI model developers. If that holds at larger parameter counts, it changes who can afford to iterate at scale - and how fast.

The Psyche network is clearly an active investment: Hermes 4.3 being explicitly the first model trained this way suggests Nous Research is treating Psyche as a production training infrastructure path, not a one-off experiment.

The 50x corpus expansion and what it changes for agents

The post-training corpus for Hermes 4.3 expanded substantially: approximately 5 million samples covering around 60 billion tokens, compared to roughly 1.2 billion tokens for Hermes 4 - a roughly 50x expansion in training token count. The intended effect was improved reasoning quality across math, code, STEM, logic, and creative writing tasks.

Two of those improvements matter most in agentic pipelines.

Schema adherence. Hermes 4.3 was specifically trained to produce valid JSON when given a schema. For agent frameworks that depend on structured tool calls, schema adherence failures are a significant source of runtime errors - and this improvement addresses one of the more common pain points practitioners reported with Hermes 3 in complex multi-tool workflows.

Refusal rate. RefusalBench measures whether a model follows user constraints precisely - whether it does what you said, in the way you specified, without adding unsolicited caveats or refusing legitimate requests. For automated agent workflows, this is not a secondary metric: an agent that softens your constraints mid-run or declines a step it deems "sensitive" breaks the automation logic as surely as a syntax error would.

Hermes 4.3 scored 74.6% on RefusalBench - meaning it answered 74.6% of questions that other aligned models refuse - compared to 59.5% for Hermes 4 70B.

The Hermes 4 405B scores 57.1%, well above GPT-4o's 17.67% and Claude Sonnet 4's 17%.

One honest caveat: RefusalBench is Nous's own benchmark, built by classifying 32 categories of requests that typically result in refusals from frontier models, then hand-crafting 166 prompts across those categories, with Claude Sonnet 4 used as an LLM-as-a-judge to identify refusals. The methodology is transparent and the prompts are published, but it is not a neutral third-party eval. Treat the comparative scores as directional signals, not settled fact.

74.6%Hermes 4.3 RefusalBenchvs 17.67% for GPT-4o
60B tokensfine-tuning corpus50× larger than Hermes 4's baseline
93.8%MATH-500 score (36B)outperforms the larger Hermes 4 70B

What the numbers look like in a real deployment

Here is where the comparison gets concrete. Hermes 4 70B costs $0.13 per million input tokens and $0.40 per million output tokens via OpenRouter. Running Hermes Agent via cloud API costs approximately $0.30 per million tokens - roughly 50x less than GPT-5 Medium.

Against that, what do you give up? Size. Hermes 4.3 is 36B parameters. Benchmark scores from the model card: MATH-500 at 93.8%, MMLU at 87.7%, BBH at 86.4%, AIME 24 at 71.9%, GPQA Diamond at 65.5%.

These are strong scores for a 36B model and outperform the larger Hermes 4 70B on several benchmarks according to the model card. But on open-ended generation and edge-case reasoning, a 405B closed model still has a ceiling advantage. The 36B is not a drop-in substitute for every task.

Dimension Hermes 4.3 36B Hermes 4 70B GPT-4o (closed)
Parameters 36B 70B ~200B (est.)
Context window 512K tokens 131K tokens 128K tokens
Hosted input price ~$0.30/M $0.13/M $2.50/M
RefusalBench 74.6% 59.5% 17.67%
MATH-500 93.8% - ~76%
Weights available Yes Yes No
Training infrastructure Psyche (decentralized) Centralized cluster Proprietary

The 512K context window on 4.3 is worth noting separately. It lets an agent hold an entire codebase, conversation history, and tool results in a single context without chunking - eliminating the retrieval overhead that degrades most frameworks on extended sessions.

Teams with data residency requirements, air-gapped compliance environments, or GPU infrastructure they want to fully utilize are the natural fit for open-weight deployments.

For teams without those constraints and comfortable with API-only providers, the open-weight model operational overhead is a cost that does not always pay off.

Beagle in action#dev-tools Slack channel, 10:52am
The ask
engineer asks 'which model should we use for the doc-to-JSON extraction agent - we keep getting refusals on the contract clauses'
Beagle drafts
pulls the Hermes 4.3 RefusalBench and schema-adherence notes, drafts a reply comparing it against the team's current GPT-4o setup on cost and compliance fit
You approve
you review and approve; the recommendation posts in-thread with source links, no context-switching required
Do this in your workspace →
Running a multi-step extraction agent
Without Beagle
the model refuses a contract clause it flags as sensitive, the pipeline throws, someone manually patches the output and re-runs
With Beagle
Hermes 4.3 follows the schema and the instruction; the structured output lands clean, no manual recovery step

When to take Hermes 4.3 seriously - and when to skip it

The Psyche story is genuinely new. A 36B model post-trained on a decentralized internet network matching or exceeding a centrally trained 70B on core benchmarks is not a headline you could have written a year ago. Whether it holds at larger base model scales is the open question Nous has implicitly committed to answering.

For teams evaluating the model right now, three conditions make Hermes 4.3 worth a serious pilot:

  • Your agents have over-refusal problems. If a closed model is declining legitimate steps in your pipeline, the RefusalBench gap is large enough to be worth testing - not as proof of production quality, but as a starting point.
  • You need the full weights. Data residency, air-gap requirements, or BYOC infrastructure all push toward open-weight. Hermes 4.3 ships on Hugging Face; you can self-host or run via a provider like Together AI or via Nous Chat.
  • You are running long-context agent sessions. The 512K window at 36B parameters is a practical advantage most 70B alternatives do not offer at this price point.

Where to be skeptical: RefusalBench is Nous's own eval. The post-training corpus expansion is large, but Nous does not publish the corpus itself. And "decentralized training" on Solana is a real architectural choice with real tradeoffs around reproducibility and auditability that enterprise teams should probe before committing.

In April 2025, Nous raised a $50 million Series A led by Paradigm, reported at a $1 billion token valuation, to fund its broader push into decentralized AI training, including the Psyche network and DisTrO and DeMo distributed-training research. The company has capital and a clear thesis. Hermes 4.3 is the first production evidence that the thesis is not purely theoretical.


Hermes 4.3 open-weight model: common questions

What is Hermes 4.3 and how does it differ from Hermes 4?

Hermes 4.3 is a 36B hybrid-reasoning model from Nous Research built on the ByteDance Seed 36B base, released December 2025. It differs from Hermes 4 in three concrete ways: it was trained on the Psyche decentralized network, its fine-tuning corpus is 50× larger (~60B vs ~1.2B tokens), and its context window is 512K tokens versus Hermes 4's 131K.

What is RefusalBench and should I trust it?

RefusalBench is Nous Research's internal benchmark covering 166 hand-crafted prompts across 32 refusal categories, judged by Claude Sonnet 4. Hermes 4.3 scores 74.6%; GPT-4o scores 17.67%. The methodology is public and the results are published, but it is not a neutral third-party eval. Use the scores as directional signal, not final proof of production behavior on your data.

How much does it cost to run Hermes 4.3?

Via hosted API, Hermes 4.3 runs at roughly $0.30 per million tokens. The 70B sibling is $0.13 input / $0.40 output per million via OpenRouter. Self-hosting on GPU cloud infrastructure typically lands between $0.10-$0.50 per million tokens at reasonable utilization - but that figure does not include GPU reservation, ops time, or the engineering cost of running inference infrastructure.

What is Nous Research's Psyche network?

Psyche is a distributed training network that coordinates GPU nodes across different data centers over the open internet using the DisTrO optimizer, with training runs recorded on the Solana blockchain. Hermes 4.3 is the first production model Nous trained entirely this way, and the Psyche-trained version outperformed the centrally trained equivalent on downstream evals.

Who should consider Hermes 4.3 for agent workflows?

Teams with data residency or air-gap requirements, those running into over-refusal failures with closed models in agentic pipelines, and those handling long-context sessions where 512K tokens matters. Teams without those constraints and no existing GPU infrastructure should weigh the operational overhead against simply routing to a managed API - the cost savings are real, but they come with an ops bill.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle