Best Model for Hermes Agent: What 49 Trillion Tokens Actually Run (2026)

Hermes ships with no default model, so everyone picks one. Verified September 2026 prices, what 49 trillion tokens of real traffic actually run, and the monthly math.

•

Published on

•

Best Model for Hermes Agent: What 49 Trillion Tokens Actually Run (2026)
Do not index
The best model for Hermes agent in late 2026 is Claude Sonnet 5 ($2 in / $10 out per million tokens) if you want the most reliable tool calling, and GPT-6 Luna ($0.10 / $0.50) or DeepSeek Flash if you want a price that makes an always-on agent sustainable. Real usage leans hard toward budget: Hermes Agent is the biggest app on OpenRouter, its public page shows 49.3 trillion tokens in the last 30 days, and the top ten models behind that traffic are all cheap flash-class options, not flagships. This guide covers verified September 2026 prices, what the community actually runs, the new official plugin that runs Hermes on a Claude subscription, and the honest monthly math. And if you need somewhere always-on to run whichever model you pick, Agent37 hosts a managed Hermes instance from $3.99/mo.

The best model for Hermes agent, by use case

Hermes ships with no default model. A fresh install has an empty model: "" sentinel in its config, so everyone answers this question at setup. Here is the short version, with September 2026 API prices per million tokens:
You want
Run this
Price (in / out)
Why
The most reliable daily driver
Claude Sonnet 5
$2 / $10
Best tool-calling consensus; price locked (the planned Sept 1 increase was cancelled)
The cheapest that still works
GPT-6 Luna
$0.10 / $0.50
Half the price of the GPT-5.6 Luna it replaced; needs its tool-calling quirks configured around
What most people actually run
DeepSeek Flash
$0.30 / $1.20 peak, half off-peak
Flash-class models dominate real Hermes traffic
Hard reasoning and coding
Claude Opus 5.5
$4 / $20
Launched Sept 22, 20% cheaper than Opus 5
Zero marginal cost
Claude Subscription DirectSDK plugin
Your existing Claude plan
Official Nous plugin, runs turns through the Claude Code CLI
Free
Nous Portal free routes
$0
Real volume, real tool-calling quirks
Private and local
Gemma 4 31B or a Qwen-class 30B+
Your GPU
The docs' current local pick with working tool calls
Every model needs at least 64,000 tokens of context. Hermes checks at startup and rejects anything smaller, which quietly disqualifies most small local setups at their default settings.

What real Hermes traffic runs

OpenRouter publishes per-app usage, and Hermes Agent tops its app rankings. The app page shows 49.3 trillion tokens in the last 30 days, spread across 486 models. The top ten by volume: DeepSeek V4 Flash 0731, GLM 5.3 Flash, DeepSeek V4.1 Flash, Solar Pro 4, DeepSeek V4 Flash 0423, MiniMax M3, Laguna S 2.1, DeepSeek V4 Pro, Nemotron 3 Ultra, and GPT-5.6 Luna.
Notice what is missing. Not one Anthropic model. Not one GPT-6 model. The DeepSeek Flash family alone accounts for roughly 17 trillion of those tokens.
The reason is structural. An always-on agent that checks your email, runs scheduled jobs, and sits in your group chats makes model calls all day. One community-built dashboard measured about 13,900 tokens of fixed overhead per call just for tool schemas and the system prompt, roughly 73% of a typical request. When every call starts with a 14k-token toll, the input price per million is what decides your bill, and flash-class pricing wins.

The quality ceiling: Sonnet 5 and Opus 5.5

If your agent approves things, moves money, or chains many tool calls, model quality stops being abstract. The r/hermesagent consensus puts Claude at the top for tool-calling reliability, and the Hermes docs' own Nous Portal quick-pick labels the Sonnet line the "best general-purpose agentic model".
Claude Sonnet 5 costs $2 in / $10 out, and Anthropic confirmed the scheduled September price increase will not happen, so that number is dependable. Claude Opus 5.5 landed on September 22 at $4 / $20, a 20% cut from Opus 5, and is the pick for genuinely hard reasoning or coding work.
The catch is volume. Run a flagship as an always-on daily driver and you are looking at a three-figure monthly bill before caching. The pattern that works: a flagship as the main brain, cheap models everywhere else. Hermes has 11 auxiliary model slots (conversation compression, vision, title generation, approval scoring, and more) that default to your main model. Point them at a flash-class model and reserve the flagship for the main loop.

The budget picks that carry real workloads

GPT-6 Luna ($0.10 / $0.50, launched September 22) is the cheapest current-generation model from a major lab. One documented catch from the Hermes issue tracker: on Chat Completions style provider routes, Luna rejects tool calls outright when reasoning is enabled. Working setups either turn reasoning off or run it through a Responses API route, so expect a few minutes of provider fiddling before it behaves.
DeepSeek Flash ($0.30 / $1.20 at peak, half price off-peak, cache hits at $0.006) is the volume king of the leaderboard above. It is fast and cheap, with two honest caveats from people running it daily: it hallucinates more than the flagships on complex tasks, and the peak/off-peak windows make bills less predictable than a flat rate.
GLM 5.3 Flash sits second on the leaderboard at 6.38 trillion tokens, and is the strongest single answer if you want one inexpensive model for everything.
GPT-6 Sol ($2 / $10) matches Sonnet 5 on price and is the natural pick if you would rather stay in the OpenAI ecosystem.

Run Hermes on your Claude subscription

The biggest model news for Hermes in months arrived on September 20: Nous shipped an official plugin, Claude Subscription DirectSDK (experimental, v0.3.0 at the time of writing), that runs Hermes turns on a Claude Pro or Max subscription through the official Claude Code CLI's Agent SDK path. No API key, no per-token bill.
Requirements are minimal: Hermes Agent 0.21.4 or newer, plus the Claude Code CLI installed and logged in. Then:
hermes plugins install claude-subscription-directsdk
Set model.provider: claude-subscription-directsdk-experimental and pick sonnet. It exposes Sonnet 5, Haiku 4.5, and Opus 5.5 (Opus needs a recent CLI build).
Two honest notes. First, this is different from Hermes's older built-in Anthropic OAuth path, which only works on Claude Max with purchased extra-usage credits and excludes Pro entirely; the plugin is the route that finally works on Pro. Second, it draws subscription usage at the Agent SDK rate, which the plugin docs peg at roughly 1.7x what interactive Claude Code consumes, and maintainers note an always-on agent eats an allowance faster than a human typing. It is a genuine bargain for Pro/Max holders, not infinite free compute.

Free and local: what actually works

Free routes are real. Nemotron 3 Ultra's free OpenRouter route does enough volume to sit ninth on the Hermes leaderboard, and Hermes's own model catalog lists several free routes (Nemotron 3, GLM 5.2, MiniMax M3) alongside Nous Portal's free plan. But the failure modes are documented all over the Hermes issue tracker: some OpenRouter :free routes reject tool calls outright with 404s, and free reasoning models can burn their entire output budget thinking, then loop retrying. Free works for chat; it gets flaky exactly when the agent needs to do things.
Local has one hard rule and one honest threshold. The hard rule is the 64k context floor: Ollama defaults to 2,048 tokens, so you must raise it or Hermes rejects the model at startup. The threshold is size: models under roughly 30B parameters tend to fail the agentic loop in a specific way, confirming they made a tool call without actually making it. The docs currently name Gemma 4 31B the best local option with working tool calls, and the community pattern is a split stack: a cloud model for general work, a local Qwen-class model for anything touching private data.
One surprise: Nous's own Hermes 4 70B and 405B models are available at a discount through Nous Portal, and Nous itself does not recommend them inside Hermes Agent. They are tuned for chat and reasoning, not the rapid-fire tool-calling loop the agent runs. Respect the honesty.

What a month of Hermes actually costs

Hermes Agent is free, MIT-licensed software. Your real costs are the model API and somewhere always-on to run it. Here is a modeled month for a busy personal agent: 40 turns a day, about 3 model calls per turn, 14k input tokens per call (the measured overhead above), 500 output tokens per call. That is roughly 50M input and 2M output tokens a month:
Model
Modeled month
GPT-6 Luna
~$6
DeepSeek Flash
~$9 off-peak, ~$17 peak
Claude Sonnet 5 / GPT-6 Sol
~$120
Claude Opus 5.5
~$240
Treat these as ceilings, not quotes. Prompt caching changes the picture a lot (Anthropic cache reads cost $0.20 per million, DeepSeek cache hits $0.006), and a quieter agent burns far less. But the shape of the table explains the leaderboard: flash-class models turn an always-on agent into a single-digit line item, flagships turn it into a subscription you have to think about.
For the infrastructure side of the bill, a capable VPS runs $5 to $24 a month depending on region and specs, or managed Hermes hosting on Agent37 starts at $3.99/mo. We wrote up the full DIY-vs-managed math in our Hermes VPS guide.

How to set the model in Hermes

Switching takes under a minute. The canonical way is the interactive wizard:
hermes model
Or edit ~/.hermes/config.yaml directly:
model:
  provider: "anthropic"
  default: "claude-sonnet-5"
In a running session, /model switches between already-configured providers, with --once for a single turn and --global to persist. Auxiliary slots live under auxiliary.<task> and default to auto, meaning they use your main model until you override them.
One gotcha worth knowing: if you configure a provider but never pick a model, Hermes falls back to the catalog default, currently z-ai/glm-5.2, which the docs describe as "deliberately a capable low-cost model, never the priciest flagship". Sensible, but it means you may be running a model you never chose.

Where to run it always-on

A model choice only pays off when the agent is reachable. On a laptop, Hermes sleeps when the lid closes, and messages queue until you are back. The whole point of a personal agent in WhatsApp or Telegram is that it answers at 3 a.m.
That leaves a VPS you maintain, or managed hosting. Agent37 runs Hermes agents as isolated always-on containers, deploys in one click, and starts at $3.99/mo with BYOK: your API keys go straight from your container to the provider, so every model in this guide works exactly as described. The Plus tier ($9.99/mo) bundles free usage of GPT-6 Luna, DeepSeek V4 Flash, and Mercury 2: that includes the exact DeepSeek Flash family dominating the real-usage leaderboard, plus the successor to the GPT-5.6 Luna in its top ten. Bundled usage is capped per rolling 5-hour window and per week, so heavy agents should still bring a key. There is also a free tier for one personal agent that sleeps after 15 minutes idle and wakes when you message it; a truly always-on agent starts at $3.99/mo.

FAQ

What is the default model for Hermes agent?

There is none. A fresh install has an empty model setting, and onboarding makes you pick a provider and model. If you configure a provider without choosing, the current catalog fallback is z-ai/glm-5.2.

Can I run Hermes agent on Nous's own Hermes 4 models?

You can: Hermes-4-70B and Hermes-4-405B are on Nous Portal at heavily discounted rates. But Nous itself recommends against using them inside Hermes Agent, because they are tuned for chat and reasoning rather than the agent's tool-calling loop.

Does the Claude subscription route work on Claude Pro?

Yes, via the official Claude Subscription DirectSDK plugin (September 2026, experimental). The older built-in Anthropic OAuth path is stricter: Claude Max with purchased extra-usage credits only, and Pro is excluded.

What is the cheapest model that runs Hermes well?

GPT-6 Luna at $0.10 / $0.50 per million tokens, once you configure around its reasoning-plus-tools quirk. DeepSeek Flash is the community's other budget staple, especially off-peak. Genuinely free routes exist through Nous Portal but get unreliable at tool calling.

How much does running Hermes cost in total?

The software is free. On a modeled 40-turn/day workload, a flash-class model costs $6 to $17 a month in API fees, a flagship $120 or more. Add somewhere always-on to run it: a $5 to $24 VPS, or managed hosting from $3.99/mo. Our is Hermes agent free post breaks down the full cost story.
Vishnu

Written by

Vishnu

Founder at Agent37, which runs managed hosting for OpenClaw and Hermes agents for 1,000+ users. Writes about what actually breaks when you leave an AI agent running.