Source: raw/reddit-1ufrtsf.md — r/hermesagent “Models, Providers & Plans Megathread — June 2026” (OP u/Jonathan_Rivera, score 61, last updated 2026-06-25; community-aggregated from 31+ r/hermesagent threads + 290+ comments, building on the May 2026 megathread by u/digitalnomadpdx). Also cites raw/reddit-1uknhk4.md (u/AnticitizenPrime) for the recursive-skill-path context-tax reproduction. Refreshed 2026-07-10 with raw/x-account-nousresearch-2075270903458373640.md (first-party @NousResearch: GPT-5.6 now available via Nous Portal). Refreshed 2026-07-12 with raw/x-account-nousresearch-2074260103469892045.md (first-party @NousResearch: Hy3 free in Nous Portal for two weeks).
This is the cloud/paid counterpart to the local-models guide (hermes-apple-silicon-local-models) — which API, subscription plan, and model to point Hermes at when you are not self-hosting. Everything below is community-reported, not verified fact. Every dollar figure, token allowance, plan name, and “best model” verdict is the consensus of a community megathread as of June 2026, attributed to the thread — not benchmarked or priced by this wiki. The thread is mod-maintained and regenerates monthly, so prices and plans decay fast; treat this as a dated snapshot and re-verify against each provider before committing. ^[the “community pick” / “best” framing throughout is the megathread’s editorial consensus, not a test this wiki ran]
Key Takeaways
- The most-recommended budget stack (community): OpenCode Go (60 in API credits”) + a Minimax 20/mo near-unlimited setup that the thread says covers ~90% of workloads.
- Never use one model for everything — route. The thread’s headline claim: smart two-tier routing (cheap orchestrator + expensive powerhouse, invoked only when needed) “buys you 5-10x more capability per dollar than any single subscription.”
- DeepSeek direct API is the cost anchor. Community-reported as 4-5x cheaper than the same model via OpenRouter resellers, with automatic prompt caching (cache reads ~0.30-2-6/day (Pro).
- Community “best” picks (June 2026): GPT-5.5 (via OpenAI Codex) for hard coding; DeepSeek v4 Pro as the value powerhouse / daily driver; DeepSeek v4 Flash or GPT-5.4-mini as the cheap-fast orchestrator; owL-alpha (free on OpenRouter) as the best free model.
- Subscriptions buy predictable billing; pay-per-token buys flexibility. The thread’s repeated horror story is the “$100 surprise day” (e.g. Claude Sonnet via OpenRouter at reseller markup) — caps + subscriptions are the fix.
- 64K tokens is the reported context-window floor for Hermes — below that the tool/skill/memory injection overflows on the first complex task. Prompt caching (DeepSeek/Anthropic/OpenAI direct APIs) reportedly cuts 50-90% of repeated-schema token cost.
- Local 24GB Macs hit a reliability wall past basic chat/bash — including a model-agnostic
write_fileconfabulation bug. A ~20-quant field test (Qwen3.5-9B, Gemma 4 26B/12B, gpt-oss-20b, Qwen3.6-35B) found both Qwen and Gemma fabricating plausible “protected file” deny errors on writes that never actually failed — indistinguishable from a real deny without checking disk. The operator’s verdict: switch primary to a cloud model for reliability. See the field signal below.
First-party model update — GPT-5.6 available via Nous Portal (added 2026-07-10)
[X signal — @NousResearch, first-party, 2026-07-09] Source: raw/x-account-nousresearch-2075270903458373640.md. Unlike the rest of this article, this is a first-party confirmed fact, not a community-reported claim: GPT-5.6 is now available for use in Hermes Agent, accessible via Nous Portal. The announcement landed one day after Hermes Agent Cloud launched (2026-07-08), in the same first-party release week. GPT-5.6 is OpenAI’s three-tier model family — Sol / Terra / Luna (see Luna) for the full release writeup) — but the announcement tweet does not specify which tier(s) are exposed inside Hermes or under what routing/pricing. This slots in next to the existing GPT-5.5 / GPT-5.4-mini picks in the tables below; expect the community megathread to re-rank once GPT-5.6 gets real-world usage reports.
First-party model update — Hy3 free in Nous Portal for two weeks (added 2026-07-12)
[X signal — @NousResearch, first-party, 2026-07-06] Source: raw/x-account-nousresearch-2074260103469892045.md. A second first-party confirmed fact from the same release week as the GPT-5.6 addition above: Hy3 — a new 295B-parameter MoE model from Tencent Hunyuan — is available free in Nous Portal for two weeks (announced 2026-07-06). Nous positions it for cost-effective agentic use: coding, reliable tool-calling, reasoning, and 256K long-context performance — squarely at the cheap-orchestrator / value-daily-driver slots the community tables below currently give to DeepSeek v4 and GPT-5.4-mini. The announcement does not state post-promo pricing, and the free window is time-boxed (ending roughly 2026-07-20^[inferred — derived from “two weeks” + the announcement date, not stated explicitly]). A follow-up post covers Nous Portal’s model access, discounts, and billing features. Expect the megathread to re-rank once Hy3 gets real-world usage reports.
TL;DR — community picks (as reported, June 2026)
The thread’s own decision table. All picks are community consensus, not verified recommendations.
| Decision | Community pick | Runner-up | Thread note |
|---|---|---|---|
| Best overall paid API | DeepSeek v4 Pro (direct) | DeepSeek v4 Flash | Pro for orchestrator, Flash for workers. “$60 got one user 8B tokens.” |
| Best value subscription | OpenCode Go ($10/mo) | Minimax $10 token plan | Go = “~10”; Minimax = “virtually unlimited” background work |
| Best premium subscription | Nous Portal ($20) | OpenAI Codex ($20 ChatGPT) | Fixed monthly cost, no billing surprises |
| Best free model (no catch) | owL-alpha (OpenRouter) | Nemotron 3 Super 120B (free) | owL-alpha: “absolute beast at tool usage” |
| Best coding model | GPT-5.5 (via Codex) | DeepSeek v4 Pro | GPT-5.5 the “undisputed king” for complex coding |
| Best orchestrator model | GPT-5.4-mini | DeepSeek v4 Flash | Fast, cheap, handles ~90% of routing/web search/light tasks |
| Most predictable billing | OpenCode Go + Minimax stack | Nous Portal | Subscriptions avoid surprise $100 days |
Providers & plans (community-reported pricing)
Prices, free tiers, and token allowances below are as stated in the megathread — verify each against the provider’s own page before relying on it.
Tier 1 — community favorites
| Provider | Reported price | Key models | Best for (per thread) | Watch for |
|---|---|---|---|---|
| DeepSeek (direct) | Pay-per-token. Flash ~0.20/M out (cache reads 0.44/M in, 0.30-2-6/day Pro | v4 Pro, v4 Flash, Coder | Primary orchestrator, heavy coding, cost-sensitive work | 4-5x cheaper direct vs OpenRouter; 503/throttling at US peak hours; China-based (data-sovereignty concern for some) |
| OpenCode Go | 5 first month); ”~$60 of credits” | DeepSeek Flash/Pro, Minimax M3, MiMo 2.5 Pro, GLM, Kimi | Best-value subscription; ~90% of agent workload | Lacks some multimodal (e.g. Gemma 4); weekly/monthly caps |
| Nous Portal | $20/mo | Hermes models, DeepSeek, Qwen, GPT-5.6 (added 2026-07-09) + routing | All-in-one, predictable billing; first-party (Nous builds Hermes) | Non-free models exhaust the $20 quickly; “Hermes Plus” Opus routing reportedly not identical to native Claude Code |
| OpenAI Codex | $20/mo (ChatGPT Plus, OAuth/BYOK) | GPT-5.5, GPT-5.4-mini | Complex coding, deep reasoning — best as “senior fixer,” not daily driver | OAuth only (no API key); burns weekly rate limits fast; shares ChatGPT limits |
Tier 2 — strong alternatives
| Provider | Reported price | Key models | Best for (per thread) | Watch for |
|---|---|---|---|---|
| Minimax | $10/mo token plan (“virtually unlimited”; 15K req/week high-speed, 1.5K/5hr) | M3, M2.7 | Background/auxiliary agent work, stable everyday use | M2.7 reliable but uncreative (M3 better); prone to looping without guardrails |
| Xiaomi MiMo | 13/mo annual (2.4B tokens/yr) | MiMo 2.5, MiMo 2.5 Pro | Agentic intelligence, vision, coding — “steal at current price” | No caching (burns tokens faster than DeepSeek); sometimes over-eager |
| Kimi/Moonshot | Pay-per-token (OpenRouter or direct) | K2.6, K2.7 | Best open-source Hermes main model; strong tool calling | Strict quotas; occasional Chinese chars in output; tends to overthink coding |
| OpenRouter | Pay-per-token (variable); free tier with $10 credit | 200+ models incl. owL-alpha (free) | Model experimentation, fallback chains, free-tier models | 4-5x markup vs direct on DeepSeek; avoid silent auto-routing; pin model IDs |
| Gemini (Google) | Pay-per-token / free tier | Flash 2.5 (free), Pro 2.5 | Free tier for light tasks; strong vision | Free-tier rate limits; OAuth-sub risky as BYOK |
| Ollama Cloud | 22 credits), $100 tier | Free/open models only | Hassle-free hosted local-style models | 3 concurrent-connection limit crashes cron jobs; reportedly degraded; no frontier models |
Tier 3 — budget / niche
| Provider | Reported price | Best for (per thread) | Watch for |
|---|---|---|---|
| GLM 5.1 / 5.2 (NeuralWatt) | Free $5 credit, then PAYG | Deep reasoning when speed doesn’t matter; stable | 5.2 pricier; painfully slow (reported 18hrs vs 1hr for GPT-5.5); prone to looping |
| Anthropic Claude (sub) | $20/mo | Opus 4.7 / Sonnet 4.5 — high-quality reasoning, code review | Agentic use explicitly discouraged by Anthropic; token hog; community says route via OpenRouter, don’t risk the account |
| NVIDIA NIM | Free tier | Nemotron 3 Super 120B — best emergency fallback | Smaller ecosystem; genuinely free |
| Grok / superGrok | 30/mo X Premium | Multi-modality, voice, tool calling | $30/mo X Premium reportedly ~2hrs agent work; API gives much more; weak at coding |
| GitHub Copilot | $10/mo | GPT-5.4 via Copilot (ACP transport) | Separate from ChatGPT Plus |
| OpenCode Zen | $10/mo | Curated model selection | Smaller selection than Go; Go is the better value |
| NanoGPT | $12/mo | Uncensored models only | ”Sketchy AF”; slow, low limits, verbose; not recommended as primary |
| Qwen OAuth | Subscription / PAYG | Qwen 3.6 / 3.5 direct | Newer; fewer community data points |
| Stepfun AI | Free voucher ($100) | Stepfun 3.7 Flash | Voucher may no longer be offered |
| Dappnode Nexus | ~$22/mo (€20) | Private/anonymous models (Kimi K2.6, MiniMax 2.7, DS 3.2, GLM 5, Qwen) | For privacy-conscious users |
Model-for-task (community consensus)
- Complex coding / deep reasoning: GPT-5.5 (Codex) or DeepSeek v4 Pro.
- Orchestrator / routing / web search / light tasks: GPT-5.4-mini, DeepSeek v4 Flash, or Kimi K2.6.
- Background / cron agents: Minimax M3 (“virtually unlimited” $10 plan) or DeepSeek v4 Flash via OpenCode Go — never expensive models.
- Best free: owL-alpha (OpenRouter) for tool use/coding; Nemotron 3 Super 120B (NVIDIA NIM/OpenRouter) as emergency fallback.
- Highest-quality reasoning / code review: Claude Opus 4.7 — but reported as a token hog, not for daily driving.
Routing strategies (orchestrator vs worker splits)
The thread’s central thesis: split a cheap-fast orchestrator from an expensive-smart powerhouse, and add a free fallback chain. Patterns, ranked roughly by sophistication:
- Gold standard — two-tier + fallback. Tier 1 orchestrator (GPT-5.4-mini / DeepSeek v4 Flash / Kimi K2.6) handles ~90% of chat, routing, web search, light tasks. Tier 2 powerhouse (GPT-5.5 via Codex / DeepSeek v4 Pro) is invoked only for complex coding, deep research, multi-step synthesis. Fallback chain when limits hit:
GLM-5.1 → Nemotron 3 Super 120B (free) → owL-alpha. - Multi-provider stack. Primary (DeepSeek v4 Pro direct or GPT-5.5/Codex) → orchestrator/auxiliary (DeepSeek v4 Flash or Minimax M3) → workers (MiMo 2.5 Pro or Kimi K2.6 via OpenCode Go) → fallback (owL-alpha / Nemotron on OpenRouter free tier). Spreading across providers avoids hitting any single rate limit.
- OpenRouter as fallback pool. Pin specific model IDs to fixed providers; use free tier (owL-alpha, Stepfun 3.7 Flash) for pennies/day. Never leave auto-routing on for anything that mutates state — the thread warns the model can silently switch mid-task and “you won’t know until it breaks.”
- Profile-level routing (advanced). A root coordinator profile routes to specialized profiles: coder → GPT-5.5 (Codex), researcher → Gemini 3.1 Pro, pm → Minimax M3 or DS v4 Pro. (“Now my profiles talk to each other.”) See hermes-profiles-multi-instance.
- Event-driven (lowest cost). Lightweight watchers poll cheaply and wake Hermes only when a filter matches, instead of cron-based fixed schedules — saves tokens on idle polling. (Community project: Watchline.)
- LiteLLM proxy (max resilience).
Hermes → LiteLLM → tiered provider poolfor provider-level failover; more setup, maximum resilience.
Subscription vs pay-per-token (decision guide)
- Use a subscription ($10-20/mo fixed) if you want predictable billing, use Hermes daily, prefer “set and forget,” or don’t yet know your usage patterns.
- Use pay-per-token (DeepSeek direct, OpenRouter) if usage is bursty, you’re extremely cost-sensitive and will monitor burn, you run multiple worker profiles on cheap models, or you self-manage fallback chains.
- Hybrid (most common in the community): subscription for the primary model → cheap pay-per-token for workers/auxiliary. (“OpenAI $20/mo + DeepSeek v4 Flash for workers. Pro for main orchestrator. Multiple providers so I never hit a single rate limit.”)
Cost-saving tips (community consensus, ranked by reported impact)
- Use DeepSeek direct, not via OpenRouter (4-5x cheaper; cache reads $0.004/M).
- Two-tier routing — Flash for ~90%, Pro/Codex only for hard work (reported 5-10x savings).
- Offload auxiliary tasks (compression, title generation, session search) to Flash or Minimax.
- Trim skills and toolsets — every enabled tool adds schemas to every prompt (one user: “50K+ tokens on every prompt”; trimming saved 60%).
- Use prompt caching (DeepSeek/Anthropic/OpenAI direct APIs) — 50-90% off repeated tool-schema tokens.
- Event-driven over cron polling.
- Minimax $10 token plan for background workers (“virtually unlimited”).
- Set spending caps — every provider supports them; the $100 surprise day is preventable.
Field signal — the context tax, measured (added 2026-06-29). [r/hermesagent, u/Beautiful-Elk-587, score 16, raw/reddit-1uiyk8i.md] An operator put a local inspection proxy between Hermes and Ollama and watched a 40-token direct prompt balloon to 20,538 tokens through Hermes (~513x) for the same input. The cause is the bootstrap Hermes injects on every call with no relevance-gating: full system prompt + user profile + rules + memory + skills list + computer-use instructions + all tool schemas + max_tokens: 65536 — even “say hi in one word” ships browser/computer-use tool definitions. This is the quantified version of cost-saving tip #4 (trim tools), and it explains both the cloud rate-limit/credit burn and the local-model choke (a small model spends minutes processing the agent bootstrap before it answers; CPU pegged ~5 min vs near-instant on a direct call). The operator’s architectural ask is pre-call context/tool selection — no tools for plain chat, load only the toolset the task needs, inject only relevant memory/skill snippets; until the framework does that automatically, trimming skills/tools/memory and running minimal profiles is the manual lever. ^[single-operator field report; the 513x figure is one measurement, reproducible via the proxy method described] A second, concrete reproduction (u/AnticitizenPrime, raw/reddit-1uknhk4.md, score 17) pins down a specific, falsifiable failure mode behind this tax: auditing OpenRouter logs exposed recursively duplicated skill-tree entries in the system prompt — the same skill re-listed under ever-deeper nested paths (general/flight-search → general/flight-search/flight-search → general/flight-search/flight-search/flight-search → …, repeating for pages) — and de-duplicating the recursive skill paths cut token usage by ~35% on every API call, corroborating the “trim skills/tools” lever with a named, reproducible root cause. ^[community-reported single-operator log audit; not independently verified]
Mixture of Agents (MoA) — first-party ensemble (added 2026-06-28)
A newer first-party answer to “which model?” — don’t pick one; fan out to several. Hermes now exposes Mixture of Agents (MoA) presets as virtual models: a prompt is sent to two or three frontier models in parallel and an aggregator model synthesizes them into one final answer (one response, several models cross-checking each other). You select it from the model dropdown like any model and can toggle it mid-task with /moa. (Source: raw/reddit-1uh9zpt.md; first-party @NousResearch via raw/x-bookmarks-recent-digest-2026-06-28.md.)
- Claimed quality (vendor benchmark, unverified): Nous reports MoA scores ~8% higher than Opus 4.8 and ~11% higher than GPT-5.5 on its own upcoming HermesBench — framed as “capabilities beyond the publicly available frontier.” Treat as a vendor claim, not an independent result. ^[the benchmark deltas are Nous’s own, not verified here]
- The economics are the surprise. A community-reported setup fans every prompt to gpt-5.5 + deepseek-v4-pro + sonnet-4.6 with opus-4.8 as the aggregator: a single Opus call costs ~0.15 — because most token cost is the shared system prompt / tool schemas, so the extra reference models run on stripped-down context and add almost nothing.
- How it relates to the routing strategies above. MoA is a parallel ensemble (breadth — many models on one prompt, then aggregate), whereas the two-tier orchestrator/worker split is conditional routing (depth — cheap model handles most, escalate to a powerhouse only when needed). MoA reframes the question from “which model should I use?” to “what’s the best team of models for this task?” — complementary to, not a replacement for, the routing patterns above.
Field signal — local 24GB reliability caveat, incl. write_file confabulation (added 2026-07-02)
[Reddit signal — r/hermesagent 2026-07-02] Source: raw/reddit-1ulkiqc.md (28 score, 50 comments, OP u/alexgranford, “MODELS” flair). A detailed local-inference field report that complicates “just self-host” as an alternative to the paid stack this article covers — worth reading alongside self-hosted counterpart guide before committing to either path.
- Stack tested. Apple Silicon M4 Pro 24GB running oMLX (after trying Ollama and Rapid-MLX); agent hosted on a UGREEN NAS DXP4800 Pro via Docker (separate containers for the agent and SearXNG); clients mostly Telegram.
- ~20 quant combos tested across Qwen3.5-9B (OptiQ 4-bit, MLX 4/8-bit, a Deckard/Heretic merge strictly worse than OptiQ), Gemma 4 26B-A4B (nvfp4, mxfp4, two QAT 4-bit converters, community oQ), Gemma 4 12B (4/6/8-bit + OptiQ), gpt-oss-20b, and Qwen3.6-35B-A3B (oQ3/oQ4). Notable throughput result: the 26B Gemma decoded faster than the 9B Qwen on single-request (~64 vs ~38 tok/s) — interleaved local/global attention won the bandwidth game; Qwen only pulled ahead under batching.
- Tooling issues hit: an oQ-on-QAT-checkpoint bug (
_TrackedTensor object has no attribute 'swapaxes', filed upstream, closed the day of posting); a tight ~22GB ceiling causing prefill OOM-kills on large prompts (the agent idles right after announcing intent, with no error surfaced); Gemma 12B 8-bit OOMing prefill at ~12K context on 24GB. - The finding that matters — model-agnostic
write_fileconfabulation. Both Qwen and Gemma would report awrite_filefailure, quoting a plausible “protected system/credential file” deny error, and mark the step failed — but the file write was never actually attempted or failed; the error was fabricated. Gemma used the exact internal deny-string wording, making the fake indistinguishable from a real deny without an actual disk check. The poster verified this at the source level — reproducible across both model families and multiple quants, not a single-quant or model-size artifact. Practical consequence: once the agent’s own report of a file operation can’t be trusted, delegating real work becomes unreliable; the poster’s posture became disk-verifying every local file op and routing must-succeed writes through a cloud delegation lane. ^[single-operator field report; the confabulation was verified at the source level by the poster, not independently reproduced by this wiki] - Verdict. Past basic chat and bash, every local model tested was “a liability” — looping, rabbit holes, breaking working skills. The poster switched to GPT-5.5 via Codex as primary: roughly 10 skills rationalized/fixed within about 30% of a week, running reliably with no confabulation. Explicit trade-off acknowledged: going cloud-primary gives up the local-sovereignty rationale that motivated the local build in the first place.
This is a data point for the cloud/paid stack this article covers, from an operator who tried the local path first and hit a specific, named trust failure rather than just a cost or speed problem — complements the local guide’s own “hybrid local + cloud fallback” recommendation and this article’s cost-saving tip #4 (trim tools) and the context-tax field signal above, but the failure mode here is reliability/trust rather than token cost.
Two first-party provider updates (2026-08-07 addition)
1. Local inference — Hermes works natively with Actual (2026-08-06). Source: raw/x-account-nousresearch-2085184069302935999.md (@NousResearch, first-party). Hermes Agent now works out of the box with Actual, a local inference stack from @actualinc, positioned as optimized for personal compute with low CPU impact and multi-platform support.
This adds a first-party-blessed local route to a table that is otherwise hosted/API providers. Read it against the field signal above rather than as a replacement for it: the 24GB local reliability report — including the model-agnostic write_file confabulation — is about model reliability at small sizes, not about the serving stack, so a better local runtime does not by itself resolve it. See self-hosted guide for the fuller picture. Actual’s specifications here are vendor claims relayed through a first-party announcement; none are independently verified.
2. DeepSeek V4 Flash 0731 promo, and a bold benchmark claim (2026-08-04). Source: raw/x-account-nousresearch-2083953441571742191.md (@NousResearch, first-party). DeepSeek V4 Flash 0731 offered at 90% off on Nous Portal for a limited period, in partnership with @novita_labs; other models at 20% off, some at 50% (GPT-5.6 Terra and Luna named).
The promotional pricing is transient and should not be written into the table above. The claim attached to it is worth recording precisely because it is falsifiable:
DeepSeek V4 Flash 0731 is “>1000x cheaper than Fable 5 on comparable tasks while beating it on Terminal-Bench 2.1.”
Treat as unverified vendor marketing. ^[first-party promotional claim; no methodology, no linked benchmark run, and the “>1000x” figure is not reconcilable with the per-token pricing in the table above without knowing what “comparable tasks” means — likely a task-completion-cost comparison rather than a token-price ratio]
Two same-week corroborating signals, both weak individually and both pointing the same way: an r/hermesagent report of the Hermes + DeepSeek V4 Flash 0731 pairing feeling markedly faster than GPT-5.6 Sol at high/medium effort (raw/reddit-1vhffj4.md, score 251 — subjective, triaged skip), and the Codex-community “escape hatch” finding of DeepSeek V4 Flash at max reasoning running ~100–200 (see Orchestrator + Cheap-Worker Routing). The cost direction is corroborated from three independent places; the Terminal-Bench 2.1 claim is corroborated from none.
Try It
- Verify before committing. Treat every price/plan/allowance here as a June-2026 community claim — open the provider’s own pricing page and confirm before subscribing. The thread is explicit: “snapshot of community consensus, not official advice.”
- Start cheap, route from day one. The community default is OpenCode Go (10 token plan; set a cheap-fast orchestrator (DeepSeek v4 Flash / GPT-5.4-mini) and reserve a powerhouse (GPT-5.5 via Codex / DeepSeek v4 Pro) for hard tasks only.
- Wire a free fallback chain so rate limits don’t stop you:
GLM-5.1 → Nemotron 3 Super 120B (free) → owL-alpha. Configure viahermes config→fallback_providers, credential pools, or a LiteLLM proxy. - Pin model IDs and set spending caps. Never leave OpenRouter on auto-routing for anything that mutates state; pin to fixed providers and cap spend in every dashboard.
- Prefer direct APIs over resellers and enable prompt caching where supported (DeepSeek/Anthropic/OpenAI) to cut repeated-schema token cost.
- For local/self-hosted instead, see hermes-apple-silicon-local-models — the hybrid pattern (fast local model + cheap cloud fallback) pairs directly with the routing strategies above.
Open Questions
- Fast decay — this is a dated snapshot. The megathread regenerates monthly and explicitly invites corrections; all pricing, token allowances, and “best model” picks are community-reported as of June 2026 and likely stale within weeks. None of the dollar figures here have been independently verified by this wiki — re-check each provider’s own page before relying on it.
- No official Nous pricing comparison in scope. The thread’s provider verdicts (e.g. DeepSeek v4 Pro “smarter than Claude Sonnet,” Minimax M3 “as good as GPT-5.5-low”) are anecdotal claims from individual users, not benchmarks. Treat model-quality rankings as opinion until corroborated.
- Hermes config specifics are summarized, not tested. Setup hints (
hermes auth add ...,.envAPI-key vars,fallback_providers, LiteLLM proxy) come from the thread; verify against current Nous docs before wiring them. - GPT-5.6 tier and pricing inside Hermes (added 2026-07-10). The first-party announcement confirms GPT-5.6 availability via Nous Portal but doesn’t state which Sol/Terra/Luna tier(s) are exposed, or how they’re priced/rate-limited relative to GPT-5.5.
Related
- hermes-apple-silicon-local-models — the local/self-hosted counterpart; the hybrid (local + cheap cloud fallback) pattern is the bridge between the two.
- nous-portal — the first-party $20/mo subscription backend listed as a “best premium” pick here.
- Luna) — OpenAI’s three-tier family; GPT-5.6 became available in Hermes via Nous Portal on 2026-07-09.
- Hermes Agent Cloud — same first-party release week (2026-07-08 to 2026-07-10) as the GPT-5.6 addition above.
- hermes-memory-providers — provider choice interacts with the context tax that drives token cost.
- hermes-grok-sub-setup — the Grok/superGrok subscription path discussed in Tier 3.
- hermes-profiles-multi-instance — profile-level routing (one model per profile) from the advanced routing patterns.
- glm-5-series-zai — background on the GLM-5 / 5.2 open-weight models the thread lists as a budget reasoning option.
- _index — Hermes Agent topic hub.