Hermes provider and account routing
Use when picking which subscription Hermes bills. Use when the user has more than one account/subscription for the same provider and wants a specific one used, when a provider is rate-limited and work must move to another credential, or when someone asks whether a role/profile/session can be pinned to a particular subscription.
The installed CLI and agent/credential_pool.py are authoritative. Discover flags with --help; do not invent them.
Core model
- Provider (
anthropic,openai-codex,openrouter,copilot,custom:*) has a credential pool: an ordered list of entries in$HERMES_HOME/auth.jsonundercredential_pool.<provider>. - Each entry has
label,auth_type(oauth/api_key),source(manual:hermes_pkce,claude_code,env:VAR,gh_cli, ...),priority(ascending; 0 is tried first), and status fields (last_status,last_error_reason,last_error_reset_at,request_count). - Selection strategy is per provider, default
fill_first. Supported:fill_first,round_robin,least_used,random, set viacredential_pool_strategies.<provider>in config. - Selection is pool-wide, not per-session. There is no supported flag that pins one credential to one session, subagent, or role. Saying "use account B" in a prompt does not switch authentication. Tell the user this plainly instead of implying per-role billing control.
- Separate subscriptions do not pool capacity. More profiles or roles do not create more limit headroom.
Inspect first
hermes auth list # pool per provider; "←" marks the active entry; shows rate-limit state + reset window
hermes auth status <provider> # logged-in / logged-out for that provider
hermes config get model # default model + aliases
hermes auth list is the fastest way to see a 429: an entry renders like
#1 device_code oauth device_code rate-limited usage_limit_reached (429) (5d 17h left).
Check this before dispatching any long provider-pinned cron/manager run — a scheduled job on an exhausted route burns wall-clock and returns RuntimeError: HTTP 429: The usage limit has been reached after tens of minutes, with zero progress.
For entry-level detail (priority, source, counters), read auth.json directly but scrub access_token/refresh_token/*_key/secret before printing. Never echo token material into chat or a report.
Changing which account is primary
hermes auth exposes add, list, remove, reset, status, logout — there is no reorder/priority subcommand. To promote an existing pooled credential:
- Confirm the exact labels with
hermes auth list. - Back up: copy
auth.jsonto a dated.bakin the same directory. - Rewrite only the
priorityintegers for the affected provider's entries, re-sort the list ascending, and write atomically (temp file +replace). - Read back with
hermes auth list— the←and#1position must reflect the intent. - Run one real probe (
hermes chat -q 'Reply with exactly: ROUTE_OK') to prove the reordered pool still authenticates. - Tell the user it was a hand-edit, not a supported command, and that a later
hermes auth addfor that provider may re-normalize order.
Adding a new pooled OAuth credential from an agent session
hermes auth add <provider> --type oauth --no-browser --label <name> starts the PKCE flow. On a headless box use --no-browser so it prints the authorization URL instead of trying to open a browser; hand that URL to the user and collect the code they paste back. Drive it as a PTY background process — a plain pipe hangs the readline prompt.
Exact working sequence (validated 2026-09):
terminal(background=True, pty=True)running thehermes auth add ... --no-browser --timeout 900.process_manage(action='poll')after ~6s; extract thehttps://claude.ai/oauth/authorize?...URL fromoutput_previewand give it to the user verbatim.- Submit the returned code with
process_manage(action='submit', data='<code>')— the parameter isdata, NOTinput. Usinginputsilently sends only a newline (bytes_written: 1/0), the CLI reads an empty line, printsNo code entered, and exits 1. This wasted several attempts before the fix. submitappends Enter itself;writesends raw bytes with no newline. Do not hand-write to/proc/<pid>/fd/0— it appends to the display but the app's controlling-terminal readline never registers the submit.- Poll again; on success
hermes auth listshows a second entry for that provider. Verify it is NOT rate-limited before firing dependent jobs.
OAuth codes are single-use and PKCE-bound: each hermes auth add invocation prints a different code_challenge/state. A code from an earlier/aborted flow will not validate against a fresh prompt — the pasted code's #<state> suffix must match the live prompt's state= query param. If they diverge, the user re-sent an old code; restart the flow and have them re-authorize at the current URL.
A fresh login lands only in the auth store of the profile/$HERMES_HOME that ran hermes auth add. Cron jobs and this session may run the default profile; if the user logs in elsewhere the new credential will not appear in hermes auth list here. Confirm auth.json updated_at bumped AND a new pool entry exists — an updated_at change alone can just be a last_status write from a failed fire, not a new credential.
Once a healthy second credential is in the pool, fill_first auto-skips a 429'd entry to the next healthy one with no per-job change needed — repinning individual jobs is unnecessary and, if you pin them to the still-exhausted provider, actively wrong.
When EVERY primary provider is 429'd at once
Adding a same-provider account is the clean fix, but when both anthropic AND openai-codex are exhausted and the user wants work to continue now, the lever is a different provider entirely. OpenRouter (OPENROUTER_API_KEY) reaches nearly every model and is usually not rate-limited when the first-party routes are — repin coordinator/worker cron jobs there with hermes cron edit <id> --provider openrouter --model <slug> (e.g. deepseek/deepseek-v4.1-flash), and smoke-test the route first with hermes -z 'Reply with exactly: OK' --provider openrouter --model <slug> before trusting it. Keep role→model bindings, not provider pins: name the role, bind the model in config, swap the binding on the owner's word. After all healthy jobs funnel through ONE account, that single account becomes the new bottleneck — stagger heavy recurring jobs onto offset, slower cadences (e.g. every 20m/25m/30m instead of all every 10m) with hermes cron edit <id> --schedule 'every 25m' so their concurrent requests fit one account's rate window; read back the persisted schedule (cron edit can exit 0 with failure text).
Other supported levers, prefer them when they fit:
hermes config set credential_pool_strategies.<provider> round_robin— alternate/share load across accounts instead of strict primary+failover.hermes auth remove <provider> <index|id|label>— drop an account from the pool entirely.hermes auth reset <provider>— clear exhaustion status after a limit window passes.
Order normalization (why a hand-edit can move)
_normalize_pool_order in agent/credential_pool.py sorts manual entries (manual:hermes_pkce and friends) by priority first, then appends seeded entries ranked by source: env:ANTHROPIC_TOKEN < env:CLAUDE_CODE_OAUTH_TOKEN < hermes_pkce < claude_code < env:ANTHROPIC_API_KEY. Hand-set priorities among manual OAuth logins are preserved relative to each other; seeded/env credentials get re-ranked by source on the next upsert.
Probe the execution plane you will actually use
A provider/model is not proven unavailable merely because a sibling client fails. Standalone coding CLIs and Hermes can use different credential stores, pools, account scopes, and rate-limit state even when both display the same provider and model names.
For any exact-route assignment:
- Identify the intended execution plane: Hermes-managed provider, standalone CLI, gateway/profile, or scheduled job.
- Probe the exact
provider/modelthrough that same plane with a no-tools sentinel response. - Treat a sibling client's result only as evidence about that client's credential plane.
- If the intended plane succeeds, launch through it; do not route the real task through the failed sibling client.
- Record both the exact route and execution plane in the attempt receipt. Never silently substitute a model or reset attempt lineage.
Example Hermes probe:
hermes chat --provider openai-codex --model <exact-model> \
--ignore-user-config --ignore-rules --safe-mode \
-q 'Reply exactly ROUTE_OK and do not use tools.'
For a write-capable continuation, remove --safe-mode, bind a narrow worktree and prompt, and explicitly select the needed toolsets/approval policy. Preserve prior partial changes and failed-attempt evidence before relaunching.
Validated session details and the client-vs-Hermes interpretation are in references/execution-plane-route-probes.md.
Profiles vs pools
Profiles ($HERMES_HOME/profiles/<name>/) are the persistent-configuration boundary — own config, skills, memory, and potentially own auth.json. They are not automatically a strict account boundary: a profile without local credentials can fall back to the root profile's. If the user wants "this profile only ever bills account X," authenticate that account inside the profile and verify effective resolution — do not assume the profile name enforces it, and never hand-copy OAuth token files between profiles (refresh collisions).
Honesty rules
- Report what you verified: pool order plus a successful probe prove the route works, not that a specific account was billed for a specific past turn.
request_count/last_statuscan remain0/nullfor gateway-served turns. - A provider rate limit on one subscription is not a defect in another; state which credential/provider the 429 belongs to.
- When the user asks "which model are you using", answer with the live resolved model id, and separate that from whatever routes the background fleet/cron jobs use.
Pitfalls
- When a role maps to an interchangeable model pool (e.g. coordinator = Fable
anthropic∨ Astraopenai-codex), pick the pool member whose PROVIDER is actually healthiest — readhermes auth listreset windows, do not assume. It is easy to move a stuck job onto the more rate-limited provider: verify both providers' 429 state and remaining windows before repinning, then read back the persisted route. Repinning to a route that is also exhausted just relocates the outage. hermes chat -qhas no--no-streamflag; an invented flag exits 2 with the full subcommand list. Checkhermes <cmd> --helpfirst.- Broad recursive greps across
$HERMES_HOMEplus large work trees time out (60s+). Scope to exact files (auth.json,config.yaml) or a single package directory. hermes config get providermay printConfig key not setwhilehermes config get modelstill showsprovider: anthropicinside the model block — read the model block, don't conclude the provider is unset.- A coding CLI's own login store (e.g. Claude Code's
~/.claude.jsonoauthAccount, or an alternateCLAUDE_CONFIG_DIR) is independent of Hermes's pool. Changing one does not change the other; check both when a user asks "which account is in use". - Never hand-edit
config.yaml; usehermes config set KEY VALUE.auth.jsonordering has no setter, which is the narrow exception documented above.
Session-specific recipe and exact observed output: references/anthropic-multi-subscription.md.
Supporting files: this skill's supporting files are held in the docsite at
docs/15-skills/_support/autonomous-ai-agents/hermes-provider-account-routing/— fetch them fresh fromjknash/docsitemain alongside this page. Source:jknash/hermes-shared-skills· branchhermes-jkdev001@1d0d545c3970·skills/autonomous-ai-agents/hermes-provider-account-routing/· view source · Imported 2026-10-04. Supporting files (references, scripts) remain in the source repository.
version 1.0.0 · author Hermes Curator · license MIT.
Published by Muse · 2026-10-04.