Skip to main content

Hermes provider and account routing

Use when picking which subscription Hermes bills. Use when the user has more than one account/subscription for the same provider and wants a specific one used, when a provider is rate-limited and work must move to another credential, or when someone asks whether a role/profile/session can be pinned to a particular subscription.

The installed CLI and agent/credential_pool.py are authoritative. Discover flags with --help; do not invent them.

Core model​

  • Provider (anthropic, openai-codex, openrouter, copilot, custom:*) has a credential pool: an ordered list of entries in $HERMES_HOME/auth.json under credential_pool.<provider>.
  • Each entry has label, auth_type (oauth/api_key), source (manual:hermes_pkce, claude_code, env:VAR, gh_cli, ...), priority (ascending; 0 is tried first), and status fields (last_status, last_error_reason, last_error_reset_at, request_count).
  • Selection strategy is per provider, default fill_first. Supported: fill_first, round_robin, least_used, random, set via credential_pool_strategies.<provider> in config.
  • Selection is pool-wide, not per-session. There is no supported flag that pins one credential to one session, subagent, or role. Saying "use account B" in a prompt does not switch authentication. Tell the user this plainly instead of implying per-role billing control.
  • Separate subscriptions do not pool capacity. More profiles or roles do not create more limit headroom.

Inspect first​

hermes auth list # pool per provider; "←" marks the active entry; shows rate-limit state + reset window
hermes auth status <provider> # logged-in / logged-out for that provider
hermes config get model # default model + aliases

hermes auth list is the fastest way to see a 429: an entry renders like #1 device_code oauth device_code rate-limited usage_limit_reached (429) (5d 17h left). Check this before dispatching any long provider-pinned cron/manager run — a scheduled job on an exhausted route burns wall-clock and returns RuntimeError: HTTP 429: The usage limit has been reached after tens of minutes, with zero progress.

For entry-level detail (priority, source, counters), read auth.json directly but scrub access_token/refresh_token/*_key/secret before printing. Never echo token material into chat or a report.

Changing which account is primary​

hermes auth exposes add, list, remove, reset, status, logout — there is no reorder/priority subcommand. To promote an existing pooled credential:

  1. Confirm the exact labels with hermes auth list.
  2. Back up: copy auth.json to a dated .bak in the same directory.
  3. Rewrite only the priority integers for the affected provider's entries, re-sort the list ascending, and write atomically (temp file + replace).
  4. Read back with hermes auth list — the ← and #1 position must reflect the intent.
  5. Run one real probe (hermes chat -q 'Reply with exactly: ROUTE_OK') to prove the reordered pool still authenticates.
  6. Tell the user it was a hand-edit, not a supported command, and that a later hermes auth add for that provider may re-normalize order.

Adding a new pooled OAuth credential from an agent session​

hermes auth add <provider> --type oauth --no-browser --label <name> starts the PKCE flow. On a headless box use --no-browser so it prints the authorization URL instead of trying to open a browser; hand that URL to the user and collect the code they paste back. Drive it as a PTY background process — a plain pipe hangs the readline prompt.

Exact working sequence (validated 2026-09):

  1. terminal(background=True, pty=True) running the hermes auth add ... --no-browser --timeout 900.
  2. process_manage(action='poll') after ~6s; extract the https://claude.ai/oauth/authorize?... URL from output_preview and give it to the user verbatim.
  3. Submit the returned code with process_manage(action='submit', data='<code>') — the parameter is data, NOT input. Using input silently sends only a newline (bytes_written: 1/0), the CLI reads an empty line, prints No code entered, and exits 1. This wasted several attempts before the fix.
  4. submit appends Enter itself; write sends raw bytes with no newline. Do not hand-write to /proc/<pid>/fd/0 — it appends to the display but the app's controlling-terminal readline never registers the submit.
  5. Poll again; on success hermes auth list shows a second entry for that provider. Verify it is NOT rate-limited before firing dependent jobs.

OAuth codes are single-use and PKCE-bound: each hermes auth add invocation prints a different code_challenge/state. A code from an earlier/aborted flow will not validate against a fresh prompt — the pasted code's #<state> suffix must match the live prompt's state= query param. If they diverge, the user re-sent an old code; restart the flow and have them re-authorize at the current URL.

A fresh login lands only in the auth store of the profile/$HERMES_HOME that ran hermes auth add. Cron jobs and this session may run the default profile; if the user logs in elsewhere the new credential will not appear in hermes auth list here. Confirm auth.json updated_at bumped AND a new pool entry exists — an updated_at change alone can just be a last_status write from a failed fire, not a new credential.

Once a healthy second credential is in the pool, fill_first auto-skips a 429'd entry to the next healthy one with no per-job change needed — repinning individual jobs is unnecessary and, if you pin them to the still-exhausted provider, actively wrong.

When EVERY primary provider is 429'd at once​

Adding a same-provider account is the clean fix, but when both anthropic AND openai-codex are exhausted and the user wants work to continue now, the lever is a different provider entirely. OpenRouter (OPENROUTER_API_KEY) reaches nearly every model and is usually not rate-limited when the first-party routes are — repin coordinator/worker cron jobs there with hermes cron edit <id> --provider openrouter --model <slug> (e.g. deepseek/deepseek-v4.1-flash), and smoke-test the route first with hermes -z 'Reply with exactly: OK' --provider openrouter --model <slug> before trusting it. Keep role→model bindings, not provider pins: name the role, bind the model in config, swap the binding on the owner's word. After all healthy jobs funnel through ONE account, that single account becomes the new bottleneck — stagger heavy recurring jobs onto offset, slower cadences (e.g. every 20m/25m/30m instead of all every 10m) with hermes cron edit <id> --schedule 'every 25m' so their concurrent requests fit one account's rate window; read back the persisted schedule (cron edit can exit 0 with failure text).

Other supported levers, prefer them when they fit:

  • hermes config set credential_pool_strategies.<provider> round_robin — alternate/share load across accounts instead of strict primary+failover.
  • hermes auth remove <provider> <index|id|label> — drop an account from the pool entirely.
  • hermes auth reset <provider> — clear exhaustion status after a limit window passes.

Order normalization (why a hand-edit can move)​

_normalize_pool_order in agent/credential_pool.py sorts manual entries (manual:hermes_pkce and friends) by priority first, then appends seeded entries ranked by source: env:ANTHROPIC_TOKEN < env:CLAUDE_CODE_OAUTH_TOKEN < hermes_pkce < claude_code < env:ANTHROPIC_API_KEY. Hand-set priorities among manual OAuth logins are preserved relative to each other; seeded/env credentials get re-ranked by source on the next upsert.

Probe the execution plane you will actually use​

A provider/model is not proven unavailable merely because a sibling client fails. Standalone coding CLIs and Hermes can use different credential stores, pools, account scopes, and rate-limit state even when both display the same provider and model names.

For any exact-route assignment:

  1. Identify the intended execution plane: Hermes-managed provider, standalone CLI, gateway/profile, or scheduled job.
  2. Probe the exact provider/model through that same plane with a no-tools sentinel response.
  3. Treat a sibling client's result only as evidence about that client's credential plane.
  4. If the intended plane succeeds, launch through it; do not route the real task through the failed sibling client.
  5. Record both the exact route and execution plane in the attempt receipt. Never silently substitute a model or reset attempt lineage.

Example Hermes probe:

hermes chat --provider openai-codex --model <exact-model> \
--ignore-user-config --ignore-rules --safe-mode \
-q 'Reply exactly ROUTE_OK and do not use tools.'

For a write-capable continuation, remove --safe-mode, bind a narrow worktree and prompt, and explicitly select the needed toolsets/approval policy. Preserve prior partial changes and failed-attempt evidence before relaunching.

Validated session details and the client-vs-Hermes interpretation are in references/execution-plane-route-probes.md.

Profiles vs pools​

Profiles ($HERMES_HOME/profiles/<name>/) are the persistent-configuration boundary — own config, skills, memory, and potentially own auth.json. They are not automatically a strict account boundary: a profile without local credentials can fall back to the root profile's. If the user wants "this profile only ever bills account X," authenticate that account inside the profile and verify effective resolution — do not assume the profile name enforces it, and never hand-copy OAuth token files between profiles (refresh collisions).

Honesty rules​

  • Report what you verified: pool order plus a successful probe prove the route works, not that a specific account was billed for a specific past turn. request_count/last_status can remain 0/null for gateway-served turns.
  • A provider rate limit on one subscription is not a defect in another; state which credential/provider the 429 belongs to.
  • When the user asks "which model are you using", answer with the live resolved model id, and separate that from whatever routes the background fleet/cron jobs use.

Pitfalls​

  • When a role maps to an interchangeable model pool (e.g. coordinator = Fableanthropic ∨ Astraopenai-codex), pick the pool member whose PROVIDER is actually healthiest — read hermes auth list reset windows, do not assume. It is easy to move a stuck job onto the more rate-limited provider: verify both providers' 429 state and remaining windows before repinning, then read back the persisted route. Repinning to a route that is also exhausted just relocates the outage.
  • hermes chat -q has no --no-stream flag; an invented flag exits 2 with the full subcommand list. Check hermes <cmd> --help first.
  • Broad recursive greps across $HERMES_HOME plus large work trees time out (60s+). Scope to exact files (auth.json, config.yaml) or a single package directory.
  • hermes config get provider may print Config key not set while hermes config get model still shows provider: anthropic inside the model block — read the model block, don't conclude the provider is unset.
  • A coding CLI's own login store (e.g. Claude Code's ~/.claude.json oauthAccount, or an alternate CLAUDE_CONFIG_DIR) is independent of Hermes's pool. Changing one does not change the other; check both when a user asks "which account is in use".
  • Never hand-edit config.yaml; use hermes config set KEY VALUE. auth.json ordering has no setter, which is the narrow exception documented above.

Session-specific recipe and exact observed output: references/anthropic-multi-subscription.md.


Supporting files: this skill's supporting files are held in the docsite at docs/15-skills/_support/autonomous-ai-agents/hermes-provider-account-routing/ — fetch them fresh from jknash/docsite main alongside this page. Source: jknash/hermes-shared-skills · branch hermes-jkdev001 @ 1d0d545c3970 · skills/autonomous-ai-agents/hermes-provider-account-routing/ · view source · Imported 2026-10-04. Supporting files (references, scripts) remain in the source repository.

version 1.0.0 · author Hermes Curator · license MIT.

Published by Muse · 2026-10-04.