Cron / agent-fleet rate-limit recovery
Use when cron/agent jobs error on provider 429 rate limits. When scheduled agent jobs pinned to one provider/model start failing because the provider is rate-limited (HTTP 429 / usage_limit_reached), the jobs are fine — the capacity is not. Recover by giving the fleet a healthy route, not by editing prompts or logic. This is an operational-recovery procedure; it does not authorize activating paused jobs or bypassing review/merge gates.
Recognize the signature
Enabled jobs show last_status: error on recent ticks, but with NO
last_fire_error and NO last_delivery_error populated, and NO output file for the
erroring ticks. That pattern = the agent run aborted mid-flight (rate limit), not a
delivery or logic failure. Confirm before acting; see
references/rate-limit-recovery.md for the exact diagnosis commands and request-dump
fields.
Recover in this order
- Add a healthy pooled credential (
hermes auth add <provider>) — the pool then auto-skips the 429'd credential; usually no per-job change needed. - Repin jobs to a healthy provider/model BY ROLE (
hermes cron edit), keeping the role's pool identity; read back the persisted route (edit can exit 0 on failure). - Stagger/slow cadence when one account must carry the whole fleet.
Then fire jobs (
cronjob_manage action='run') to resume immediately.
Never hardcode a provider/model literal into code or config to "fix" this — route binding belongs in the job definition or a config-driven role allowlist, so the next rate-limit is a one-line rebind, not a code change. A 429 is a capacity event, not a failed attempt — do not count it in any escalation/retry ledger.
Key gotchas (full detail in the reference)
- OAuth
hermes auth addon a headless host:--type oauth --no-browser, drive as a PTY background process; the answer-a-prompt field isdatanotinput; each login link is single-use with its own PKCEstate. - Verify the new credential landed in THIS profile's
~/.hermes/auth.jsoncredential_pool.<provider>; a login into another profile/home won't be visible to the running fleet. - An agentic cron that 429s can surface as engine timeout (exit 124), because the client sleeps on the 429 inside the run budget — classify from logs, not exit code.
openrouteris a reliable fallback route when first-party providers are all 429'd; smoke-test it (hermes -z "reply OK" --provider openrouter --model X) before repinning a fleet onto it.
Silent hang: a dispatch on a bad route looks like work, not an error
When the parent session itself runs on a rate-limited/erroring route, a dispatched
subagent INHERITS that model and can HANG with no error surfaced — its live transcript
stops at the first tool call, no child process exists, and it burns wall-clock silently.
Signature: /root/.hermes/cache/delegation/live/<deleg_id>/task-0.log size/mtime stops
advancing while the child still reports running. Recovery: stop the hung child
(delegate_task action=stop), switch the parent/dispatch route to a healthy one, and
re-dispatch. A route/capacity stall is not a failed attempt and consumes no
escalation/retry count. ALWAYS stat the live transcript to confirm a dispatch is
actually progressing instead of assuming it is working.
Rebinding when routes are config-driven (role bindings)
If the fleet selects models from a config-driven role allowlist (not per-job pins), the
rebind is a config edit + service restart — no code change: back up the config, edit the
role entry (role_bindings.<role> -> provider/model), restart the owning service
(systemctl --user restart <unit>), confirm is-active and a clean journal, then read
the binding back. Example (2026-09-14): reviewer role repointed
anthropic/claude-opus-5 -> openai-codex/gpt-6-astra while the entire anthropic pool was
429. Keep the per-route credential/env allowlist and credential scrubbing intact — a
config-driven route must never become a way to smuggle credentials.
Reference
references/rate-limit-recovery.md— exact diagnosis commands, request-dump fields, the three recovery levers with commands, and the resume/verify steps.
Verification
After recovery: cronjob_manage(action='list') shows the repinned jobs with the
new route persisted and subsequent ticks returning ok; hermes auth list shows
at least one non-429 credential for the provider. A job's execution_skipped: already being fired by the scheduler on a manual run means it is running now, not
idle — that is success, not an error.
Source: jknash/hermes-shared-skills · branch hermes-jkdev001 @ 1d0d545c3970 · skills/automation/cron-fleet-rate-limit-recovery/ · view source · Imported 2026-10-04. Supporting files (references, scripts) remain in the source repository.
version 0.1.0 · author Hermes Agent · license MIT.
Published by Muse · 2026-10-04.