Skip to main content

Cron / agent-fleet rate-limit recovery

Use when cron/agent jobs error on provider 429 rate limits. When scheduled agent jobs pinned to one provider/model start failing because the provider is rate-limited (HTTP 429 / usage_limit_reached), the jobs are fine — the capacity is not. Recover by giving the fleet a healthy route, not by editing prompts or logic. This is an operational-recovery procedure; it does not authorize activating paused jobs or bypassing review/merge gates.

Recognize the signature​

Enabled jobs show last_status: error on recent ticks, but with NO last_fire_error and NO last_delivery_error populated, and NO output file for the erroring ticks. That pattern = the agent run aborted mid-flight (rate limit), not a delivery or logic failure. Confirm before acting; see references/rate-limit-recovery.md for the exact diagnosis commands and request-dump fields.

Recover in this order​

  1. Add a healthy pooled credential (hermes auth add <provider>) — the pool then auto-skips the 429'd credential; usually no per-job change needed.
  2. Repin jobs to a healthy provider/model BY ROLE (hermes cron edit), keeping the role's pool identity; read back the persisted route (edit can exit 0 on failure).
  3. Stagger/slow cadence when one account must carry the whole fleet. Then fire jobs (cronjob_manage action='run') to resume immediately.

Never hardcode a provider/model literal into code or config to "fix" this — route binding belongs in the job definition or a config-driven role allowlist, so the next rate-limit is a one-line rebind, not a code change. A 429 is a capacity event, not a failed attempt — do not count it in any escalation/retry ledger.

Key gotchas (full detail in the reference)​

  • OAuth hermes auth add on a headless host: --type oauth --no-browser, drive as a PTY background process; the answer-a-prompt field is data not input; each login link is single-use with its own PKCE state.
  • Verify the new credential landed in THIS profile's ~/.hermes/auth.json credential_pool.<provider>; a login into another profile/home won't be visible to the running fleet.
  • An agentic cron that 429s can surface as engine timeout (exit 124), because the client sleeps on the 429 inside the run budget — classify from logs, not exit code.
  • openrouter is a reliable fallback route when first-party providers are all 429'd; smoke-test it (hermes -z "reply OK" --provider openrouter --model X) before repinning a fleet onto it.

Silent hang: a dispatch on a bad route looks like work, not an error​

When the parent session itself runs on a rate-limited/erroring route, a dispatched subagent INHERITS that model and can HANG with no error surfaced — its live transcript stops at the first tool call, no child process exists, and it burns wall-clock silently. Signature: /root/.hermes/cache/delegation/live/<deleg_id>/task-0.log size/mtime stops advancing while the child still reports running. Recovery: stop the hung child (delegate_task action=stop), switch the parent/dispatch route to a healthy one, and re-dispatch. A route/capacity stall is not a failed attempt and consumes no escalation/retry count. ALWAYS stat the live transcript to confirm a dispatch is actually progressing instead of assuming it is working.

Rebinding when routes are config-driven (role bindings)​

If the fleet selects models from a config-driven role allowlist (not per-job pins), the rebind is a config edit + service restart — no code change: back up the config, edit the role entry (role_bindings.<role> -> provider/model), restart the owning service (systemctl --user restart <unit>), confirm is-active and a clean journal, then read the binding back. Example (2026-09-14): reviewer role repointed anthropic/claude-opus-5 -> openai-codex/gpt-6-astra while the entire anthropic pool was 429. Keep the per-route credential/env allowlist and credential scrubbing intact — a config-driven route must never become a way to smuggle credentials.

Reference​

  • references/rate-limit-recovery.md — exact diagnosis commands, request-dump fields, the three recovery levers with commands, and the resume/verify steps.

Verification​

After recovery: cronjob_manage(action='list') shows the repinned jobs with the new route persisted and subsequent ticks returning ok; hermes auth list shows at least one non-429 credential for the provider. A job's execution_skipped: already being fired by the scheduler on a manual run means it is running now, not idle — that is success, not an error.


Source: jknash/hermes-shared-skills · branch hermes-jkdev001 @ 1d0d545c3970 · skills/automation/cron-fleet-rate-limit-recovery/ · view source · Imported 2026-10-04. Supporting files (references, scripts) remain in the source repository.

version 0.1.0 · author Hermes Agent · license MIT.

Published by Muse · 2026-10-04.