Skip to main content

Durable Fleet Delivery

Use when making agent fleets deliver durably. Design and operate recoverable delivery workflows, not merely persistent worker processes. This is a procedural skill, not an installed workflow engine. Version 0.2.0 incorporates fixture-tested phase, journal, handoff, reserve, recovery, review, release, and disposable-drill procedures from the staged Assessor implementation. It does not certify a deployed controller, real product integration/release, or unattended reliability. Systemd examples belong in deployment-specific skills and require Linux.

When to use​

  • A fleet stays alive but does not reliably complete, review, integrate or release work.
  • The user repeatedly has to refill workers or recover interrupted handoffs.
  • Building reusable orchestration for another authorized fleet.
  • Do not use this to authorize new repositories, model fallback, production releases or credential access.

Prerequisites​

Use read_file and read-only terminal discovery to establish the live board, repository, provider routes, process owners, existing controller and release requirements. Load the platform's orchestration skill before changing its configuration. For Hermes-specific behavior, use the hermes-agent skill and official documentation. Discover CLI flags rather than inventing them.

Complete templates/fleet-contract.md before deployment. Use the existing task authority; do not create a parallel backlog in runtime JSON. Record all assumptions and unsupported adapters. Scope credentials to roles; environment filtering is not a filesystem sandbox.

Procedure​

  1. Define delivered. Bind release acceptance IDs to tests, artifact destinations, target platforms and authorized integration/release owners. Separate implemented, reviewed, integrated and delivered. Completion: every release criterion has an evidence obligation.
  2. Preserve and census. Inspect tracked jobs, OS child identities/CWDs, board claims, current Git/PR state and dirty/untracked work. Completion: sole known mutation ownership or an explicit safety block; useful files preserved.
  3. Checkpoint the workflow. Persist task/phase/revision/candidate, evidence, owner generation, next action/due time and resume phase. Use fenced compare-and-swap transitions and an action intent/receipt journal. Completion: restart resumes the unfinished step, not the entire assignment.
  4. Prepare replacement work early. Keep a small route-compatible reserve on the existing board. Include exact acceptance gaps, scope, dependency/overlap evidence and prepared brief. Revalidate at reservation and worker execution. Completion: eligible vacancy can dispatch without fresh global LLM planning.
  5. Separate responsibilities. Ready preparation, dispatch, handoff assessment, review/integration and release verification are independent durable obligations. Shared contracts and heavy gates remain serialized. Completion: one difficult handoff cannot starve unrelated eligible work.
  6. Classify exits. Worker output is advisory. Inspect intended tracked/untracked artifacts and acceptance gates; choose missing implementation, validation, exact-candidate review or explicit block. A complete/no-edit report is not a failed writer to respawn blindly. Completion: each exit has an owned next phase.
  7. Recover by cause. Apply references/failure-recovery.md. Retries bind canonical task/candidate/phase, not cosmetic prompt hashes or caller-provided scope labels; a material new shape needs a deduplicated authority record. An ambiguous external action needs one leased recovery owner before readback or retry. Completion: every blocked/due item has a fenced owner, reason and retry/escalation policy.
  8. Integrate exact evidence. Independent reviews bind immutable candidate hashes. Candidate changes invalidate relevant evidence; unrelated main movement does not automatically invalidate approval. Check current conflicts/protections/CI and authorization before integration. Read back remote actions and verify intended content, including squash merges. Completion: remote integrated state is independently confirmed.
  9. Verify the deliverable. Validate the approved artifact and platform/user-flow acceptance; distinguish fixtures from live-environment proof. Bind source commit/checksum/environment/gates and authorized destination. Completion: artifact delivery readback and release criteria, not a successful process.
  10. Cut over and prove recovery. Follow references/verification.md: shadow mode, one dispatch authority, safe migration, worker survival, rollback, failure drills and unattended soak. Completion: measured evidence meets the fleet contract; no blanket perpetual-autonomy claim.
  11. Report outcomes. Lead with accepted/integrated work and remaining release blockers. Include phase age, reserve health, next automatic action and decisions required. Worker counts are secondary. Completion: totals derive from canonical records, unknown state is visible, notices are deduplicated.
  12. Generalize only proven pieces. Keep local paths/models/board identities in fleet adapters. Add scripts only after execution tests; test a second disposable fleet before claiming portability. Completion: versioned evidence and honest limitations accompany each skill update.

Reference implementation boundaries​

The staged Assessor source demonstrates strict JSON state schemas, atomic replace plus directory fsync, file locks, owner generations, semantic retry keys, exact-role review records, artifact manifests, and explicit fixture/live evidence classes. Those Python modules are examples, not a stable cross-project API. Do not copy live configs or credential-bearing runtime into a drill. A reserve record is cached intent, never authority: re-read board/dependencies/owner/overlap/route immediately before the controller-only launch. Wrap every external action in an intent/receipt journal; after an action-before-receipt crash, exact target readback must decide already-applied, retryable, or unknown.

Expired ownership is not proof of death. Reclaim an obligation or resource only after identity/CWD/process readback proves the old owner quiescent, then advance a fence. A stale process must be unable to complete under its old generation. Provider probes use the same exact provider/model and never authorize fallback.

Pitfalls​

  • Re-read the latest canonical task comments inside the same launch fence immediately before any model assignment or phase handoff. A concurrent owner can publish a stricter plan/authority requirement after an earlier census; a process launched without that newest prerequisite remains evidence-only even when its route, candidate, and locks are otherwise exact. Preserve a surviving read-only run, disclose the admission failure, and do not retroactively treat a later plan as authorization.
  • On Hermes boards with auto-decomposition, TRIAGE is active input, not a safe parking state. A triage program can be decomposed into assigned children and dispatched automatically. kanban block may fail for triage/todo yet append a misleading BLOCKED comment; read canonical status/events. Prepare a supported dispatch guard before import, verify actual topology, and never run an external staging writer beside a spawned gateway child. If this occurs, preserve files, block the actionable root step, stop only the verified duplicate PID and reuse/correct the existing graph (auto-generated phase labels may misread the plan).
  • A successful cutover can still leave production adapters and acceptance incomplete. Reconcile the original program against newest external receipts, not immutable staging status alone; keep completion ownership through implementation, exact-role reviews and gated soak. A recovery watchdog must remain enabled while its child runs (observe-only under verified ownership), not pause merely on transfer and strand the next phase.
  • Restart=always proves service lifetime, not delivery throughput.
  • A successful planner receipt can hide zero useful transitions.
  • Broad repeated backlog inspection can consume the entire reconciliation budget.
  • Single-row microtasks can spend more time on rediscovery than output; choose bounded acceptance-complete slices, not arbitrary mutation quotas.
  • Fresh content changes are activity signals, not acceptance; no changes can mean work is already complete.
  • Exit/wrapper failure does not prove children are dead. Fence by boot/start identity and CWD, not PID alone.
  • When a launch wrapper acquires an inheritable flock before exec, the worker's own Hermes ancestor is the expected lock owner. Ownership census must distinguish that ancestry from unrelated writers; treating its own parent as a competitor can cause a no-op exit. State inherited ownership explicitly in worker briefs and retain the lock across exec.
  • A timeout after remote action can mean the action succeeded; query before retrying. A durable not_applied readback must be consumed by a fsynced action-start event before external execution; an exception or malformed result cannot reuse that observation even under the same recovery generation. Success requires a nonempty durable evidence reference. These are acceptance requirements from independent review, not claims that the current staged implementation satisfies them.
  • A supervisor may report an accepted/active unit before its wrapper's create-once active receipt appears. Treat a missing receipt inside a short polling window as action ambiguity, not not_applied: read back the exact unit invocation, PID/start identity, CWD/cgroup and later receipt under the same action generation, then persist an observed-by-readback recovery receipt. Never launch a successor merely because the first local receipt read raced its atomic creation.
  • Multi-platform releases need artifact-specific IDs, bytes/digests, architecture and destination bindings; a single digest shared across Windows/macOS packages cannot prove either artifact.
  • Health boundaries must catch domain-specific durable-state exceptions and emit an unhealthy/unknown observation rather than crashing or fabricating outcome totals.
  • Module/fixture coverage is not workflow completion: verify that production controller paths actually invoke handoff stores, reserve claims, action journals and review/recovery adapters before calling the fleet implemented.
  • Holding a local reserve lock across an external action does not itself make the action replay-safe; the durable intent must precede the action and ambiguous restart must read back the target.
  • Four busy writers cannot compensate for stalled reviews/integration. Do not increase occupancy by violating safety or inventing work.
  • Merge authority and release authority are separate from implementation authority.
  • Host reboot recovery also depends on boot enablement, mounts, secrets availability and process supervision configuration.
  • Do not claim a newly written skill is a working orchestrator.

Verification​

Use tool-backed evidence for every claimed gate. Required matrix, cutover and soak are in references/verification.md. Keep planned, fixture-tested, live-exercised and production-soak-verified capabilities distinct. Run default discovery without live dispatch; real systemd and delivery drills require their documented explicit opt-ins and disposable paths. For planning-only requests, save a concrete plan without changing live services; explicit skill-authoring requests may create/update this procedural package.


Supporting files: this skill's supporting files are held in the docsite at docs/15-skills/_support/engineering/durable-fleet-delivery/ — fetch them fresh from jknash/docsite main alongside this page. Source: jknash/hermes-shared-skills · branch hermes-jkdev001 @ 1d0d545c3970 · skills/engineering/durable-fleet-delivery/ · view source · Imported 2026-10-03. Supporting files (references, scripts) remain in the source repository.

version 0.2.0 · author jknash, Hermes Agent · license MIT.

Published by Muse · 2026-10-03.