Durable Fleet Delivery
Use when making agent fleets deliver durably. Design and operate recoverable delivery workflows, not merely persistent worker processes. This is a procedural skill, not an installed workflow engine. Version 0.2.0 incorporates fixture-tested phase, journal, handoff, reserve, recovery, review, release, and disposable-drill procedures from the staged Assessor implementation. It does not certify a deployed controller, real product integration/release, or unattended reliability. Systemd examples belong in deployment-specific skills and require Linux.
When to use
- A fleet stays alive but does not reliably complete, review, integrate or release work.
- The user repeatedly has to refill workers or recover interrupted handoffs.
- Building reusable orchestration for another authorized fleet.
- Do not use this to authorize new repositories, model fallback, production releases or credential access.
Prerequisites
Use read_file and read-only terminal discovery to establish the live board, repository, provider routes, process owners, existing controller and release requirements. Load the platform's orchestration skill before changing its configuration. For Hermes-specific behavior, use the hermes-agent skill and official documentation. Discover CLI flags rather than inventing them.
Complete templates/fleet-contract.md before deployment. Use the existing task authority; do not create a parallel backlog in runtime JSON. Record all assumptions and unsupported adapters. Scope credentials to roles; environment filtering is not a filesystem sandbox.
Procedure
- Define delivered. Bind release acceptance IDs to tests, artifact destinations, target platforms and authorized integration/release owners. Separate implemented, reviewed, integrated and delivered. Completion: every release criterion has an evidence obligation.
- Preserve and census. Inspect tracked jobs, OS child identities/CWDs, board claims, current Git/PR state and dirty/untracked work. Completion: sole known mutation ownership or an explicit safety block; useful files preserved.
- Checkpoint the workflow. Persist task/phase/revision/candidate, evidence, owner generation, next action/due time and resume phase. Use fenced compare-and-swap transitions and an action intent/receipt journal. Completion: restart resumes the unfinished step, not the entire assignment.
- Prepare replacement work early. Keep a small route-compatible reserve on the existing board. Include exact acceptance gaps, scope, dependency/overlap evidence and prepared brief. Revalidate at reservation and worker execution. Completion: eligible vacancy can dispatch without fresh global LLM planning.
- Separate responsibilities. Ready preparation, dispatch, handoff assessment, review/integration and release verification are independent durable obligations. Shared contracts and heavy gates remain serialized. Completion: one difficult handoff cannot starve unrelated eligible work.
- Classify exits. Worker output is advisory. Inspect intended tracked/untracked artifacts and acceptance gates; choose missing implementation, validation, exact-candidate review or explicit block. A complete/no-edit report is not a failed writer to respawn blindly. Completion: each exit has an owned next phase.
- Recover by cause. Apply
references/failure-recovery.md. Retries bind canonical task/candidate/phase, not cosmetic prompt hashes or caller-provided scope labels; a material new shape needs a deduplicated authority record. An ambiguous external action needs one leased recovery owner before readback or retry. Completion: every blocked/due item has a fenced owner, reason and retry/escalation policy. - Integrate exact evidence. Independent reviews bind immutable candidate hashes. Candidate changes invalidate relevant evidence; unrelated main movement does not automatically invalidate approval. Check current conflicts/protections/CI and authorization before integration. Read back remote actions and verify intended content, including squash merges. Completion: remote integrated state is independently confirmed.
- Verify the deliverable. Validate the approved artifact and platform/user-flow acceptance; distinguish fixtures from live-environment proof. Bind source commit/checksum/environment/gates and authorized destination. Completion: artifact delivery readback and release criteria, not a successful process.
- Cut over and prove recovery. Follow
references/verification.md: shadow mode, one dispatch authority, safe migration, worker survival, rollback, failure drills and unattended soak. Completion: measured evidence meets the fleet contract; no blanket perpetual-autonomy claim. - Report outcomes. Lead with accepted/integrated work and remaining release blockers. Include phase age, reserve health, next automatic action and decisions required. Worker counts are secondary. Completion: totals derive from canonical records, unknown state is visible, notices are deduplicated.
- Generalize only proven pieces. Keep local paths/models/board identities in fleet adapters. Add scripts only after execution tests; test a second disposable fleet before claiming portability. Completion: versioned evidence and honest limitations accompany each skill update.
Reference implementation boundaries
The staged Assessor source demonstrates strict JSON state schemas, atomic replace plus directory fsync, file locks, owner generations, semantic retry keys, exact-role review records, artifact manifests, and explicit fixture/live evidence classes. Those Python modules are examples, not a stable cross-project API. Do not copy live configs or credential-bearing runtime into a drill. A reserve record is cached intent, never authority: re-read board/dependencies/owner/overlap/route immediately before the controller-only launch. Wrap every external action in an intent/receipt journal; after an action-before-receipt crash, exact target readback must decide already-applied, retryable, or unknown.
Expired ownership is not proof of death. Reclaim an obligation or resource only after identity/CWD/process readback proves the old owner quiescent, then advance a fence. A stale process must be unable to complete under its old generation. Provider probes use the same exact provider/model and never authorize fallback.
Pitfalls
- Re-read the latest canonical task comments inside the same launch fence immediately before any model assignment or phase handoff. A concurrent owner can publish a stricter plan/authority requirement after an earlier census; a process launched without that newest prerequisite remains evidence-only even when its route, candidate, and locks are otherwise exact. Preserve a surviving read-only run, disclose the admission failure, and do not retroactively treat a later plan as authorization.
- On Hermes boards with auto-decomposition, TRIAGE is active input, not a safe parking state. A triage program can be decomposed into assigned children and dispatched automatically.
kanban blockmay fail for triage/todo yet append a misleading BLOCKED comment; read canonical status/events. Prepare a supported dispatch guard before import, verify actual topology, and never run an external staging writer beside a spawned gateway child. If this occurs, preserve files, block the actionable root step, stop only the verified duplicate PID and reuse/correct the existing graph (auto-generated phase labels may misread the plan). - A successful cutover can still leave production adapters and acceptance incomplete. Reconcile the original program against newest external receipts, not immutable staging status alone; keep completion ownership through implementation, exact-role reviews and gated soak. A recovery watchdog must remain enabled while its child runs (observe-only under verified ownership), not pause merely on transfer and strand the next phase.
- Restart=always proves service lifetime, not delivery throughput.
- A successful planner receipt can hide zero useful transitions.
- Broad repeated backlog inspection can consume the entire reconciliation budget.
- Single-row microtasks can spend more time on rediscovery than output; choose bounded acceptance-complete slices, not arbitrary mutation quotas.
- Fresh content changes are activity signals, not acceptance; no changes can mean work is already complete.
- Exit/wrapper failure does not prove children are dead. Fence by boot/start identity and CWD, not PID alone.
- When a launch wrapper acquires an inheritable flock before exec, the worker's own Hermes ancestor is the expected lock owner. Ownership census must distinguish that ancestry from unrelated writers; treating its own parent as a competitor can cause a no-op exit. State inherited ownership explicitly in worker briefs and retain the lock across exec.
- A timeout after remote action can mean the action succeeded; query before retrying. A durable not_applied readback must be consumed by a fsynced action-start event before external execution; an exception or malformed result cannot reuse that observation even under the same recovery generation. Success requires a nonempty durable evidence reference. These are acceptance requirements from independent review, not claims that the current staged implementation satisfies them.
- A supervisor may report an accepted/active unit before its wrapper's create-once active receipt appears. Treat a missing receipt inside a short polling window as action ambiguity, not
not_applied: read back the exact unit invocation, PID/start identity, CWD/cgroup and later receipt under the same action generation, then persist an observed-by-readback recovery receipt. Never launch a successor merely because the first local receipt read raced its atomic creation. - Multi-platform releases need artifact-specific IDs, bytes/digests, architecture and destination bindings; a single digest shared across Windows/macOS packages cannot prove either artifact.
- Health boundaries must catch domain-specific durable-state exceptions and emit an unhealthy/unknown observation rather than crashing or fabricating outcome totals.
- Module/fixture coverage is not workflow completion: verify that production controller paths actually invoke handoff stores, reserve claims, action journals and review/recovery adapters before calling the fleet implemented.
- Holding a local reserve lock across an external action does not itself make the action replay-safe; the durable intent must precede the action and ambiguous restart must read back the target.
- Four busy writers cannot compensate for stalled reviews/integration. Do not increase occupancy by violating safety or inventing work.
- Merge authority and release authority are separate from implementation authority.
- Host reboot recovery also depends on boot enablement, mounts, secrets availability and process supervision configuration.
- Do not claim a newly written skill is a working orchestrator.
Verification
Use tool-backed evidence for every claimed gate. Required matrix, cutover and soak are in references/verification.md. Keep planned, fixture-tested, live-exercised and production-soak-verified capabilities distinct. Run default discovery without live dispatch; real systemd and delivery drills require their documented explicit opt-ins and disposable paths. For planning-only requests, save a concrete plan without changing live services; explicit skill-authoring requests may create/update this procedural package.
Supporting files: this skill's supporting files are held in the docsite at
docs/15-skills/_support/engineering/durable-fleet-delivery/— fetch them fresh fromjknash/docsitemain alongside this page. Source:jknash/hermes-shared-skills· branchhermes-jkdev001@1d0d545c3970·skills/engineering/durable-fleet-delivery/· view source · Imported 2026-10-03. Supporting files (references, scripts) remain in the source repository.
version 0.2.0 · author jknash, Hermes Agent · license MIT.
Published by Muse · 2026-10-03.