External Coding Agent Orchestration
Use when orchestrating autonomous coding CLI workers. Use this skill when dispatching coding CLIs such as Reasonix, Codex, Claude Code, or separately-profiled Hermes workers, especially for long-running work in parallel worktrees.
Operating model
- Inspect the CLI's current help and diagnostics; live-smoke the exact provider/model identifier that will appear in the worker command. A different identifier from the same model family is not an alias.
- Start from a clean fetched baseline and give each independent task its own branch and real Git worktree.
- Before dispatch, verify the task carries an explicit external route marker plus
fallback=forbidden, its authoritative issue/lock is current, dependencies are merged, and no open PR overlaps its files. - Prompt with task ID, acceptance criteria, TDD expectations, allowed scope, full verification commands, and commit/push policy.
- Run workers as tracked background processes with completion notification.
- Verify both the process registry and OS process table. External workers may not appear in delegation lists.
- Treat task-sized waves as continuous delivery; do not use one immortal agent for an entire program or overlap shared migrations/contracts.
- When the user requires continuous progress, enforce a real writer floor: after a merge, worker exit, or gateway recovery, immediately preserve/rebase/resume a safe implementation lane rather than waiting for the next status cycle. Pair a frequent local execution supervisor with an explicit origin-facing path for ready PRs and genuinely human-only blockers; local receipts alone do not notify the user.
- Treat a user-approved concurrency change as a durable policy mutation, not a one-time launch. Inspect the active supervisor/queue prompt for its configured floor, update it through the scheduler's supported API, read it back, and only then launch or retire workers. Otherwise a supervisor still saying “maintain one” will undo or fail to refill a newly authorized second lane. Encode both the target count and refill condition (“if fewer than N qualifying writers are alive”), plus gate-serialization and overlap rules.
- Distinguish a minimum floor (“keep at least N alive”) from a target/cap (“run exactly N” or “no more than N”). Do not silently interpret a minimum as authorization for arbitrary fan-out. Before adding a supplemental named-issue writer while the floor is already met, reconcile the user's capacity intent, host/gate limits, and supervisor policy; prefer an eligible review lane when no extra writer slot is authorized. Under an exact-N supervisor, do not repeatedly launch an N+1 foreground writer: the supervisor may correctly terminate it before mutation. Persist the requested issue order in the supervisor, preserve productive incumbents, and rotate the priority issue into the next safe vacancy. Also remove stale merged-task priority text before it can steer a later tick. See
references/exact-cap-priority-steering.md. - Before foreground recovery launches, read whether a supervisor tick is already
running. Do not let the foreground and an unleased in-flight tick refill the same vacancy: wait for the bounded dispatch phase or establish a shared recovery lease/hold. If they race, preserve all residue, select canonical owners by worktree, stop duplicates, and verify the floor again after the tick exits. A transient process snapshot during simultaneous dispatch is not a stable capacity claim. Seereferences/supervisor-orchestrator-race-control.mdandreferences/supervisor-owned-sigterm-batch-recovery.mdfor the repeated-SIGTERM attribution and stable post-tick refill procedure.
- Distinguish a minimum floor (“keep at least N alive”) from a target/cap (“run exactly N” or “no more than N”). Do not silently interpret a minimum as authorization for arbitrary fan-out. Before adding a supplemental named-issue writer while the floor is already met, reconcile the user's capacity intent, host/gate limits, and supervisor policy; prefer an eligible review lane when no extra writer slot is authorized. Under an exact-N supervisor, do not repeatedly launch an N+1 foreground writer: the supervisor may correctly terminate it before mutation. Persist the requested issue order in the supervisor, preserve productive incumbents, and rotate the priority issue into the next safe vacancy. Also remove stale merged-task priority text before it can steer a later tick. See
- On every worker completion notification, recount the entire target fleet from the OS process table and tracked registry before refilling. More than the reported worker may have exited since the last observation, so replacing only the named vacancy can silently leave the fleet below its floor. Treat a completion batch as historical evidence rather than an atomic fleet snapshot: processes can exit while reports are read and review/refill actions are launched. Inspect and clean the completed handoff, update its durable board receipt, launch enough safe disjoint lanes to restore all missing slots, then recount again after the action wave and verify every child survives before asserting the floor. See
references/continuous-writer-floor-completion-refill.mdfor the bounded callback procedure andreferences/completion-batch-races-and-immutable-handoffs.mdfor mid-callback exits, wrapper/child counting, staged-residue attribution, and stash-shaped candidate verification. - Enforce a single mutation-capable owner per worktree, including untracked supervisor children. A wrapper exit, process-manager receipt, or terse worker summary does not release the lease: inspect all OS agent processes by exact CWD and include nested workers that detached or were reparented to a daemon. If a worker says it “kicked off a background agent,” treat that as a live writer until the actual child exits. Before direct orchestrator edits or replacement launches, enumerate coding processes by CWD, stop collisions, classify/preserve residue, and restore only incomplete writer-owned files. Treat orchestrator-side
git add,commit,commit --amend, andmergeas mutation leases too: exclude the worktree from an in-flight supervisor, snapshot HEAD/index immediately before the command, and read them back afterward. A supervisor can commit between two foreground commands, turning an intended amend into an extra commit or invalidating the candidate identity. Never clean generated residue while any detached child can still write it. Seereferences/orchestrator-git-mutation-leases.mdfor the bounded protocol. Before proposing bulk cleanup, state the exact worktree and paths, prove whether they are tracked, explain why they are disposable, and obtain command-specific approval; a blocked cleanup is not permission to retry another deletion shape. Seereferences/detached-child-and-cleanup-approval.md. For descriptor-owned renderers, pair this lease discipline with real regular-file/symlink displacement tests and retained-handle streaming. Seereferences/single-writer-worktree-lease-and-renderer-streaming.md. When such a security gate differs by OS, preserve assertion-order evidence, distinguishENOENTfromreadlink/EINVAL, and dispatch a bounded investigator with explicit target-OS instrumentation rather than calling it flaky; seereferences/cross-platform-filesystem-gate-investigation.md.
Named role identity and stalled-PR recovery
- Treat human-readable reviewer names as literal route contracts. Persist the role’s exact provider, exact model, and fallback policy; live-smoke that route; verify the live child command line; and record the route in the durable exact-head verdict.
- A wrapper/profile/role label is not identity evidence. If a report claims a named role but records another provider/model, invalidate the role approval, preserve its technical findings only as leads, and rerun the exact head through the contracted route.
Action-coupled review cascade
A required review is an action trigger, not a status sentence. Once an immutable candidate satisfies its pre-review gates, launch the contracted reviewer in the same callback/tick unless that exact candidate already has a live or durable current verdict. Do not report “review needed” and defer dispatch to a later status cycle.
-
Prefer event-driven dispatch: a writer/gate completion should durably transition an exact-candidate review task to ready and let the owning execution supervisor launch it. A slower unified status reporter is a watchdog, not the primary dispatcher.
-
Treat provider/model overrides as phase-scoped routing state, not permanent task identity. Before every implementation→review or changes-requested→implementation dispatch, atomically restore the destination role's exact provider/model, assign that role, dispatch, and verify the spawned command line. A same-card handoff can otherwise inherit the previous phase's route. See
references/kanban-phase-route-reset.md. -
When a unified reporter nevertheless discovers an eligible missing review, it must invoke the repository execution supervisor immediately in that same run, passing the exact candidate identity and route contract. A scheduler acceptance handle,
okwrapper receipt, claimed tick, timeout, orunknownresult is not proof that the reviewer started; verify the literal child and arm a short, bounded, deduplicated recovery check when launch cannot yet be proven. Reporting-only prompts must explicitly permit this narrow supervisor-to-supervisor escalation while still forbidding direct product mutation or review. -
If prerequisites are missing, immediately assign and start the exact prerequisite work (cleanup, focused/full gates, formatting, immutable local head), persist the missing-gate receipt, and launch the reviewer at the first eligible completion callback without waiting for user prompting.
-
Read the durable reviewer report as soon as the process exits.
APPROVEimmediately triggers the contracted final verifier when its prerequisites are satisfied;REQUEST CHANGESimmediately triggers a bounded remediation owner with the report attached. Do not wait for the next periodic supervisor tick merely to route the verdict. -
After remediation, create a new immutable candidate and relaunch the same exact-role review; never reuse an older-head verdict.
-
Treat
kanban_request_reviewas the writer-lease cutoff: enumerate processes by worktree CWD, stop any surviving mutation child, and recompute the candidate identity before reviewer dispatch. A baseline-plus-diff digest can support an explicitly local pre-commit review, but it cannot support claims about the PR head or remote CI; commit and push first whenever the gate promises exact-head PR evidence. Seereferences/event-driven-kanban-delivery-cascade.md. -
Keep review/final-verification lanes separate from writer-floor counts. They may run alongside disjoint implementation writers when heavy/stateful gates are serialized, but do not run a dependent final verifier concurrently with its prerequisite reviewer. Until the required exact-candidate reviewer returns a parseable
APPROVE, a second model is advisory cross-review only; any later remediation invalidates it. Seereferences/reviewer-final-verifier-gate-ordering.md. -
Preflight dependency-released profiles before completing the parent. A parent completion can promote and spawn its final-verifier child within seconds, so verify in advance that the child profile resolves every declared
--skillsentry and can live-smoke the exact provider/model/fallback route from the intended workspace. A model-only smoke is insufficient when dispatcher startup preloads skills. A successful provider smoke proves transport/inference health only; it does not prove that the dispatcher profile can load its skills, obey the worktree identity, or perform the required terminal Kanban transition. Track provider, profile-startup, workflow-protocol, and child-survivorship health separately. If startup still fails, inspect the worker log, repair the profile capability, smoke the exact full invocation shape, then unblock and redispatch the existing child; do not create a replacement card. Read back a fresh claimed/spawned event, surviving PID, exact command line, first heartbeat, and eventual terminal transition. Seereferences/profile-skill-preflight-and-recovery.md. -
Before every launch, verify exact candidate identity, literal provider/model, fallback prohibition, CWD, duplicate absence, real-child survivorship, and credential minimization. Inspect the complete descendant process tree: removing credential variables from the main agent does not prove a configured MCP/helper subprocess lacks them. Read back the report before classifying approval.
-
Treat
--ignore-user-configas a configuration-isolation request, not proof that the launched agent has no MCP/helper processes. When the task requires a narrow tool surface, pass an explicit toolset allowlist supported by the live CLI (for example terminal/file/coding/web/skills as needed), scrub unrelated credential variables in the launcher, then inspect direct descendants immediately after launch. Record the literal route, CWD, PID, session handle, and tool-surface expectation in durable status. If unexpected helpers remain, classify the launch as broader than intended and correct the next bounded handoff rather than claiming isolation from the flag alone. Seereferences/bounded-hermes-worker-launch.md. -
When coherent remediation is still uncommitted and a normal commit is not yet authorized,
git stash createcan produce a temporary immutable tracked-file snapshot without changing the branch, index, or source worktree. Bind isolated detached review worktrees to its commit and tree IDs; reject intended untracked content and do not confuse the dangling snapshot with a delivery commit. Seereferences/immutable-uncommitted-review-candidates.md. -
If an exact reviewer exits with empty output, preserve the immutable candidate and classify it as reviewer execution failure. Retry once through the normal route and once through an independently configured invocation path; if both are empty, open a candidate-scoped reviewer-outage circuit instead of thrashing. Persist the exact identity and retry only after provider-health change or a later bounded tick. Never launch the dependent final verifier or weaken the named role.
-
A supervisor
okreceipt is not operational proof. Read the complete saved output’s actual response section (long files may begin with injected skill text), then independently verify claimed PIDs/CWDs, task-scoped mutations, candidate identities, and authoritative label/board readback. -
See
references/action-coupled-review-cascade.mdfor the callback state machine and supervisor acceptance contract. -
See
references/event-driven-kanban-delivery-cascade.mdfor the concrete same-card implementation→review topology, dependency-released final verifier, safe guarded graph construction, notifications, and acceptance verification. -
See
references/empty-reviewer-output-and-supervisor-readback.mdfor the bounded outage circuit and post-tick verification checklist. -
Repeated blocked status with no live child owning the next action is a recovery failure. Re-read the PR and worktree, identify the single exact blocker, and start/resume actionable engineering or review work in the same run. Report the next-action owner, active recovery evidence, and last concrete progress—not merely “still blocked.”
-
For mutex/capacity recovery, a bounded watcher may release an existing transiently blocked card after proving the owner task and its real child have exited. Treat generic
kanban dispatch --max 1as non-targeted: parsespawned[].task_id, then require the intended card's own claimed/spawned events and live child before reporting dispatch. Canonicalstatusoutranks stale auxiliary fields such asblock_kindor an older summary. Match external workers byopencode runanywhere in argv plus exact CWD/route; absolute executable paths make prefix-only checks unreliable. Before exact-head review, remove known generated coordination artifacts and prove the worktree is clean including untracked files; if the named reviewer profile is occupied, preserve canonicalreviewstate rather than launching Sol or a fallback reviewer. Seereferences/targeted-kanban-capacity-recovery.md. -
If replacement workers repeatedly rerun discovery or gates without producing a reviewable handoff, stop the restart churn. Preserve and inspect the residue, run one bounded continuation with a literal readiness postcondition, then create a local unpushed commit from the coherent reviewed scope so the opposite-family reviewer receives one immutable exact head. A local review head is not permission to push; final verification and lease-protected remote update still follow.
-
Treat mutation as the writer-liveness checkpoint: PID/CPU/tool activity with a clean worktree is stalled discovery. After one no-edit continuation, stop enlarging that transcript and launch a fresh, minimal edit-first session; in a credential-scrubbed isolated no-push worktree,
opencode run --automay be the narrowest practical mutation-capable mode. Pause duplicate dispatch until the first expected diff appears. -
Apply a bounded restart budget scoped to the exact
(issue, worktree, route/model, task shape)attempt—not as a fleet-wide stop latch. If the same route/model misses the declared mutation checkpoint again after the minimal edit-first retry, stop relaunching that task shape. Requeue it with preserved evidence, switch to a demonstrably mutation-capable route/task, or perform only deterministic mechanical synthesis from verified inputs. Exhausting one attempt tuple must not leave a requested writer floor at zero while other safe, dependency-clear issues or remediation findings exist. Do not count alive-but-reading processes toward a requested writer floor. For large mapping tasks, split source-bound identity extraction/validation from semantic enrichment rather than repeatedly asking one worker to complete the entire research-and-artifact pipeline. Seereferences/mutation-checkpoint-and-restart-budget.md. -
For
opencode run, inspect current help and beware that array-valued-f/--filecan consume a trailing positional message as another filename. Put the positional instruction before-f brief.md, poll immediately, and treat an argument-parsing exit as a worker that never started. -
A non-fast-forward push to a live PR branch containing real product commits is not a stale claim-only lock. Merge and preserve both histories, prove the resolved tree, rerun gates, and obtain an exact-new-head review rather than force-pushing.
-
When retiring branches behind merged PRs, separate the remote PR head from any local-only residue, preserve and disposition every divergent commit/artifact, and prove squash-merged content by patch or per-file content rather than commit ancestry. See
references/merged-lock-branch-recovery-and-content-proof.mdfor the verified patch-ID fallback, binary/truncation caveats, bundle sequence, and exact-target read-back. -
Prefer deterministic concurrency barriers over timing sleeps in regression tests. Prove that the competing operation was actually enqueued before releasing the held operation, then assert the final state.
-
For re-entrant singleton-lease deadlocks, make the regression call the real public reader/writer from inside the callback that previously held the lease; a test that merely asserts the whole callback stays locked can codify the bug. Use a bounded completion guard, prove RED against the exact parent implementation, restore production byte-for-byte, and then prove GREEN plus the adjacent service suite before review.
-
When one lifecycle/security finding spans several output formats, remediate one format at a time. Require a RED event trace proving render/write begins before gate acquisition, then a GREEN trace proving reservation, render, descriptor-based write, and durable registration all occur inside the gate. Validate the pattern before cloning it to the next format, and do not close the parent finding until every format plus combined serialized gates pass. See
references/format-specific-lifecycle-gate-remediation.md. -
Treat renderer memory safety as more than byte-length counting. Stream through the retained descriptor where the library permits it. When an API must return a whole archive, reject before renderer construction using a budget that covers every actually rendered formatted string and fixed per-record/per-slide object cost; otherwise many tiny rows bypass a character-only cap. Hash retained artifacts incrementally in bounded chunks, verify complete reads/EOF, and never allocate the entire output again merely to record identity. Real-renderer displacement tests remain required. See
references/single-writer-worktree-lease-and-renderer-streaming.md. -
In retention and cleanup code, only
ENOENTproves absence. Permission, I/O, short-read, and directory-enumeration errors are uncertain state and must remain retryable/fail closed; they must not advance a purge marker or clear a durable obligation. -
When reviews discover adjacent failures one round at a time, stop patching only the cited line. Model the entire ownership/state transition matrix—reserve, write, promote, durable re-point, rollback, restore, cleanup/unlink, and successful return—then inject every failure and race window. Require byte preservation, durable discoverability, retention exclusion of caller-owned data, and no silent orphan before requesting another exact-head review.
-
Apply the same matrix discipline at API and asynchronous UI boundaries: distinguish absent from present-malformed values, exact origin from hostname-only checks, transport success from schema validity, and stable resource tuples from unique request invocations. Test revisit collisions and the render-before-passive-effect boundary before synthesizing the next candidate. See
references/trust-boundary-and-request-identity-review-matrix.md. -
When a fail-closed security change intentionally reclassifies legacy records as unowned, run the complete adjacent service suite before advancing to the next finding. Older success-path fixtures may still use an identity-less convenience API and should be migrated to the current identity-bearing production registration path; preserve separate legacy tests that prove identity-less rows remain untouched. Never weaken the production check merely to restore stale fixture expectations.
-
See
references/named-review-role-and-stall-recovery.mdfor the validated role-correction, stall-recovery, divergent-head reconciliation, and deterministic-lock-test workflow. Seereferences/state-machine-failure-matrix-review.mdfor converging repeated lifecycle/artifact review rounds into one complete transition-matrix pass. Seereferences/fail-closed-fixture-contract-migration.mdfor reconciling adjacent test suites after ownership or provenance semantics tighten. -
See
references/lifecycle-quarantine-and-singleton-leases.mdwhen directory occupants cannot use hard-link restoration, when a mutable singleton database needs a full operation lease rather than a switch lock, when zero-byte artifacts participate in ownership, or when mutation tests must restore production files byte-for-byte after interruption.
Expanding a live mixed-model fleet
Before adding a second writer, review the complete live issue/PR/branch surface rather than selecting from labels alone:
- Fetch every open issue with bodies/comments, every open PR and exact head, issue branches, and recent CI.
- Compare issue branches to current default and distinguish claim-only, real implementation, PR-active, and stale. A
startingcomment or remote lock without a live process is not active work. - Map candidate issue files/contracts against active PRs and uncommitted worktrees. Treat shared schemas, migrations, IPC contracts, central services, and broad renderer surfaces as overlap even when issue titles differ.
- If implementation candidates overlap, use the added capacity for an opposite-family exact-head review or bounded read-only research instead of forcing another writer. A reviewer lane can shorten the critical path while preserving single ownership.
- Live-smoke every exact model route through the integration that will actually launch it; OAuth acceptance in Hermes does not prove a standalone Codex/OpenCode CLI accepts the same identifier, and a standalone CLI rejection does not prove the contracted Hermes provider route is unavailable. Smoke the real launch path, check that fallback is absent or forbidden, then verify the actual child command line and fresh model/tool activity after launch.
- When replacing a durable fleet slot with a subscription-backed model, mutate and read back the supervisor policy before process changes, reconcile any already-in-flight tick that captured the old prompt, and replace productive old-route workers only at a safe completion boundary. Latch weekly fallback only on an explicit subscription/quota response; distinguish temporary or persistent throttling from proven exhaustion. See
references/subscription-route-replacement-and-quota-fallback.md. - When an owner combines a successor-route change with “start another if safe,” separate future replacement policy from immediate capacity demand: grandfather productive old-route workers, persist/read back the new route, then recount literal writers and immediately fill a vacancy or explicitly authorized added lane with a verified disjoint task. Direct CLI workers that lack Kanban tools stay behind a duplicate-dispatch hold, and the supervisor owns their post-exit review transition. See
references/route-transition-and-immediate-capacity.md. - When repository policy mandates one issue-lock namespace, a supplemental model must reuse a proven stale claim-only branch rather than create a model-specific competing lock. Preserve and locally rebase the empty claim onto API-verified current default, remove GitHub credentials, apply a process-local disabled push URL, and leave remote replacement to the orchestrator under an exact lease.
- Keep supplemental subscription models role-honest: bounded research, documentation, tests, or advisory review do not inherit a contracted Terra/Sol-style reviewer or final-verifier identity.
- Check recent OOM evidence as well as current free memory. After package-install OOMs, serialize dependency installs and heavy full-suite gates even when two lightweight agent processes can safely coexist.
Supplemental operations workers beside a live fleet
When this user authorizes another worker “next to all the others,” preserve productive incumbents and use the explicitly added capacity for the current task. For an isolated operations-build worker, do not require unrelated product writers, reviewers, or the product planner to exit: scope exclusion to shared worktrees, candidate mutation, dispatch authority, locks, and heavy gates. This does not implicitly enlarge the product controller's slot count or authorize unlimited workers. Persist the scoped exception in the owning supervisor and canonical card before launch; remove contradictory global-idle wording rather than merely appending an exception. Retain single operations ownership and independent exact-role reviews. Reconcile checkpoints and briefs too, so a later recovery tick cannot resurrect the obsolete global stop condition.
Keep completion recovery enabled in observe-only mode while its canonical child runs; transfer to a child is not completion of the program. Require fresh child identity, CWD, inherited-lock and credential-name readback after every corrected launch. A live process without fresh task edits is only started, not implementation progress.
See references/supplemental-operations-worker-isolation.md for the scope checklist, launch-recovery evidence limits, and completion handoff rules.
Issue-wide remediation escalation
For Assessor, the user's correction supersedes the former Terra-as-remediator rule: more than three distinct Terra/Sol change-request rounds on the same GitHub issue requires first-party Sonnet as the next remediation writer, not DeepSeek. Count cumulatively across cards/candidates/PRs and deduplicate verdict events; row continuations do not reset the count. Preserve productive incumbents through checkpoint/exit, retain literal Terra review then literal Sol verification, and record the evidence and next route on the authoritative delivery card. A saved planner instruction is policy, not proof of enforcement: verify the next actual writer route.
Phase-aware backlog eligibility and named-issue refill
When the user asks for work regardless of priority, inspect exact delivery-card bodies, latest summaries/comments, and candidate state BEFORE saying an issue is unblocked. A GitHub ready label and empty parent list prove neither dispatch readiness nor acceptance readiness. A capability hold can be a duplicate-dispatch guard or carry a concrete gate/review blocker; read the reason before calling it administrative. Present implementation, gate recovery, review pending, and external acceptance separately.
For this user's named-issue refill requests, prioritize the selected set in the durable manager and board, preserve productive incumbents and dedicated lane roles, then actually launch eligible work through the authorized controller. Policy edits and queued comments are not launch evidence. Reuse existing candidates rather than assigning redundant writers. Report which issues started, which wait for a compatible slot, and which need review. Gate-only recovery may occupy capacity but does not establish an implementation-writer floor. Do not claim all selected issues are underway when some only have next-action notes.
For this user's “review the backlog, plan eligible work, then implement to completion” requests, an analyst report and worker launch are intermediate steps. Supply exact-main source and full selected-card histories, parent-validate recommendations, run real gates on preserved candidates before assigning authorship, and carry the validated plan through review/integration or an evidenced external blocker. Verify all issue classifications/counts in code; re-fetch named dependency states rather than trusting old blocked text. A canonical repository directory can be checked out at an old commit and must not be called current main.
See references/phase-aware-backlog-refill.md for the eligibility evidence checklist and truthful refill reporting. See references/backlog-plan-grounding-and-candidate-recovery.md for stale-source audit correction, gate-first remediation, meaningful untracked-file progress, and blob-proven alternate-index assembly onto current main.
Kanban activation, freshness, and takeover safety
-
A filtered capability-held queue can omit newly eligible product successors and stale needs-input cards. Reconcile current merged prerequisites and exact baseline content before concluding no work is available. Classify preserved candidates rather than reimplementing them blindly. Independent source-manifest or private compatibility foundations can proceed only with explicit scope exclusions and later integration gates; they do not complete the parent issue.
-
See
references/queue-and-supervisor-recovery-readback.mdfor post-failure card readback, wrapper-exit races, independent supervision health, and precise artifact progress reporting. -
Never bulk-dispatch a newly imported backlog. Mirror only executable children, hold every imported card behind a dispatcher-recognized guard, reconcile dependencies/routes/workspaces, then release one bounded item.
-
Treat the product tracker and execution board as different layers when the project contract says so: GitHub Issues/PRs may be authoritative for product state while Hermes Kanban preserves execution attempts, workspaces, reviews, blockers, and receipts. Do not mechanically force every card to match its issue state.
-
Make board freshness an explicit supervisor acceptance criterion. At the start and end of every durable tick, reconcile active cards against live issue/PR state, exact heads, branch locks, worktrees, worker processes, and persisted review artifacts; correct known drift in the same run and read changed cards back.
-
Preserve immutable run history while removing stale operational state. Terminalize superseded attempts with exact closure/merge receipts, and create idempotent dedicated cards for live PR review or human action when the historical implementation card no longer represents the current next step.
-
A card created with
--initial-status blockedcan be promoted immediately if it has no dispatcher-recognized unresolved guard. Apply a typed human/capability/dependency block with a precise reason and next-action owner, then verify the latest event and status after the dispatcher has had a chance to act. -
Treat review-card transitions as dispatch-capable mutations. Unblocking a review card merely to convert it into a human merge hold can immediately spawn a duplicate reviewer. Prefer a dedicated idempotent human-action card; otherwise suppress dispatch, use the supported review-to-ready transition, apply the typed hold immediately, and inspect both card events and OS processes. Reclaim only a duplicate whose immutable exact-head review already exists.
-
When an owner resolves a blocker by splitting scope, reconcile the authoritative issue, PR body/draft state, hosted workflow activation, residual backlog issue, board cards, and processes as one transaction. Preserve quality gates unless the owner explicitly changes them; enabling a workflow is not permission to weaken it. See
references/owner-scope-decision-reconciliation.md. -
Merge or otherwise establish the governing role/routing contract before new implementation dispatch. If productive workers were already running when drift was found, preserve their work while preventing further dispatch.
-
A Hermes Kanban wrapper marked
runningis not proof that implementation is running. Confirm a live external CLI child with the exact route; otherwise report coordination, probing, or blockage rather than implementation. -
When a writer is intentionally launched outside the Kanban dispatcher, a
blocked/held card may be the correct duplicate-prevention state rather than evidence that engineering is blocked. Keep the dispatcher guard in place, append the current external process handle/PID, exact route, worktree, next owner, and gate-serialization rule, then read the card back. Report the lane as externally implementation-active while describing the card itself as held against duplicate dispatch; do not unblock it merely to make board status look active. -
When recovering accidental fan-out, preserve all worktree residue, stop stale-route or no-writer coordinators, write the canonical route back to each affected card, and keep cards blocked until explicitly released. Do not delete residue or restart everything merely to make board counts look clean.
-
When converting stale on-disk plans into a live fleet backlog, first audit each obligation against current code and record delivered work as epic baseline. Build a stable-key dependency graph, make every unresolved identity/retention/provider/schedule/publication decision a human-owned leaf, independently review the graph, then create issues topologically from a resumable manifest. Future migration tasks use
<next>and graph predecessors—never predicted migration numbers. -
Read back every Kanban idempotency hit before applying holds: it may be an existing root with runs or decomposed lifecycle children. Repeated same-kind blocking can trigger loop protection, so preserve existing topology and verify the exact event/status rather than forcing the board to match a planned count.
-
See
references/continuous-kanban-reconciliation.mdfor the validated start/end reconciliation loop, typed-block read-back check, stale-attempt terminalization, and cron acceptance contract. -
See
references/plan-to-issue-fleet-activation.mdfor current-state plan auditing, independently reviewed issue graphs, resumable bulk creation, migration serialization, held Kanban imports, and one-lane release.
Reconciling a stale claim-only lock branch
A remote issue lock can remain claim-only while useful implementation exists on a different local history based on newer main. Do not merge the two claim commits blindly and do not delete/recreate the lock.
Before replacing the stale remote head:
- Read the issue state and labels, all open PRs for the branch, the current remote branch SHA, and the remote comparison against default in the same turn.
- Prove the remote branch contains only the stale claim commit—no product files or unreviewed commits—and that no live worker owns it.
- Prove the local candidate starts from current approved default, contains only the intended reviewed change, and has current focused/full gate evidence.
- Commit the reviewed scope locally, then update the existing lock branch with an explicit exact-head lease, for example
--force-with-lease=refs/heads/claude/issue-N:<observed-remote-sha>. Never use an unqualified force push. - If the lease fails, stop: another actor changed the branch. Re-read issue/PR/branch state rather than retrying.
- Read back the remote SHA before opening the PR, then verify the PR base, head ref, exact head SHA, body, and exact-head CI runs.
This preserves the one-lock namespace while preventing a stale claim commit from forcing useful reviewed work onto an obsolete base.
Permission and duration policy
- Edit-accepting modes may still block shell, network, package-manager, or test commands.
- A worker sandbox may bind-mount manifests or governance files read-only even when the host paths are writable. For an immutable-SHA integration blocked by
unable to unlink ...: Device or resource busy, preserve the attempt, create a fresh detached worktree, perform only the exact no-commit merge in the orchestrator/host context, verifyMERGE_HEADand staged state, then hand the integrated candidate back to the restricted worker for tests and bounded repairs. Do not grant mount capabilities or weaken the sandbox globally. Seereferences/protected-file-sandbox-merge-recovery.md. - Prefer the ordinary tracked background-process or durable dispatcher path and verify that the real child survives after the launch call returns. On Linux, when that backend does not retain the child, a user-systemd transient service can provide an independently managed lifetime; use an absolute verified CLI path, credential-scrubbed launcher, durable logs, and post-return PID/mutation checks. See
references/durable-linux-systemd-workers.md. - In a disposable isolated worktree with no push authority, use the narrowest permission mode that can execute the required gates; otherwise run blocked verification through the orchestrator.
- For an approval-protected named-role review, prepare the final invocation before requesting approval: reconcile concurrent supervisor receipts/processes, require a clean immutable exact head, generate the SHA-bound brief, and remove unrelated credential variables in the launch command itself. Do not launch unsanitized and inspect credentials afterward.
- Treat approval as command-specific. Adding credential sanitization or otherwise changing a protected launch may require a fresh approval; if it times out, stop and request approval for that exact sanitized command rather than trying alternate shapes or falling back to an unsafe launch.
- Immediately after launch, verify the real child command, CWD, route/fallback marker, and absence of unrelated credential-variable names. If the worktree drifted or a supervisor already reviewed/fixed the prior head, terminate duplicate/stale review work and regenerate the brief for one committed exact head.
- Avoid arbitrary tool-round caps for delivery work. Prefer resumable task-sized runs with no fixed cap or a deliberately high cap when persistence is known.
- For multi-source research or documentation workers, do not let discovery consume the entire turn budget before the deliverable exists. Require durable per-source notes or incremental target-file edits after each verified batch, reserve a final tranche for writing and checks, and inspect the worktree at checkpoints. An exit-zero
iteration budget reachedmessage with no required artifact is an incomplete run. - Resume only the missing bounded phase. If resuming a very large transcript triggers expensive preflight compaction or repeats discovery, prefer a fresh continuation supplied with the durable notes and a literal artifact postcondition; do not rely on unpersisted model context as the only copy of completed research.
- Stop after one bounded continuation and one evidence-packet synthesis attempt fail before mutation. Terminate the worker to prevent a race, let the orchestrator perform only the mechanical synthesis from verified evidence, and report research contribution, artifact authorship, and delivery state separately. Never credit the worker with completing an artifact it did not write.
- See
references/research-worker-artifact-recovery.mdfor artifact-first checkpoints, evidence-packet construction, retry limits, citation verification, and honest handoff attribution. - Use the task worktree's exact lockfile. Never validate against a sibling checkout's
node_modules. If a writer discovers missing dependencies, stop before it substitutes a sibling symlink or letsnpxauto-fetch implicitly; have the orchestrator perform one explicit, serialized lockfile install in that same worktree (for examplenpm ci --ignore-scriptswhen the repository permits it), record the install/audit result, then resume only the missing implementation or gate phase. - Do not parallelize two gates that execute the same stateful integration or recovery harness. For example, a full test suite and coverage suite may both invoke Docker-backed recovery drills and race over temporary canaries/cleanup. Serialize those gates unless the repository proves per-run isolation.
Completion is evidence, not prose
A worker summary and exit code are advisory. The same rule applies to reviewers: if stdout contains only a terse verdict while the invocation requested a durable report, read and validate that exact report before classifying findings or launching remediation. Recheck the supposedly read-only worktree afterward; reviewer bootstrap attempts can leave package-manager metadata or other residue. Reviewers, verifiers, discovery processes, and gate-only workers never count toward an implementation-writer floor. See references/reviewer-report-readback-and-residue.md.
The orchestrator must inspect:
- live/exited process state;
git status, full diff, changed-file scope, andgit diff --check;- security-sensitive changed files;
- focused tests after the latest write;
- required full gates;
- tooling residue and duplicate draft tests;
- independent review for high-risk changes.
Exit code zero is not completion when the worker admits implementation or tests are pending. It is also not completion when the transcript ends on a denied permission request, repository/credential discovery, directory listing, malformed edit/script, or other detour without the required verdict, artifact, or gate receipts. Define explicit postconditions before launch (for example: named report/verdict, changed-file set, and every required command result) and validate them independently after exit. For source/test-writing workers, run the smallest parser/compiler/focused import test before preserving the handoff. If the only delta is uncommitted syntactically corrupt residue from a recorded immutable head and no coherent production artifact exists, restore only those files to the head, prove clean state, then retry with incremental syntax checks; never preserve corruption merely because the process exited zero. See references/malformed-worker-output-recovery.md for the rollback boundary, route-policy preflight, and fleet-refill procedure.
For structured artifacts, never infer “no mutation” from the tail of an interrupted transcript. A worker may have written a malformed or half-populated row before returning to discovery. After every exit—including SIGTERM—inspect the actual diff, validate schema width and source-bound totals, and compare the worker's writable slice to the last independently green checkpoint. Keep mechanical identity extraction separate from semantic enrichment; preserve the complete validated ledger, roll back only the incomplete enrichment slice, and rerun the biting validator before refill. See references/partial-artifact-checkpoint-recovery.md.
Treat delayed completion receipts as historical evidence, not instructions to restore old state. Before reporting or refilling, re-read live processes by CWD, current HEAD, staged/unstaged/untracked files, and newly created supervisor commits; separate inherited delta from worker-authored delta, and absorb superseded receipts without reviving stale PIDs or conclusions. See references/superseded-receipts-and-delta-provenance.md.
A green checkpoint validator is not proof that the newly assigned enrichment slice landed. The validator may still encode the prior accepted mapped set and correctly pass while the target row remains blank. Before accepting a structured-slice completion, inspect the exact target identity, require every newly mandatory field to be populated, confirm the validator's pinned mapped set advanced by exactly the intended identities, and prove all previously accepted rows remain byte-exact. If the target was already populated and the worker only strengthened a mutation check, report it as validator hardening—not as a newly mapped row or semantic progress. See references/slice-acceptance-vs-checkpoint-validation.md.
For constrained read-only reviews, prefer a complete final report on stdout; have the orchestrator persist it afterward. Asking the worker to write directly to a report path outside its worktree can trigger an external-directory denial even though local review is permitted. Supply the exact base/head and complete task brief up front, explicitly forbid network, credential, GitHub-config, parent-directory, and unrelated repository discovery, and treat any missing report as an incomplete attempt rather than approval.
A nonzero exit is not necessarily code failure: delivery presets may reject the final response because a post-mutation verification receipt or review report is missing even when edits and tests succeeded. In that case, inspect and verify the work instead of blindly restarting.
User-directed writer replacement and artifact-first recovery
For this user's “start it now” / “have Fable fix it” corrections, a manager invocation or verified task comment is not the requested worker launch. Verify the real implementation child, exact route, CWD and write/test capabilities; distinguish dispatch requested, worker started, actual edits, tested and accepted. Do not repeatedly substitute another manager handoff for execution.
When the owner replaces a writer, reconcile both current dispatch policy and pending automatic successors so stale Astra-only or reviewer-only assignments cannot return. Preserve work; stop only identity-verified superseded workers. Coordinate with any in-flight manager under the existing shared ownership fence before attempting a takeover. A lock refusal requires incumbent readback, never a new lock path or deletion.
The user's objection to long no-output runs requires early saved artifacts and bounded progress checkpoints. Compare content hashes to the staged baseline, not file existence alone. A copied plan is not authored output; a changed plan is not code implementation. Inspect real activity and preserve partial edits before diagnosing/restarting a stalled or failed worker. Do not impose an arbitrary wall-clock kill on productive work.
When the owner explicitly makes low-severity findings nonblocking, track them as deferred without claiming resolution or starting another low-only correction loop. Keep this scoped to the authorized repair; medium-or-higher defects, mandatory security/acceptance checks and independent final verification remain. Count validated review verdicts separately from provider/launcher failures, retain history, and honor an explicit subsequent owner decision releasing a threshold hold.
For timed pauses, bind the resolved timezone/instant and check it immediately before every model launch, including automatic successors. Do not retroactively admit a review that ran before the hold expired.
See references/artifact-first-owner-recovery.md for the observed duplicate-launch race, progress attribution and evidence-closure lessons.
Recovery sequence
- Preserve useful uncommitted edits.
- Identify exactly what is missing: implementation, RED/GREEN evidence, full gates, cleanup, or independent review.
- Rerun the missing commands directly or resume with a prompt naming only the missing work.
- Relaunch only when code work remains; do not relaunch merely to obtain polished final prose.
- Send security-sensitive diffs to the designated opposite-family reviewer after the latest mutation.
- If mistaken completion deleted a scratch candidate, keep the old card as failure/review evidence and create a replacement in a durable worktree or explicit directory. Reconstruct from current main and rerun every gate; historical green output is not proof for rebuilt code.
- Do not trust
--initial-status blockedas a dispatcher hold unless the card has an unresolved dependency or recognized guard. With an active gateway, prepare the durable workspace, exact lock, route smoke, and complete body before creation, then read back events immediately.
Live status
When asked what is running, check tracked processes, delegated children, OS-level agent/profile processes, cron, containers, repository/PR state, and repository-hosted automation. Local process absence does not prove nothing is building: GitHub Actions schedules, manual dispatches, pushes, and PR events can run entirely off-host. Report running, exited, verified, blocked, and scheduled separately.
For every running workstream, distinguish the orchestration wrapper from the actual implementation child. Name a task as implementation-active only when the external CLI child is alive on the approved exact route and has met the task's declared mutation checkpoint; a CPU-active child that only reads or probes while the worktree stays clean is stalled discovery, not code writing. A wrapper that is merely reading configuration, smoke-testing, waiting, or recovering is not an active writer. Treat completed local commands and failed watcher processes as exited evidence, not current work.
- For cron-owned execution, verify child survivorship again after the cron tick completes. A successful cron receipt proves the tick completed, not that a shell child remains alive; use a durable dispatcher/queue or transfer ownership to a parent-managed process when the backend scopes children to the tick.
- Keep each supervisor tick well below the scheduler's stale-claim allowance: launch/reconcile, persist receipts, and return without waiting for workers, reviews, suites, builds, or CI. A tick that outlives the allowance can strand the in-flight guard and skip later refills. Add an explicit runtime budget to the durable prompt. Disable run continuity when historical output repeatedly resurrects merged PRs or stale branches; delayed receipts remain evidence, never instructions.
When the user says a repository or fleet is paused, enforce the boundary at both layers:
- hard-scope supervisors so they cannot inspect, enter, fetch, build, test, modify, or dispatch work for paused product repositories;
- allow system-level inspection only to detect and safely stop out-of-scope product work without deleting residue;
- inspect hosted workflow runs and their trigger events before attributing the source;
- inspect workflow activation state, because a scheduled workflow can restart after the current run exits;
- if the instruction clearly means no builds at all, disable every build-capable workflow in the paused repository, not only the currently visible scheduled workflow;
- map workflow IDs to names/paths before mutation, change one intended target at a time, and read back all workflow states plus queued/in-progress runs after the final change.
Do not mistake monitoring for building: repository status enumeration is read-only, but it should still be excluded when a hard scope says not to inspect unrelated repositories. State whether observed work came from a local agent, local scheduler, or repository-hosted automation.
If gh pr checks --watch or gh run watch fails on restricted Checks/annotation access, do not report CI failure. Read every Actions run for the exact head SHA through REST and require all required workflow conclusions to be successful.
Reference
- See
references/exhaustive-read-only-project-audits.mdfor exact-route whole-repository audits, complete tracker-packet compaction, large-prompt handling, credential-scrubbed read-only launches, and audit-to-sliced-delivery handoff. - See
references/kanban-dispatch-activation-safety.mdfor safe board import, durable activation holds, exact external-route metadata, and accidental-fan-out recovery. - See
references/live-status-evidence-recovery.mdfor wrapper/child process classification and unified status evidence. - See
references/paused-repository-build-suppression.mdfor diagnosing off-host scheduled builds, hard-scoping supervisors, disabling paused-repository workflows safely, and proving the repository is quiet. - See
references/paused-repository-reactivation.mdbefore re-authorizing a paused repository or cloning another fleet's setup; it covers stale/credentialized checkouts, old cron deny-lists, governance-first activation, product-specific drift checks, held Kanban import, and Projects-v2 access fallback. - See
references/github-actions-rest-monitoring.mdfor restricted-token CI monitoring. - See
references/write-capability-and-evidence-artifact-recovery.mdfor mutation-capability preflights, evidence handoffs, and protectedAGENTS.mdgovernance writes that require a dedicated foreground approval rather than mixed parallel batches. - See
references/reasonix-recovery-patterns.mdfor concrete Reasonix permission, step-limit, readiness-gate, and exact-lockfile recovery patterns. - See
references/kanban-accidental-completion-recovery.mdfor rebuilding a lost scratch candidate into a durable replacement card without falsifying the original review history. - See
references/durable-roadmap-cron-supervision.mdfor action-coupled status ticks, exact-head review continuity, writer-floor enforcement, duplicate prevention, and recurring-schedule verification. - See
references/mixed-opencode-second-lane-selection.mdfor whole-backlog branch classification, overlap-aware second-lane selection, exact-head reviewer dispatch, and resource-safe mixed-model launches. - See
references/supplemental-subscription-model-lanes.mdfor assigning spare OAuth/subscription model quota beside a primary writer, preserving stale claim-only locks, credential-isolated launch guards, and proof of real model activity. - See
references/opposite-family-review-correction-loop.mdfor exact-head review, validated finding correction, delta re-review, post-rebase gates, lease-protected branch updates, and merged-main verification. - See
references/protected-review-launch-preflight.mdfor stable-head reconciliation, credential-minimized protected launches, command-specific approval handling, and live child verification. - See
references/stalled-pr-handoff-and-opencode-attachment.mdfor converting repeated worker restarts into an immutable local review head, reliable OpenCode attachment argument ordering, and credential-isolated live-process verification. - See
references/opencode-no-edit-recovery.mdfor mutation-based writer liveness, fresh edit-first--autorecovery, cron child-lifetime checks, and digest-bound review of staged candidates. - See
references/bounded-slice-checkpoint-recovery.mdfor splitting source-bound ledgers from semantic enrichment, recovering partial structured artifacts after any exit, restoring writer floors, and adversarially checking race fixes for residual check-then-act windows. - See
references/retiring-external-agent-runtime.mdfor removing a retired agent's executable, credentials, services, fleet state, cron routing, and durable role assignments while preserving inert project evidence. - See
references/owner-scope-decision-reconciliation.mdfor propagating a human scope decision across issue acceptance, PR body/readiness, workflow activation, split backlog, Kanban holds, and duplicate-dispatch cleanup. - See
references/completion-batch-races-and-immutable-handoffs.mdfor double-recounting completion batches, distinguishing reviewer wrappers from duplicate writers, attributing staged residue, and validatinggit stash createcandidates against their first parent. - See
references/github-branch-claim-reconciliation-safety.mdfor first-commit claim identity, exact-head CAS deletion, PR/label projection semantics, concurrency convergence, and immutable workflow-action pinning.
Supporting files: this skill's supporting files are held in the docsite at
docs/15-skills/_support/engineering/external-coding-agent-orchestration/— fetch them fresh fromjknash/docsitemain alongside this page. Source:jknash/hermes-shared-skills· branchhermes-jkdev001@1d0d545c3970·skills/engineering/external-coding-agent-orchestration/· view source · Imported 2026-10-03. Supporting files (references, scripts) remain in the source repository.
version 1.0.0 · author Hermes Curator · license MIT.
Published by Muse · 2026-10-03.