Skip to main content

Multi-Agent Program Delivery

Use for phased software delivery across multiple agents. Use this skill for large software initiatives that span multiple phases, security boundaries, modules, migrations, or pull requests and assign distinct models/agents to planning, implementation, repetitive work, cross-review, and final verification.

This skill coordinates the program. Load the domain skills for security, architecture, TDD, Git, CI, and module implementation as needed.

Core operating model​

Assign explicit roles before coding:

  • Planner/verifier: owns capability boundaries, specifications, dependencies, acceptance gates, integration decisions, and final truth claims.
  • Implementation workers: own focused S/M-sized changes in isolated worktrees.
  • Opposite-family reviewer: reviews each implementation independently so an agent never certifies its own work.
  • Mechanical/high-volume worker: handles fixture matrices, repetitive file moves, contract inventories, dependency-graph checks, and other bounded work after architecture is fixed.

For this user's personal-dashboard program, default to Sol for planning/final verification, GPT-5.6 Terra and Reasonix DeepSeek Flash for implementation and reciprocal review, and GPT-5.6 Luna for high-volume repetitive work.

Owner corrections: execution, capacity, and named takeover​

This user expects a request to review a completion plan and implement it to continue into verified implementation, not end at an audit, revised plan, scheduled job, or worker launch. Validate and consume review findings into the existing task authority; assign the next missing acceptance step and preserve completion ownership. Reports must distinguish review complete, implementation running, tested, integrated, and delivered.

When the user says “just have Astra do the work,” treat that as a task-scoped executor correction: assign actual implementation to literal Astra, not another planning pass or a handoff to a different writer. If a separate Astra process is launched, say so rather than implying the foreground assistant made the code changes. Preserve completed work and failed-attempt lineage; refresh source, ownership, and candidate state before takeover. A newly authorized attempt does not erase the old failure or waive independent acceptance gates. Apply current user role instructions rather than historical reviewer defaults in this skill.

When the user authorizes more workers for speed, identify independent executable slices before increasing capacity. Use disjoint file allowlists, isolated workspaces, per-assignment ownership, and serialized shared/heavy gates. Extra implementation workers do not confer extra deployment, merge, or product-launch authority. Keep the existing coordinator and canonical backlog; do not make another manager just to supervise the new workers. More workers cannot shorten an evidence-mandated soak or replace missing production wiring.

Approved split means verified workers, not another manager request​

When this user approves a proposed split and says to add workers, carry the request through actual worker dispatch and identity/tool-use readback. A manager prompt edit, canonical-card comment, or deduplicated cron fire is coordination evidence only. If the scheduler already owns a fire, reconcile that owner; do not infer a worker launch or start a competing dispatch. Under explicit supplemental-worker authorization, reserve the exact named lanes with the existing coordinator before foreground dispatch, and publish their actual launch handles afterward.

For findings that overlap central files, split by implementation interface rather than finding number alone: one primary writer owns shared controller/boundary wiring; supplemental workers own standalone modules and their tests. Bind APIs, exclusive paths, baseline hashes and integration order before launch. Adding new module paths is a named staging-scope amendment, not silent permission to exceed an earlier allowlist or deploy additional files. Keep one integration owner and independent exact-candidate acceptance. Distinguish workers started, initial checkpoints, actual code changes, tested modules and integrated fixes; standalone module tests do not establish end-to-end closure.

For build-impact questions, distinguish application-source changes from worker execution, acceptance-evidence and rollout changes. Give evidence-backed qualitative risk rather than invented percentages, and name the real-build proof still needed. A functioning trusted operations worker does not prove a confined product worker can bootstrap, authenticate, access Git or produce an accepted artifact.

See references/parallel-remediation-interface-splits.md for the verified supplemental-lane dispatch pattern and its acceptance limits. See references/completion-ownership-and-named-takeover.md for audit handoff, parallel-slice boundaries, terminal-result validation, and session-derived evidence limits.

Workflow​

1. Establish the live baseline​

Before planning or dispatch:

  1. Inspect the active branch, clean/dirty status, divergence from origin/main, open PRs, and current test state.
  2. Preserve unmerged work. Put planning in a separate worktree/branch rather than mixing it into a feature branch.
  3. Verify each named worker is actually configured with the intended provider/model using a one-shot inference marker.
  4. Verify external-agent CLI flags from its current help before launching a long run.

Do not treat a configured profile name as proof that the intended model can execute.

2. Create and approve a capability map​

For an initiative with multiple independently testable areas:

  1. Define stable module IDs, responsibilities, dependencies, trust boundaries, data classes, and build order.
  2. Record explicit assumptions: user model, interfaces, deployment style, technology constraints, deferred scope, and migration prerequisites.
  3. Define module boundary rules before implementation details.
  4. Open a reviewable capability-map PR.
  5. Obtain explicit human approval before writing the master spec or implementation code.

A capability map is a control against scope drift and accidental cyclic architecture, not merely an outline.

3. Write the master spec and canonical task registry​

The planner owns:

  • target architecture and dependency direction;
  • canonical module IDs;
  • security/privacy contracts;
  • transport-neutral command/query/event contracts;
  • phase gates and success criteria;
  • a dependency-ordered roadmap;
  • a canonical task file with stable IDs, owner, dependencies, Acceptance, and Verification for every card.

Use recursive planning. Near-term cards may be executable; distant cards may remain bounded epics, but explicitly require a module spec and S/M decomposition before assignment.

Avoid duplicating the complete task-ID registry in multiple documents. Name one canonical file and have overview plans reference it.

4. Adversarially review the plan before approval​

Run three complementary passes:

  • Architecture/security reviewer: finds dependency cycles, late foundations, transport coupling, authority gaps, privacy omissions, and unsafe parallelism.
  • Independent second reviewer: challenges the first review and verifies findings against the actual repository.
  • Mechanical validator: checks duplicate/missing task IDs, nonexistent dependencies, phase cycles, naming drift, owner ambiguity, and missing Acceptance/Verification sections.

Resolve findings in the artifacts, then rerun targeted second passes. The planning owner decides readiness only after independently checking the revisions.

Common plan defects to catch:

  • a feature phase depends on UI/platform work scheduled later;
  • encryption/key management appears after sensitive-data modules;
  • domain modules register directly with transport adapters despite a transport-neutral architecture;
  • data retention/export/deletion is deferred until release instead of required in each module spec;
  • multiple agents are allowed to modify one migration chain or central contract concurrently;
  • a monorepo/web-host restructure is assumed but not represented as a task;
  • dependency-remediation work combines unrelated package families into one rollback unit;
  • AI/provider egress is less constrained than connector egress.

5. Execute in small, reviewable slices​

For each task:

  1. Start from current origin/main in an isolated worktree.
  2. Give one agent the approved scope, allowed files, acceptance criteria, required tests, and explicit commit/push policy.
  3. Use tests first for behavioral changes.
  4. Keep database privilege, identity, audit, encryption metadata, outbox, and other shared migration changes sequential.
  5. Parallelize only genuinely independent files/contracts after their shared interface has merged.
  6. Have the opposite agent family inspect the actual diff and test evidence.
  7. Return required findings to the implementer, then rerun the review.

A reviewer finding is not closed because the implementer says it is fixed. The reviewer or planner must inspect the revised diff.

6. Independently verify agent work​

Agent summaries are advisory. The planner/verifier must:

  1. Inspect git status, diff, changed-file scope, and git diff --check.
  2. Read security-sensitive changed files directly.
  3. Run focused tests after the latest write.
  4. Run the required full repository gates.
  5. Rebase/update against current origin/main without losing work.
  6. Rerun relevant gates after the rebase.
  7. Verify the remote branch, PR body, merge state, merged commit, and clean repository state.

If an external worker cannot execute verification because of its approval layer, do not retry indefinitely and do not accept its self-review. Preserve the edits, run the commands through the orchestrator, and feed real results to the cross-reviewer.

Before restarting an interrupted external worker, reconcile both the tracked process registry and OS-level child processes. A wrapper exiting with 143, an empty receipt, or a stale tracked status does not prove the coding child stopped. Never launch another mutator into that worktree until duplicate or orphaned children are identified and only one owner remains.

Treat golden fixtures as executable contracts, but reconcile them against canonical application helpers and declared inputs before forcing exact equality. If a fixture contradicts production compliance math, timezone rules, ordering, or lacks source data for editorial prose/grouping, stop the downstream lane: record the concrete mismatch, keep the failing test honest, and obtain a source-of-truth decision. Never weaken the assertion or hard-code sample/customer content merely to make the fixture green.

After authority is decided, preserve both purposes of the artifact: the machine fixture must be fully derivable and exact, while a rich visual/reference example must remain representative enough to exercise layout, long prose, evidence, remediation, grouping, and pagination. Prefer enriching general-purpose input metadata or maintaining a clearly labeled visual fixture over replacing a polished reference wholesale with generic output. Any broad fixture rewrite requires independent review before downstream renderer work starts.

See references/worker-process-reconciliation-and-fixture-conflicts.md for the recovery, contradiction-gate, and post-decision fixture-preservation procedure. For reporting-specific effective-status, scope filtering, deterministic fixtures, and permission-causation rules, see references/report-contract-fixture-and-causation.md.

7. Use the high-volume worker safely​

Good uses:

  • expected privilege matrices;
  • recurrence/adversarial fixtures;
  • task-ID/dependency validation;
  • approved mechanical file moves;
  • documentation indexes and inventories.

Do not delegate architecture, cryptography, authorization models, destructive migrations, or release decisions to the mechanical worker.

8. Maintain explicit stop/go gates​

Every phase gate should link:

  • exact commands or evidence artifacts;
  • required cross-reviews;
  • pass/fail criteria;
  • residual-risk decisions;
  • rollback/restore evidence where relevant.

Do not begin downstream sensitive work while an upstream security or delivery gate is open.

9. Answer interruption/status questions from live state​

When the user asks what is running or where work stands:

  1. Inspect the tracked background-process registry.
  2. Inspect live OS processes for relevant workers/tests/servers.
  3. Inspect Git branch divergence and open PR state.
  4. Separate running, exited, verified, and unfinished remote action clearly.
  5. Never describe a completed background process as still running because it remains in the registry.

Give the concise operational answer first; include only the next unfinished action and blockers.

10. Archive program knowledge​

When the user asks to preserve the work in a semantic memory provider:

  • archive the active user/assistant conversation;
  • archive durable artifacts such as review reports, capability maps, specs, plans, and task registries;
  • add a concise initiative summary with current gate, review decisions, and artifact/PR references;
  • mine into the intended stable user/project namespace;
  • verify with at least two semantic searches and provider status;
  • update the summary after the current code/PR state changes materially.

Do not claim persistence from provider configuration alone; verify retrieval.

Pull-request discipline​

  • Keep planning, behavior, refactors, migrations, and dependency families in separate PRs when they have independent rollback paths.
  • Target 100–300 changed lines; split work approaching 1,000 lines.
  • Every PR body records scope, real verification output, cross-review disposition, and roadmap task ID.
  • Rebase/merge current planning changes before force-updating an older feature branch; use lease-protected history updates.
  • Final merge is a planner/verifier decision, never an implementing agent's claim.

Verification checklist​

  • Live baseline and unmerged work preserved
  • Named worker models live-tested
  • Capability map approved
  • Master spec and canonical task registry approved
  • Plan received architecture, independent, and mechanical reviews
  • Shared migrations/contracts serialized
  • Each implementation received opposite-family review
  • Planner inspected diff and reran focused/full gates after latest write/rebase
  • Remote PR/merge state independently verified
  • Phase gate evidence and residual risks recorded
  • Semantic archive retrieved successfully when requested

Reference​

See references/personal-workspace-program-example.md for the session-derived example that motivated the cycle, encryption, agent-review, safe-integer, and MemPalace verification checks.


Supporting files: this skill's supporting files are held in the docsite at docs/15-skills/_support/engineering/multi-agent-program-delivery/ — fetch them fresh from jknash/docsite main alongside this page. Source: jknash/hermes-shared-skills · branch hermes-jkdev001 @ 1d0d545c3970 · skills/engineering/multi-agent-program-delivery/ · view source · Imported 2026-10-03. Supporting files (references, scripts) remain in the source repository.

version 1.0.0 · author Hermes Curator · license MIT.

Published by Muse · 2026-10-03.