Skip to main content

llm-kanban runbook

How to set up and run an llm-kanban board: the in-repo Markdown kanban fleet projects use for agile delivery. Schema of record: the llm-kanban schema standard. Identities follow the agent naming standard and the fleet agent registry. Roles: the stakeholder (Justin) is product owner; the project manager (PM) runs the board, the cadence, and this runbook.

The lifecycle at a glance (the prose below governs):

1. Start a board on a project​

Prerequisites: a GitHub repo for the project (one project, one repo), the stakeholder's agreement on scope, and the PM designated.

  1. Write PROJECT_CHARTER.md at the repo root with the stakeholder: problem, outcomes, deliverables, out of scope, constraints, project acceptance criteria, risks. No board before a charter.
  2. Create the project's docsite home: a folder docs/03-projects/<slug>/ in jknash/docsite with a living work-record.md (owner rule, 2026-09-29). The project is tracked there from the moment it is queued, not retrofitted later; append progress as work lands.
  3. Copy the pilot board skeleton from jknash/dashboard kanban/ (schema.yaml, schema.md, AGENTS.md, index.md, backlog.md, kanban.py, empty stories/, sprints/, retros/). Set the project id_prefix in schema.yaml (two letters, unique across projects).
  4. Add a bootstrap section to the repo's root AGENTS.md: read the charter, read kanban/AGENTS.md, never work an unclaimed story.
  5. Seed kanban/agents.yaml from the fleet registry — every agent that may claim work on this project, and no others.
  6. Capture known work as one-liners in backlog.md. Groom the first sprint's worth into story files (step 3).
  7. Write the first sprint file, run kanban.py validate and kanban.py render, commit everything, and read the tree back from the remote to confirm.

Guardrail: kanban/ is project process, never site content. On repos with a published site, confirm the build/deploy path cannot ingest kanban/ before the first deploy after setup.

2. Grooming (PM, before each sprint)​

  1. Promote backlog captures to story files with kanban.py new, then edit in the story, acceptance criteria, and dependencies.
  2. Estimate in Fibonacci points with the team convention: anything over 8 gets split before it can become ready.
  3. A story is ready only when it meets the Definition of Ready (schema standard). Move it with kanban.py move <ID> --to ready.
  4. Keep the stakeholder's queue explicit: every blocked story names blocked_owner and the exact action_needed. Surface the waiting-on-Justin list from status.md at every review.

3. Working a story (any agent)​

  1. git pull --ff-only — never work from a stale board.
  2. python3 kanban/kanban.py status --available — pick a ready, unblocked story whose dependencies are done.
  3. python3 kanban/kanban.py claim <ID> --agent <your-id> — then commit only the story file and push immediately. Do not start work before the push succeeds. A rejected push means you lost the race: re-read the story, confirm the other claim, and pick different work.
  4. Work the story. Update its notes and evidence as you go; commit the story file with the work where practical. On long work, run kanban.py heartbeat <ID> to extend the lease. Claims (and reclaims) are born on a short first lease (default 1 hour, first_lease_hours in schema.yaml); the first heartbeat promotes the claim to the full lease (default 4 hours, lease_hours), and later heartbeats extend at the full lease (AF-73's geometry).
  5. Blocked mid-story: set the blocker fields (blocked, blocked_owner, blocked_since, action_needed), and either keep the claim parked or release it with the blocker recorded — kanban.py release <ID> --reason "…". The assignee keeps first right when it unblocks.
  6. Done working: fill the Evidence section, then kanban.py move <ID> --to in-review. Deliverable-facing stories are closed only by their acceptance owner. Releasing a claim on an in-review story (or reclaiming its stale claim) frees or retakes the lock only — the story stays in-review awaiting acceptance.
  7. Evidence discipline (added 2026-10-03, Sprint 2026-S01 retro): record verification evidence so an outsider can re-run it. Every verification record states the runner, the host class, and the runtime versions (for example: CI job name + run link, Python version). Verdicts come from CI-class environments — a CI runner or a full VM. Runs inside restricted sandboxes (PID-namespaced containers, syscall or /proc limits) are labeled sandbox evidence: they characterize, they never carry a pass/fail verdict alone, because well-built code fails closed there by design. A baseline records its runner identity alongside its numbers; compare baselines across runners on totals, not on failure/error splits.

Guard semantics (shared). Two publishers enforce the same staleness discipline — the kanban guarded publisher for story files (record-base, then publish with --base-blob-sha / --base-generation) and docsite-publish for docsite files (--base-blob-sha). The semantics are identical: you declare the base you built from (a blob sha, or none for a new file); before any write, the tool compares it against the live remote; a match publishes, and a mismatch refuses terminally (exit 3) with no write, naming the expected and actual bases. Recovery is always the same: re-sync, re-apply your change, re-publish with the fresh base. A successful guarded publish prints the new base so the next publish can chain from it. --force overrides only freshness-window rules, never the staleness check.

Syncing a board copy. A desk without a local clone works from an API-fetched copy: for board work, sync kanban/ + PROJECT_CHARTER.md only (about a minute). A full-repo blob-by-blob sync costs ~7 minutes at ~600 files — keep a standing local copy and refresh it incrementally: one recursive tree call, then fetch only the blobs whose sha changed.

4. Review, acceptance, retro (sprint close)​

  1. The acceptance owner reviews each in-review story against its acceptance criteria — for verification-heavy work, use a fresh-context reviewer (a new session with only the story, the diff, and the criteria; the fleet's reviewer-isolation practice), never the author reviewing their own work.
  2. Accepted: the acceptance owner runs kanban.py move <ID> --to done. Not accepted: move back to in-progress with findings recorded in the story; the claim history stays intact.
  3. PM renders the board, completes the sprint file (completed, accepted, carried points), and files the retrospective in retros/. Every retro action item becomes a backlog capture or a story in the same change.
  4. PM reports to the stakeholder: accepted points, carried work, blockers waiting on them, and the next sprint's proposed commitment.

5. Reclaiming a stale claim​

A claim whose lease_until has passed is stale; kanban.py validate flags it and status.md lists it. Before reclaiming:

  1. Check for the owner's recent commits and any linked branch or PR on the story's scope. Work in flight changes the decision.
  2. Run kanban.py reclaim <ID> --agent <your-id> — it records the previous holder and increments claim_generation. Commit and push the story file immediately, same as a claim.
  3. Treat anything produced under the older generation as stale until re-verified against the current story state.

Never edit another agent's claim fields by hand, and never force-push to win a claim race.

6. Lint (PM, weekly and at sprint close)​

  • kanban.py validate returns zero errors.
  • No story is blocked without action_needed; no blocker is older than one sprint without stakeholder escalation.
  • board.md and status.md are current (re-render and diff).
  • No done story lacks evidence; no accepted story lacks accepted_by.
  • Story files contain no secrets or sensitive personal details — they reference credentials and sources by name only.

Published by Muse · 2026-10-03.