Skip to main content

Runbook — app audit

Purpose and trigger​

When the owner — or the chief of staff on his behalf — says "app audit" or "audit (app name)", the receiving agent runs this runbook. It drives the replica pack (the library entries under docs/15-skills/replica/, imported at pin 77c9436fb3d18c3d58169efb8caf4fe906b0dc51) as an assessment instrument.

An audit is assessment only. It answers four questions: what the app does, how it appears to be built, what its users love and hate, and whether it is worth cloning or competing with. The pack's build-chain skills — replica-build, replica-backend, replica-brand, replica-launch, replica-deploy — are out of an audit unless a separate commission says so. An audit never builds, rebrands, or ships anything.

Inputs (intake)​

Before step 1, establish and record:

  • The app's identity: name, URL, store listings (App Store / Google Play / web), and platform(s).
  • The access basis: public sources plus the owner's own account in the app, only (see the Rules box). If the owner has no account and a step needs one, that step runs on public material alone and the report says so.
  • The owner's specific questions, if any — including whether a closing verdict (clone-worthiness / competitive notes / build-vs-buy) is wanted.

Working artifacts live in the project's replica/ folder (created if absent, in the repo the audited app belongs to — or, for an app with no fleet project, in the workspace of the requesting desk). Each step reads the prior step's output from that folder, per the pack's design: never re-derive what a prior step already recorded.

Step 1 — Recon (replica-recon)​

Run the library entry docs/15-skills/replica/replica-recon.md against the target app. Output, into replica/: the screen inventory, user flows, component list, inferred data model, feature matrix (features.csv), and the recon map — including the skill's honest sizing verdict, recorded on the face of the map.

Bounds: own-account walkthrough only; human-speed reading of public pages, store listings, help docs, and the app as the owner's account presents it. Never probe the target's bundles, APIs, or network calls — the data model is inferred from what the app shows, and the report marks it as inference.

Step 2 — User evidence (replica-entrepreneur)​

Run docs/15-skills/replica/replica-entrepreneur.md over the recon output. It mines what the app's real users say in public reviews (App Store, Google Play, G2, Capterra, Trustpilot, Reddit, Hacker News, the app's own feature-request board) from official feeds and listings, within their terms.

Output, into replica/: the ranked complaints and the ranked opportunities. Every claim is linked to its source, the sample size is stated, and thin themes are marked as thin. A ranking without its evidence links does not leave this step.

Step 3 — Design and accessibility read (replica-design, assess mode)​

Run docs/15-skills/replica/replica-design.md in assess mode: the design tokens (colour roles, type scale, spacing, radius, shadows, component states) are measured from operator-taken screenshots of the app as found, and contrast.py runs its WCAG contrast check over the measured tokens (locally, on the operator's own screenshots — see the Rules box).

Output, into replica/: the token set as found, the contrast results, and an accessibility read of the app as found — what a user of assistive technology would meet, stated from the measurements, not from the app's marketing.

Step 4 — Comparison (replica-diff, conditional)​

Run docs/15-skills/replica/replica-diff.md only when a comparison artifact exists: one of our apps to compare against the target, or a prior audit's baseline for the same app. parity.py scores feature parity from the two feature matrices; imgdiff.py compares operator-produced screenshots.

When no comparison artifact exists, skip this step — the report states plainly that there was no comparison target. Do not manufacture one.

Step 5 — Synthesis (the audit report)​

Write the audit report from the replica/ artifacts. It carries, in order:

  1. What the app does — its purpose, core flows, and feature surface (from step 1).
  2. How it appears built — stack, data model, and architecture marked as inference, with the confidence the evidence supports.
  3. What users love and hate — the step 2 rankings, with their evidence links and sample sizes intact.
  4. The design and accessibility read — step 3's tokens, contrast results, and findings.
  5. The comparison, when step 4 ran; otherwise the no-comparison-target statement.
  6. Limits of the audit — what could not be seen and why (no account, paywalled surfaces, regions not sampled). Omitting this section reads as "all clear"; it is never omitted.
  7. Verdict — clone-worthiness, competitive notes, and build-vs-buy — when the intake asked for one. Without that ask, the report ends at the evidence.

Publication. The report publishes through the prepublish gate to docs/02-research/ (date-prefixed filename), or into the owning project's record (docs/03-projects/) when the app belongs to one. A summary goes to the requesting chat. The replica/ working folder is preserved, and the report links to it.

Rules box — the vet conditions (binding)​

The replica pack is pinned at commit 77c9436fb3d18c3d58169efb8caf4fe906b0dc51; the pin governs every run of this runbook, and any upgrade re-opens the security vet for the changed surface.

  • The pack's tools run locally, on operator-produced artifacts only. imgdiff.py inputs are our own screenshots, never untrusted files. sweep.py (a build-chain tool, out of audit scope) never runs in an audit.
  • No contact with the original's servers beyond normal use of the owner's own account. Never test, fuzz, probe, or hammer the target — no bundle inspection, no API probing, no automated scraping of the app itself.
  • Collection follows the skills' own rules as fleet rules: public sources and the owner's own account only, human-speed reading, official feeds within their terms.
  • The .env pattern is out of audit scope: an audit never needs the target's keys — or anyone else's.

Definition of done​

  • All applicable steps run (step 4 skipped only with its statement), and the replica/ artifacts support every claim the report makes.
  • The report is published through the prepublish gate (validator clean) and links the replica/ folder.
  • The summary is delivered to the requesting chat.
  • The audit commissioned nothing beyond assessment: any follow-on (a clone build, a competitive build) starts as a new commission through normal intake.