Skip to main content

SEO Audit

The fleet's SEO audit method: a standing, local-only audit of a website against technical-SEO best practice, run on a cadence and reported diff-first. The method was commissioned for x-centric.com (Framer-hosted, 319 sitemap URLs) on the research report of record (docs/02-research/2026-10-08-seo-skills-and-tooling.md); the stack generalizes to any fleet surface by changing the target configuration only.

The library previously held zero SEO skills. This skill is fleet-authored: it borrows its check taxonomy openly from the claude-seo and marketingskills seo-audit frameworks and encodes it against three local tools the fleet vets and pins. No hosted SEO service, API, or validator is used at any step — fleet URLs are never submitted to third parties.

The stack (pinned; upgrades re-open the security vet)​

Installs live in the SEO goal workspace (~/workspace/goals/seo-agent-for-x-centric-com/tools/), nothing system-wide. Each tool was security-vetted at its pinned version (vet memo of record, 2026-10-09); a version upgrade of any of the three re-opens the vet for the changed surface and is never auto-applied.

  1. SiteOne Crawler v2.6.1 (single Rust binary, built from the release tag; sha256 recorded at install). Primary crawl: site-wide technical SEO, metadata, Open Graph, and link structure, emitted as JSON.
  2. Unlighthouse 0.19.1 (Node, lockfile-pinned). Lighthouse SEO category + Core Web Vitals in the fleet Chromium: a fixed sample weekly, the full sitemap monthly.
  3. extruct 0.18.0 (Python, pinned dependencies). Structured-data extraction (JSON-LD, microdata, RDFa, OpenGraph) over HTML the crawl legs already fetched — library API only.

Run rules (the vet conditions, restated as hard constraints)​

  • SiteOne AI analysis stays OFF. The run configuration pins ai_enabled=false explicitly (the tool pre-populates AI actions, so the flag is the single switch — never rely on absence), and no AI key, endpoint, or provider setting is present. If AI analysis is ever wanted, the endpoint must be a local OpenAI-compatible server, and enabling it is a new decision, not a config flip.
  • Raw-HTML mode. Browser rendering stays disabled (browser_enabled=false). The target's content lives in raw HTML; rendering buys nothing for the crawl leg.
  • robots.txt is enforced (ignore_robots_txt=false). Never crawl against a target's robots directives.
  • No mailer, no report server. The SMTP exporter is unconfigured; reports are local files. One-shot local viewing bound to 127.0.0.1 at most.
  • Unlighthouse never downloads a browser. The audit configuration pins both puppeteerOptions.executablePath to the fleet Chromium (/opt/meta-chromium/chrome) and chrome.useDownloadFallback=false — a resolution failure errors, it never downloads. Installs are registry-only with PUPPETEER_SKIP_DOWNLOAD=1.
  • CI provider only (unlighthouse-ci): static filesystem output, no interactive server, and no update-check call.
  • Filesystem output only, everywhere. LHCI temporary-public-storage is never a target, under any configuration, ever. No report leaves the fleet's machines.
  • Hosted validators are not used. Google's Rich Results Test and validator.schema.org receive no fleet URLs; structured-data validation is the local vocabulary check below.
  • extruct's CLI fetcher is never used — the library parses HTML already fetched by the crawl legs.

Cadence​

  • Weekly: full SiteOne crawl (the target's URL count makes a full crawl minutes of work) + Unlighthouse on the fixed sample (home, one location page, one service page, the blog index, top glossary entries).
  • Monthly: Unlighthouse across the full sitemap; the monthly run becomes the pinned baseline (below).
  • On deploy: an extra crawl + sample run when the target site deploys or restructures. Framer deploys are outside fleet CI, so this trigger is a standing instruction from whoever learns of the deploy — it is not a webhook.

Procedure​

  1. Run SiteOne with the pinned run configuration; keep the JSON output as the run's substrate.
  2. Run Unlighthouse (CI provider) on the sample or the full set, per the cadence; keep the static output with the run.
  3. Feed the fetched HTML through extruct for every crawled URL; record extracted structured data per URL (today's expected result on x-centric.com: none found — the check fires until Organization/LocalBusiness and BlogPosting schema exist).
  4. Run the standing checks (below) over the three legs' outputs.
  5. Diff against the previous run and against the pinned monthly baseline; write the report diff-first (below).
  6. Store the run's raw outputs with the run; never edit a stored run — a correction is a new run.

Standing checks​

  • Technical SEO: status codes, redirect chains (including known multi-hop chains), canonicals (present, self-referencing, matching the sitemap), crawl traps, indexability.
  • Metadata and Open Graph: title, meta description, canonical, OG/Twitter tags on every crawled URL — present, unique, non-empty.
  • Structured data: extruct extraction across the full URL set, then local validation: JSON syntax, schema.org vocabulary, and required properties per type (Organization/LocalBusiness: name, url, address where applicable; BlogPosting: headline, datePublished, author). Presence without validity is a finding; absence on a business site is the headline finding until fixed.
  • Sitemap and robots hygiene: the sitemap parses; every sitemap URL returns 200 and matches its canonical; sitemap↔crawl set difference in both directions (orphans either way); robots.txt parsed with RFC 9309 semantics (longest match, not stdlib first-match). Missing lastmod values are a standing hygiene note, not a per-run finding.
  • Internal linking: broken internal links, orphan pages, click depth from the crawl graph.
  • Core Web Vitals and page weight: Lighthouse performance + SEO categories on the sample (lab data), page weight tracked as a named metric run over run.
  • Mobile: Lighthouse mobile emulation on the sample (viewport, tap targets, legible font sizes — its SEO category carries these).

Out of scope by design: accessibility (the fleet's AccessLint method owns it; cross-reference intersecting findings such as non-descriptive anchor text) and content quality / search intent (an agent judgment under this framework, never a tool verdict).

Reporting: diff-first​

Every report opens with a one-line health summary (finding counts by severity). Findings follow in three sections — new, fixed since last run, standing — each finding carrying the check name, the URL, and the evidence (the SiteOne JSON record, the Lighthouse audit id, or the extracted snippet). Scores (Lighthouse SEO category, SiteOne quality scores) appear as context, never as the verdict, and never bare. The first run of a target is its baseline run: reported in full, with no diff sections.

Data-blocked items (recorded, never faked)​

  • Backlinks: known only to hosted indexes; blocked under the local-only rule unless the owner approves a named source.
  • Search Console / GA4: behind the owner's own Google account — an owner decision plus a credential step, not a tooling gap.
  • Log-file analysis: no data source exists for a Framer-hosted site (no raw log export).
  • Rank tracking / field CWV (CrUX): hosted data by construction; lab data is the local substitute.

Provenance​

Fleet-authored 2026-10-09 by the sweeper desk (maverick-muse-sweeper-001) on the chief of staff's dispatch, from the research report of record (docs/02-research/2026-10-08-seo-skills-and-tooling.md) and the security vet memo of the same date. Check taxonomy borrowed from AgriciDaniel/claude-seo and coreyhaines31/marketingskills (seo-audit), both MIT, neither imported.