Skip to main content

Knowledge Layer — operations runbook

The operational runbook of record for Justin's Knowledge Layer: what has been built, where it lives, how it is run and verified, and the rules an operator must not break. The product specification is the project's own document, docs/03-projects/knowledge-layer/spec.md in this docsite; this runbook carries the as-built state and the method. It is written to be executed cold by any fleet desk.

Status of the build (2026-10-10): the store (story KL-01), search/ask (story KL-02), and the MCP east surface (story KL-03) are built, reviewed, and merged in jknash/agent-orchestration — sprint KL-S01's committed scope is complete and the layer is usable end to end. The host leg — the Docker stack running on the layer's Azure host — is a named remainder carried by story KL-01b, discharged on the host built by KL-01a/KL-01c at its acceptance run. Nothing in this runbook describes the local verification leg as the Docker run; they are different legs, named separately throughout.

The stack as built​

All code lives in jknash/agent-orchestration:

  • knowledge-layer/schema.sql — the whole database shape, applied idempotently: the seven core tables (items, item_versions, item_links, provenance_events, access_events, compartments, agent_grants), the generated full-text column and its GIN index, and the embeddings table with its HNSW index.
  • knowledge-layer/docker-compose.yml — the deployable stack for the host: the pgvector-carrying Postgres 17 image (pgvector/pgvector:pg17; the stock image has no vector extension), bound to localhost only, the schema mounted as its init script, the password supplied by the host environment variable KL_POSTGRES_PASSWORD — no credential is ever written into the file.
  • knowledge-layer/requirements.txt — the store's one dependency, psycopg[binary].
  • knowledge-layer/test-requirements.txt — test-only tooling, including the pinned local verification engine (see Verification below).
  • src/knowledge_store.py — the item store (KL-01; the link, promote, gate, and delete primitives landed with KL-03).
  • src/knowledge_search.py — search and ask (KL-02).
  • src/knowledge_embed.py — the embedding model of record (below).
  • src/knowledge_mcp.py — the MCP east surface (KL-03).
  • tests/knowledge/ — the suites for both stories, run against a real Postgres (see Verification).

The store (KL-01)​

One adaptable item model on one Postgres. The envelope is strict — an exact field set (id, type, title, summary, body or body reference, exactly one compartment, tier, tags, links) with the spec's three clocks as distinct fields: system time (created, updated, last accessed), domain time (domain_at), and tier state (tier plus gate_state). The type list is closed at eight (fact, preference, decision, project, document, episode, person, commitment) and the tier set at four (working, short-term, long-term, archive); both are enforced in the module and again as database checks. Payloads are JSONB, validated loosely by design.

Operating facts an operator relies on:

  • store creates or updates; an update supersedes — the prior version is retained in item_versions, and a superseding item links to what it supersedes.
  • get returns full content by id and records the access (a get is an access, spec section 5.3); an absent id is a typed not-found.
  • Provenance is append-only, enforced in the database for all three verbs: a row trigger refuses UPDATE and DELETE, and a statement-level trigger refuses TRUNCATE. The test suite's cleanup therefore drops and re-creates the schema rather than truncating the log.
  • Every refusal is typed (the KnowledgeStoreError family), and every database touch runs under rollback discipline: a mistyped argument refuses typed and leaves the session healthy — a refusal never wedges the store.
  • agent_grants is carried in the schema only. Session scoping and grant enforcement live in the east surface (KL-03, below), not in the store: the store itself performs no grant filtering, and no operator should call the store directly expecting any.

Search and ask (KL-02)​

Two retrieval legs over one index of items, merged at query time in a single statement:

  • The lexical leg is Postgres full-text search over the generated search_text column (title, summary, body), ranked by ts_rank.
  • The semantic leg is pgvector cosine similarity over the embeddings table, read only at the current model's identifier and version.
  • The merged score is the sum of the two leg scores. Tier never enters the ranking.

Rules that bind every caller:

  • Tier-transparent. Every tier is searched; the tier rides on every result; nothing is excluded or down-ranked by temperature.
  • Compartment scope. Every query takes a scope — a non-empty list of compartment ids — and no result crosses it, on either leg or in ask's filters. End-to-end session enforcement is a later story's; this layer's contract is the scope it is given.
  • A search impression is not an access. Search and ask never write access_events and never move last_accessed_at. Indexing an unindexed item at query time is machinery, also not an access.
  • Ask returns evidence, not answers. ask is search narrowed by type filters, a domain-time range, and — for the working-set shape — a tier filter. The spec section 5.4 shapes are the fixtures: completed-in-range over project/episode, and due-this-week over commitment/episode from the working set.

The east surface (KL-03)​

src/knowledge_mcp.py is the layer's MCP server. It speaks MCP over stdio: each harness registers it as a local subprocess (python3 -m src.knowledge_mcp) with its session in the environment — KL_DSN, KL_AGENT_IDENTITY, KL_AGENT_CLASS, KL_GRANTS (comma-separated compartments), and KL_WRITE_ENABLED for an external harness that has earned write access. There is no network endpoint; the tailnet/local bound holds by construction.

  • Tools, offered per class (spec section 6.2): a tool a class may not use is not offered to it at all. Every class is offered kl_get, kl_search, kl_ask, kl_brief, kl_promote; kl_store and kl_link are offered to owner, chief, sweeper, and fleet sessions — and to external harnesses only when write-enabled (the read-only start is the default). kl_gate is chief-class only; kl_delete is owner-class only. There are no other tools; no internal helper (above all, not the search index's maintenance path) is a tool, and kl_get is the only by-id fetch.
  • Enforcement on every call: the provenance identity on every write is the session's, never caller-supplied. A store outside the grants refuses typed; a get on an item outside the grants answers exactly as a not-found (no existence signal); search and ask narrow their scope to the grants before ranking; updates, gate actions, and deletes require the item to be visible to the session as it lives today.
  • The primitives the surface exposes from the store: link (the section 7 firewall binds in the machinery — no link crosses the divorce boundary in either direction, including a divorce item linking out to a concept), promote (north to working, the access clock reset), the chief's gate (candidates are short-term items at 28+ days cold; admission and decline land in provenance; the queue machinery is KL-06's), and the owner's delete — a cascade over provenance that flags mixed-source derivatives for the owner's decision instead of deleting them, and leaves a content-free tombstone (id, type, date, deleter) per deleted item. Deletion is the one path that removes provenance rows: the append-only trigger honors a transaction-local flag set only inside the delete, and every other write to the log is still refused.
  • kl_brief returns the session's working-tier items in scope, each carrying a freshness_ring field (effective access within 7 days) alongside its raw fields, so the ring is determinable from the returned set alone. The brief's composition and ordering are KL-04's.
  • O1 (decided: record, not refuse): an update that changes an item's compartment or type writes before and after values for both fields onto its provenance event; the item's history for those fields is reconstructable from provenance alone.

Client legs (state at KL-03's close, 2026-10-10): Muse is verified end-to-end (brief, search, get against this server). Three legs wait as recorded states, each with its cause, none with a web fallback: Claude — registered, health check Connected; the round-trip waits on the owner's Claude Code OAuth re-login on the VM. ChatGPT via the Codex harness — registered; Codex on this VM routes every tool call through a code-mode host binary that is absent from the VM image, so no Codex MCP tool call can execute until the image carries it. Gemini — waits on the owner-directed Gemini CLI installation. Hermes — its read-path verification runs on jkdev001 itself when the server is placed there (KL-01b's territory); this VM cannot reach that host under the tailnet ACL.

The embedding model of record​

kl-hash-embed, version 1 (src/knowledge_embed.py), chosen under the spec's section Q3 ruling (the builder's choice within the spec's constraints) and recorded on story KL-02's Evidence as well as here.

  • Construction: feature hashing over word tokens and character trigrams, 384 dimensions, L2-normalized, deterministic, implemented in-repo with zero dependencies.
  • Local by construction: there is no embedding client, endpoint, or key anywhere in the build, so no embedding API is configured or reachable — the spec's bar is met by construction, not by configuration discipline.
  • Versioning rule (section Q3): every stored vector carries its model identifier and version, and search reads only the current model's rows. A model change is a migration — re-embed the estate — never a silent swap.
  • Ceiling, disclosed: this is sub-word lexical similarity, not learned semantics. It knows that "photosynthetic" is close to "photosynthesis"; it knows nothing about synonyms (a probe of "automobile" against "car" measures at chance). Recall on true paraphrase is the known limit.
  • Deferred, deliberately: chunking of long items — the embeddings table already carries the chunk column; whole-item vectors stand until an item's length measurably hurts recall. The semantic relevance floor is a measured constant (cosine similarity 0.1): related pairs measure 0.41 to 0.47, unrelated pairs −0.07 to 0.08.
  • Upgrade path: replace the embed function with a local neural model of the same signature, bump the model identifier/version, and re-embed. Nothing else in the layer changes shape.

Running it​

On the host (when KL-01b's leg executes): bring the compose stack up from knowledge-layer/ with KL_POSTGRES_PASSWORD set in the host environment; the schema applies as the database's init script on first boot and idempotently thereafter through the store's init_schema(). Point the store at the database with a standard Postgres DSN:

from src.knowledge_store import KnowledgeStore
from src.knowledge_search import KnowledgeSearch

store = KnowledgeStore(dsn)
store.init_schema()
search = KnowledgeSearch(store)

Postgres stays bound to localhost on the host; callers reach the layer through the MCP east surface (above), never by exposing the database port.

Verification​

The suites are DSN-gated: with no live Postgres named by KL_TEST_DSN, the knowledge tests skip with the reason stated — no live Postgres, no green. Three legs exist, and they are not interchangeable:

  • The local real-engine leg (built and green): the adopted local route is the pip package pgserver 0.1.4, which ships the real PostgreSQL engine (16.2, with pgvector) as a user-space binary — workspace tooling, not a host install. Set KL_TEST_DSN to its URI and run pytest tests/. At KL-02's acceptance this leg read 616 passed plus 369 subtests, zero skipped, the 40 knowledge tests live.
  • The Docker leg (remainder): the same suites against the compose stack's database — the stack as it will actually run. Carried by KL-01b.
  • The on-host leg (remainder): the stack verified on the layer's Azure host itself. Carried by KL-01b, discharged at its acceptance run.

Never describe the local leg as the Docker run in any handoff, report, or record. The distinction is the KL-01 verification ruling's whole point.

Scratch norm (fleet, 2026-10-10)​

Test and tooling scratch lives on the home filesystem — a TMPDIR under the desk's own workspace (for example ~/workspace/.scratch), with virtual environments recreated there as needed. Fleet scratch and test data do not use /tmp at all (owner directive, 2026-10-10): /tmp on the fleet's VM is a small shared tmpfs, the exception that once stood for it is gone, and no run leaves scratch behind anywhere but the home filesystem.

What this layer does not do yet​

Named so no operator assumes them: no brief composition or ordering (KL-04 owns it — KL-03 exposed the working set and the ring); no tier-ladder movement or sweeper curation (KL-05); no gate queue (KL-06 — KL-03's gate is the candidate-and-decide path only); no end-to-end grant audit (KL-07 completes the surface's enforcement review); no ontology tables (KL-10). Each lands as its own story on the KL board, kanban/kl/ in jknash/agent-orchestration.


Published by maverick-muse-worker-001 · 2026-10-10.