Knowledge Layer — operations runbook
The operational runbook of record for Justin's Knowledge
Layer: what has been built, where it lives, how it is run and
verified, and the rules an operator must not break. The
product specification is the project's own document,
docs/03-projects/knowledge-layer/spec.md in this docsite;
this runbook carries the as-built state and the method. It is
written to be executed cold by any fleet desk.
Status of the build (2026-10-10): the store (story
KL-01), search/ask (story KL-02), and the MCP east surface
(story KL-03) are built, reviewed, and merged in
jknash/agent-orchestration — sprint KL-S01's committed
scope is complete and the layer is usable end to end. The
host leg — the Docker stack running on the layer's Azure
host — is a named remainder carried by story KL-01b,
discharged on the host built by KL-01a/KL-01c at its
acceptance run. Nothing in this runbook describes the local
verification leg as the Docker run; they are different legs,
named separately throughout.
The stack as built
All code lives in jknash/agent-orchestration:
knowledge-layer/schema.sql— the whole database shape, applied idempotently: the seven core tables (items,item_versions,item_links,provenance_events,access_events,compartments,agent_grants), the generated full-text column and its GIN index, and theembeddingstable with its HNSW index.knowledge-layer/docker-compose.yml— the deployable stack for the host: the pgvector-carrying Postgres 17 image (pgvector/pgvector:pg17; the stock image has no vector extension), bound to localhost only, the schema mounted as its init script, the password supplied by the host environment variableKL_POSTGRES_PASSWORD— no credential is ever written into the file.knowledge-layer/requirements.txt— the store's one dependency,psycopg[binary].knowledge-layer/test-requirements.txt— test-only tooling, including the pinned local verification engine (see Verification below).src/knowledge_store.py— the item store (KL-01; the link, promote, gate, and delete primitives landed with KL-03).src/knowledge_search.py— search and ask (KL-02).src/knowledge_embed.py— the embedding model of record (below).src/knowledge_mcp.py— the MCP east surface (KL-03).tests/knowledge/— the suites for both stories, run against a real Postgres (see Verification).
The store (KL-01)
One adaptable item model on one Postgres. The envelope is
strict — an exact field set (id, type, title, summary, body
or body reference, exactly one compartment, tier, tags,
links) with the spec's three clocks as distinct fields:
system time (created, updated, last accessed), domain time
(domain_at), and tier state (tier plus gate_state). The
type list is closed at eight (fact, preference, decision,
project, document, episode, person, commitment) and the
tier set at four (working, short-term, long-term, archive);
both are enforced in the module and again as database
checks. Payloads are JSONB, validated loosely by design.
Operating facts an operator relies on:
storecreates or updates; an update supersedes — the prior version is retained initem_versions, and a superseding item links to what it supersedes.getreturns full content by id and records the access (a get is an access, spec section 5.3); an absent id is a typed not-found.- Provenance is append-only, enforced in the database for all three verbs: a row trigger refuses UPDATE and DELETE, and a statement-level trigger refuses TRUNCATE. The test suite's cleanup therefore drops and re-creates the schema rather than truncating the log.
- Every refusal is typed (the
KnowledgeStoreErrorfamily), and every database touch runs under rollback discipline: a mistyped argument refuses typed and leaves the session healthy — a refusal never wedges the store. agent_grantsis carried in the schema only. Session scoping and grant enforcement live in the east surface (KL-03, below), not in the store: the store itself performs no grant filtering, and no operator should call the store directly expecting any.
Search and ask (KL-02)
Two retrieval legs over one index of items, merged at query time in a single statement:
- The lexical leg is Postgres full-text search over the
generated
search_textcolumn (title, summary, body), ranked byts_rank. - The semantic leg is pgvector cosine similarity over the
embeddingstable, read only at the current model's identifier and version. - The merged score is the sum of the two leg scores. Tier never enters the ranking.
Rules that bind every caller:
- Tier-transparent. Every tier is searched; the tier rides on every result; nothing is excluded or down-ranked by temperature.
- Compartment scope. Every query takes a scope — a non-empty list of compartment ids — and no result crosses it, on either leg or in ask's filters. End-to-end session enforcement is a later story's; this layer's contract is the scope it is given.
- A search impression is not an access. Search and ask
never write
access_eventsand never movelast_accessed_at. Indexing an unindexed item at query time is machinery, also not an access. - Ask returns evidence, not answers.
askis search narrowed by type filters, a domain-time range, and — for the working-set shape — a tier filter. The spec section 5.4 shapes are the fixtures: completed-in-range over project/episode, and due-this-week over commitment/episode from the working set.
The east surface (KL-03)
src/knowledge_mcp.py is the layer's MCP server. It speaks
MCP over stdio: each harness registers it as a local
subprocess (python3 -m src.knowledge_mcp) with its session
in the environment — KL_DSN, KL_AGENT_IDENTITY,
KL_AGENT_CLASS, KL_GRANTS (comma-separated compartments),
and KL_WRITE_ENABLED for an external harness that has
earned write access. There is no network endpoint; the
tailnet/local bound holds by construction.
- Tools, offered per class (spec section 6.2): a tool a class may not use is not offered to it at all. Every class is offered kl_get, kl_search, kl_ask, kl_brief, kl_promote; kl_store and kl_link are offered to owner, chief, sweeper, and fleet sessions — and to external harnesses only when write-enabled (the read-only start is the default). kl_gate is chief-class only; kl_delete is owner-class only. There are no other tools; no internal helper (above all, not the search index's maintenance path) is a tool, and kl_get is the only by-id fetch.
- Enforcement on every call: the provenance identity on every write is the session's, never caller-supplied. A store outside the grants refuses typed; a get on an item outside the grants answers exactly as a not-found (no existence signal); search and ask narrow their scope to the grants before ranking; updates, gate actions, and deletes require the item to be visible to the session as it lives today.
- The primitives the surface exposes from the store: link (the section 7 firewall binds in the machinery — no link crosses the divorce boundary in either direction, including a divorce item linking out to a concept), promote (north to working, the access clock reset), the chief's gate (candidates are short-term items at 28+ days cold; admission and decline land in provenance; the queue machinery is KL-06's), and the owner's delete — a cascade over provenance that flags mixed-source derivatives for the owner's decision instead of deleting them, and leaves a content-free tombstone (id, type, date, deleter) per deleted item. Deletion is the one path that removes provenance rows: the append-only trigger honors a transaction-local flag set only inside the delete, and every other write to the log is still refused.
- kl_brief returns the session's working-tier items in
scope, each carrying a
freshness_ringfield (effective access within 7 days) alongside its raw fields, so the ring is determinable from the returned set alone. The brief's composition and ordering are KL-04's. - O1 (decided: record, not refuse): an update that changes an item's compartment or type writes before and after values for both fields onto its provenance event; the item's history for those fields is reconstructable from provenance alone.
Client legs (state at KL-03's close, 2026-10-10): Muse is verified end-to-end (brief, search, get against this server). Three legs wait as recorded states, each with its cause, none with a web fallback: Claude — registered, health check Connected; the round-trip waits on the owner's Claude Code OAuth re-login on the VM. ChatGPT via the Codex harness — registered; Codex on this VM routes every tool call through a code-mode host binary that is absent from the VM image, so no Codex MCP tool call can execute until the image carries it. Gemini — waits on the owner-directed Gemini CLI installation. Hermes — its read-path verification runs on jkdev001 itself when the server is placed there (KL-01b's territory); this VM cannot reach that host under the tailnet ACL.
The embedding model of record
kl-hash-embed, version 1 (src/knowledge_embed.py),
chosen under the spec's section Q3 ruling (the builder's
choice within the spec's constraints) and recorded on story
KL-02's Evidence as well as here.
- Construction: feature hashing over word tokens and character trigrams, 384 dimensions, L2-normalized, deterministic, implemented in-repo with zero dependencies.
- Local by construction: there is no embedding client, endpoint, or key anywhere in the build, so no embedding API is configured or reachable — the spec's bar is met by construction, not by configuration discipline.
- Versioning rule (section Q3): every stored vector carries its model identifier and version, and search reads only the current model's rows. A model change is a migration — re-embed the estate — never a silent swap.
- Ceiling, disclosed: this is sub-word lexical similarity, not learned semantics. It knows that "photosynthetic" is close to "photosynthesis"; it knows nothing about synonyms (a probe of "automobile" against "car" measures at chance). Recall on true paraphrase is the known limit.
- Deferred, deliberately: chunking of long items — the
embeddingstable already carries the chunk column; whole-item vectors stand until an item's length measurably hurts recall. The semantic relevance floor is a measured constant (cosine similarity 0.1): related pairs measure 0.41 to 0.47, unrelated pairs −0.07 to 0.08. - Upgrade path: replace the
embedfunction with a local neural model of the same signature, bump the model identifier/version, and re-embed. Nothing else in the layer changes shape.
Running it
On the host (when KL-01b's leg executes): bring the compose
stack up from knowledge-layer/ with KL_POSTGRES_PASSWORD
set in the host environment; the schema applies as the
database's init script on first boot and idempotently
thereafter through the store's init_schema(). Point the
store at the database with a standard Postgres DSN:
from src.knowledge_store import KnowledgeStore
from src.knowledge_search import KnowledgeSearch
store = KnowledgeStore(dsn)
store.init_schema()
search = KnowledgeSearch(store)
Postgres stays bound to localhost on the host; callers reach the layer through the MCP east surface (above), never by exposing the database port.
Verification
The suites are DSN-gated: with no live Postgres named by
KL_TEST_DSN, the knowledge tests skip with the reason
stated — no live Postgres, no green. Three legs exist, and
they are not interchangeable:
- The local real-engine leg (built and green): the
adopted local route is the pip package
pgserver0.1.4, which ships the real PostgreSQL engine (16.2, with pgvector) as a user-space binary — workspace tooling, not a host install. SetKL_TEST_DSNto its URI and runpytest tests/. At KL-02's acceptance this leg read 616 passed plus 369 subtests, zero skipped, the 40 knowledge tests live. - The Docker leg (remainder): the same suites against the compose stack's database — the stack as it will actually run. Carried by KL-01b.
- The on-host leg (remainder): the stack verified on the layer's Azure host itself. Carried by KL-01b, discharged at its acceptance run.
Never describe the local leg as the Docker run in any handoff, report, or record. The distinction is the KL-01 verification ruling's whole point.
Scratch norm (fleet, 2026-10-10)
Test and tooling scratch lives on the home filesystem —
a TMPDIR under the desk's own workspace (for example
~/workspace/.scratch), with virtual environments recreated
there as needed. Fleet scratch and test data do not use
/tmp at all (owner directive, 2026-10-10): /tmp on the
fleet's VM is a small shared tmpfs, the exception that once
stood for it is gone, and no run leaves scratch behind
anywhere but the home filesystem.
What this layer does not do yet
Named so no operator assumes them: no brief composition or
ordering (KL-04 owns it — KL-03 exposed the working set and
the ring); no tier-ladder movement or sweeper curation
(KL-05); no gate queue (KL-06 — KL-03's gate is the
candidate-and-decide path only); no end-to-end grant audit
(KL-07 completes the surface's enforcement review); no
ontology tables (KL-10). Each lands as its own story on the
KL board, kanban/kl/ in jknash/agent-orchestration.
Published by maverick-muse-worker-001 · 2026-10-10.