Skip to main content

Removing a live single-model gate

Use when one model blocks a fleet. Remove the gate. Companion to the design rule in role-based-model-binding (bind roles, not models). That skill says how to BUILD role-based routing; this one is how to DISMANTLE a single-model gate that is already deployed and blocking work — and make the removal stick.

Owner rule this serves: "I don't want a single model blocking anything." A review/gate pinned to one literal provider/model with fallback_allowed:false turns one 429 or one REQUEST_CHANGES into a fleet-wide freeze. When you see a blocker like *_direct_sol_policy_not_accepted or any "fresh literal <one model>, fallback forbidden" requirement, that is removable policy, not a real blocker — say so and remove it.

The gate lives on FOUR coupled surfaces​

Fixing only one leaves the others to re-assert it. Fix them together, in this order:

  1. Config binding. Turn the role's single {provider, model, fallback_allowed:false} into an approved_routes LIST + preferred_order + fallback_allowed:true. fallback_allowed:true here means "fall over to the next ALLOWLISTED first-party route on 429/unavailable" — never an unbound catalog default. Back up config first (config.json.bak-<UTCstamp>).

  2. Code enforcement points. The gate is usually pinned in 2–3 places, typically a resolver (_load_review_routes-style) that RAISES when fallback_allowed is true, and a receipt validator that requires provider+model to equal the ONE bound route. Change: resolver accepts an allowlist; validator accepts a receipt whose provider/model matches ANY allowlist entry. An EMPTY allowlist must STILL fail closed (a role can never run unbound). Keep every OTHER gate byte-for-byte: exact-head fencing, exactly-one parsed VERDICT, evidence-ref, candidate-SHA match, independent ephemeral reviewer, no auto-approve without a real review. Grep the codebase for the literal model string and the fallback_allowed field to find all enforcement sites before editing.

  3. Stored blocker state. The blocker text is often PERSISTED in delivery/ledger records (a blocker_class/blocker_reason on each record), NOT re-derived each tick. Proof: grep the codebase for the exact blocker string — if it is absent from the code, it is stored state. Restarting the service will NOT clear it. Clear it through the supported transition path (e.g. integration.py delivery-transition <request.json>), resuming each blocked record to its resume_phase. In a well-built contract, transitioning a blocked record to its source/resume phase clears blocker_class/resume_phase to null automatically (verify by reading the returned record). Read each record's LIVE revision + owner_generation + candidate_sha first; the transition is a compare-and-swap and will StaleTransition if you use a stale revision.

  4. Cron / manager prompts. Autonomous crons re-assert the old policy from their OWN prompt text on the next tick, undoing your work in minutes. Grep each relevant cron prompt for the model literal / "fallback forbidden" / "direct-SOL". Re-pin every manager that reads the policy via cronjob_manage action=update with a surgically amended prompt (preserve the rest of the contract), AND write an owner-authority record (activation addendum + a superseding action-intent that names the new route_policy) that the managers consume. Note which managers are paused vs enabled — only enabled ones can re-block.

Decouple genuine safety fixes from the routing gate​

A fused blocker like auxiliary_launcher_AND_direct_sol_policy_not_accepted bundles a real engineering fix (e.g. a launcher fail-open bug) with the single-model policy. Split them: the safety fix may still gate the component's INSTALL, but must NOT gate product-candidate REVIEW routing. Candidates review now through the allowlist; they do not wait for the launcher to ship. State this decoupling explicitly in the authority record.

Updating tests that encode the old gate​

Tests asserting provider=='X' and model=='Y', or that fallback_allowed:true must RAISE, ARE the single-model gate in test form — update them to the allowlist policy; they are not regressions to defend. Assertions to keep: the previously-pinned model is STILL an accepted route (assertIn), the allowlist has >1 member (no single-model block), an unlisted model is rejected, an empty allowlist fails closed. Distinguish these from genuinely unrelated pre-existing failures: run the FULL suite, and for any leftover failure grep the failing test for reviewer|ROUTES|role_binding|approved_route — zero hits means the routing change cannot be the cause. Confirm the same failures exist against the pristine config backup. NEVER weaken a real security/isolation assertion (env-allowlist floor, credential scrub) to make a routing change pass.

Verify and read back​

  • Config loads through the controller's own loader (not just json.load).
  • Resolver returns the allowlist; a receipt from EACH allowlisted model validates; an unlisted model is rejected.
  • Restart the service (systemd user unit: systemctl --user restart <unit>; confirm new PID active).
  • After transitions, re-read each record: phase advanced, blocker_class == null.
  • The re-pinned crons carry the new policy (read the job back).

Overlap note​

This is the operational sibling of role-based-model-binding and touches the same territory as multi-agent-delivery-verification (reviewer routing) and assessor-lane-controller (the concrete controller). Those three are user-owned; if they get adopted for curation, the removal procedure here should be merged in rather than kept separate.


Source: jknash/hermes-shared-skills · branch hermes-jkdev001 @ 1d0d545c3970 · skills/engineering/removing-single-model-gates/ · view source · Imported 2026-10-04. Supporting files (references, scripts) remain in the source repository.

Published by Muse · 2026-10-04.