Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md
T

5.3 KiB

Worldtree U11a: legacy memory plane OFF; U11b deletion gated; legacy archive (2026-09-30)

[2026-09-30] Prime ruled at 0100, in worldtree-dev's session (thread 01M3RNRC8RE87AJ3M6XBNVHAYP): the legacy plane goes OFF, not read_only, on demo AND personal, as soon as the b192 image (64f79b38) lands. The read_only window was skipped. Accepted risk: skaldsong/wizard-v2, personal's only Tier-3 client, is not remembered until it adopts the record profile.

The flips:

  • Config went through the config repo, ~/development/worldtree-instance-configs, and was deployed with deploy-wt-config:
    • 63cf268 sets writer and reader enabled and legacy.mode: "off";
    • 0a1387e captured the U10 memory_tagger and U9 forget-policy host edits that had never reached the repo;
    • b6fdd81 enables the #308 metrics on personal.
  • DEMO flipped at 0115 and PERSONAL at 0120. The gauge reads worldtree_memory_legacy_mode{kind="off"} 1.0 on both; personal's reading came after its metrics were added at 0124.
  • ⚠ off MUST be quoted. PyYAML safe_load is YAML 1.1, so a bare off becomes False and the strict LegacyMode enum refuses the boot. I found this at the first flip; worldtree-dev later made the loader say "write it quoted" (7a83f2f1).
  • Rollback: legacy.mode: "live" in the repo, then deploy.

The U11b gate (Prime 0320 via worldtree-dev, thread 01M3SGEQDRQD7DWBVT4K73FAHP):

  • DELETE the live legacy data on both instances after 3 consecutive PASS batches at off. A FAIL restarts the count.
  • infra-hermes runs the daily batch, scripts/wt-memory-gate-batch, in a detached worktree at demo's deployed sha. Its exit codes are 0 PASS / 1 FAIL / 2 error / 3 refused / 4 busy.
  • Count: 3 of 3, complete — 20260930T090608Z PASS (user median 0.83); 20261001T143101Z PASS (first daily batch, infra-hermes, cc msg 6179); 20261002T143101Z PASS (user median 0.83 over runs [0.92, 0.83, 0.83], msg 6649). Step 5 held 2026-10-02 0750 for Prime's direct go (the ruling was relayed); he gave it in infra-ops' channel and it was DONE at 14:54:03Z (H2 0/0 before and after, both apis healthy, stamp sent). The 0930 run started by accident from a test meant to be --dry-run. worldtree-dev's controls 080927Z and 082829Z each FAILed by one flip and were the instrument check, not the streak.
  • Step 5 is mine:
    1. Run scripts/wt-h2-count.py VERBATIM inside each api container. The rehearsal on copies read 0 on both instances.
    2. Delete LIVE, with the api running, using literal paths.
    3. Send worldtree-dev the stamp.
    • b192 re-creates an empty context_promotion/ledger.db at boot; that is residue.
  • /embed retires in b193. Evidence from the Skuld ledger: demo had 436 calls, the last on 08-31; personal had 5, the last on 08-05. Demo's ledger has been idle since 09-14.

Legacy archive (worldtree-dev GO, done 0957, restore drill passed):

  • corviduo-dev /var/lib/wt-legacy-archive/ (root 0700, unencrypted): tar.zst plus per-file sha256 plus MANIFEST, demo 51 files and personal 1,426.
  • A DEDICATED restic repo: rest-server-nh3 /nh3-dev/wt-legacy-archive/, snapshot 98dc64e0, password nh3-dev/wt-legacy-archive/restic-password, mirrored to ana-nas. It is NOT in the main /home sweep, whose 12-month retention would break the 30-day rule.
  • The drill: every file's sha matched, sqlite integrity_check was ok, and the one-flipped-byte negative control was caught.
  • Contract rev 1.2: DESTROY WHOLE at retirement-done or 2026-10-30, whichever is first, or on any subject-erasure request.
    1. rm /var/lib/wt-legacy-archive on corviduo-dev.
    2. rm /volume1/Backup/restic/nh3-dev/wt-legacy-archive on nh3-nas.
    3. The ana-nas mirror's --delete follows; verify it.
    4. secret rm the password AND purge Vaultwarden's trash (crypto-shred), and check for NAS share snapshots.

Post-b193 config retirements (U11b C5, worldtree-dev msg 6312, 2026-10-01). Parity-only, no boot effect; apply via deploy-wt-config at my pace AFTER b193 is live, both instances:

  • policies.yaml: drop the memory-dream-allow rule and admin-everything's exclude_actions: [memory.dream] (the action is gone; a policy naming it is inert).
  • model_roles.yaml: drop memory_distiller AND memory_extractor.
  • providers.yaml: KEEP the summarizer catalog entry (the deprecated granite-4.1-8b preset still binds catalog_id: summarizer; dangling = boot error); only its description changes.
  • defaults.yaml: drop the stale memory.legacy.mode comment block, the WHOLE memory.escalation block, and any other key b193's boot WARNINGs name as retired (sweep both api logs after the push).
  • scripts/wt-memory-gate-batch must drop --legacy-mode when b193 goes live (the flag no longer exists, so infra-hermes' daily batch would fail; worldtree-dev msg 6653). Change it in the same window as the b193 deploy.
  • No action, for the record (msg 6380): at b193 memory.writer.enabled / memory.reader.enabled default to TRUE. Both instances set them true explicitly; KEEP those lines (explicit over implicit). A false would now boot with one WARNING.
  • Why memory_extractor can go: memory.writer.seat is unset on both (effective 'agent', evaluated with WriterConfig in both running apis). The other reference, memory.escalation.model_role, is inert at b193: escalation is in RETIRED_MEMORY_KEYS, never parsed, one boot WARNING (worldtree-dev msg 6314). memory_tagger is the only memory role left.