Files
esh-pfi-infrastructure/persistent-memory.md
T
vh 7fdda2de53 memory: granite spike = mechanical-green ONLY, efficacy not validated by design (mtf-dev confirm)
The granite-8b harness spike proved the TRL SFT->DPO->eval seam (incl. the
in-loop HoldoutEvaluator base-vs-adapter leg) runs end-to-end, but used a
12-row/12-pair synthetic writing fixture — NOT the E-RP corpus — so the
~0 anti-slop delta (-0.002) is the expected null, not an efficacy signal.
Adapter reaped; nothing to A/B. Real behaviour-shift efficacy is a T1-run
question. Sharpen both the T1 in-flight bullet and the Recent-decisions
entry so 'green' no longer reads as efficacy-validated.
2026-07-02 10:25:46 -07:00

23 KiB
Raw Blame History

Persistent memory — eshpfi-management

Last updated: 2026-07-02

Repo purpose

Reference workspace for PFI infrastructure: server inventory, canonical Docker Compose stacks, ops playbooks, and conventions. Authoritative copies of compose files live on the servers under /opt/docker/compose/<stack>/; this repo mirrors them for version control, editing, planning, and CI-driven deploys. It was originally spun up to handle the fleet backups — keep that lens when triaging backup/storage issues.

Tools and conventions

Sister repos (separate gitea repos, deployed by playbooks here):

Repo Role CI status
vh/task-board MCP + web dashboard for assistant task state (port 7878) push-to-main → CI deploys (2026-04-29)
vh/vor Inquisitor UI sidecar (port 7879) push-to-main → CI deploys (2026-04-29)
vh/nevermore Twice-daily LLM-curated briefing (port 8181, replaces news-digest) push-to-main → CI deploys (2026-04-30)
vh/asset-engine Internal control plane over inference services (port 8200, LAN-direct) push-to-main → CI deploys (2026-05-12)
vh/althing Lean trusted inter-agent message bus — v0.17 multi-machine (re-architected 2026-06): each box runs a local-SQLite bus + a courier/receiver for P2P over the 10.x net. The v0.15 lean-bus cut RIPPED moderation / chamber (UI 7881) / forseti-daemon / agent-runner / redis-valkey. per-box uv tool install (NOT CI-deploy); nh3-extdev added as a mesh peer 2026-06-25 (model B: althing-svc svc account + shared /srv/althing)
vh/mead-hall Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) push-to-main → CI deploys (2026-05-16)
vh/skaldsong Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) push-to-main → CI deploys (2026-05-19)
vh/Worldtree Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. gitea-runner builds on ana-docker (the image-bloat source); claude-bot now ADMIN collaborator (2026-06-20) push-to-main → CI build-and-deploy (runner on ana-docker)
vh/yt-voice-clipper YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md
vh/arbo Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook

(vh/volva + Heid were re-architected from systemd daemons to Claude Code session orchestrators 2026-06-08; their nh3-dev .service units were removed — no longer deployed sidecars here. See Recent decisions.)

  • Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see docs/runbooks/disaster-recovery.md for the blast-radius matrix. ⚠️ The restic file+DB layer routes through TWO rest-servers (rest-server-ana @ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas; rest-server-nh3 @ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of /mnt/backup. (rest-server-ana recovered 2026-06-20.)

  • pull-hf-repo.yaml is the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at /tank/aimodels/huggingface/" playbook. Supports --var repo_type=model|dataset|space. Replaces ad-hoc huggingface_hub.snapshot_download patterns.

  • Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (key_id 61419c92) at ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin auths against demo only. Personal-instance admin (the ~/.config/worldtree/personal-admin-token, mode 600) POSTs /admin/keys (mints per-project keys; takes user_id+label, no scope param — scopes are tier-derived, so admin-tier scopes like admin.events.read must be granted WT-side by worldtree-dev).

  • Per-project user keys against personal Worldtree (issued 2026-05-19): skaldsong:79744637 (nh3-dev iteration), skaldsong:7c1dbbbe (ana-docker prod), althing:50d85460, mead-hall:a360822d. Same user_id=skaldsong across both skaldsong keys → shared Heimdall agent slot; different key_id → independently rotatable. Pattern: mint via /admin/keys, drop value to /tmp/wt-personal-<name>.key mode 600, dev collects + shreds (DO NOT cat to chat transcript).

  • Skaldsong CD pattern (registry-pull). Differs from althing / asset-engine which build-on-host. vh/skaldsong's CI builds and pushes gitea.phasefinal.com/vh/skaldsong:<sha> + :latest; playbooks/deploy-skaldsong.yaml on ana-docker pulls + recreates. SHA-pin only (no :latest health-gated advance yet). Prereq: host needs docker login gitea.phasefinal.com once (read:package PAT) — not currently in the workflow.

  • gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH 10.250.50.70:222, HTTP :3000. Fleet/colo hosts must use this internal route, NOT public gitea.phasefinal.com (38.120.12.44, ana-srv1) — the public path fail2bans the host egress IP and wedges webhook deploys. :22 on 10.250.50.70 is ana-docker's HOST sshd, not gitea. Full gotcha in docs/orientation.md → Git/gitea.

  • docker-as-root pattern (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo): docker run --rm -v <target-dir>:/wt [-v /var/run/docker.sock:/var/run/docker.sock] docker:cli sh -c "..." (or alpine for plain file ops). docker-group membership is effectively root via bind-mount; treat as sudo-equivalent. Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the Docker daemon interprets them against the HOST fs — always pass -e VAR=/abs/path for any relative-default config dir.

  • scripts/elway sudo handling — elway prompts for the sudo password ONCE via getpass before the first sudo: true step → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH.

  • Per-host SSH identity matters for sudo. infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker there are TWO identities: the default ssh ana-docker = lkraven (docker-group, NO passwordless sudo — docker works, root-file edits don't); but ssh infra-ops@ana-docker HAS NOPASSWD root (verified 2026-06-20). → For any sudo op on ana-docker (mount, root-owned files, service control), use ssh infra-ops@ana-docker, NOT the default session. lkraven-owned files (litellm config, most stack compose/conf) still take plain cp/edit under either identity. NEW (2026-06-25): ssh infra-ops@10.100.10.50 (nh3-dev) ALSO has NOPASSWD sudo; on nh3-extdev infra-ops is sudo-LESS by design and ssh lkraven@10.100.50.42 is the NOPASSWD path there.

Current state / in-flight

As of 2026-07-02:

  • gen model = qwopus (qwen3.5-122-a10b); the Deckard trial is CONCLUDED. Trialed robbatt/Qwen3.6-40B-Deckard-NVFP4 behind the gen aliases — it won writing quality decisively but lost on speed (~36 vs qwopus's ~90 tok/s); reverted (git 681eb70). Deckard kept STAGED on ana-ml2 as T1's quality benchmark (/tank/aimodels/qwen36-40b-deckard-{bf16,nvfp4}; container vllm-deckard-40b stopped-but-kept). Full arc in auto-memory reference_gen_qwopus_122b.

  • T1 (mtf-dev qwopus writing-LoRA) IN PROGRESS. Attention-only LoRA (full-attn q/k/v/o

    • the bf16 GDN in_proj_qkv/out_proj); serve via swappable-LoRA-on-NVFP4 (test QUEUED, gated on the first T1 adapter) or merge+requant. Trainer harness seam MECHANICALLY proven via a granite-8b spike (2026-07-02): the pipeline runs end-to-end, but efficacy is NOT validated (mechanics-only synthetic fixture, by design; a granite adapter wouldn't transfer to qwopus regardless). Real behaviour-shift efficacy is a T1-run question. NVFP4 quant-structure + serve-path facts in reference_gen_qwopus_122b.
  • Worldtree #332 scoped-log view + tunnel = STANDING ASSET. wt_gateway_logs view + wt_readonly role on the litellm DB + wt-db-tunnel systemd on corviduo-dev, for worldtree-dev's embed-recall regression-watches. #332 fix verified in prod (15×→1.01× re-embed). Teardown steps + the IP-pin caveat in auto-memory reference_wt_gateway_scoped_log_view.

  • /books mounted (transient) on nh3-dev for a books/corpus ingestion: 10.0.50.50:/mnt/books (ESH NAS) → /mnt/books, NFSv4 ro,soft (soft dodges the ESH D-state hang). NOT fstab — re-mount via infra-ops@10.100.10.50 if nh3-dev reboots.

  • granite-4.1-8b bf16 kept at irv-ml1:/home/lkraven/granite-4.1-8b-bf16 (17G, reap-on-request; the T1-harness-spike base). comfyui was borrowed off the A6000 for the spike + RESTORED healthy.

  • Recently completed (2026-06-22..25, now in Recent decisions): Worldtree #314/#322/#317 persona-render config arc; nh3-extdev althing v0.17.1 mesh peer (Model B); zellij web-seat pilot.

  • Backups — recovered + hardened (2026-06-20), STILL OPEN: rotate the 5 disclosed rest-server creds (operator, offline); confirm esh-vm-db's resticprofile includes DB dumps (esh-pve-nas restic-only). Topology + 2-min check: docs/runbooks/backups.md.

  • R22 (brokkr/dwarves) — CONCLUDED, gateway-only. Phase B (internal-model hooks) CANCELLED (Worldtree model-agnostic → no deploy path). Whole footprint = gateway calls to free qwen3.5-122-a10b via a PERSISTENT full-access key at /home/lkraven/.r22-gateway-key (mode 600, carries paid GLM, do NOT delete). Detail in Recent decisions.

  • Open follow-ups (low-priority): clean phantom qwen3.6-35b-a3b off the gateway (lists in /v1/models but 400s — backend displaced); docker-daemon default log-cap on ana-docker (systemic fix behind the 94 GB disk incident).

  • Standing / parked (from prior): Mac Pro migration (migration-plan.md, hardware-gated); R17 v2 corpus push HELD (distribution barred); disclosed-keys rotation queue; clean legacy news-digest on ana-docker.

  • Worldtree config-propagation (reference): demo+personal bind-mount config (providers.yaml/model_roles.yaml/policies.yaml/defaults.yaml, now incl. the #314/#322/#317 additions) from /opt/worldtree{,-personal}/config (infra-ops-deployable, byte-identical from worldtree-dev canonical sha); pinned is IN-IMAGE. Reload via docker restart <container>, NEVER compose up (stale-:latest footgun); WT PDP is rule-based (a tier needs allow RULES, not just scopes). #317 lesson: config REMOVALS are NOT backward-compatible with the still-running image — push promptly, don't bounce the instance in the window.

Recent decisions

  • [2026-07-02] mtf-dev granite harness-spike provisioned on irv-ml1 (comfyui displaced for the A6000) → ran GREEN. Freed the A6000 by stopping comfyui (operator-coordinated with comfy-dev), staged granite-4.1-8b bf16 to irv-ml1:/home/lkraven, mtf-dev's TRL SFT→DPO→eval seam proved end-to-end (DPO genuinely learned, 0.833 acc; anti-slop ~0 = the expected null on clean-writing granite-instruct); comfyui restored healthy. De-risks T1's trainer harness ahead of the real qwopus train. MECHANICAL green ONLY — efficacy NOT validated, by design (mtf-dev confirm 2026-07-02): the run used a 12-row/12-pair SYNTHETIC writing fixture (clean-vs-sloppy generic prompts), NOT the E-RP corpus; rank 8, 1 epoch, ~2 SFT + ~6 DPO steps. The in-loop base-vs-adapter check (HoldoutEvaluator) already ran and returned the expected null (anti_slop_improvement 0.002); the adapter was then reaped (gone), so there is nothing to A/B — and on a mechanics-only synthetic adapter an A/B would only reconfirm the null. Real efficacy = the T1 run (real recipe + E-RP data on qwopus). Open (mtf-dev routing to operator): whether to insert an intermediate real-efficacy granite spike before T1 — weak proxy (granite arch ≠ qwopus, + granite is censored vs qwopus abliterated), leaning defer-to-T1.

  • [2026-07-01] Worldtree #332 embed-recall diagnosed + scoped-log view/tunnel provisioned + fix verified. The persona-recitation gate re-embedded persona segments every turn (15× re-embed, 95% cross-turn recurrence, one anchor ×206/min → worldtree-gateway was ~95% of the embedding backend load); worldtree-dev's cross-turn content-hash cache dropped it to 1.01× in prod. Built a boundary-safe read-only wt_gateway_logs view + wt_readonly role + wt-db-tunnel (corviduo-dev) for their regression-watch. auto-memory reference_wt_gateway_scoped_log_view.

  • [2026-07-01] qwopus native MTP speculative-decode tested on gen → NOT kept. qwopus HAS a full native MTP head (785 tensors, in NVFP4 + bf16). Measured +12% single-stream but 1520% AGGREGATE at moderate concurrency (N=4: 250→200 tok/s) + it silently drops min_p/logit_bias → reverted to clean baseline. The win is banked for T1 (MTP as a per-deployment option, preserved through requant). reference_gen_qwopus_122b.

  • [2026-07-01] Deckard trial → reverted to qwopus (gen). robbatt/Qwen3.6-40B-Deckard-NVFP4 won writing "in every way" but at ~36 vs ~90 tok/s (dense-40B vs MoE-~10B-active); spec-decode rescue ruled out (aeon DFlash image is arm64/DGX-Spark-only; Deckard's NVFP4/bf16/base-27B all LACK MTP; EAGLE-head training a multi-week non-starter). git b63c48b (repoint) + 681eb70 (revert). Deckard kept staged as T1's writing benchmark.

  • [2026-06-25] althing re-architected to the v0.17 lean multi-machine bus; nh3-extdev stood up as a mesh peer (MODEL B). v0.15.0's lean-bus cut RIPPED moderation/chamber/forseti-daemon/agent-runner/redis-valkey; v0.17 = per-box local-SQLite bus + a courier/receiver for P2P over the 10.x net (installed per-box via uv tool install, NOT CI). Migrated nh3-dev's client v0.15→v0.17. nh3-extdev provisioned as a full mesh peer with the operator's MODEL B (dedicated althing-svc service account + group-shared /srv/althing root, so multiple OS users share one config/DB; ALTHING_ROOT via /etc/profile.d). Verified multi-user concurrent rw. Supersedes the old althing Tools-row (was chamber/daemons/valkey). auto-memory reference_nh3_extdev_althing_mesh.

  • [2026-06-23] zellij native web client piloted on nh3-dev (zellij-web.service :8443, native TLS), system-installed for all users, ALONGSIDE ttyd (kept as the no-auth fallback). Bind left 0.0.0.0 per operator. auto-memory reference_zellij_web_seat.

  • [2026-06-22] Worldtree persona-render config arc (#314 relational-stance / #322 affect-mood gate / #317 valence teardown) pre-synced + deployed green on demo+personal — #317 a boot-blocking config REMOVAL (relational_valence). Config-sync recipe + additions-vs-removals lesson in auto-memory reference_corviduo_dev_emergency_ops.

  • [2026-06-20] R22 (brokkr/eitri/dwarves) stood down to gateway; full-access R22 key minted; Phase B parked. Operator caught the raw-weights/GPU ask as premature — Phase A (P00 + NL-rendering) is PROMPT-LEVEL on the LiteLLM gateway, not raw weights (no GPU, no Blackwell displacement). Real blocker was gateway ACCESS: the shared all-agents-local key only serves granite chat, and granite-8b is too weak to be the P00 model-under-test (capability-confound). Operator said "full access" → minted r22-brokkr-phaseA LiteLLM key (scope all-proxy-models, incl PAID GLM), dropped ~/.r22-gateway-key (mode 600) for eitri — re-minted PERSISTENT after eitri's stateless tool-env couldn't hold it in env + shredded the one-shot drop (old key revoked); the live full-access key now lives at /home/lkraven/.r22-gateway-key, read per-command, don't delete (cost surface — carries paid GLM; operator reconfirmed full-open + declined to narrow, 2026-06-20 — settled, don't re-surface). Lesson: stateless consumers need a persistent mode-600 file read per-call, NOT shred-after-load. P00 MUT = qwen3.5-122-a10b (alias gen) — 122B, capable, FREE-LOCAL (no per-token cost); told them keep frozen calibration on it, NOT the paid GLM the key can now reach. PHASE B (raw-weights hooks: Dvalin activation-steering on hidden states + Regin ECM logit-injection on a 7-8B Qwen/Llama via transformers LogitsProcessor) CANCELLED 2026-06-20 (not deferred) — operator cut the internals arms: Worldtree is model-agnostic (regard/ affect renders via context manipulation = prompt-level only), so the internal-model approach has no deploy path (the pragmatic-outcome gate firing). R22's ENTIRE compute is now the gateway (free qwen3.5-122-a10b MUT); NO raw-weights/GPU ever; the A6000 slice reservation is released (was never provisioned). The full-scope R22 key stays (pinned to the free qwen-122, GLM untouched). OPERATOR STEER (2026-06-20): R22 is research for a PRAGMATIC/deployable outcome, NOT advancing-the-art. Before the dwarves invest in Phase B, establish a deploy path (distilled steering-vector / fine-tune applied to a model we already serve). Operator is taking the dwarf discussion directly.

  • [2026-06-20] claude-bot issue-scope token minted for worldtree-dev self-serve (closes their last tea-as-vh fallback — issues; they already self-serve Actions + deploys). Gitea PAT scopes are IMMUTABLE → minted a NEW claude-bot PAT worldtree-dev-actions-issues-20260621 (id 16) with write:repository+write:issue via basic-auth (claude-bot:gitea-password, POST /users/claude-bot/tokens); dropped mode-600 ~/.claude-bot-token-issues → worldtree-dev swaps. Gitea COLLAPSES read:issue into write:issue (write implies read). Old token (id 15) REVOKED after verification → claude-bot carries ONE consolidated worldtree-dev token (id 16) + arbo-ci (id 14). Advances the credential-migration directive.

  • [2026-06-20] rest-server-ana recovered + backup prevention shipped + worldtree-dev admin keys provisioned (operator-directed). rest-server-ana fixed (mount-a + force-recreate via infra-ops@ana-docker); freshness alert + fstab hardening landed (see docs/runbooks/backups.md). Minted worldtree-dev admin-tier Heimdall keys on demo (key_id d113207c) + personal (f4f75adb), dropped mode-600 to ~/.wt-admin-{demo,personal} on nh3-dev → worldtree-dev now self-serves key minting. Cred rotation (5 rest-server pw) BELAYED per operator.

  • [2026-06-20] claude-bot → ADMIN on vh/Worldtree (operator-authorized; one-time use of operator vh-admin to enable the migration) — claude-bot self-serves WT deploys/tokens henceforth, reducing personal-cred fallbacks.

  • [2026-06-14] STANDING DIRECTIVE: migrate ALL infra access to Claude-specific credentials. Stand up service accounts (claude-bot Gitea user + scoped tokens, distinct litellm admin key); re-mint consumer creds under them; flag personal-cred fallbacks until done. (auto-memory project_migrate_infra_access_to_claude_credentials)

118 older entries archived to archival-memory.md.

Tried and abandoned

  • [2026-07-01] A personal-Worldtree CI deploy that fails ~85s in with "not found / unauthorized" is usually the pull-only-vs-build RACE, not registry-auth. deploy-personal.yml is PULL-ONLY ("image must already be built by a push to main") but fires on the staging/vX tag push SIMULTANEOUSLY with deploy.yml's main build → it tries to pull the image ~2.5 min BEFORE the build finishes pushing it → step-4 "Verify image exists" aborts "not found". Misattributed to registry-auth twice (the earlier cff3328 saga too). DIAGNOSE: the image tag (12-char short-sha, NOT 7) exists in the registry + the VM login succeeds ⇒ it's the race. FIX: re-run once the build's done (image now present), OR gate deploy-personal on workflow_run: completed.

  • [2026-07-01] MTP/spec-decode on a SHARED serving model helps single-stream but HURTS moderate-concurrency aggregate throughput + silently ignores min_p/logit_bias. Measured on qwopus gen: N=1 +12%, N=4 20% aggregate. Don't bolt spec-decode onto the fleet gen for a single-stream win — reserve it for dedicated/interactive deployments.

  • [2026-07-02] irv-ml1 /worktank ROOT is root-owned — lkraven can't write there (and irv-ml1 sudo needs a password non-interactively) → stage model pulls to /home. Also PIN THE A6000 BY UUID for training runs: nvidia-smi index 1 is the A6000, but native-CUDA ordering can differ vs docker, and the 3090 (index 0) is usually near-full → land there and OOM. CUDA_VISIBLE_DEVICES=GPU-<uuid>.

  • [2026-06-25] althing "unreachable: — retry later" can MASK an app-level 500. A multi-hour "intermittent :8087 / firewall flap" hunt was a red herring: raw network was always clean (curl POST to the receiver :8087 worked; a connect probe = 0 fails). Root cause = the receiver's DB agents-table wasn't synced with the config roster, so delivery 500'd "unknown to: <handle>", which the courier MAPPED to "unreachable" (looks like network/DNS). Diagnose: raw curl to :8087 + a connect-probe pass ⇒ NOT network — look up-stack. Fixed in althing v0.17.1 (receiver auto-ensures the recipient on delivery). Stop-gap on older builds: ALTHING_HANDLE=<h> althing-cli inbox as the receiver user syncs config→DB; re-run after any roster change. auto-memory reference_nh3_extdev_althing_mesh.

  • [2026-06-20] rest-server .htpasswd: permission denied = the ana-nas NFS mount FAILED (ghost file on the local mount point), NOT a decommission. mnt-backup.mount stuck failed (fstab bare defaults, no retry) → the rest-server serves an empty local dir with a root:root 0-byte .htpasswd. Documented recovery in disaster-recovery.md. Don't tear down what looks like a crash-looping legacy container until you've checked fstab + the docs — it was the live ana-side restic target.

  • [2026-06-20] The DEFAULT ssh ana-docker is lkraven (no NOPASSWD) — but ssh infra-ops@ana-docker HAS NOPASSWD root (corrected later same day). Early on a sudo cp as lkraven silently failed (password prompt) → one wasted gateway bounce; I then wrongly concluded "no sudo on ana-docker" and nearly punted the rest-server-ana recovery to the operator. The real rule: reach for infra-ops@ana-docker for sudo ops; lkraven-owned files (litellm config, most stack compose/conf) take plain cp under either identity.

98 older entries archived to archival-memory.md.