# Persistent memory — eshpfi-management _Last updated: 2026-07-05_ ## Repo purpose Reference workspace for PFI infrastructure: server inventory, canonical Docker Compose stacks, ops playbooks, and conventions. Authoritative copies of compose files live on the servers under `/opt/docker/compose//`; this repo mirrors them for version control, editing, planning, and CI-driven deploys. **It was originally spun up to handle the fleet backups** — keep that lens when triaging backup/storage issues. ## Tools and conventions Sister repos (separate gitea repos, deployed by playbooks here): | Repo | Role | CI status | |---|---|---| | `vh/task-board` | MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) | | `vh/vor` | Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) | | `vh/nevermore` | Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) | | `vh/asset-engine` | Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) | | `vh/althing` | Lean trusted inter-agent message bus — **v2 "email model" (v2.0.0b2, 2026-07)**: per-box local-SQLite bus + courier/receiver for P2P over the 10.x net; pillars = open-loops / per-box herald + wake-listener / roaming owner API `/owner/*` / `althing-mcp` stdio surface. The v0.15 lean-bus cut RIPPED moderation / chamber / forseti-daemon / agent-runner / redis-valkey. | per-box `uv tool install` (NOT CI-deploy); **nh3-dev = the DEV box** (editable install of `~/development/althing`, gets new versions first); **nh3-extdev** a mesh peer (model B: althing-svc + shared `/srv/althing`) | | `vh/mead-hall` | Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) | | `vh/skaldsong` | Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) | | `vh/Worldtree` | Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. **gitea-runner builds on ana-docker**; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. | push-to-main → CI build-and-deploy (runner on ana-docker) | | `vh/yt-voice-clipper` | YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → **gitea-webhook auto-deploy** to irv-ml1 (2026-06-03) — see `docs/runbooks/ytvc-autodeploy.md` | | `vh/arbo` | Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) | push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook | | `model-training-forge` (mtf-dev) | Fine-tuning recipe forge; **T1 = qwopus E-RP writing LoRA** (SFT→DPO, LitBench-RM reward) | training runs, not a deployed sidecar | (`vh/volva` + Heid were re-architected from systemd daemons to Claude Code session orchestrators 2026-06-08; their nh3-dev `.service` units were removed — no longer deployed sidecars here. See Recent decisions.) - **Two-layer backups** — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see `docs/runbooks/disaster-recovery.md` for the blast-radius matrix. **⚠️ The restic file+DB layer routes through TWO rest-servers** (`rest-server-ana` @ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas; `rest-server-nh3` @ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of `/mnt/backup`. (rest-server-ana recovered 2026-06-20.) - **`pull-hf-repo.yaml`** is the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at `/tank/aimodels/huggingface/`" playbook. Supports `--var repo_type=model|dataset|space`. Replaces ad-hoc `huggingface_hub.snapshot_download` patterns. - **Worldtree admin auth — per-instance.** Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (`key_id 61419c92`) at `ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin` auths against **demo only**. Personal-instance admin (the `~/.config/worldtree/personal-admin-token`, mode 600) POSTs `/admin/keys` (mints per-project keys; takes `user_id`+`label`, **no scope param** — scopes are tier-derived). **On-instance mint recipe (cleaner than DB-manip):** `docker exec worldtree-worldtree-api-1` POST `/admin/keys` with the in-container `WORLDTREE_BOOTSTRAP_ADMIN_KEY`; cleartext once in `.key`=`wt_live_+16hex`. auto-memory `reference_worldtree_demo_key_mint`. - **Per-project user keys against personal Worldtree** (issued 2026-05-19): `skaldsong:79744637`, `skaldsong:7c1dbbbe`, `althing:50d85460`, `mead-hall:a360822d`. Mint via `/admin/keys`, drop value to `/tmp/wt-personal-.key` mode 600, dev collects + shreds (DO NOT cat to chat transcript). - **Skaldsong CD pattern (registry-pull).** vh/skaldsong's CI builds and pushes `gitea.phasefinal.com/vh/skaldsong:` + `:latest`; `playbooks/deploy-skaldsong.yaml` on ana-docker pulls + recreates. SHA-pin only. Prereq: host needs `docker login gitea.phasefinal.com` once. - **gitea internal route for fleet hosts.** gitea is a container on **ana-docker** — git-SSH `10.250.50.70:222`, HTTP `:3000`. Fleet/colo hosts must use this internal route, NOT public `gitea.phasefinal.com` (`38.120.12.44`) — the public path fail2bans the host egress IP. Full gotcha in `docs/orientation.md` → Git/gitea. - **docker-as-root pattern** (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo): `docker run --rm -v :/wt docker:cli sh -c "..."`. docker-group membership is effectively root via bind-mount. **Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass `-e VAR=/abs/path` for any relative-default config dir.** - **`scripts/elway` sudo handling** — elway prompts for the sudo password ONCE via `getpass` before the first `sudo: true` step → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH. - **Per-host SSH identity matters for sudo.** infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: **default `ssh ana-docker` = `lkraven`** (docker-group, NO passwordless sudo); **`ssh infra-ops@ana-docker` HAS NOPASSWD root**. **→ For any sudo op on ana-docker, use `ssh infra-ops@ana-docker`.** `ssh infra-ops@10.100.10.50` (nh3-dev) ALSO NOPASSWD sudo; on **nh3-extdev** infra-ops is sudo-LESS by design (`ssh lkraven@10.100.50.42` is the NOPASSWD path). **irv-ml1: `ssh irv-ml1` = lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD)** — stage model pulls to `/home`, not root-owned `/worktank`. ## Current state / in-flight _As of 2026-07-05:_ - **T1 TRAIN — venue decision live; NEXT ACTION = the ana-ml2 SMOKE.** T1 (qwopus E-RP writing LoRA) datasets + recipe + smoke set are all staged (mtf-dev, on `nh3-dev:data/export/qwen-3.5-122b-erp-lora/` incl. `smoke/`); bf16 base is on ana-ml2 `/tank/aimodels/qwopus3.5-122b-a10b-bf16` (233G, PUBLIC HF repo, zero-auth pull). On-prem RULED OUT except a full ana-ml2 shutdown → CPU-offload → ~1-day fleet outage; cloud (Vast.ai 8×80GB, ~3-6h, ~$60-500, zero impact) is the clean alt. **Operator chose smoke-first on ana-ml2** to measure real samples/sec before committing. ⚠️ the smoke ITSELF needs both GPUs + the ~421G vLLM RAM freed (192G VRAM < 244G base → CPU offload) = a brief full ana-ml2 serving pause; keeping any serving up drops it into the ~week NVMe regime. Full runbook + both paths in `reference_t1_cloud_train_plan`. Post-train serve-path (swappable-LoRA-on-NVFP4 test) still queued: `reference_gen_qwopus_122b`. - **LitBench-RM UP (irv-ml1 A6000); comfyui DISPLACED.** Respun the R19 creative-writing reward (`vllm-litbench-rm` :8202, ~19.6G) for T1's reward ensemble; stopped comfyui to free the A6000 (operator is sole comfyui consumer + holding image-gen until LitBench done). Score scale re-verified from nh3-dev (literary 0.95 > flat 0.44 > slop 0.12). **TEARDOWN on "litbench done":** `ssh irv-ml1 'docker rm -f vllm-litbench-rm && docker start comfyui'` — ⚠️ then watch the mmartial boot trap (root-owned venv → crash-loop; `chown -R 1000:1000 /comfy/mnt/venv`). `reference_litbench_rm_irv_ml1`. - **character-rp role — SHIPPED; one QUEUED spot-check.** Proved per-request `extra_body` (top_k/repetition_penalty) forwards through the `gen-reasoning` alias to the vLLM sampler (no gateway cap needed); pre-staged the `character-rp` role in demo+personal bind-mount `model_roles.yaml` (byte-verified live on b18; caught the cached-registry ordering so the deploy's own restart activates it). worldtree-dev shipped **#344 (v1.0.0b19)** fixing the durable-agent override-drop (role now applies the RP profile). QUEUED: ratatoskr-dev pings a turn-timestamp → I spot-check the gateway spend_logs to confirm the RP sampler params (temp 0.75 / top_p 0.85 / presence 1.5 / top_k 20 / rep 1.0) reach vLLM. - **althing v2 (herald + receiver) = systemd-supervised on nh3-dev (the althing DEV box).** Formalized this session: `althing-herald.service` (Restart=always, **Environment=PATH incl ~/.cargo/bin** — the fix for the silent pane-dispatch outage) + `althing-receiver.service` (v2 → pillar-3 `/owner/*` live). Stale forseti unit removed; both on v2.0.0b2, canonicalized by forseti. `reference_nh3_dev_althing_herald`. - **glm-5.2 = z.ai 1M-in / 128K-out, NO gateway cap** (probed z.ai; recorded in the config comment + `reference_litellm_gateway`; commit 624a07e). All GLM entries are pure z.ai passthrough — the effective ceiling is z.ai's canonical, not a gateway limit. - **`/books` mounted (transient) on nh3-dev** (`10.0.50.50:/mnt/books` → NFSv4 ro,soft; re-mount via `infra-ops@10.100.10.50` if it reboots). - **Backups — recovered + hardened (2026-06-20), STILL OPEN:** rotate the 5 disclosed rest-server creds (operator, offline); confirm esh-vm-db's resticprofile includes DB dumps. `docs/runbooks/backups.md`. - **Open follow-ups (low-priority):** clean phantom `qwen3.6-35b-a3b` off the gateway (lists in /v1/models but 400s); docker-daemon default log-cap on ana-docker; drop the now-inert `mood.decay_rate`/`mood.stale_hours` keys from the deployed `/opt/worldtree*/config` bind-mounts (harmless-present per worldtree-dev; edit-only/no-restart at a no-deploy window). - **Standing / parked:** Mac Pro migration (hardware-gated); R17 v2 corpus HELD; disclosed-keys rotation queue; clean legacy `news-digest`; R22 gateway-only full-access key at `/home/lkraven/.r22-gateway-key` (mode 600, carries paid GLM, do NOT delete); granite-4.1-8b bf16 at `irv-ml1:/home/lkraven/granite-4.1-8b-bf16` (reap-on-request); Deckard staged on ana-ml2 as T1's writing benchmark. - **Worldtree config-propagation (reference):** demo+personal bind-mount config from `/opt/worldtree{,-personal}/config` (infra-ops-deployable, byte-identical from canonical); reload via `docker restart `, NEVER `compose up` (stale-`:latest` footgun). **The role registry is loaded ONCE + CACHED at startup** (`ConversationService.role_registry`) → a bind-mount `model_roles.yaml` change needs a container restart to take effect; pre-stage the bind-mount BEFORE the activating deploy's restart (else it needs a second bounce). config REMOVALS are NOT backward-compatible with the still-running image. ## Recent decisions - `[2026-07-05]` **T1 training venue: CLOUD recommended; operator chose smoke-first on ana-ml2.** On-prem ruled out (ana-ml2 full — both 96G GPUs ~93G used): keep-serving = NVMe offload ~6-8 DAYS; full ana-ml2 shutdown = CPU offload ~1 DAY but a whole-fleet outage. Cloud Vast.ai 8×80GB (no offload → ~3-6h, ~$60-500, zero fleet impact) is the clean alt (mtf-dev + infra-ops both rec; Vast for its no-content-AUP marketplace + likely-existing VastBlue account). Operator's next step = the ana-ml2 CPU-offload SMOKE (~60 steps) to get real samples/sec before the full-outage-vs-cloud call. HF base verified public (zero-auth pull). Runbook + gotchas in `reference_t1_cloud_train_plan`. - `[2026-07-05]` **glm-5.2 canonical limits recorded** (probed live vs z.ai): **1,048,576 (1M) input context / 131,072 (128K) max output**; NO gateway-side cap (pure passthrough → z.ai's limits are effective). Written to the config comment (commit `624a07e`) + `reference_litellm_gateway`. - `[2026-07-04]` **character-rp: gateway-forwarding proven + role pre-staged + #344 shipped.** Empirically confirmed per-request `extra_body` (top_k/repetition_penalty) forwards through the `gen-reasoning` LiteLLM alias to vLLM + standard params override the alias defaults — no gateway cap needed (I over-built a dedicated alias, operator corrected, reverted with zero fleet impact). Pre-staged the `character-rp` role into demo+personal bind-mount `model_roles.yaml` (byte-verified on b18; caught the cached-registry ordering). worldtree-dev shipped **#344 (v1.0.0b19)** for the durable-agent override-drop. spend_logs spot-check queued (ratatoskr's timestamp ping). - `[2026-07-04]` **althing v2 herald+receiver formalized as systemd on nh3-dev.** `althing-herald.service` (Restart=always, **Environment=PATH incl ~/.cargo/bin** — the pane-dispatch fix) + `althing-receiver.service` (v2 → pillar-3 `/owner/*` live); stale forseti unit removed; both on v2.0.0b2, canonicalized by forseti. `reference_nh3_dev_althing_herald`. - `[2026-07-04]` **LitBench-RM respun (irv-ml1 A6000, comfyui displaced)** for T1's reward ensemble; operator sole comfyui consumer, holding image-gen until LitBench done. `reference_litbench_rm_irv_ml1`. - `[2026-07-03]` **ratatoskr-dev DEMO Heimdall key provisioned (R30 φ0).** Minted a tier-user key on the demo via `POST /admin/keys` (bootstrap admin key), mirroring their personal base consumer (no character-binding); base-agent affect reads work ungated. `reference_worldtree_demo_key_mint`. - `[2026-07-02]` **mtf-dev granite harness-spike ran GREEN — MECHANICAL only, efficacy DEFERRED to the T1 run.** Trainer TRL SFT→DPO→eval seam proven end-to-end on a synthetic fixture (not the E-RP corpus); operator DECIDED no intermediate real-efficacy granite spike (uninterpretable proxy — arch gap + abliteration axis). `reference_gen_qwopus_122b`. - `[2026-07-01]` **Worldtree #332 embed-recall diagnosed + scoped-log view/tunnel provisioned + fix verified** (15×→1.01× re-embed). `reference_wt_gateway_scoped_log_view`. - `[2026-07-01]` **qwopus native MTP speculative-decode tested on `gen` → NOT kept** (+12% single-stream, −15–20% aggregate at concurrency, silently drops min_p/logit_bias). Banked for T1. `reference_gen_qwopus_122b`. - `[2026-07-01]` **Deckard trial → reverted to qwopus (`gen`)** (won writing "in every way" but ~36 vs ~90 tok/s; spec-decode rescue ruled out). git `b63c48b`→`681eb70`. Deckard kept staged as T1's writing benchmark. - `[2026-06-25]` **althing re-architected to the lean multi-machine bus; nh3-extdev stood up as a MODEL B mesh peer** (dedicated `althing-svc` + group-shared `/srv/althing`). `reference_nh3_extdev_althing_mesh`. - `[2026-06-23]` **zellij native web client piloted on nh3-dev** (`zellij-web.service` :8443) alongside ttyd. `reference_zellij_web_seat`. - `[2026-06-22]` **Worldtree persona-render config arc (#314/#322/#317) pre-synced + deployed green on demo+personal** — #317 a boot-blocking config REMOVAL. `reference_corviduo_dev_emergency_ops`. - `[2026-06-20]` **R22 (brokkr/dwarves) stood down to gateway-only; full-access R22 key minted; Phase B CANCELLED** (Worldtree model-agnostic → no deploy path). Key at `/home/lkraven/.r22-gateway-key` (persistent mode-600, carries paid GLM, don't delete). MUT = free `qwen3.5-122-a10b` (`gen`). Operator steer: R22 research is gated on a pragmatic/deployable outcome, not advancing-the-art. - `[2026-06-20]` **claude-bot issue-scope token minted for worldtree-dev self-serve** (id 16, `write:repository`+`write:issue`); old token revoked. Advances the credential-migration directive. - `[2026-06-20]` **rest-server-ana recovered + backup prevention shipped + worldtree-dev admin keys provisioned** (demo d113207c / personal f4f75adb). Cred rotation (5 rest-server pw) BELAYED. - `[2026-06-20]` **claude-bot → ADMIN on vh/Worldtree** (operator-authorized) — self-serves WT deploys/tokens henceforth. - `[2026-06-14]` **STANDING DIRECTIVE: migrate ALL infra access to Claude-specific credentials.** (auto-memory `project_migrate_infra_access_to_claude_credentials`) _118 older entries archived to archival-memory.md._ ## Tried and abandoned - `[2026-07-04]` **LiteLLM (this gateway version) mutates the SHARED deployment config in-place on per-request sampler-param merge** → my deliberately-invalid `top_k=-5` forwarding-probe bled into a param-less character-rp request (vLLM 400, ONE-OFF, self-cleared by a later valid probe). NOT caching (none configured), NOT a config change. **Never fire invalid/distinctive sampler values at a SHARED gateway alias with live consumers** — use a throwaway alias, or a `docker restart litellm` flushes residual carryover. `feedback_litellm_shared_param_mutation`. - `[2026-07-04]` **A systemd `--user` daemon that shells out to `~/.cargo/bin`/`~/.local/bin` tools needs an explicit `Environment=PATH`** — the minimal `--user` default silently drops them. The althing herald lost `zellij` → silent `pane-miss` for ALL config-backed TUI/pane agents; CC + FIFO routes were unaffected, so it was invisible from a CC session. `reference_nh3_dev_althing_herald`. - `[2026-07-04]` **On-prem T1 train that keeps ANY ana-ml2 serving up = ~6-8 DAYS** (1-GPU + NVMe ZeRO-Infinity offload; MoE ~10B-active cuts FLOPs but NOT the 244G base's param I/O). The only fast on-prem path is a FULL ana-ml2 shutdown (both GPUs + the ~421G vLLM RAM freed → base fits in the 566G CPU RAM) → CPU offload → ~1-day full-fleet outage. Cloud (no offload) = hours. `reference_t1_cloud_train_plan`. - `[2026-07-01]` **A personal-Worldtree CI deploy that fails ~85s in with "not found / unauthorized" is usually the pull-only-vs-build RACE, not registry-auth.** `deploy-personal.yml` is PULL-ONLY but fires on the `staging/vX` tag simultaneously with `deploy.yml`'s build → pulls before the push finishes. FIX: re-run once built, or gate on `workflow_run: completed`. - `[2026-07-01]` **MTP/spec-decode on a SHARED serving model helps single-stream but HURTS moderate-concurrency aggregate + silently ignores `min_p`/`logit_bias`** (qwopus `gen`: N=1 +12%, N=4 −20%). Reserve for dedicated/interactive deployments. - `[2026-07-02]` **irv-ml1 `/worktank` ROOT is root-owned — lkraven can't write there (irv-ml1 sudo needs a password) → stage model pulls to `/home`.** PIN THE A6000 BY UUID for training (native-CUDA ordering differs vs docker; the 3090 index 0 is usually near-full → OOM). `CUDA_VISIBLE_DEVICES=GPU-`. - `[2026-06-25]` **althing "unreachable: " can MASK an app-level 500.** Raw network was clean; root cause = receiver DB agents-table not synced with the config roster → delivery 500'd "unknown to: ", MAPPED to "unreachable". Diagnose: raw curl to :8087 + connect-probe ⇒ NOT network. Fixed in althing v0.17.1. `reference_nh3_extdev_althing_mesh`. - `[2026-06-20]` **rest-server `.htpasswd: permission denied` = the ana-nas NFS mount FAILED (ghost file on the local mount point), NOT a decommission.** `mnt-backup.mount` stuck `failed` (fstab bare `defaults`) → rest-server serves an empty local dir. Recovery in disaster-recovery.md. - `[2026-06-20]` **The DEFAULT `ssh ana-docker` is `lkraven` (no NOPASSWD) — but `ssh infra-ops@ana-docker` HAS NOPASSWD root.** A `sudo cp` as lkraven silently failed → nearly punted the rest-server recovery. Reach for `infra-ops@ana-docker` for sudo ops. _98 older entries archived to archival-memory.md._