# Persistent memory — eshpfi-management _Last updated: 2026-06-25_ ## Repo purpose Reference workspace for PFI infrastructure: server inventory, canonical Docker Compose stacks, ops playbooks, and conventions. Authoritative copies of compose files live on the servers under `/opt/docker/compose//`; this repo mirrors them for version control, editing, planning, and CI-driven deploys. **It was originally spun up to handle the fleet backups** — keep that lens when triaging backup/storage issues. ## Tools and conventions Sister repos (separate gitea repos, deployed by playbooks here): | Repo | Role | CI status | |---|---|---| | `vh/task-board` | MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) | | `vh/vor` | Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) | | `vh/nevermore` | Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) | | `vh/asset-engine` | Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) | | `vh/althing` | Lean trusted inter-agent message bus — **v0.17 multi-machine** (re-architected 2026-06): each box runs a local-SQLite bus + a courier/receiver for P2P over the 10.x net. The v0.15 lean-bus cut RIPPED moderation / chamber (UI 7881) / forseti-daemon / agent-runner / redis-valkey. | per-box `uv tool install` (NOT CI-deploy); **nh3-extdev** added as a mesh peer 2026-06-25 (model B: althing-svc svc account + shared `/srv/althing`) | | `vh/mead-hall` | Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) | | `vh/skaldsong` | Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) | | `vh/Worldtree` | Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. **gitea-runner builds on ana-docker** (the image-bloat source); claude-bot now ADMIN collaborator (2026-06-20) | push-to-main → CI build-and-deploy (runner on ana-docker) | | `vh/yt-voice-clipper` | YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → **gitea-webhook auto-deploy** to irv-ml1 (2026-06-03) — see `docs/runbooks/ytvc-autodeploy.md` | | `vh/arbo` | Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) | push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook | (`vh/volva` + Heid were re-architected from systemd daemons to Claude Code session orchestrators 2026-06-08; their nh3-dev `.service` units were removed — no longer deployed sidecars here. See Recent decisions.) - **Two-layer backups** — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see `docs/runbooks/disaster-recovery.md` for the blast-radius matrix. **⚠️ The restic file+DB layer routes through TWO rest-servers** (`rest-server-ana` @ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas; `rest-server-nh3` @ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of `/mnt/backup`. (rest-server-ana recovered 2026-06-20.) - **`pull-hf-repo.yaml`** is the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at `/tank/aimodels/huggingface/`" playbook. Supports `--var repo_type=model|dataset|space`. Replaces ad-hoc `huggingface_hub.snapshot_download` patterns. - **Worldtree admin auth — per-instance.** Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (`key_id 61419c92`) at `ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin` auths against **demo only**. Personal-instance admin (the `~/.config/worldtree/personal-admin-token`, mode 600) POSTs `/admin/keys` (mints per-project keys; takes `user_id`+`label`, **no scope param** — scopes are tier-derived, so admin-tier scopes like `admin.events.read` must be granted WT-side by worldtree-dev). - **Per-project user keys against personal Worldtree** (issued 2026-05-19): `skaldsong:79744637` (nh3-dev iteration), `skaldsong:7c1dbbbe` (ana-docker prod), `althing:50d85460`, `mead-hall:a360822d`. Same `user_id=skaldsong` across both skaldsong keys → shared Heimdall agent slot; different `key_id` → independently rotatable. Pattern: mint via `/admin/keys`, drop value to `/tmp/wt-personal-.key` mode 600, dev collects + shreds (DO NOT cat to chat transcript). - **Skaldsong CD pattern (registry-pull).** Differs from althing / asset-engine which build-on-host. vh/skaldsong's CI builds and pushes `gitea.phasefinal.com/vh/skaldsong:` + `:latest`; `playbooks/deploy-skaldsong.yaml` on ana-docker pulls + recreates. SHA-pin only (no `:latest` health-gated advance yet). Prereq: host needs `docker login gitea.phasefinal.com` once (read:package PAT) — not currently in the workflow. - **gitea internal route for fleet hosts.** gitea is a container on **ana-docker** — git-SSH `10.250.50.70:222`, HTTP `:3000`. Fleet/colo hosts must use this internal route, NOT public `gitea.phasefinal.com` (`38.120.12.44`, ana-srv1) — the public path fail2bans the host egress IP and wedges webhook deploys. `:22` on `10.250.50.70` is ana-docker's HOST sshd, not gitea. Full gotcha in `docs/orientation.md` → Git/gitea. - **docker-as-root pattern** (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo): `docker run --rm -v :/wt [-v /var/run/docker.sock:/var/run/docker.sock] docker:cli sh -c "..."` (or `alpine` for plain file ops). docker-group membership is effectively root via bind-mount; treat as sudo-equivalent. **Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the Docker daemon interprets them against the HOST fs — always pass `-e VAR=/abs/path` for any relative-default config dir.** - **`scripts/elway` sudo handling** — elway prompts for the sudo password ONCE via `getpass` before the first `sudo: true` step → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH. - **Per-host SSH identity matters for sudo.** infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker there are TWO identities: the **default `ssh ana-docker` = `lkraven`** (docker-group, NO passwordless sudo — docker works, root-file edits don't); **but `ssh infra-ops@ana-docker` HAS NOPASSWD root** (verified 2026-06-20). **→ For any sudo op on ana-docker (mount, root-owned files, service control), use `ssh infra-ops@ana-docker`, NOT the default session.** lkraven-owned files (litellm config, most stack compose/conf) still take plain `cp`/edit under either identity. **NEW (2026-06-25):** `ssh infra-ops@10.100.10.50` (nh3-dev) ALSO has NOPASSWD sudo; on **nh3-extdev** infra-ops is sudo-LESS by design and **`ssh lkraven@10.100.50.42` is the NOPASSWD path** there. ## Current state / in-flight _As of 2026-06-25:_ - **Worldtree persona-render config arc COMPLETE (demo + personal).** Pre-synced + deployed three bind-mount config deltas on corviduo-dev demo+personal: #314 relational-stance (relational_extractor role + relational_stance block + stance/ dir), #322 affect/mood gate (affect-render-baseline-allow policy rule + mood_tier_deployment_cap), #317 valence teardown (REMOVED relational_valence — a **boot-blocking removal**). All deployed green; the full two-layer persona render (relational + mood) is live on demo+personal; Worldtree at v0.37.38. Recipe + the additions-vs-removals lesson in auto-memory `reference_corviduo_dev_emergency_ops`. - **nh3-extdev (10.100.50.42) stood up as an althing v0.17.1 mesh PEER — MODEL B.** Full provision in one session: system zellij + althing (`/usr/local/bin`, both users), CLAUDE.md, and the mesh (handles `mailman`+`ldp-dev`, ↔nh3-dev). MODEL B (operator's multi-user choice): receiver runs as the **`althing-svc`** service account + a group-shared **`/srv/althing`** root (group `althing`; infra-ops+lkraven members; `ALTHING_ROOT=/srv/althing` via `/etc/profile.d/althing.sh`). Live + verified (two OS users sharing ONE config/DB). Layout + access facts in auto-memory `reference_nh3_extdev_althing_mesh`. Open: forseti's multi-user ADR. - **zellij native web client piloted on nh3-dev** (`zellij-web.service` :8443, native TLS) ALONGSIDE the ttyd seats (ttyd KEPT as the no-auth fallback). Verdict: wins on clipboard/auth/TLS/simplicity, loses device-independent fonts; bind left `0.0.0.0` per operator. auto-memory `reference_zellij_web_seat`. - **Backups — recovered + hardened (2026-06-20), STILL OPEN:** rotate the 5 disclosed rest-server creds (operator, offline); confirm esh-vm-db's resticprofile includes DB dumps (esh-pve-nas restic-only). Topology + 2-min check: `docs/runbooks/backups.md`. - **R22 (brokkr/dwarves) — CONCLUDED, gateway-only.** Phase B (internal-model hooks) CANCELLED (Worldtree model-agnostic → no deploy path). Whole footprint = gateway calls to free `qwen3.5-122-a10b` via a PERSISTENT full-access key at `/home/lkraven/.r22-gateway-key` (mode 600, carries paid GLM, do NOT delete). Detail in Recent decisions. - **Open follow-ups (low-priority):** clean phantom `qwen3.6-35b-a3b` off the gateway (lists in /v1/models but 400s — backend displaced); docker-daemon default log-cap on ana-docker (systemic fix behind the 94 GB disk incident). - **Standing / parked (from prior):** Mac Pro migration (`migration-plan.md`, hardware-gated); R17 v2 corpus push HELD (distribution barred); disclosed-keys rotation queue; clean legacy `news-digest` on ana-docker. - **Worldtree config-propagation (reference):** demo+personal bind-mount config (`providers.yaml`/`model_roles.yaml`/`policies.yaml`/`defaults.yaml`, now incl. the #314/#322/#317 additions) from `/opt/worldtree{,-personal}/config` (infra-ops-deployable, byte-identical from worldtree-dev canonical sha); pinned is IN-IMAGE. Reload via `docker restart `, NEVER `compose up` (stale-`:latest` footgun); WT PDP is rule-based (a tier needs allow RULES, not just scopes). **#317 lesson: config REMOVALS are NOT backward-compatible with the still-running image — push promptly, don't bounce the instance in the window.** ## Recent decisions - `[2026-06-25]` **althing re-architected to the v0.17 lean multi-machine bus; nh3-extdev stood up as a mesh peer (MODEL B).** v0.15.0's lean-bus cut RIPPED moderation/chamber/forseti-daemon/agent-runner/redis-valkey; **v0.17 = per-box local-SQLite bus + a courier/receiver for P2P over the 10.x net** (installed per-box via `uv tool install`, NOT CI). Migrated nh3-dev's client v0.15→v0.17. nh3-extdev provisioned as a full mesh peer with the operator's **MODEL B** (dedicated `althing-svc` service account + group-shared `/srv/althing` root, so multiple OS users share one config/DB; `ALTHING_ROOT` via /etc/profile.d). Verified multi-user concurrent rw. Supersedes the old althing Tools-row (was chamber/daemons/valkey). auto-memory `reference_nh3_extdev_althing_mesh`. - `[2026-06-23]` **zellij native web client piloted on nh3-dev** (`zellij-web.service` :8443, native TLS), system-installed for all users, ALONGSIDE ttyd (kept as the no-auth fallback). Bind left `0.0.0.0` per operator. auto-memory `reference_zellij_web_seat`. - `[2026-06-22]` **Worldtree persona-render config arc (#314 relational-stance / #322 affect-mood gate / #317 valence teardown) pre-synced + deployed green on demo+personal** — #317 a boot-blocking config REMOVAL (relational_valence). Config-sync recipe + additions-vs-removals lesson in auto-memory `reference_corviduo_dev_emergency_ops`. - `[2026-06-20]` **R22 (brokkr/eitri/dwarves) stood down to gateway; full-access R22 key minted; Phase B parked.** Operator caught the raw-weights/GPU ask as premature — Phase A (P00 + NL-rendering) is PROMPT-LEVEL on the LiteLLM gateway, not raw weights (no GPU, no Blackwell displacement). Real blocker was gateway ACCESS: the shared `all-agents-local` key only serves granite chat, and granite-8b is too weak to be the P00 model-under-test (capability-confound). Operator said "full access" → minted `r22-brokkr-phaseA` LiteLLM key (scope **all-proxy-models**, incl PAID GLM), dropped `~/.r22-gateway-key` (mode 600) for eitri — **re-minted PERSISTENT** after eitri's stateless tool-env couldn't hold it in env + shredded the one-shot drop (old key revoked); the live full-access key now lives at `/home/lkraven/.r22-gateway-key`, read per-command, **don't delete** (cost surface — carries paid GLM; operator reconfirmed full-open + declined to narrow, 2026-06-20 — settled, don't re-surface). Lesson: stateless consumers need a persistent mode-600 file read per-call, NOT shred-after-load. **P00 MUT = `qwen3.5-122-a10b` (alias `gen`)** — 122B, capable, FREE-LOCAL (no per-token cost); told them keep frozen calibration on it, NOT the paid GLM the key can now reach. **PHASE B** (raw-weights hooks: Dvalin activation-steering on hidden states + Regin ECM logit-injection on a 7-8B Qwen/Llama via transformers LogitsProcessor) **CANCELLED 2026-06-20 (not deferred)** — operator cut the internals arms: **Worldtree is model-agnostic** (regard/ affect renders via context manipulation = prompt-level only), so the internal-model approach has no deploy path (the pragmatic-outcome gate firing). R22's ENTIRE compute is now the gateway (free qwen3.5-122-a10b MUT); NO raw-weights/GPU ever; the A6000 slice reservation is released (was never provisioned). The full-scope R22 key stays (pinned to the free qwen-122, GLM untouched). **OPERATOR STEER (2026-06-20): R22 is research for a PRAGMATIC/deployable outcome, NOT advancing-the-art.** Before the dwarves invest in Phase B, establish a deploy path (distilled steering-vector / fine-tune applied to a model we already serve). Operator is taking the dwarf discussion directly. - `[2026-06-20]` **claude-bot issue-scope token minted for worldtree-dev self-serve** (closes their last `tea`-as-vh fallback — issues; they already self-serve Actions + deploys). Gitea PAT scopes are IMMUTABLE → minted a NEW claude-bot PAT `worldtree-dev-actions-issues-20260621` (id 16) with `write:repository`+`write:issue` via basic-auth (`claude-bot:gitea-password`, POST `/users/claude-bot/tokens`); dropped mode-600 `~/.claude-bot-token-issues` → worldtree-dev swaps. **Gitea COLLAPSES `read:issue` into `write:issue`** (write implies read). Old token (id 15) REVOKED after verification → claude-bot carries ONE consolidated worldtree-dev token (id 16) + arbo-ci (id 14). Advances the credential-migration directive. - `[2026-06-20]` **rest-server-ana recovered + backup prevention shipped + worldtree-dev admin keys provisioned** (operator-directed). rest-server-ana fixed (mount-a + force-recreate via `infra-ops@ana-docker`); freshness alert + fstab hardening landed (see `docs/runbooks/backups.md`). Minted worldtree-dev **admin-tier Heimdall keys** on demo (key_id d113207c) + personal (f4f75adb), dropped mode-600 to `~/.wt-admin-{demo,personal}` on nh3-dev → worldtree-dev now self-serves key minting. Cred rotation (5 rest-server pw) BELAYED per operator. - `[2026-06-20]` **claude-bot → ADMIN on vh/Worldtree** (operator-authorized; one-time use of operator vh-admin to enable the migration) — claude-bot self-serves WT deploys/tokens henceforth, reducing personal-cred fallbacks. - `[2026-06-14]` **STANDING DIRECTIVE: migrate ALL infra access to Claude-specific credentials.** Stand up service accounts (claude-bot Gitea user + scoped tokens, distinct litellm admin key); re-mint consumer creds under them; flag personal-cred fallbacks until done. (auto-memory `project_migrate_infra_access_to_claude_credentials`) _118 older entries archived to archival-memory.md._ ## Tried and abandoned - `[2026-06-25]` **althing "unreachable: — retry later" can MASK an app-level 500.** A multi-hour "intermittent :8087 / firewall flap" hunt was a red herring: raw network was always clean (curl POST to the receiver `:8087` worked; a connect probe = 0 fails). Root cause = the receiver's DB agents-table wasn't synced with the config roster, so delivery 500'd `"unknown to: "`, which the courier MAPPED to "unreachable" (looks like network/DNS). **Diagnose: raw curl to :8087 + a connect-probe pass ⇒ NOT network — look up-stack.** Fixed in althing **v0.17.1** (receiver auto-ensures the recipient on delivery). Stop-gap on older builds: `ALTHING_HANDLE= althing-cli inbox` as the receiver user syncs config→DB; re-run after any roster change. auto-memory `reference_nh3_extdev_althing_mesh`. - `[2026-06-20]` **rest-server `.htpasswd: permission denied` = the ana-nas NFS mount FAILED (ghost file on the local mount point), NOT a decommission.** `mnt-backup.mount` stuck `failed` (fstab bare `defaults`, no retry) → the rest-server serves an empty local dir with a root:root 0-byte `.htpasswd`. Documented recovery in disaster-recovery.md. Don't tear down what looks like a crash-looping legacy container until you've checked fstab + the docs — it was the live ana-side restic target. - `[2026-06-20]` **The DEFAULT `ssh ana-docker` is `lkraven` (no NOPASSWD) — but `ssh infra-ops@ana-docker` HAS NOPASSWD root** (corrected later same day). Early on a `sudo cp` as lkraven silently failed (password prompt) → one wasted gateway bounce; I then wrongly concluded "no sudo on ana-docker" and nearly punted the rest-server-ana recovery to the operator. The real rule: reach for `infra-ops@ana-docker` for sudo ops; lkraven-owned files (litellm config, most stack compose/conf) take plain `cp` under either identity. _98 older entries archived to archival-memory.md._