Files
esh-pfi-infrastructure/persistent-memory.md
T

292 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Persistent memory — eshpfi-management
_Last updated: 2026-07-05_
## Repo purpose
Reference workspace for PFI infrastructure: server inventory, canonical
Docker Compose stacks, ops playbooks, and conventions. Authoritative
copies of compose files live on the servers under
`/opt/docker/compose/<stack>/`; this repo mirrors them for version
control, editing, planning, and CI-driven deploys. **It was originally
spun up to handle the fleet backups** — keep that lens when triaging
backup/storage issues.
## Tools and conventions
Sister repos (separate gitea repos, deployed by playbooks here):
| Repo | Role | CI status |
|---|---|---|
| `vh/task-board` | MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) |
| `vh/vor` | Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
| `vh/nevermore` | Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
| `vh/asset-engine` | Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
| `vh/althing` | Lean trusted inter-agent message bus — **v2 "email model" (v2.0.0b2, 2026-07)**: per-box local-SQLite bus + courier/receiver for P2P over the 10.x net; pillars = open-loops / per-box herald + wake-listener / roaming owner API `/owner/*` / `althing-mcp` stdio surface. The v0.15 lean-bus cut RIPPED moderation / chamber / forseti-daemon / agent-runner / redis-valkey. | per-box `uv tool install` (NOT CI-deploy); **nh3-dev = the DEV box** (editable install of `~/development/althing`, gets new versions first); **nh3-extdev** a mesh peer (model B: althing-svc + shared `/srv/althing`) |
| `vh/mead-hall` | Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
| `vh/skaldsong` | Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
| `vh/Worldtree` | Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. **gitea-runner builds on ana-docker**; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. | push-to-main → CI build-and-deploy (runner on ana-docker) |
| `vh/yt-voice-clipper` | YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → **gitea-webhook auto-deploy** to irv-ml1 (2026-06-03) — see `docs/runbooks/ytvc-autodeploy.md` |
| `vh/arbo` | Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) | push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook |
| `model-training-forge` (mtf-dev) | Fine-tuning recipe forge; **T1 = qwopus E-RP writing LoRA** (SFT→DPO, LitBench-RM reward) | training runs, not a deployed sidecar |
(`vh/volva` + Heid were re-architected from systemd daemons to Claude Code
session orchestrators 2026-06-08; their nh3-dev `.service` units were removed —
no longer deployed sidecars here. See Recent decisions.)
- **Two-layer backups** — Backrest orchestrates restic for file+DB (5
fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for
VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore +
cross-site restic targets — see `docs/runbooks/disaster-recovery.md`
for the blast-radius matrix. **⚠️ The restic file+DB layer routes
through TWO rest-servers** (`rest-server-ana` @ ana-docker:8000 →
ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas; `rest-server-nh3` @
nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS
export of `/mnt/backup`. (rest-server-ana recovered 2026-06-20.)
- **`pull-hf-repo.yaml`** is the canonical "get a HuggingFace
model/dataset onto ana-ml2's shared cache at
`/tank/aimodels/huggingface/`" playbook. Supports `--var repo_type=model|dataset|space`. Replaces ad-hoc `huggingface_hub.snapshot_download` patterns.
- **Worldtree admin auth — per-instance.** Each Worldtree deployment
(demo :8080, personal :8081, pinned :8082) has its own Heimdall
registry and its own bootstrap admin key. Infra-ops's stored
long-lived admin key (`key_id 61419c92`) at
`ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin`
auths against **demo only**. Personal-instance admin (the
`~/.config/worldtree/personal-admin-token`, mode 600) POSTs
`/admin/keys` (mints per-project keys; takes `user_id`+`label`, **no
scope param** — scopes are tier-derived). **On-instance mint recipe
(cleaner than DB-manip):** `docker exec worldtree-worldtree-api-1` POST
`/admin/keys` with the in-container `WORLDTREE_BOOTSTRAP_ADMIN_KEY`; cleartext
once in `.key`=`wt_live_+16hex`. auto-memory `reference_worldtree_demo_key_mint`.
- **Per-project user keys against personal Worldtree** (issued
2026-05-19): `skaldsong:79744637`, `skaldsong:7c1dbbbe`,
`althing:50d85460`, `mead-hall:a360822d`. Mint via `/admin/keys`, drop
value to `/tmp/wt-personal-<name>.key` mode 600, dev collects +
shreds (DO NOT cat to chat transcript).
- **Skaldsong CD pattern (registry-pull).** vh/skaldsong's CI builds and
pushes `gitea.phasefinal.com/vh/skaldsong:<sha>` + `:latest`;
`playbooks/deploy-skaldsong.yaml` on ana-docker pulls + recreates.
SHA-pin only. Prereq: host needs `docker login gitea.phasefinal.com` once.
- **gitea internal route for fleet hosts.** gitea is a container on
**ana-docker** — git-SSH `10.250.50.70:222`, HTTP `:3000`. Fleet/colo
hosts must use this internal route, NOT public `gitea.phasefinal.com`
(`38.120.12.44`) — the public path fail2bans the host egress IP. Full
gotcha in `docs/orientation.md` → Git/gitea.
- **docker-as-root pattern** (for ops with no admin API, or to edit
deploy-owned/root-owned files without sudo): `docker run --rm -v
<target-dir>:/wt docker:cli sh -c "..."`. docker-group membership is
effectively root via bind-mount. **Foot-gun: relative paths in compose.yaml
resolve against the sandbox CWD but the daemon interprets them against the
HOST fs — always pass `-e VAR=/abs/path` for any relative-default config dir.**
- **`scripts/elway` sudo handling** — elway prompts for the sudo password
ONCE via `getpass` before the first `sudo: true` step → can't run
unattended from a non-TTY tool if any step needs sudo. Sudo-free
playbooks run fully non-interactive over key SSH.
- **Per-host SSH identity matters for sudo.** infra-ops has NOPASSWD sudo
on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On
ana-docker: **default `ssh ana-docker` = `lkraven`** (docker-group, NO
passwordless sudo); **`ssh infra-ops@ana-docker` HAS NOPASSWD root**. **
For any sudo op on ana-docker, use `ssh infra-ops@ana-docker`.** `ssh
infra-ops@10.100.10.50` (nh3-dev) ALSO NOPASSWD sudo; on **nh3-extdev** infra-ops
is sudo-LESS by design (`ssh lkraven@10.100.50.42` is the NOPASSWD path). **irv-ml1:
`ssh irv-ml1` = lkraven, docker-group (plain docker) but sudo needs a PASSWORD
(no NOPASSWD)** — stage model pulls to `/home`, not root-owned `/worktank`.
## Current state / in-flight
_As of 2026-07-05:_
- **T1 TRAIN — venue decision live; NEXT ACTION = the ana-ml2 SMOKE.** T1 (qwopus E-RP writing
LoRA) datasets + recipe + smoke set are all staged (mtf-dev, on `nh3-dev:data/export/qwen-3.5-122b-erp-lora/`
incl. `smoke/`); bf16 base is on ana-ml2 `/tank/aimodels/qwopus3.5-122b-a10b-bf16` (233G, PUBLIC
HF repo, zero-auth pull). On-prem RULED OUT except a full ana-ml2 shutdown → CPU-offload → ~1-day
fleet outage; cloud (Vast.ai 8×80GB, ~3-6h, ~$60-500, zero impact) is the clean alt. **Operator
chose smoke-first on ana-ml2** to measure real samples/sec before committing. ⚠️ the smoke ITSELF
needs both GPUs + the ~421G vLLM RAM freed (192G VRAM < 244G base → CPU offload) = a brief full
ana-ml2 serving pause; keeping any serving up drops it into the ~week NVMe regime. Full runbook +
both paths in `reference_t1_cloud_train_plan`. Post-train serve-path (swappable-LoRA-on-NVFP4
test) still queued: `reference_gen_qwopus_122b`.
- **LitBench-RM UP (irv-ml1 A6000); comfyui DISPLACED.** Respun the R19 creative-writing reward
(`vllm-litbench-rm` :8202, ~19.6G) for T1's reward ensemble; stopped comfyui to free the A6000
(operator is sole comfyui consumer + holding image-gen until LitBench done). Score scale
re-verified from nh3-dev (literary 0.95 > flat 0.44 > slop 0.12). **TEARDOWN on "litbench done":**
`ssh irv-ml1 'docker rm -f vllm-litbench-rm && docker start comfyui'` — ⚠️ then watch the mmartial
boot trap (root-owned venv → crash-loop; `chown -R 1000:1000 /comfy/mnt/venv`). `reference_litbench_rm_irv_ml1`.
- **character-rp role — SHIPPED; one QUEUED spot-check.** Proved per-request `extra_body`
(top_k/repetition_penalty) forwards through the `gen-reasoning` alias to the vLLM sampler (no
gateway cap needed); pre-staged the `character-rp` role in demo+personal bind-mount
`model_roles.yaml` (byte-verified live on b18; caught the cached-registry ordering so the deploy's
own restart activates it). worldtree-dev shipped **#344 (v1.0.0b19)** fixing the durable-agent
override-drop (role now applies the RP profile). QUEUED: ratatoskr-dev pings a turn-timestamp → I
spot-check the gateway spend_logs to confirm the RP sampler params (temp 0.75 / top_p 0.85 /
presence 1.5 / top_k 20 / rep 1.0) reach vLLM.
- **althing v2 (herald + receiver) = systemd-supervised on nh3-dev (the althing DEV box).**
Formalized this session: `althing-herald.service` (Restart=always, **Environment=PATH incl
~/.cargo/bin** — the fix for the silent pane-dispatch outage) + `althing-receiver.service` (v2 →
pillar-3 `/owner/*` live). Stale forseti unit removed; both on v2.0.0b2, canonicalized by forseti.
`reference_nh3_dev_althing_herald`.
- **glm-5.2 = z.ai 1M-in / 128K-out, NO gateway cap** (probed z.ai; recorded in the config comment
+ `reference_litellm_gateway`; commit 624a07e). All GLM entries are pure z.ai passthrough — the
effective ceiling is z.ai's canonical, not a gateway limit.
- **`/books` mounted (transient) on nh3-dev** (`10.0.50.50:/mnt/books` → NFSv4 ro,soft; re-mount via
`infra-ops@10.100.10.50` if it reboots).
- **Backups — recovered + hardened (2026-06-20), STILL OPEN:** rotate the 5 disclosed rest-server
creds (operator, offline); confirm esh-vm-db's resticprofile includes DB dumps. `docs/runbooks/backups.md`.
- **Open follow-ups (low-priority):** clean phantom `qwen3.6-35b-a3b` off the gateway (lists in
/v1/models but 400s); docker-daemon default log-cap on ana-docker; drop the now-inert
`mood.decay_rate`/`mood.stale_hours` keys from the deployed `/opt/worldtree*/config` bind-mounts
(harmless-present per worldtree-dev; edit-only/no-restart at a no-deploy window).
- **Standing / parked:** Mac Pro migration (hardware-gated); R17 v2 corpus HELD; disclosed-keys
rotation queue; clean legacy `news-digest`; R22 gateway-only full-access key at
`/home/lkraven/.r22-gateway-key` (mode 600, carries paid GLM, do NOT delete); granite-4.1-8b bf16
at `irv-ml1:/home/lkraven/granite-4.1-8b-bf16` (reap-on-request); Deckard staged on ana-ml2 as
T1's writing benchmark.
- **Worldtree config-propagation (reference):** demo+personal bind-mount config from
`/opt/worldtree{,-personal}/config` (infra-ops-deployable, byte-identical from canonical); reload
via `docker restart <container>`, NEVER `compose up` (stale-`:latest` footgun). **The role registry
is loaded ONCE + CACHED at startup** (`ConversationService.role_registry`) → a bind-mount
`model_roles.yaml` change needs a container restart to take effect; pre-stage the bind-mount
BEFORE the activating deploy's restart (else it needs a second bounce). config REMOVALS are NOT
backward-compatible with the still-running image.
## Recent decisions
- `[2026-07-05]` **T1 training venue: CLOUD recommended; operator chose smoke-first on ana-ml2.**
On-prem ruled out (ana-ml2 full — both 96G GPUs ~93G used): keep-serving = NVMe offload ~6-8 DAYS;
full ana-ml2 shutdown = CPU offload ~1 DAY but a whole-fleet outage. Cloud Vast.ai 8×80GB (no
offload → ~3-6h, ~$60-500, zero fleet impact) is the clean alt (mtf-dev + infra-ops both rec;
Vast for its no-content-AUP marketplace + likely-existing VastBlue account). Operator's next step
= the ana-ml2 CPU-offload SMOKE (~60 steps) to get real samples/sec before the full-outage-vs-cloud
call. HF base verified public (zero-auth pull). Runbook + gotchas in `reference_t1_cloud_train_plan`.
- `[2026-07-05]` **glm-5.2 canonical limits recorded** (probed live vs z.ai): **1,048,576 (1M) input
context / 131,072 (128K) max output**; NO gateway-side cap (pure passthrough → z.ai's limits are
effective). Written to the config comment (commit `624a07e`) + `reference_litellm_gateway`.
- `[2026-07-04]` **character-rp: gateway-forwarding proven + role pre-staged + #344 shipped.**
Empirically confirmed per-request `extra_body` (top_k/repetition_penalty) forwards through the
`gen-reasoning` LiteLLM alias to vLLM + standard params override the alias defaults — no gateway
cap needed (I over-built a dedicated alias, operator corrected, reverted with zero fleet impact).
Pre-staged the `character-rp` role into demo+personal bind-mount `model_roles.yaml` (byte-verified
on b18; caught the cached-registry ordering). worldtree-dev shipped **#344 (v1.0.0b19)** for the
durable-agent override-drop. spend_logs spot-check queued (ratatoskr's timestamp ping).
- `[2026-07-04]` **althing v2 herald+receiver formalized as systemd on nh3-dev.** `althing-herald.service`
(Restart=always, **Environment=PATH incl ~/.cargo/bin** — the pane-dispatch fix) + `althing-receiver.service`
(v2 → pillar-3 `/owner/*` live); stale forseti unit removed; both on v2.0.0b2, canonicalized by
forseti. `reference_nh3_dev_althing_herald`.
- `[2026-07-04]` **LitBench-RM respun (irv-ml1 A6000, comfyui displaced)** for T1's reward ensemble;
operator sole comfyui consumer, holding image-gen until LitBench done. `reference_litbench_rm_irv_ml1`.
- `[2026-07-03]` **ratatoskr-dev DEMO Heimdall key provisioned (R30 φ0).** Minted a tier-user key on
the demo via `POST /admin/keys` (bootstrap admin key), mirroring their personal base consumer (no
character-binding); base-agent affect reads work ungated. `reference_worldtree_demo_key_mint`.
- `[2026-07-02]` **mtf-dev granite harness-spike ran GREEN — MECHANICAL only, efficacy DEFERRED to
the T1 run.** Trainer TRL SFT→DPO→eval seam proven end-to-end on a synthetic fixture (not the E-RP
corpus); operator DECIDED no intermediate real-efficacy granite spike (uninterpretable proxy —
arch gap + abliteration axis). `reference_gen_qwopus_122b`.
- `[2026-07-01]` **Worldtree #332 embed-recall diagnosed + scoped-log view/tunnel provisioned + fix
verified** (15×→1.01× re-embed). `reference_wt_gateway_scoped_log_view`.
- `[2026-07-01]` **qwopus native MTP speculative-decode tested on `gen` → NOT kept** (+12% single-stream,
1520% aggregate at concurrency, silently drops min_p/logit_bias). Banked for T1. `reference_gen_qwopus_122b`.
- `[2026-07-01]` **Deckard trial → reverted to qwopus (`gen`)** (won writing "in every way" but ~36 vs
~90 tok/s; spec-decode rescue ruled out). git `b63c48b``681eb70`. Deckard kept staged as T1's
writing benchmark.
- `[2026-06-25]` **althing re-architected to the lean multi-machine bus; nh3-extdev stood up as a
MODEL B mesh peer** (dedicated `althing-svc` + group-shared `/srv/althing`). `reference_nh3_extdev_althing_mesh`.
- `[2026-06-23]` **zellij native web client piloted on nh3-dev** (`zellij-web.service` :8443)
alongside ttyd. `reference_zellij_web_seat`.
- `[2026-06-22]` **Worldtree persona-render config arc (#314/#322/#317) pre-synced + deployed green
on demo+personal** — #317 a boot-blocking config REMOVAL. `reference_corviduo_dev_emergency_ops`.
- `[2026-06-20]` **R22 (brokkr/dwarves) stood down to gateway-only; full-access R22 key minted;
Phase B CANCELLED** (Worldtree model-agnostic → no deploy path). Key at `/home/lkraven/.r22-gateway-key`
(persistent mode-600, carries paid GLM, don't delete). MUT = free `qwen3.5-122-a10b` (`gen`).
Operator steer: R22 research is gated on a pragmatic/deployable outcome, not advancing-the-art.
- `[2026-06-20]` **claude-bot issue-scope token minted for worldtree-dev self-serve** (id 16,
`write:repository`+`write:issue`); old token revoked. Advances the credential-migration directive.
- `[2026-06-20]` **rest-server-ana recovered + backup prevention shipped + worldtree-dev admin keys
provisioned** (demo d113207c / personal f4f75adb). Cred rotation (5 rest-server pw) BELAYED.
- `[2026-06-20]` **claude-bot → ADMIN on vh/Worldtree** (operator-authorized) — self-serves WT
deploys/tokens henceforth.
- `[2026-06-14]` **STANDING DIRECTIVE: migrate ALL infra access to Claude-specific credentials.**
(auto-memory `project_migrate_infra_access_to_claude_credentials`)
_118 older entries archived to archival-memory.md._
## Tried and abandoned
- `[2026-07-04]` **LiteLLM (this gateway version) mutates the SHARED deployment config in-place on
per-request sampler-param merge** → my deliberately-invalid `top_k=-5` forwarding-probe bled into a
param-less character-rp request (vLLM 400, ONE-OFF, self-cleared by a later valid probe). NOT
caching (none configured), NOT a config change. **Never fire invalid/distinctive sampler values at
a SHARED gateway alias with live consumers** — use a throwaway alias, or a `docker restart litellm`
flushes residual carryover. `feedback_litellm_shared_param_mutation`.
- `[2026-07-04]` **A systemd `--user` daemon that shells out to `~/.cargo/bin`/`~/.local/bin` tools
needs an explicit `Environment=PATH`** — the minimal `--user` default silently drops them. The
althing herald lost `zellij` → silent `pane-miss` for ALL config-backed TUI/pane agents; CC + FIFO
routes were unaffected, so it was invisible from a CC session. `reference_nh3_dev_althing_herald`.
- `[2026-07-04]` **On-prem T1 train that keeps ANY ana-ml2 serving up = ~6-8 DAYS** (1-GPU + NVMe
ZeRO-Infinity offload; MoE ~10B-active cuts FLOPs but NOT the 244G base's param I/O). The only fast
on-prem path is a FULL ana-ml2 shutdown (both GPUs + the ~421G vLLM RAM freed → base fits in the
566G CPU RAM) → CPU offload → ~1-day full-fleet outage. Cloud (no offload) = hours. `reference_t1_cloud_train_plan`.
- `[2026-07-01]` **A personal-Worldtree CI deploy that fails ~85s in with "not found / unauthorized"
is usually the pull-only-vs-build RACE, not registry-auth.** `deploy-personal.yml` is PULL-ONLY but
fires on the `staging/vX` tag simultaneously with `deploy.yml`'s build → pulls before the push
finishes. FIX: re-run once built, or gate on `workflow_run: completed`.
- `[2026-07-01]` **MTP/spec-decode on a SHARED serving model helps single-stream but HURTS
moderate-concurrency aggregate + silently ignores `min_p`/`logit_bias`** (qwopus `gen`: N=1 +12%,
N=4 20%). Reserve for dedicated/interactive deployments.
- `[2026-07-02]` **irv-ml1 `/worktank` ROOT is root-owned — lkraven can't write there (irv-ml1 sudo
needs a password) → stage model pulls to `/home`.** PIN THE A6000 BY UUID for training (native-CUDA
ordering differs vs docker; the 3090 index 0 is usually near-full → OOM). `CUDA_VISIBLE_DEVICES=GPU-<uuid>`.
- `[2026-06-25]` **althing "unreachable: <machine>" can MASK an app-level 500.** Raw network was
clean; root cause = receiver DB agents-table not synced with the config roster → delivery 500'd
"unknown to: <handle>", MAPPED to "unreachable". Diagnose: raw curl to :8087 + connect-probe ⇒ NOT
network. Fixed in althing v0.17.1. `reference_nh3_extdev_althing_mesh`.
- `[2026-06-20]` **rest-server `.htpasswd: permission denied` = the ana-nas NFS mount FAILED (ghost
file on the local mount point), NOT a decommission.** `mnt-backup.mount` stuck `failed` (fstab bare
`defaults`) → rest-server serves an empty local dir. Recovery in disaster-recovery.md.
- `[2026-06-20]` **The DEFAULT `ssh ana-docker` is `lkraven` (no NOPASSWD) — but `ssh
infra-ops@ana-docker` HAS NOPASSWD root.** A `sudo cp` as lkraven silently failed → nearly punted
the rest-server recovery. Reach for `infra-ops@ana-docker` for sudo ops.
_98 older entries archived to archival-memory.md._