Files
esh-pfi-infrastructure/persistent-memory.md
T

20 KiB
Raw Blame History

Persistent memory — eshpfi-management

Last updated: 2026-07-05

Repo purpose

Reference workspace for PFI infrastructure: server inventory, canonical Docker Compose stacks, ops playbooks, and conventions. Authoritative copies of compose files live on the servers under /opt/docker/compose/<stack>/; this repo mirrors them for version control, editing, planning, and CI-driven deploys. It was originally spun up to handle the fleet backups — keep that lens when triaging backup/storage issues.

Tools and conventions

Sister repos (separate gitea repos, deployed by playbooks here):

Repo Role CI status
vh/task-board MCP + web dashboard for assistant task state (port 7878) push-to-main → CI deploys (2026-04-29)
vh/vor Inquisitor UI sidecar (port 7879) push-to-main → CI deploys (2026-04-29)
vh/nevermore Twice-daily LLM-curated briefing (port 8181, replaces news-digest) push-to-main → CI deploys (2026-04-30)
vh/asset-engine Internal control plane over inference services (port 8200, LAN-direct) push-to-main → CI deploys (2026-05-12)
vh/althing Lean trusted inter-agent message bus — v2 "email model" (v2.0.0b2, 2026-07): per-box local-SQLite bus + courier/receiver for P2P over the 10.x net; pillars = open-loops / per-box herald + wake-listener / roaming owner API /owner/* / althing-mcp stdio surface. The v0.15 lean-bus cut RIPPED moderation / chamber / forseti-daemon / agent-runner / redis-valkey. per-box uv tool install (NOT CI-deploy); nh3-dev = the DEV box (editable install of ~/development/althing, gets new versions first); nh3-extdev a mesh peer (model B: althing-svc + shared /srv/althing)
vh/mead-hall Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) push-to-main → CI deploys (2026-05-16)
vh/skaldsong Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) push-to-main → CI deploys (2026-05-19)
vh/Worldtree Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. gitea-runner builds on ana-docker; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. push-to-main → CI build-and-deploy (runner on ana-docker)
vh/yt-voice-clipper YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md
vh/arbo Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook
model-training-forge (mtf-dev) Fine-tuning recipe forge; T1 = qwopus E-RP writing LoRA (SFT→DPO, LitBench-RM reward) training runs, not a deployed sidecar

(vh/volva + Heid were re-architected from systemd daemons to Claude Code session orchestrators 2026-06-08; their nh3-dev .service units were removed — no longer deployed sidecars here. See Recent decisions.)

  • Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see docs/runbooks/disaster-recovery.md for the blast-radius matrix. ⚠️ The restic file+DB layer routes through TWO rest-servers (rest-server-ana @ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas; rest-server-nh3 @ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of /mnt/backup. (rest-server-ana recovered 2026-06-20.)

  • pull-hf-repo.yaml is the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at /tank/aimodels/huggingface/" playbook. Supports --var repo_type=model|dataset|space. Replaces ad-hoc huggingface_hub.snapshot_download patterns.

  • Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (key_id 61419c92) at ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin auths against demo only. Personal-instance admin (the ~/.config/worldtree/personal-admin-token, mode 600) POSTs /admin/keys (mints per-project keys; takes user_id+label, no scope param — scopes are tier-derived). On-instance mint recipe (cleaner than DB-manip): docker exec worldtree-worldtree-api-1 POST /admin/keys with the in-container WORLDTREE_BOOTSTRAP_ADMIN_KEY; cleartext once in .key=wt_live_+16hex. auto-memory reference_worldtree_demo_key_mint.

  • Per-project user keys against personal Worldtree (issued 2026-05-19): skaldsong:79744637, skaldsong:7c1dbbbe, althing:50d85460, mead-hall:a360822d. Mint via /admin/keys, drop value to /tmp/wt-personal-<name>.key mode 600, dev collects + shreds (DO NOT cat to chat transcript).

  • Skaldsong CD pattern (registry-pull). vh/skaldsong's CI builds and pushes gitea.phasefinal.com/vh/skaldsong:<sha> + :latest; playbooks/deploy-skaldsong.yaml on ana-docker pulls + recreates. SHA-pin only. Prereq: host needs docker login gitea.phasefinal.com once.

  • gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH 10.250.50.70:222, HTTP :3000. Fleet/colo hosts must use this internal route, NOT public gitea.phasefinal.com (38.120.12.44) — the public path fail2bans the host egress IP. Full gotcha in docs/orientation.md → Git/gitea.

  • docker-as-root pattern (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo): docker run --rm -v <target-dir>:/wt docker:cli sh -c "...". docker-group membership is effectively root via bind-mount. Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass -e VAR=/abs/path for any relative-default config dir.

  • scripts/elway sudo handling — elway prompts for the sudo password ONCE via getpass before the first sudo: true step → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH.

  • Per-host SSH identity matters for sudo. infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: default ssh ana-docker = lkraven (docker-group, NO passwordless sudo); ssh infra-ops@ana-docker HAS NOPASSWD root. → For any sudo op on ana-docker, use ssh infra-ops@ana-docker. ssh infra-ops@10.100.10.50 (nh3-dev) ALSO NOPASSWD sudo; on nh3-extdev infra-ops is sudo-LESS by design (ssh lkraven@10.100.50.42 is the NOPASSWD path). irv-ml1: ssh irv-ml1 = lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD) — stage model pulls to /home, not root-owned /worktank.

Current state / in-flight

As of 2026-07-05:

  • T1 TRAIN — venue decision live; NEXT ACTION = the ana-ml2 SMOKE. T1 (qwopus E-RP writing LoRA) datasets + recipe + smoke set are all staged (mtf-dev, on nh3-dev:data/export/qwen-3.5-122b-erp-lora/ incl. smoke/); bf16 base is on ana-ml2 /tank/aimodels/qwopus3.5-122b-a10b-bf16 (233G, PUBLIC HF repo, zero-auth pull). On-prem RULED OUT except a full ana-ml2 shutdown → CPU-offload → ~1-day fleet outage; cloud (Vast.ai 8×80GB, ~3-6h, ~$60-500, zero impact) is the clean alt. Operator chose smoke-first on ana-ml2 to measure real samples/sec before committing. ⚠️ the smoke ITSELF needs both GPUs + the ~421G vLLM RAM freed (192G VRAM < 244G base → CPU offload) = a brief full ana-ml2 serving pause; keeping any serving up drops it into the ~week NVMe regime. Full runbook + both paths in reference_t1_cloud_train_plan. Post-train serve-path (swappable-LoRA-on-NVFP4 test) still queued: reference_gen_qwopus_122b.

  • LitBench-RM UP (irv-ml1 A6000); comfyui DISPLACED. Respun the R19 creative-writing reward (vllm-litbench-rm :8202, ~19.6G) for T1's reward ensemble; stopped comfyui to free the A6000 (operator is sole comfyui consumer + holding image-gen until LitBench done). Score scale re-verified from nh3-dev (literary 0.95 > flat 0.44 > slop 0.12). TEARDOWN on "litbench done": ssh irv-ml1 'docker rm -f vllm-litbench-rm && docker start comfyui'⚠️ then watch the mmartial boot trap (root-owned venv → crash-loop; chown -R 1000:1000 /comfy/mnt/venv). reference_litbench_rm_irv_ml1.

  • character-rp role — SHIPPED; one QUEUED spot-check. Proved per-request extra_body (top_k/repetition_penalty) forwards through the gen-reasoning alias to the vLLM sampler (no gateway cap needed); pre-staged the character-rp role in demo+personal bind-mount model_roles.yaml (byte-verified live on b18; caught the cached-registry ordering so the deploy's own restart activates it). worldtree-dev shipped #344 (v1.0.0b19) fixing the durable-agent override-drop (role now applies the RP profile). QUEUED: ratatoskr-dev pings a turn-timestamp → I spot-check the gateway spend_logs to confirm the RP sampler params (temp 0.75 / top_p 0.85 / presence 1.5 / top_k 20 / rep 1.0) reach vLLM.

  • althing v2 (herald + receiver) = systemd-supervised on nh3-dev (the althing DEV box). Formalized this session: althing-herald.service (Restart=always, Environment=PATH incl ~/.cargo/bin — the fix for the silent pane-dispatch outage) + althing-receiver.service (v2 → pillar-3 /owner/* live). Stale forseti unit removed; both on v2.0.0b2, canonicalized by forseti. reference_nh3_dev_althing_herald.

  • glm-5.2 = z.ai 1M-in / 128K-out, NO gateway cap (probed z.ai; recorded in the config comment

    • reference_litellm_gateway; commit 624a07e). All GLM entries are pure z.ai passthrough — the effective ceiling is z.ai's canonical, not a gateway limit.
  • /books mounted (transient) on nh3-dev (10.0.50.50:/mnt/books → NFSv4 ro,soft; re-mount via infra-ops@10.100.10.50 if it reboots).

  • Backups — recovered + hardened (2026-06-20), STILL OPEN: rotate the 5 disclosed rest-server creds (operator, offline); confirm esh-vm-db's resticprofile includes DB dumps. docs/runbooks/backups.md.

  • Open follow-ups (low-priority): clean phantom qwen3.6-35b-a3b off the gateway (lists in /v1/models but 400s); docker-daemon default log-cap on ana-docker; drop the now-inert mood.decay_rate/mood.stale_hours keys from the deployed /opt/worldtree*/config bind-mounts (harmless-present per worldtree-dev; edit-only/no-restart at a no-deploy window).

  • Standing / parked: Mac Pro migration (hardware-gated); R17 v2 corpus HELD; disclosed-keys rotation queue; clean legacy news-digest; R22 gateway-only full-access key at /home/lkraven/.r22-gateway-key (mode 600, carries paid GLM, do NOT delete); granite-4.1-8b bf16 at irv-ml1:/home/lkraven/granite-4.1-8b-bf16 (reap-on-request); Deckard staged on ana-ml2 as T1's writing benchmark.

  • Worldtree config-propagation (reference): demo+personal bind-mount config from /opt/worldtree{,-personal}/config (infra-ops-deployable, byte-identical from canonical); reload via docker restart <container>, NEVER compose up (stale-:latest footgun). The role registry is loaded ONCE + CACHED at startup (ConversationService.role_registry) → a bind-mount model_roles.yaml change needs a container restart to take effect; pre-stage the bind-mount BEFORE the activating deploy's restart (else it needs a second bounce). config REMOVALS are NOT backward-compatible with the still-running image.

Recent decisions

  • [2026-07-05] T1 training venue: CLOUD recommended; operator chose smoke-first on ana-ml2. On-prem ruled out (ana-ml2 full — both 96G GPUs ~93G used): keep-serving = NVMe offload ~6-8 DAYS; full ana-ml2 shutdown = CPU offload ~1 DAY but a whole-fleet outage. Cloud Vast.ai 8×80GB (no offload → ~3-6h, ~$60-500, zero fleet impact) is the clean alt (mtf-dev + infra-ops both rec; Vast for its no-content-AUP marketplace + likely-existing VastBlue account). Operator's next step = the ana-ml2 CPU-offload SMOKE (~60 steps) to get real samples/sec before the full-outage-vs-cloud call. HF base verified public (zero-auth pull). Runbook + gotchas in reference_t1_cloud_train_plan.

  • [2026-07-05] glm-5.2 canonical limits recorded (probed live vs z.ai): 1,048,576 (1M) input context / 131,072 (128K) max output; NO gateway-side cap (pure passthrough → z.ai's limits are effective). Written to the config comment (commit 624a07e) + reference_litellm_gateway.

  • [2026-07-04] character-rp: gateway-forwarding proven + role pre-staged + #344 shipped. Empirically confirmed per-request extra_body (top_k/repetition_penalty) forwards through the gen-reasoning LiteLLM alias to vLLM + standard params override the alias defaults — no gateway cap needed (I over-built a dedicated alias, operator corrected, reverted with zero fleet impact). Pre-staged the character-rp role into demo+personal bind-mount model_roles.yaml (byte-verified on b18; caught the cached-registry ordering). worldtree-dev shipped #344 (v1.0.0b19) for the durable-agent override-drop. spend_logs spot-check queued (ratatoskr's timestamp ping).

  • [2026-07-04] althing v2 herald+receiver formalized as systemd on nh3-dev. althing-herald.service (Restart=always, Environment=PATH incl ~/.cargo/bin — the pane-dispatch fix) + althing-receiver.service (v2 → pillar-3 /owner/* live); stale forseti unit removed; both on v2.0.0b2, canonicalized by forseti. reference_nh3_dev_althing_herald.

  • [2026-07-04] LitBench-RM respun (irv-ml1 A6000, comfyui displaced) for T1's reward ensemble; operator sole comfyui consumer, holding image-gen until LitBench done. reference_litbench_rm_irv_ml1.

  • [2026-07-03] ratatoskr-dev DEMO Heimdall key provisioned (R30 φ0). Minted a tier-user key on the demo via POST /admin/keys (bootstrap admin key), mirroring their personal base consumer (no character-binding); base-agent affect reads work ungated. reference_worldtree_demo_key_mint.

  • [2026-07-02] mtf-dev granite harness-spike ran GREEN — MECHANICAL only, efficacy DEFERRED to the T1 run. Trainer TRL SFT→DPO→eval seam proven end-to-end on a synthetic fixture (not the E-RP corpus); operator DECIDED no intermediate real-efficacy granite spike (uninterpretable proxy — arch gap + abliteration axis). reference_gen_qwopus_122b.

  • [2026-07-01] Worldtree #332 embed-recall diagnosed + scoped-log view/tunnel provisioned + fix verified (15×→1.01× re-embed). reference_wt_gateway_scoped_log_view.

  • [2026-07-01] qwopus native MTP speculative-decode tested on gen → NOT kept (+12% single-stream, 1520% aggregate at concurrency, silently drops min_p/logit_bias). Banked for T1. reference_gen_qwopus_122b.

  • [2026-07-01] Deckard trial → reverted to qwopus (gen) (won writing "in every way" but ~36 vs ~90 tok/s; spec-decode rescue ruled out). git b63c48b681eb70. Deckard kept staged as T1's writing benchmark.

  • [2026-06-25] althing re-architected to the lean multi-machine bus; nh3-extdev stood up as a MODEL B mesh peer (dedicated althing-svc + group-shared /srv/althing). reference_nh3_extdev_althing_mesh.

  • [2026-06-23] zellij native web client piloted on nh3-dev (zellij-web.service :8443) alongside ttyd. reference_zellij_web_seat.

  • [2026-06-22] Worldtree persona-render config arc (#314/#322/#317) pre-synced + deployed green on demo+personal — #317 a boot-blocking config REMOVAL. reference_corviduo_dev_emergency_ops.

  • [2026-06-20] R22 (brokkr/dwarves) stood down to gateway-only; full-access R22 key minted; Phase B CANCELLED (Worldtree model-agnostic → no deploy path). Key at /home/lkraven/.r22-gateway-key (persistent mode-600, carries paid GLM, don't delete). MUT = free qwen3.5-122-a10b (gen). Operator steer: R22 research is gated on a pragmatic/deployable outcome, not advancing-the-art.

  • [2026-06-20] claude-bot issue-scope token minted for worldtree-dev self-serve (id 16, write:repository+write:issue); old token revoked. Advances the credential-migration directive.

  • [2026-06-20] rest-server-ana recovered + backup prevention shipped + worldtree-dev admin keys provisioned (demo d113207c / personal f4f75adb). Cred rotation (5 rest-server pw) BELAYED.

  • [2026-06-20] claude-bot → ADMIN on vh/Worldtree (operator-authorized) — self-serves WT deploys/tokens henceforth.

  • [2026-06-14] STANDING DIRECTIVE: migrate ALL infra access to Claude-specific credentials. (auto-memory project_migrate_infra_access_to_claude_credentials)

118 older entries archived to archival-memory.md.

Tried and abandoned

  • [2026-07-04] LiteLLM (this gateway version) mutates the SHARED deployment config in-place on per-request sampler-param merge → my deliberately-invalid top_k=-5 forwarding-probe bled into a param-less character-rp request (vLLM 400, ONE-OFF, self-cleared by a later valid probe). NOT caching (none configured), NOT a config change. Never fire invalid/distinctive sampler values at a SHARED gateway alias with live consumers — use a throwaway alias, or a docker restart litellm flushes residual carryover. feedback_litellm_shared_param_mutation.

  • [2026-07-04] A systemd --user daemon that shells out to ~/.cargo/bin/~/.local/bin tools needs an explicit Environment=PATH — the minimal --user default silently drops them. The althing herald lost zellij → silent pane-miss for ALL config-backed TUI/pane agents; CC + FIFO routes were unaffected, so it was invisible from a CC session. reference_nh3_dev_althing_herald.

  • [2026-07-04] On-prem T1 train that keeps ANY ana-ml2 serving up = ~6-8 DAYS (1-GPU + NVMe ZeRO-Infinity offload; MoE ~10B-active cuts FLOPs but NOT the 244G base's param I/O). The only fast on-prem path is a FULL ana-ml2 shutdown (both GPUs + the ~421G vLLM RAM freed → base fits in the 566G CPU RAM) → CPU offload → ~1-day full-fleet outage. Cloud (no offload) = hours. reference_t1_cloud_train_plan.

  • [2026-07-01] A personal-Worldtree CI deploy that fails ~85s in with "not found / unauthorized" is usually the pull-only-vs-build RACE, not registry-auth. deploy-personal.yml is PULL-ONLY but fires on the staging/vX tag simultaneously with deploy.yml's build → pulls before the push finishes. FIX: re-run once built, or gate on workflow_run: completed.

  • [2026-07-01] MTP/spec-decode on a SHARED serving model helps single-stream but HURTS moderate-concurrency aggregate + silently ignores min_p/logit_bias (qwopus gen: N=1 +12%, N=4 20%). Reserve for dedicated/interactive deployments.

  • [2026-07-02] irv-ml1 /worktank ROOT is root-owned — lkraven can't write there (irv-ml1 sudo needs a password) → stage model pulls to /home. PIN THE A6000 BY UUID for training (native-CUDA ordering differs vs docker; the 3090 index 0 is usually near-full → OOM). CUDA_VISIBLE_DEVICES=GPU-<uuid>.

  • [2026-06-25] althing "unreachable: " can MASK an app-level 500. Raw network was clean; root cause = receiver DB agents-table not synced with the config roster → delivery 500'd "unknown to: ", MAPPED to "unreachable". Diagnose: raw curl to :8087 + connect-probe ⇒ NOT network. Fixed in althing v0.17.1. reference_nh3_extdev_althing_mesh.

  • [2026-06-20] rest-server .htpasswd: permission denied = the ana-nas NFS mount FAILED (ghost file on the local mount point), NOT a decommission. mnt-backup.mount stuck failed (fstab bare defaults) → rest-server serves an empty local dir. Recovery in disaster-recovery.md.

  • [2026-06-20] The DEFAULT ssh ana-docker is lkraven (no NOPASSWD) — but ssh infra-ops@ana-docker HAS NOPASSWD root. A sudo cp as lkraven silently failed → nearly punted the rest-server recovery. Reach for infra-ops@ana-docker for sudo ops.

98 older entries archived to archival-memory.md.