20 KiB
Persistent memory — eshpfi-management
Last updated: 2026-07-05
Repo purpose
Reference workspace for PFI infrastructure: server inventory, canonical
Docker Compose stacks, ops playbooks, and conventions. Authoritative
copies of compose files live on the servers under
/opt/docker/compose/<stack>/; this repo mirrors them for version
control, editing, planning, and CI-driven deploys. It was originally
spun up to handle the fleet backups — keep that lens when triaging
backup/storage issues.
Tools and conventions
Sister repos (separate gitea repos, deployed by playbooks here):
| Repo | Role | CI status |
|---|---|---|
vh/task-board |
MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) |
vh/vor |
Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
vh/nevermore |
Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
vh/asset-engine |
Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
vh/althing |
Lean trusted inter-agent message bus — v2 "email model" (v2.0.0b2, 2026-07): per-box local-SQLite bus + courier/receiver for P2P over the 10.x net; pillars = open-loops / per-box herald + wake-listener / roaming owner API /owner/* / althing-mcp stdio surface. The v0.15 lean-bus cut RIPPED moderation / chamber / forseti-daemon / agent-runner / redis-valkey. |
per-box uv tool install (NOT CI-deploy); nh3-dev = the DEV box (editable install of ~/development/althing, gets new versions first); nh3-extdev a mesh peer (model B: althing-svc + shared /srv/althing) |
vh/mead-hall |
Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
vh/skaldsong |
Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
vh/Worldtree |
Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. gitea-runner builds on ana-docker; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. | push-to-main → CI build-and-deploy (runner on ana-docker) |
vh/yt-voice-clipper |
YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md |
vh/arbo |
Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) | push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook |
model-training-forge (mtf-dev) |
Fine-tuning recipe forge; T1 = qwopus E-RP writing LoRA (SFT→DPO, LitBench-RM reward) | training runs, not a deployed sidecar |
(vh/volva + Heid were re-architected from systemd daemons to Claude Code
session orchestrators 2026-06-08; their nh3-dev .service units were removed —
no longer deployed sidecars here. See Recent decisions.)
-
Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see
docs/runbooks/disaster-recovery.mdfor the blast-radius matrix. ⚠️ The restic file+DB layer routes through TWO rest-servers (rest-server-ana@ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas;rest-server-nh3@ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of/mnt/backup. (rest-server-ana recovered 2026-06-20.) -
pull-hf-repo.yamlis the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at/tank/aimodels/huggingface/" playbook. Supports--var repo_type=model|dataset|space. Replaces ad-hochuggingface_hub.snapshot_downloadpatterns. -
Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (
key_id 61419c92) atana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-adminauths against demo only. Personal-instance admin (the~/.config/worldtree/personal-admin-token, mode 600) POSTs/admin/keys(mints per-project keys; takesuser_id+label, no scope param — scopes are tier-derived). On-instance mint recipe (cleaner than DB-manip):docker exec worldtree-worldtree-api-1POST/admin/keyswith the in-containerWORLDTREE_BOOTSTRAP_ADMIN_KEY; cleartext once in.key=wt_live_+16hex. auto-memoryreference_worldtree_demo_key_mint. -
Per-project user keys against personal Worldtree (issued 2026-05-19):
skaldsong:79744637,skaldsong:7c1dbbbe,althing:50d85460,mead-hall:a360822d. Mint via/admin/keys, drop value to/tmp/wt-personal-<name>.keymode 600, dev collects + shreds (DO NOT cat to chat transcript). -
Skaldsong CD pattern (registry-pull). vh/skaldsong's CI builds and pushes
gitea.phasefinal.com/vh/skaldsong:<sha>+:latest;playbooks/deploy-skaldsong.yamlon ana-docker pulls + recreates. SHA-pin only. Prereq: host needsdocker login gitea.phasefinal.comonce. -
gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH
10.250.50.70:222, HTTP:3000. Fleet/colo hosts must use this internal route, NOT publicgitea.phasefinal.com(38.120.12.44) — the public path fail2bans the host egress IP. Full gotcha indocs/orientation.md→ Git/gitea. -
docker-as-root pattern (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo):
docker run --rm -v <target-dir>:/wt docker:cli sh -c "...". docker-group membership is effectively root via bind-mount. Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass-e VAR=/abs/pathfor any relative-default config dir. -
scripts/elwaysudo handling — elway prompts for the sudo password ONCE viagetpassbefore the firstsudo: truestep → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH. -
Per-host SSH identity matters for sudo. infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: default
ssh ana-docker=lkraven(docker-group, NO passwordless sudo);ssh infra-ops@ana-dockerHAS NOPASSWD root. → For any sudo op on ana-docker, usessh infra-ops@ana-docker.ssh infra-ops@10.100.10.50(nh3-dev) ALSO NOPASSWD sudo; on nh3-extdev infra-ops is sudo-LESS by design (ssh lkraven@10.100.50.42is the NOPASSWD path). irv-ml1:ssh irv-ml1= lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD) — stage model pulls to/home, not root-owned/worktank.
Current state / in-flight
As of 2026-07-05:
-
T1 TRAIN — venue decision live; NEXT ACTION = the ana-ml2 SMOKE. T1 (qwopus E-RP writing LoRA) datasets + recipe + smoke set are all staged (mtf-dev, on
nh3-dev:data/export/qwen-3.5-122b-erp-lora/incl.smoke/); bf16 base is on ana-ml2/tank/aimodels/qwopus3.5-122b-a10b-bf16(233G, PUBLIC HF repo, zero-auth pull). On-prem RULED OUT except a full ana-ml2 shutdown → CPU-offload → ~1-day fleet outage; cloud (Vast.ai 8×80GB, ~3-6h, ~$60-500, zero impact) is the clean alt. Operator chose smoke-first on ana-ml2 to measure real samples/sec before committing. ⚠️ the smoke ITSELF needs both GPUs + the ~421G vLLM RAM freed (192G VRAM < 244G base → CPU offload) = a brief full ana-ml2 serving pause; keeping any serving up drops it into the ~week NVMe regime. Full runbook + both paths inreference_t1_cloud_train_plan. Post-train serve-path (swappable-LoRA-on-NVFP4 test) still queued:reference_gen_qwopus_122b. -
LitBench-RM UP (irv-ml1 A6000); comfyui DISPLACED. Respun the R19 creative-writing reward (
vllm-litbench-rm:8202, ~19.6G) for T1's reward ensemble; stopped comfyui to free the A6000 (operator is sole comfyui consumer + holding image-gen until LitBench done). Score scale re-verified from nh3-dev (literary 0.95 > flat 0.44 > slop 0.12). TEARDOWN on "litbench done":ssh irv-ml1 'docker rm -f vllm-litbench-rm && docker start comfyui'— ⚠️ then watch the mmartial boot trap (root-owned venv → crash-loop;chown -R 1000:1000 /comfy/mnt/venv).reference_litbench_rm_irv_ml1. -
character-rp role — SHIPPED; one QUEUED spot-check. Proved per-request
extra_body(top_k/repetition_penalty) forwards through thegen-reasoningalias to the vLLM sampler (no gateway cap needed); pre-staged thecharacter-rprole in demo+personal bind-mountmodel_roles.yaml(byte-verified live on b18; caught the cached-registry ordering so the deploy's own restart activates it). worldtree-dev shipped #344 (v1.0.0b19) fixing the durable-agent override-drop (role now applies the RP profile). QUEUED: ratatoskr-dev pings a turn-timestamp → I spot-check the gateway spend_logs to confirm the RP sampler params (temp 0.75 / top_p 0.85 / presence 1.5 / top_k 20 / rep 1.0) reach vLLM. -
althing v2 (herald + receiver) = systemd-supervised on nh3-dev (the althing DEV box). Formalized this session:
althing-herald.service(Restart=always, Environment=PATH incl ~/.cargo/bin — the fix for the silent pane-dispatch outage) +althing-receiver.service(v2 → pillar-3/owner/*live). Stale forseti unit removed; both on v2.0.0b2, canonicalized by forseti.reference_nh3_dev_althing_herald. -
glm-5.2 = z.ai 1M-in / 128K-out, NO gateway cap (probed z.ai; recorded in the config comment
reference_litellm_gateway; commit624a07e). All GLM entries are pure z.ai passthrough — the effective ceiling is z.ai's canonical, not a gateway limit.
-
/booksmounted (transient) on nh3-dev (10.0.50.50:/mnt/books→ NFSv4 ro,soft; re-mount viainfra-ops@10.100.10.50if it reboots). -
Backups — recovered + hardened (2026-06-20), STILL OPEN: rotate the 5 disclosed rest-server creds (operator, offline); confirm esh-vm-db's resticprofile includes DB dumps.
docs/runbooks/backups.md. -
Open follow-ups (low-priority): clean phantom
qwen3.6-35b-a3boff the gateway (lists in /v1/models but 400s); docker-daemon default log-cap on ana-docker; drop the now-inertmood.decay_rate/mood.stale_hourskeys from the deployed/opt/worldtree*/configbind-mounts (harmless-present per worldtree-dev; edit-only/no-restart at a no-deploy window). -
Standing / parked: Mac Pro migration (hardware-gated); R17 v2 corpus HELD; disclosed-keys rotation queue; clean legacy
news-digest; R22 gateway-only full-access key at/home/lkraven/.r22-gateway-key(mode 600, carries paid GLM, do NOT delete); granite-4.1-8b bf16 atirv-ml1:/home/lkraven/granite-4.1-8b-bf16(reap-on-request); Deckard staged on ana-ml2 as T1's writing benchmark. -
Worldtree config-propagation (reference): demo+personal bind-mount config from
/opt/worldtree{,-personal}/config(infra-ops-deployable, byte-identical from canonical); reload viadocker restart <container>, NEVERcompose up(stale-:latestfootgun). The role registry is loaded ONCE + CACHED at startup (ConversationService.role_registry) → a bind-mountmodel_roles.yamlchange needs a container restart to take effect; pre-stage the bind-mount BEFORE the activating deploy's restart (else it needs a second bounce). config REMOVALS are NOT backward-compatible with the still-running image.
Recent decisions
-
[2026-07-05]T1 training venue: CLOUD recommended; operator chose smoke-first on ana-ml2. On-prem ruled out (ana-ml2 full — both 96G GPUs ~93G used): keep-serving = NVMe offload ~6-8 DAYS; full ana-ml2 shutdown = CPU offload ~1 DAY but a whole-fleet outage. Cloud Vast.ai 8×80GB (no offload → ~3-6h, ~$60-500, zero fleet impact) is the clean alt (mtf-dev + infra-ops both rec; Vast for its no-content-AUP marketplace + likely-existing VastBlue account). Operator's next step = the ana-ml2 CPU-offload SMOKE (~60 steps) to get real samples/sec before the full-outage-vs-cloud call. HF base verified public (zero-auth pull). Runbook + gotchas inreference_t1_cloud_train_plan. -
[2026-07-05]glm-5.2 canonical limits recorded (probed live vs z.ai): 1,048,576 (1M) input context / 131,072 (128K) max output; NO gateway-side cap (pure passthrough → z.ai's limits are effective). Written to the config comment (commit624a07e) +reference_litellm_gateway. -
[2026-07-04]character-rp: gateway-forwarding proven + role pre-staged + #344 shipped. Empirically confirmed per-requestextra_body(top_k/repetition_penalty) forwards through thegen-reasoningLiteLLM alias to vLLM + standard params override the alias defaults — no gateway cap needed (I over-built a dedicated alias, operator corrected, reverted with zero fleet impact). Pre-staged thecharacter-rprole into demo+personal bind-mountmodel_roles.yaml(byte-verified on b18; caught the cached-registry ordering). worldtree-dev shipped #344 (v1.0.0b19) for the durable-agent override-drop. spend_logs spot-check queued (ratatoskr's timestamp ping). -
[2026-07-04]althing v2 herald+receiver formalized as systemd on nh3-dev.althing-herald.service(Restart=always, Environment=PATH incl ~/.cargo/bin — the pane-dispatch fix) +althing-receiver.service(v2 → pillar-3/owner/*live); stale forseti unit removed; both on v2.0.0b2, canonicalized by forseti.reference_nh3_dev_althing_herald. -
[2026-07-04]LitBench-RM respun (irv-ml1 A6000, comfyui displaced) for T1's reward ensemble; operator sole comfyui consumer, holding image-gen until LitBench done.reference_litbench_rm_irv_ml1. -
[2026-07-03]ratatoskr-dev DEMO Heimdall key provisioned (R30 φ0). Minted a tier-user key on the demo viaPOST /admin/keys(bootstrap admin key), mirroring their personal base consumer (no character-binding); base-agent affect reads work ungated.reference_worldtree_demo_key_mint. -
[2026-07-02]mtf-dev granite harness-spike ran GREEN — MECHANICAL only, efficacy DEFERRED to the T1 run. Trainer TRL SFT→DPO→eval seam proven end-to-end on a synthetic fixture (not the E-RP corpus); operator DECIDED no intermediate real-efficacy granite spike (uninterpretable proxy — arch gap + abliteration axis).reference_gen_qwopus_122b. -
[2026-07-01]Worldtree #332 embed-recall diagnosed + scoped-log view/tunnel provisioned + fix verified (15×→1.01× re-embed).reference_wt_gateway_scoped_log_view. -
[2026-07-01]qwopus native MTP speculative-decode tested ongen→ NOT kept (+12% single-stream, −15–20% aggregate at concurrency, silently drops min_p/logit_bias). Banked for T1.reference_gen_qwopus_122b. -
[2026-07-01]Deckard trial → reverted to qwopus (gen) (won writing "in every way" but ~36 vs ~90 tok/s; spec-decode rescue ruled out). gitb63c48b→681eb70. Deckard kept staged as T1's writing benchmark. -
[2026-06-25]althing re-architected to the lean multi-machine bus; nh3-extdev stood up as a MODEL B mesh peer (dedicatedalthing-svc+ group-shared/srv/althing).reference_nh3_extdev_althing_mesh. -
[2026-06-23]zellij native web client piloted on nh3-dev (zellij-web.service:8443) alongside ttyd.reference_zellij_web_seat. -
[2026-06-22]Worldtree persona-render config arc (#314/#322/#317) pre-synced + deployed green on demo+personal — #317 a boot-blocking config REMOVAL.reference_corviduo_dev_emergency_ops. -
[2026-06-20]R22 (brokkr/dwarves) stood down to gateway-only; full-access R22 key minted; Phase B CANCELLED (Worldtree model-agnostic → no deploy path). Key at/home/lkraven/.r22-gateway-key(persistent mode-600, carries paid GLM, don't delete). MUT = freeqwen3.5-122-a10b(gen). Operator steer: R22 research is gated on a pragmatic/deployable outcome, not advancing-the-art. -
[2026-06-20]claude-bot issue-scope token minted for worldtree-dev self-serve (id 16,write:repository+write:issue); old token revoked. Advances the credential-migration directive. -
[2026-06-20]rest-server-ana recovered + backup prevention shipped + worldtree-dev admin keys provisioned (demo d113207c / personal f4f75adb). Cred rotation (5 rest-server pw) BELAYED. -
[2026-06-20]claude-bot → ADMIN on vh/Worldtree (operator-authorized) — self-serves WT deploys/tokens henceforth. -
[2026-06-14]STANDING DIRECTIVE: migrate ALL infra access to Claude-specific credentials. (auto-memoryproject_migrate_infra_access_to_claude_credentials)
118 older entries archived to archival-memory.md.
Tried and abandoned
-
[2026-07-04]LiteLLM (this gateway version) mutates the SHARED deployment config in-place on per-request sampler-param merge → my deliberately-invalidtop_k=-5forwarding-probe bled into a param-less character-rp request (vLLM 400, ONE-OFF, self-cleared by a later valid probe). NOT caching (none configured), NOT a config change. Never fire invalid/distinctive sampler values at a SHARED gateway alias with live consumers — use a throwaway alias, or adocker restart litellmflushes residual carryover.feedback_litellm_shared_param_mutation. -
[2026-07-04]A systemd--userdaemon that shells out to~/.cargo/bin/~/.local/bintools needs an explicitEnvironment=PATH— the minimal--userdefault silently drops them. The althing herald lostzellij→ silentpane-missfor ALL config-backed TUI/pane agents; CC + FIFO routes were unaffected, so it was invisible from a CC session.reference_nh3_dev_althing_herald. -
[2026-07-04]On-prem T1 train that keeps ANY ana-ml2 serving up = ~6-8 DAYS (1-GPU + NVMe ZeRO-Infinity offload; MoE ~10B-active cuts FLOPs but NOT the 244G base's param I/O). The only fast on-prem path is a FULL ana-ml2 shutdown (both GPUs + the ~421G vLLM RAM freed → base fits in the 566G CPU RAM) → CPU offload → ~1-day full-fleet outage. Cloud (no offload) = hours.reference_t1_cloud_train_plan. -
[2026-07-01]A personal-Worldtree CI deploy that fails ~85s in with "not found / unauthorized" is usually the pull-only-vs-build RACE, not registry-auth.deploy-personal.ymlis PULL-ONLY but fires on thestaging/vXtag simultaneously withdeploy.yml's build → pulls before the push finishes. FIX: re-run once built, or gate onworkflow_run: completed. -
[2026-07-01]MTP/spec-decode on a SHARED serving model helps single-stream but HURTS moderate-concurrency aggregate + silently ignoresmin_p/logit_bias(qwopusgen: N=1 +12%, N=4 −20%). Reserve for dedicated/interactive deployments. -
[2026-07-02]irv-ml1/worktankROOT is root-owned — lkraven can't write there (irv-ml1 sudo needs a password) → stage model pulls to/home. PIN THE A6000 BY UUID for training (native-CUDA ordering differs vs docker; the 3090 index 0 is usually near-full → OOM).CUDA_VISIBLE_DEVICES=GPU-<uuid>. -
[2026-06-25]althing "unreachable: " can MASK an app-level 500. Raw network was clean; root cause = receiver DB agents-table not synced with the config roster → delivery 500'd "unknown to: ", MAPPED to "unreachable". Diagnose: raw curl to :8087 + connect-probe ⇒ NOT network. Fixed in althing v0.17.1.reference_nh3_extdev_althing_mesh. -
[2026-06-20]rest-server.htpasswd: permission denied= the ana-nas NFS mount FAILED (ghost file on the local mount point), NOT a decommission.mnt-backup.mountstuckfailed(fstab baredefaults) → rest-server serves an empty local dir. Recovery in disaster-recovery.md. -
[2026-06-20]The DEFAULTssh ana-dockerislkraven(no NOPASSWD) — butssh infra-ops@ana-dockerHAS NOPASSWD root. Asudo cpas lkraven silently failed → nearly punted the rest-server recovery. Reach forinfra-ops@ana-dockerfor sudo ops.
98 older entries archived to archival-memory.md.