Files
esh-pfi-infrastructure/persistent-memory.md
T
vh 307b01e8a1 snapshot: capture skaldsong env-var-name footgun (CD wiping state)
Same lesson family as the /app/web/dist mismatch — encoding
container-internal contract (paths OR env var names) in compose
needs to be verified against the Dockerfile + app, not against
design-doc shorthand. Wrong env var names silently no-op; app
falls back to Dockerfile defaults which orthogonally miss the
bind mount, and state goes to ephemeral layer until next recreate.
2026-05-20 21:57:51 -07:00

24 KiB

Persistent memory — eshpfi-management

Last updated: 2026-05-20

Repo purpose

Reference workspace for PFI infrastructure: server inventory, canonical Docker Compose stacks, ops playbooks, and conventions. Authoritative copies of compose files live on the servers under /opt/docker/compose/<stack>/; this repo mirrors them for version control, editing, planning, and CI-driven deploys.

Tools and conventions

Sister repos (separate gitea repos, deployed by playbooks here):

Repo Role CI status
vh/task-board MCP + web dashboard for assistant task state (port 7878) push-to-main → CI deploys (2026-04-29)
vh/vor Inquisitor UI sidecar (port 7879) push-to-main → CI deploys (2026-04-29)
vh/nevermore Twice-daily LLM-curated briefing (port 8181, replaces news-digest) push-to-main → CI deploys (2026-04-30)
vh/asset-engine Internal control plane over inference services (port 8200, LAN-direct) push-to-main → CI deploys (2026-05-12)
vh/althing Inter-agent message bus (chamber UI port 7881, forseti + agent-runner daemons, valkey IPC) push-to-main → CI deploys (2026-05-14)
vh/mead-hall Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) push-to-main → CI deploys (2026-05-16)
vh/skaldsong Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) push-to-main → CI deploys (2026-05-19)
vh/worldtree Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration push-to-main → CI deploys
vh/volva Codex peer agent on althing bus (single-turn oracle, systemd daemon on nh3-dev) manual install via deploy/volva.service (2026-05-18)
  • Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see docs/runbooks/disaster-recovery.md for the blast-radius matrix.

  • pull-hf-repo.yaml is the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at /tank/aimodels/huggingface/" playbook. Supports --var repo_type=model|dataset|space. Replaces ad-hoc huggingface_hub.snapshot_download patterns.

  • Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (key_id 61419c92) at ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin auths against demo only. For personal-instance admin ops, fetch the bootstrap admin per-op via docker exec worldtree-personal-worldtree-api-1 printenv WORLDTREE_BOOTSTRAP_ADMIN_KEY on corviduo-dev. Used for POST /admin/keys, admin diagnostics (/admin/sessions/<id>/{bifrost,tools}, etc.).

  • Per-project user keys against personal Worldtree (issued 2026-05-19): skaldsong:79744637 (nh3-dev iteration), skaldsong:7c1dbbbe (ana-docker prod), althing:50d85460, mead-hall:a360822d. Same user_id=skaldsong across both skaldsong keys → shared Heimdall agent slot; different key_id → independently rotatable. Pattern: mint via /admin/keys, drop value to /tmp/wt-personal-<name>.key mode 600, dev collects + shreds (DO NOT cat to chat transcript).

  • Skaldsong CD pattern (registry-pull). Differs from althing / asset-engine which build-on-host. vh/skaldsong's CI builds and pushes gitea.phasefinal.com/vh/skaldsong:<sha> + :latest; playbooks/deploy-skaldsong.yaml on ana-docker pulls + recreates. SHA-pin only (no :latest health-gated advance yet). Prereq: host needs docker login gitea.phasefinal.com once (read:package PAT) — not currently in the workflow.

  • docker-as-root pattern (for ops that have no admin API, e.g. SqliteUserStore.set_bifrost_credentials): on hosts where the SSH user is in the docker group but lacks passwordless sudo, run docker run --rm -v <target-dir>:/wt -v /var/run/docker.sock:/var/run/docker.sock docker:cli sh -c "..." to edit deploy-owned files without sudo. Documented with security warning in servers/corviduo-dev/README.md. docker-group membership is effectively root via bind-mount; treat as a sudo-equivalent grant.

Current state / in-flight

As of 2026-05-20:

  • Worldtree CD disk-hygiene PR #184 approved, ready for merge. Three commits on infra/disk-watermark-gate-cd in vh/Worldtree: watermark gate (pre-pull recovery), eager post-deploy SHA prune (post-:latest-advance), set -e hardening in SSH blocks. Pending worldtree-dev's merge button. Follow-up extension to deploy-personal.yml + deploy-pinned.yml + Beszel disk-usage metric all queued, not bundled.

  • Skaldsong live on ana-docker:8300 (registry-pull pattern), pointed at personal Worldtree (:8081). Three first-deploy footguns surfaced + canonical-patched: SPA path mismatch (/app/web/dist/app/spa), CORS env shape (pydantic-settings expects JSON array), verify-step race vs healthcheck start_period (grep ^Up not healthy). nh3-dev hand-launched skaldsong on :8080 also still running for dev iteration — same user_id=skaldsong, different key_id, both authorized.

  • Worldtree personal (:8081) is the iteration-target instance. skaldsong + mead-hall + althing-chamber all rotated to personal via the per-project user keys minted 2026-05-19. Demo (:8080) stays as the canonical prod surface; pinned (:8082) intentionally lags. Personal's BIFROST_CLIENT_ALLOWED_HOSTS now mirrors demo's: 10.100.10.50:5173,10.100.10.50:8080,10.250.50.70:8300.

  • Disclosed-keys hygiene queue — these warrant rotation at convenience (mostly in-bus disclosure during this session):

    • /tmp/wt-personal-skaldsong-prod.key on nh3-dev — operator should shred after verify (file mode 600, lkraven-owned)
    • mead-hall's prior Worldtree bearer (wt_live_80e1570620ef2aba998dc63954cce3a6) — now superseded by personal-instance key a360822d per the 2026-05-19 rotation
    • Worldtree provider keys Z_AI_API_KEY + ZAI_API_KEY — from the 2026-05-12 corviduo-dev outage
    • chamber config.yaml's forseti.api_key + agent_runner.api_key — superseded by personal-instance key 50d85460 2026-05-19
    • Gitea runner registration token (a1135753...) — rotate via Gitea admin UI's runner-token reset
  • Still open from prior sessions: rotate MINIFLUX_PASSWORD (leaked twice); clean up legacy news-digest detritus on ana-docker (/opt/docker/compose/news-digest/, /opt/docker/data/news-digest/, image local/news-digest:v5); watch nh3-nas /volume1 (was 65%; recheck retention or expand before ~80%); the docker push 60s client-side ceiling mystery remains uninstrumented.

Recent decisions

  • [2026-05-19] Worldtree CD disk-hygiene strategy: watermark gate (env-tunable threshold + window, fail-loud on still-low post-prune)

    • eager post-deploy prune (only after :latest advance succeeds, uses docker image prune -a --filter "until=24h" which respects in-use semantic — protects pinned + personal images automatically). Combined: demo VM holds ~24h of deploy history instead of unbounded accumulation. Shipped in vh/Worldtree PR #184 (306cd61 + 613dac2 + bd91df5).
  • [2026-05-19] Skaldsong CD shape: shape (1) of three operator options — container + Gitea registry + pull-restart, matching Worldtree's pattern. Target host ana-docker (NOT nh3-dev where skaldsong-dev runs for iteration). SHA-pin only for now; health-gated :latest advance is a follow-up once /health exercises Worldtree

    • Kokoro reachability.
  • [2026-05-19] Skaldsong prod (ana-docker) switched from demo Worldtree (:8080) to personal (:8081). Same user_id=skaldsong as the nh3-dev hand-launch key — shared Heimdall agent slot (skaldsong:wizard-v2), different key_ids for independent rotation. Demo Worldtree stays for isolation; personal becomes the multi-consumer dev iteration instance.

  • [2026-05-19] mead-hall Bifrost v0.3 end-to-end smoke green. Closed task #32 (althing thread 01KRV1M2KW6N6HBEXGTH72QXCA). Wire layer (handshake + binding + dispatch) + data-flow (per-dispatch JWT claims → ctx.session_id populated → real session-scoped data) + agent-loop (LLM reads + quotes back) all proven. Resolves the "stalled mid-Worldtree" state from the 2026-05-17 snapshot.

  • [2026-05-18] Volva systemd install complete after three-stage debug. Final unit at /etc/systemd/system/volva.service runs as User=lkraven with ProtectHome=read-only + ReadWritePaths=/home/lkraven/.althing /home/lkraven/.codex carve-outs for state writes. VOLVA_ALTHING_CLI=/home/lkraven/ .local/bin/althing-cli + ALTHING_HANDLE=volva both pinned in env.sh.

  • [2026-05-17] Phase 3.1 cross-process streaming uses Valkey 8 alpine as a sibling compose service in stacks/althing-chamber/, redis-protocol pub/sub for high-volume msg_delta / msg_thinking / msg_start / msg_complete event kinds. DB bridge keeps msg_curated + floor_grant (structured / canonical). Two-channel architecture, no overlap. chamber + agent-runner depends_on: valkey: service_healthy.

  • [2026-05-17] Worldtree admin workflow shift (per vh): infra-ops gets its own permanent admin-tier key (61419c92, stored at ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin). Future admin ops route through this key, not the bootstrap admin via docker-as-root.

  • [2026-05-17] Worldtree env-var addition checklist: anytime introducing os.environ.get("FOO") in worldtree code, update BOTH .env.example AND compose.yaml's &worldtree-env anchor in the same PR. Same Z_AI_API_KEY-shape footgun bit BIFROST_CLIENT_ALLOWED_HOSTS (#170) until worldtree-dev added the passthrough line in 08f02b2.

  • [2026-05-16] althing-chamber Phase 2: added althing-agent-runner as third compose service (worldtree-driver agent dispatcher). All three althing services use the same image; command: selects entrypoint. Safe to enable preemptively (sleeps when no driver=worldtree handles declared).

  • [2026-05-14] althing-chamber stack scaffolded: chamber + forseti. Internal LAN-only at port 7881 (chamber default 7878 collides with task-board). Two-service compose, shared SQLite bind-mount, build-on-host pattern via vh/althing's gitea-workflow. Forseti is the canonical dev for this stack (galdrabok is on a different project).

  • [2026-05-13] vllm-qwen3vllm stack rename. Added vllm-reward service (Skywork-Reward-V2-Llama-3.1-8B-AWQ classifier). Three vLLM services share GPU 1 (embed 0.20, rerank 0.20, reward 0.30 utilization; 30% headroom). All use --runner pooling; classification drives via model's architectures: [LlamaForSequenceClassification] in config.json, NOT --task classify (deprecated in vLLM 0.19.1).

  • [2026-05-13] pull-hf-repo.yaml is the canonical HF-fetch playbook on ana-ml2. Supports --var repo_type=model|dataset|space. Replaces ad-hoc huggingface_hub.snapshot_download calls.

  • [2026-05-13] Selene-1-Mini-Llama-3.1-8B added to llama-swap as judge model. mradermacher i1-Q6_K imatrix quant (~6.5GB). AtlaAI reward/eval model — temp 0.01, ctx 32K, q8_0 KV cache. New JUDGE / EVAL MODELS section in stacks/llama-swap/conf/config.yaml.

  • [2026-05-13] /tend-docs first pass deletions: stacks/infinity/ removed (retired by vllm). Archived docs/asset-engine/design-brief.mddocs/archive/asset-engine/ with archival header. Fixed pfi-pve VM list to full qm list enumeration. Dropped stale weak-password section from pfi-postgres (rotation done 2026-04-23).

  • [2026-05-12] corviduo-dev (Worldtree-team dev VM, 10.250.50.152, CT 106 on pfi-pve) added to servers/ inventory. Treat like SF client hosts: PFI hosts + provides emergency-ops backstop; Worldtree team owns OS config + deploys + backup decisions.

  • [2026-05-12] Worldtree :latest tag drift bug — fixed by health-gated :latest advance in vh/worldtree's deploy workflow (architect commit 8ef3801): only tag :latest AFTER the new container's /health probe passes. Build-on-host stacks here don't have this problem because the playbook always builds the SHA-tagged image from a git reset --hard <ref> checkout.

  • [2026-05-12] asset-engine stack scaffolded LAN-direct at http://10.250.50.70:8200. Initially included Traefik labels for public hostname; user pulled them out (internal tool, no public TLS surface needed). Pattern: internal tools default LAN-direct; Traefik wiring only when external/TLS required.

  • [2026-05-12] asset-engine catalog gains lifecycle: { stack, vram_gb, gpu_device_id } per irv-ml1 service for the orchestrator feature. SSH keypair scaffolded at ana-docker:/opt/docker/conf/asset-engine/ssh/ for asset-engine container → irv-ml1 orchestration via dedicated ed25519 key.

Tried and abandoned

  • [2026-05-19] Naive docker rmi worldtree:<old-sha> --force for CD SHA cleanup — would untag pinned/personal worldtree images since all three deployments share corviduo-dev. Use docker image prune -a --filter "until=Xh" instead — respects in-use semantic (Docker won't remove an image referenced by any container on the host), so pinned/personal protected automatically.

  • [2026-05-19] Skaldsong CD first attempt: docker pull step failed with 401 unauthorized. ana-docker had no docker login for gitea.phasefinal.com. My playbook prereq note ("docker login has been done at least once") was an unverified assumption. One-time manual login persists in ~/.docker/config.json; architectural fix (workflow-side ssh ana-docker 'docker login ...' step using REGISTRY_USER/REGISTRY_TOKEN secrets) flagged as v2.

  • [2026-05-19] SKALDSONG_HOST_CORS_ORIGINS=http://10.250.50.70:8300 as a bare URL — pydantic-settings parses complex env vars via json.loads(); first-boot crashloop with SettingsError: error parsing value for field "cors_origins". Must be JSON array literal: SKALDSONG_HOST_CORS_ORIGINS=["http://..."].

  • [2026-05-19] SKALDSONG_HOST_STATIC_ASSETS_PATH=/app/web/dist in compose — mismatched Dockerfile reality. The Dockerfile COPYs SvelteKit build output flat into /app/spa (not /app/spa/dist). Lifted the path from skaldsong-dev's CD-ask message ("/app/web/dist") rather than verifying against the actual Dockerfile they shipped. Lesson: when encoding container-internal paths in compose, verify against the Dockerfile, not the design-doc.

  • [2026-05-20] SKALDSONG_DB_PATH + SKALDSONG_RUNS_DIR in compose env block — names skaldsong's app doesn't read. App reads SKALDSONG_HOST_SQLITE_PATH + SKALDSONG_HOST_RUNS_ROOT (per Dockerfile ENV defaults). Wrong names = silently no-op; app fell back to Dockerfile defaults pointing at /app/data/... which the compose's bind mount did NOT cover (mount target was /app/state/...). Result: every --force-recreate wiped the SQLite DB along with the ephemeral container layer. Caught by skaldsong-dev after operator noticed stories vanishing on each CD push (althing thread 01KS4DPF6SXTBP4Q360JZVWPNT). Fix in 52e98fa — rename env vars, bind targets unchanged. Same lesson as the /app/web/dist footgun: verify env var NAMES against the Dockerfile/app, not against design-doc shorthand.

  • [2026-05-19] Playbook verify step docker ps | grep healthy racing the container's start_period (30s in compose's healthcheck). Verify ran 0.09s after compose up -d --force-recreate — well before docker's healthcheck could flip the status from (health: starting) to (healthy). False-negative; container was operationally up (the earlier /health poll verify already confirmed). Fix: grep ^Up not healthy. /health-200 IS the liveness check; docker's (healthy) is just a delayed echo.

  • [2026-05-19] Volva daemon impersonating volva-dev (the human) for hours of debugging because cwd-based handle resolution. session_handles.json maps /home/lkraven/development/volvavolva-dev; the daemon's systemd unit sets WorkingDirectory=/home/lkraven/development/volva and env.sh didn't pin ALTHING_HANDLE, so the daemon polled volva-dev's inbox and replied AS volva-dev. Catch: cross-reference journal PIDs against message_received events to verify which voice generated which reply. Fix: pin ALTHING_HANDLE=volva in env.sh (worldtree-dev's e39be87 made this the default in env.sh.template). Real follow-on risk: hallucinated structured feedback from impersonating-daemon can bootstrap real (and correct) downstream commits — the technical artifact survives but the conversational attribution rots.

  • [2026-05-18] Volva env.sh.template $HOME in commented examples — systemd's EnvironmentFile= parser doesn't expand $HOME; uncommenting lands the literal $HOME/... string. Volva-dev's f4dda73 swapped to /home/<svc-user>/... placeholders.

  • [2026-05-18] Initial Volva systemd unit's ProtectHome=read-only without ReadWritePaths= — althing-cli's SQLite (~/.althing/ althing.db) and codex's session state (~/.codex/) both need to write. Container started but every poll failed with "db path not writable". Surgical fix: ReadWritePaths=/home/lkraven/.althing /home/lkraven/.codex (preserves the hardening intent, only carves out the specific dirs).

  • [2026-05-18] Trusting that env.sh's export VOLVA_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" template line works under systemd — EnvironmentFile= parser aborts on the first unparseable line (command substitution), and VOLVA_ALTHING_CLI declared below silently never lands. Symptom: Environment= property empty, daemon error "althing-cli not found at 'althing-cli'". Fix: replace command-substitution with literal path. Volva-dev's d436c3c dropped VOLVA_ROOT entirely upstream.

  • [2026-05-17] --task classify for Skywork in vLLM 0.19.1 — flag was deprecated. Use --runner pooling; the model's architectures: [LlamaForSequenceClassification] in config.json drives the classification head. Surfaced as vllm: error: unrecognized arguments: --task classify in container logs.

  • [2026-05-17] Trusting that .env edit alone propagates a new env var into a worldtree container — compose.yaml's &worldtree-env anchor must explicitly declare the passthrough or the value silently doesn't land. Same footgun bit Z_AI_API_KEY (2026-05-12) AND BIFROST_CLIENT_ALLOWED_HOSTS (2026-05-17). Cost ~10 min of "why is env empty?" diagnosis each time. Worldtree-side fix in vh/worldtree@08f02b2.

  • [2026-05-17] --force-recreate --pull never from the docker:cli sandbox without explicit -e WORLDTREE_IMAGE=<sha> re-pins the container to :latest, even when a newer SHA-tagged image is on disk. Symptom: container "recreated" but actually reverted to a stale image. Pass -e WORLDTREE_IMAGE=...:<sha> to the docker run invocation. Worldtree-dev's 8ef3801 health-gated :latest advance is the long-term fix.

  • [2026-05-13] Initial Voxtral default voice alloy (OpenAI-compat naming) — vLLM-Omni serving Voxtral does NOT translate aliases. Native presets are <register>_<gender> shape (neutral_female, casual_male, etc.). Always live-probe /v1/audio/voices for the exact wrapper-deployed preset names before setting a catalog default. Same caveat for Qwen3-TTS (wrapper exposes 15 voices: 9 Qwen presets + 6 OpenAI aliases) and Kyutai-TTS (NillPointer wrapper has NO voice-listing endpoint at all; voices are filesystem paths under the kyutai/tts-voices HF repo).

  • [2026-05-12] Defaulting asset-engine to Traefik-routed (asset-engine.phasefinal.com with anaprod cert resolver) on first scaffold — user pulled it back to LAN-direct. Internal tools default LAN-direct; only add Traefik when an external/TLS surface is actually needed.

  • [2026-05-12] Routing althing thread replies through galdrabok when the actual dev handle is forseti — bus rejected to=forseti initially because thread participants list was [galdrabok, infra]. Solved by starting a new thread with forseti as the direct recipient. Lesson: when the bus auto-resolves a sender handle that doesn't match the actual dev role, start a fresh thread rather than fighting the participant list.

  • [2026-05-08] Filtering Traefik's UTC access log by Gitea-local-PDT timestamp substrings (grep "2026/05/08 15:1[2-7]") returned zero matches and led to a wrong "no /v2/ traffic in 12 days" conclusion. Gitea logs in PDT, Traefik logs in UTC — same host, different timezones. Always normalize timezones (UTC) when correlating logs across services on the same box. Cost: ~30 min in the wrong direction.

  • [2026-05-08] Validating a user-proposed Traefik/Gitea timeout bump for "Traefik is dropping connections on big docker pushes" without first verifying which component was actually in the failure path. Traefik turned out to be innocent (499 / 60012ms = client closed, Traefik never timed out), Gitea's PER_WRITE_TIMEOUT governs response writes (wrong direction), and the real ceiling was a ~60s client-side timer no compose change can reach. Lesson: validate the diagnostic premise — which component is actually in the failure path? — before refining the proposed fix.

  • [2026-05-08] Bumping Gitea PER_WRITE_TIMEOUT / PER_WRITE_PER_KB_TIMEOUT to address unexpected EOF on /v2/.../blobs/uploads/ PATCH — wrong direction. Both govern response writes, not request body reads. unexpected EOF from Go's HTTP server means the client closed mid-body-upload; not a knob Gitea exposes server-side.

  • [2026-04-30] task-board workflow with container: image: debian:bookworm-slim — fails: actions/checkout@v4 needs node at runtime, slim image lacks it. Switched to node:20-bookworm-slim (has node + apt) or runner-label default. (Pattern revisited 2026-05-17 for skaldsong-dev: container override needs nodejs apt-installed unless it IS the default.)

  • [2026-04-30] Dropping the container: directive before runner re-registration with docker-schema labels — runner silently falls back to host mode (jobs run inside the alpine act_runner container itself, no apt). The :host suffix in startup logs (labels updated to: [pfi-fleet:host ana-docker:host]) is the giveaway. Fix: register with pfi-fleet:docker://<image> schema labels.

  • [2026-04-30] Updating runner labels by editing .env and bouncing — doesn't take. The .runner registration cache pins labels at first registration; env-var updates are read each start but the stored token + UUID are tied to the original label set on the gitea side. Fix: stop runner, delete .runner, generate new admin registration token, redeploy.

  • [2026-04-30] git reset --hard origin/<sha> in deploy-task-board.yaml (and the in-repo nevermore playbook before fix) — invalid syntax: origin/ prefix only works for branch refs. SHAs need git reset --hard <sha> directly. Resolved with git rev-parse --verify --quiet "origin/{{ ref }}^{commit}" first, then bare "{{ ref }}^{commit}" fallback.

  • [2026-04-30] Assuming DEPLOY_SSH_KEY was at user scope after task-board wiring — it was actually only repo-scope on vh/task-board. vor's first CI run failed with empty SSH key (printf '%s\n' "" > ~/.ssh/id_ed25519). Fix: copy secret to user scope at gitea.phasefinal.com/user/settings/actions/secrets.

  • [2026-04-30] grep -vE "^(#|$)" to inspect .env for sanity — leaked the full MINIFLUX_PASSWORD line into the transcript. Then a follow-up redaction attempt with sed -E "s/=(.{4}).*$/=\1<redacted>/" still leaked the first 4 chars. Lesson: when probing secret-bearing files, use field-by-field SELECTIVE inspection (grep -E "^(KEY1|KEY2)=") rather than negative filters; for any password line, grep -c (existence) or test -n "$(...)" (non-empty), never cat or value-printing.