19 KiB
Persistent memory — eshpfi-management
Last updated: 2026-06-01
Repo purpose
Reference workspace for PFI infrastructure: server inventory, canonical
Docker Compose stacks, ops playbooks, and conventions. Authoritative
copies of compose files live on the servers under
/opt/docker/compose/<stack>/; this repo mirrors them for version
control, editing, planning, and CI-driven deploys.
Tools and conventions
Sister repos (separate gitea repos, deployed by playbooks here):
| Repo | Role | CI status |
|---|---|---|
vh/task-board |
MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) |
vh/vor |
Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
vh/nevermore |
Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
vh/asset-engine |
Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
vh/althing |
Inter-agent message bus (chamber UI port 7881, forseti + agent-runner daemons, valkey IPC) | push-to-main → CI deploys (2026-05-14) |
vh/mead-hall |
Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
vh/skaldsong |
Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
vh/worldtree |
Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration | push-to-main → CI deploys |
vh/volva |
Codex peer agent on althing bus (single-turn oracle, systemd daemon on nh3-dev) | manual install via deploy/volva.service (2026-05-18) |
-
Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see
docs/runbooks/disaster-recovery.mdfor the blast-radius matrix. -
pull-hf-repo.yamlis the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at/tank/aimodels/huggingface/" playbook. Supports--var repo_type=model|dataset|space. Replaces ad-hochuggingface_hub.snapshot_downloadpatterns. -
Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (
key_id 61419c92) atana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-adminauths against demo only. For personal-instance admin ops, fetch the bootstrap admin per-op viadocker exec worldtree-personal-worldtree-api-1 printenv WORLDTREE_BOOTSTRAP_ADMIN_KEYon corviduo-dev. Used forPOST /admin/keys, admin diagnostics (/admin/sessions/<id>/{bifrost,tools}, etc.). -
Per-project user keys against personal Worldtree (issued 2026-05-19):
skaldsong:79744637(nh3-dev iteration),skaldsong:7c1dbbbe(ana-docker prod),althing:50d85460,mead-hall:a360822d. Sameuser_id=skaldsongacross both skaldsong keys → shared Heimdall agent slot; differentkey_id→ independently rotatable. Pattern: mint via/admin/keys, drop value to/tmp/wt-personal-<name>.keymode 600, dev collects + shreds (DO NOT cat to chat transcript). -
Skaldsong CD pattern (registry-pull). Differs from althing / asset-engine which build-on-host. vh/skaldsong's CI builds and pushes
gitea.phasefinal.com/vh/skaldsong:<sha>+:latest;playbooks/deploy-skaldsong.yamlon ana-docker pulls + recreates. SHA-pin only (no:latesthealth-gated advance yet). Prereq: host needsdocker login gitea.phasefinal.comonce (read:package PAT) — not currently in the workflow. -
docker-as-root pattern (for ops that have no admin API, e.g.
SqliteUserStore.set_bifrost_credentials): on hosts where the SSH user is in thedockergroup but lacks passwordless sudo, rundocker run --rm -v <target-dir>:/wt -v /var/run/docker.sock:/var/run/docker.sock docker:cli sh -c "..."to edit deploy-owned files without sudo. Documented with security warning inservers/corviduo-dev/README.md. docker-group membership is effectively root via bind-mount; treat as a sudo-equivalent grant. Foot-gun: when runningdocker composeinside this sandbox, any relative path in compose.yaml (e.g.${WORLDTREE_CONFIG_DIR:-./config}) resolves against the sandbox CWD, but Docker daemon interprets the resulting path against the HOST filesystem. Always pass-e VAR=/abs/pathto the docker run invocation for any relative-default config dir. -
scripts/elwaysudo handling — elway prompts for the sudo password ONCE viagetpassbefore the firstsudo: truestep. That prompt is interactive → elway can't run unattended from a non-TTY tool if any step needs sudo. For sudo-free playbooks (nosudo: truesteps) it runs fully non-interactive over key SSH. To create root-owned dirs WITHOUT host sudo, use the docker-daemon-root trick:docker run --rm -v /worktank:/mnt alpine sh -c 'mkdir -p /mnt/<x> && chown -R 1000:1000 /mnt/<x>'.
Current state / in-flight
As of 2026-06-01:
- TTS fleet on irv-ml1 reworked this session; asset-engine catalog now
17 services (CSM removed). Canonical
docs/asset-engine/services.yaml- vendored
vh/asset-engineboth pushed (eshpfi38d9e3b; asset-engine0b552bd). Per-engine live state:
- dia2-2b (:8200) + dia2-1b (:8202) — LIVE + healthy (restarted this
session;
local/dia:v2, voices + cache intact).status: ready. - Chatterbox (:8196) — LIVE. Catalog v2: routes to
/tts, exposes emotion levers (exaggeration 0.5, cfg_weight 0.5, temperature 0.8), 28 built-in predefined voices (defaultEmily.wav) + clone. Fast + competent cloner (it audibly cloned Imogen). PlainImogen.wavuploaded to its clone dropdown (/get_reference_filesnow: Imogen.wav + glados_25s.wav). - Fish-S2 (:8195) — LIVE but HEAVY: ~240s torch.compile warmup on
cold start, ~24 GB on the A6000; first call after any restart eats the
warmup. Catalog v6: full
/ttsServeTTSRequest surface (temperature/top_p/repetition_penalty/latency/seed/format) +reference_idis a 32-voice static "Voice" dropdown (default Emily) + advanced base64 clone. NOT OpenAI-compat (only/v1/tts, no/openapi.json, no/voicesAPI). Research/non-commercial license. - zonos (:8203) down (adapter built, not deployed). ComfyUI :8188 catalog-deferred link.
- vendored
- BLOCKED on asset-engine-dev — P1 "undefined" select bug (default voice
picks send the literal string
"undefined"→ Chatterbox 404s, Fish falls to default voice). ROOT-CAUSED: the form's Kokoro voice-blend widget (templates/partials/_fields.html) runsrecompute()at page load, reading the Shoelace<sl-select>.valueBEFORE hydration → writes "undefined" into the submit field. Escalated with full root cause + 3-part fix (althing msg01KT2K2SY9N7AY69R9V0B4RXSW→ asset-engine-dev). Workaround until fixed: click the dropdown (fires change) or hit the API directly. My half (blendable: falsecatalog flag) is queued — see Recent decisions. - Fish cloning verdict UNRESOLVED — re-test parked. My in-session "Fish is
a weak cloner" conclusion was confounded (bogus reference transcript + maybe
reference_id: "undefined"from the form bug + weak resemblyzer encoder) — see Tried and abandoned. Literature says Fish/OpenAudio is SOTA cloning. Whisper has the real transcript ("If the red of a second bow…"). Clean re-test (correct transcript + ECAPA scoring) ready when wanted. Future voice clones are operator-handled (pitch-shift deepening abandoned). - On-host consented voice library — ~992 real-person clips cached in the
kyutai tts-voices repo (
/worktank/kyutai-tts/.../snapshots/.../): VCTK (CC BY 4.0, accent-tagged speaker IDs), Unmute voice-donations (CC0), EARS + expresso (CC BY-NC). Source for future clone voices. British-female Southern England speakers p225/p228/p229 staged into Fish as Imogen/Eleanor/Beatrice. - Parakeet ASR (:8765) is DOWN — no container, but catalog says
status: ready(drift). Use Whisper-in-a-container (faster-whisper) for transcripts instead; it's self-contained (CTranslate2, no torch). - Worldtree healthy v0.29.13 (last-known); Skaldsong v0.32.2 (ana-docker:8300, Kokoro SSE streaming); artemis-31b-v1i live on llama-swap + worldtree personal; ttyd fleet driver seat on nh3-dev (RO :7681 / drive :7682 via Caddy) — all out of this repo, last-known.
- Disclosed-keys hygiene queue — rotate at convenience: HF token
hf_HBl…(lkraven's HF account) leaked into BuildKit logs during the CSM build attempt (logs shredded, never committed — low urgency);/tmp/wt-personal-skaldsong-prod.keyon nh3-dev; mead-hall's prior Worldtree bearer (superseded bya360822d); WorldtreeZ_AI_API_KEY/ZAI_API_KEY; chamberforseti/agent_runnerapi_keys (superseded by50d85460); Gitea runner registration token (a1135753…). - Still open from prior sessions: rotate
MINIFLUX_PASSWORD(leaked twice); clean up legacynews-digestdetritus on ana-docker; watch nh3-nas/volume1(was 65%; recheck before ~80%); thedocker push 60s client-side ceilingmystery remains uninstrumented.
Recent decisions
-
[2026-06-01]CSM (Sesame csm-1b) torn down entirely — removed from catalog,stacks/csm/,playbooks/deploy-csm.yaml, and host (c54ab13). Two reasons: (1) deep-research verdict — the acclaimed Maya/Miles demo is a fine-tuned, larger variant Sesame never open-sourced; the open 1B base ships no usable voices, can't generate text, English-only, no real-time streaming, absent from leaderboards, dominated by Kokoro/Dia2/Fish/IndexTTS. (2) The upstream image is rotted (see Tried and abandoned). asset-engine re-synced (306aae9). -
[2026-06-01]Chatterbox catalog v1→v2 (93cedd6) — re-point to the devnen wrapper's/tts(was the thin OpenAI/v1/audio/speechexposing none of the emotion knobs → poor output). Defaultexaggeration: 0.5per Resemble's README ("works well for most"); the server SHIPS demo-tuned 1.3 which over-acts — deliberately NOT adopted. 28 predefined voices via/get_predefined_voices(default Emily.wav). Corrected seedable false→true. -
[2026-06-01]Fish-S2 catalog v1→v6 (f3c260e..38d9e3b) — was a blank 2-field form; now exposes the full fish-speech ServeTTSRequest surface- a 32-voice
reference_id"Voice" dropdown (STATIC list — Fish has no /voices API).reference_id-by-name is the WORKING clone path; the path-formreferencesshape 500s (see Tried and abandoned). Corrected vram 4→24 GB, cold_start 8→240s, seedable false→true. Added 3 consenting British-female VCTK voices (Imogen=p225/Eleanor=p228/Beatrice=p229).
- a 32-voice
-
[2026-06-01]blendable: falsecatalog flag — DEFERRED, pending asset-engine-dev honoring it in the form template. Clone-voice selects (chatterbox/fish/dia2) must render as plain dropdowns, not the Kokoro blend widget (blending reference clones is nonsensical). I ship the catalog flag + schema field once they read it. Tracked: althing msg01KT2K2SY9N7AY69R9V0B4RXSWto asset-engine-dev. -
[2026-05-31]Dia2 deployed as two fixed-model instances (dia2-2b:8200,dia2-1b:8202) fromlocal/dia:v2, retiring legacy Dia 1.6B; catalogdiaentry removed → dia2-2b + dia2-1b (breaking for asset-engine). Rationale: the devnen wrapper is single-model and IGNORES the OpenAImodelfield (verified on its live OpenAPI), so the only way to offer both Dia2 models as real per-request asset-engine choices is one fixed endpoint per model.3139e81(deploy),db15638(catalog swap). -
[2026-05-31]Both dia2 catalog entries route to the wrapper's richer/ttsendpoint (not/v1/audio/speech) to expose the full control surface (cfg_scale/temperature/top_p/cfg_filter_top_k/voice_mode/clone); all defaults sourced from the wrapper'sCustomTTSRequestPydantic blessed values. Voice default isvoice_mode: clone+clone_reference_filename: Abigail.wavso a stable (non-random-gender) voice is the out-of-box behavior.55602b7,5c47843. -
[2026-05-31]Zonos REST adapter (stacks/zonos/adapter/,local/zonos-api) — thin OpenAI-ish/v1/audio/speechFastAPI in front of the Gradio-only Zonos SDK; JSON-envelope{audio, audio_format, seed}(Zonos is the fleet's first seedable TTS). Port 8203 (moved off 8201 — collided with csm). Built; NOT deployed (stack down for VRAM). Also fixed the upstream image's missing CMD (71df6f7).81efa8d. -
[2026-05-31]Catalog schema regenerated: addedCatalogLifecycle+reproducibility.seed_field(b7b2130). Resolves the stale-schema hand-off; catalog now validates clean. (asset_enginecatalog.pyPydantic already supported both — schema file was just behind.) -
[2026-05-31]TTS bench expanded withstacks/{dia,zonos,csm}(666f7f3dia+zonos,a4b8c2acsm). The bench already had Fish S2-Pro / Chatterbox-Turbo / IndexTTS-2 / CosyVoice3 / Kokoro / VibeVoice / Qwen3-TTS / Kyutai. (csm since removed 2026-06-01.) -
[2026-05-31]Remote browser/iPad/Vision-Pro driver seat for the agent-fleet zellij sessionClaudestood up on nh3-dev (ttyd behind Caddy, network-gated). Out of this repo — full architecture + the HTTP2/OSC52/Safari-auth gotchas in auto-memoryreference_ttyd_fleet_seat. -
[2026-05-30]esh-docker-vm NFS boot-ordering fix:playbooks/fix-esh-nfs-boot-ordering.yaml(c0458d9, +53157b1drop-in filename-collision fix) adds_netdev,nofailto the four 10.0.50.50 NFS mounts + a dockerAfter=remote-fs.targetdrop-in — resolves paperlessExited(255)on reboot. traefik also gainedrestart: unless-stopped. Full incident → auto-memoryincident_esh_docker_nfs_boot_race.
25 older entries archived to archival-memory.md.
Tried and abandoned
-
[2026-06-01]CSM bring-up — upstreamphildougherty/sesame_csm_openaiDockerfile pins NOhuggingface_hubversion, so it now resolves to 1.17.0 wherehuggingface-cliis fully removed (replaced byhf) → the build dies athuggingface-cli login/download. Passing the HF token as a build-arg got the token in (the model downloads at BUILD time, not runtime — our compose only wired it as runtime env, the original failure) but then hit the rotted CLI. Fixing would need vendoring + patching the Dockerfile; not worth it for a weak model → CSM removed. -
[2026-06-01]Fish path-form references ({audio:"/app/references/x.wav"}) → HTTP 500 on this build, despite the stack README showing that shape. The working clone paths arereference_id(by name → reads<name>.wav+ optional<name>.txt) or inline base64references. The catalog now usesreference_id. -
[2026-06-01]Fish reference transcript as a provenance note (not the actual spoken words) — Fish/fish-speech uses the reference transcript to disambiguate phonemes, and a misaligned transcript hurts cloning badly. I staged VCTK voices with provenance-note.txtsidecars, which CONFOUNDED my "Fish is a weak cloner" verdict (that verdict is retracted/unproven). Lesson: stage the REAL transcript (VCTK ships ground-truth; or ASR via Whisper) for any clone reference. -
[2026-06-01]Pitch-shift register control (rubberband, to deepen Imogen to contralto/mezzo) — Fish ignores small reference shifts and overshoots large ones;-3stWITHOUT formant preservation → "baritone through a tuba" (formants collapsed);-2/-3stWITH--formant→ too subtle / still bad. Abandoned at every depth; all variants deleted. Finer independent pitch/formant control needs praat (not installed). Future clones = operator's. -
[2026-06-01]resemblyzer for cloning-fidelity scoring — weak/dated encoder (2019 LSTM, English-biased) + a synthetic-vs-natural domain gap that depresses cosine regardless of true fidelity → conclusions were soft. Use ECAPA-TDNN (speechbrainspkrec-ecapa-voxceleb) for rigorous speaker-verification scoring. -
[2026-05-31]Building the dia2-capable image surfaced THREE upstream packaging quirks: (1)pip install -e nari-labs/dia2fails — no PEP 660build_editablehook; (2) plainpip installbuilds an emptyUNKNOWN-0.0.0wheel (base setuptools 59.6 < dia2's required ≥70); (3)--no-depsleavestransformers/sphn/whisper-timestampedmissing. Fix (local/dia:v2): copy the pure-pythondia2/package into site-packages + install ONLY those 3 deps; base torch/numpy already satisfy Dia2. -
[2026-05-31]Dia2 predefined voices (43, baked at/app/voices) are NOT reachable from the/ttsclone path — it resolvesclone_reference_filenameagainst the reference_audio dir ONLY. The OpenAI/v1/audio/speechvoiceparam auto-resolves them (separate code path), which masked the gap. Fix: stage/app/voices/*into/worktank/dia/reference_audio. Lesson: verify on the endpoint the catalog ACTUALLY targets. -
[2026-05-31]voice_mode=clonewith an emptyclone_reference_filename→ asset-engine serializes it as the literal string"undefined"→/tts404. First observed on dia2; worked around in the catalog (default the field to a real voice). [2026-06-01] root cause found — the Kokoro voice-blend widget reading Shoelace.valuebefore hydration (see Current state); the real fix is asset-engine-side and is escalated. -
[2026-05-31]asset-engineservices.schema.jsonis DERIVED (regen from the Pydantic model viadump_schema.py) and had DRIFTED — rejected thelifecyclefield 12/14 services use. RESOLVED: regenerated withCatalogLifecycle+reproducibility.seed_field(b7b2130). Lesson: hand-editingservices.yamlshape without regenerating re-introduces drift. -
[2026-05-31]ttyd-over-TLS forces HTTP/2 (kills ttyd's terminal WebSocket → blank screen); Safari/WebKit never sends HTTP basic-auth on WS upgrades. Both solved for the fleet seat (Caddy forces HTTP/1.1; auth → network-gating) — detail in auto-memoryreference_ttyd_fleet_seat. -
[2026-05-30]esh-docker-vm:hardNFS mounts from 10.0.50.50 froze a container worker in UNKILLABLE D-state when the NAS stalled — only a host reboot clears it. Separately,fstab defaults(no_netdev) made NFS-bind containersExited(255)on reboot. → auto-memoryincident_esh_docker_nfs_boot_race.
29 older entries archived to archival-memory.md.