Files
esh-pfi-infrastructure/persistent-memory.md
T

31 KiB
Raw Blame History

Persistent memory — eshpfi-management

Last updated: 2026-06-03

Repo purpose

Reference workspace for PFI infrastructure: server inventory, canonical Docker Compose stacks, ops playbooks, and conventions. Authoritative copies of compose files live on the servers under /opt/docker/compose/<stack>/; this repo mirrors them for version control, editing, planning, and CI-driven deploys.

Tools and conventions

Sister repos (separate gitea repos, deployed by playbooks here):

Repo Role CI status
vh/task-board MCP + web dashboard for assistant task state (port 7878) push-to-main → CI deploys (2026-04-29)
vh/vor Inquisitor UI sidecar (port 7879) push-to-main → CI deploys (2026-04-29)
vh/nevermore Twice-daily LLM-curated briefing (port 8181, replaces news-digest) push-to-main → CI deploys (2026-04-30)
vh/asset-engine Internal control plane over inference services (port 8200, LAN-direct) push-to-main → CI deploys (2026-05-12)
vh/althing Inter-agent message bus (chamber UI port 7881, forseti + agent-runner daemons, valkey IPC) push-to-main → CI deploys (2026-05-14)
vh/mead-hall Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) push-to-main → CI deploys (2026-05-16)
vh/skaldsong Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) push-to-main → CI deploys (2026-05-19)
vh/worldtree Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration push-to-main → CI deploys
vh/volva Codex peer agent on althing bus (single-turn oracle, systemd daemon on nh3-dev) manual install via deploy/volva.service (2026-05-18)
vh/yt-voice-clipper YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md
  • Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see docs/runbooks/disaster-recovery.md for the blast-radius matrix.

  • pull-hf-repo.yaml is the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at /tank/aimodels/huggingface/" playbook. Supports --var repo_type=model|dataset|space. Replaces ad-hoc huggingface_hub.snapshot_download patterns.

  • Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (key_id 61419c92) at ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin auths against demo only. For personal-instance admin ops, fetch the bootstrap admin per-op via docker exec worldtree-personal-worldtree-api-1 printenv WORLDTREE_BOOTSTRAP_ADMIN_KEY on corviduo-dev. Used for POST /admin/keys, admin diagnostics (/admin/sessions/<id>/{bifrost,tools}, etc.).

  • Per-project user keys against personal Worldtree (issued 2026-05-19): skaldsong:79744637 (nh3-dev iteration), skaldsong:7c1dbbbe (ana-docker prod), althing:50d85460, mead-hall:a360822d. Same user_id=skaldsong across both skaldsong keys → shared Heimdall agent slot; different key_id → independently rotatable. Pattern: mint via /admin/keys, drop value to /tmp/wt-personal-<name>.key mode 600, dev collects + shreds (DO NOT cat to chat transcript).

  • Skaldsong CD pattern (registry-pull). Differs from althing / asset-engine which build-on-host. vh/skaldsong's CI builds and pushes gitea.phasefinal.com/vh/skaldsong:<sha> + :latest; playbooks/deploy-skaldsong.yaml on ana-docker pulls + recreates. SHA-pin only (no :latest health-gated advance yet). Prereq: host needs docker login gitea.phasefinal.com once (read:package PAT) — not currently in the workflow.

  • docker-as-root pattern (for ops that have no admin API, e.g. SqliteUserStore.set_bifrost_credentials): on hosts where the SSH user is in the docker group but lacks passwordless sudo, run docker run --rm -v <target-dir>:/wt -v /var/run/docker.sock:/var/run/docker.sock docker:cli sh -c "..." to edit deploy-owned files without sudo. Documented with security warning in servers/corviduo-dev/README.md. docker-group membership is effectively root via bind-mount; treat as a sudo-equivalent grant. Foot-gun: when running docker compose inside this sandbox, any relative path in compose.yaml (e.g. ${WORLDTREE_CONFIG_DIR:-./config}) resolves against the sandbox CWD, but Docker daemon interprets the resulting path against the HOST filesystem. Always pass -e VAR=/abs/path to the docker run invocation for any relative-default config dir.

  • scripts/elway sudo handling — elway prompts for the sudo password ONCE via getpass before the first sudo: true step. That prompt is interactive → elway can't run unattended from a non-TTY tool if any step needs sudo. For sudo-free playbooks (no sudo: true steps) it runs fully non-interactive over key SSH. To create root-owned dirs WITHOUT host sudo, use the docker-daemon-root trick: docker run --rm -v /worktank:/mnt alpine sh -c 'mkdir -p /mnt/<x> && chown -R 1000:1000 /mnt/<x>'.

Current state / in-flight

As of 2026-06-03:

  • INFRA SESSION 2026-06-03 — NH3 egress proxy, R14 provisioning, yt-voice-clipper CD. Five threads landed (detail in the cited auto-memories + runbooks + commits):

    • NH3 egress proxy (DURABLE). dante SOCKS5 on nh3-dev socks5h://10.100.10.50:1080, internal-only (client ACL 10.100.0.0/16), systemd. NH3 residential egress (~70.230.226.88) for colo services gated on their datacenter IP. scripts/setup-nh3-egress-proxy.sh + docs/runbooks/nh3-egress-proxy.md (f27ee47); auto-memory reference_nh3_egress_proxy.
    • yt-voice-clipper on irv-ml1 (/opt/docker/compose/yt-voice-clipper, :8000 console+API + bgutil pot-provider sidecar). yt-dlp via YTVC_PROXY egress beats YouTube's datacenter bot-gate (proven end-to-end). Console + audition + trim-by-playhead deployed; Homepage label (AI Systems) in the override. Push-to-deploy LIVE: gitea webhook → HMAC listener irv-ml1:9008 → git-clone deploy dir rebuilds on push to main. docs/runbooks/ytvc-autodeploy.md (d4f180d); auto-memory reference_ytvc_autodeploy.
    • R14 VC-NVV tracer (brokkr-smithy) provisioned scope-(a) on irv-ml1: box + CUDA env + 4 engine run-envs (knn-vc/seed-vc/knn-svc/RVC, A6000-validated; one proven on a real VIVAE conversion) + /mnt/smithy nfs4 mount + VIVAE-reachability. brokkr wires arms.py + runs. Env at ~/r14; auto-memory reference_irv_ml1_gpu_r14. OPEN: irv-ml1 git-pull deploy key for vh/brokkr-smithy pending brokkr authorizing the pubkey (tracked in the R14 althing thread).
    • gitea [webhook] ALLOWED_HOST_LIST = external, 10.100.0.0/16 on the gitea container (ana-docker) — the anti-SSRF default external was denying webhooks to the private WG net, blocking the ytvc auto-deploy. Backup app.ini.bak-pre-webhook.
    • irv-ml1 disk 97%→49% — cleared all ollama models (~194 GB) + 2023-era /bootstore/llm-models/Storage/ LLMs (~472 GB; 2 llmuser-owned dirs needed operator sudo).
  • chatterbox-fast — DONE (was the primary focus). Deployed irv-ml1 :8197, cataloged ready, EXTRACTED to its own MIT repo vh/chatterbox-fast (547a557) — the eshpfi stack is now a pointer; auto-memory reference_chatterbox_fast_repo. Adaptive buffer-ratchet chunking, TTFB ~0.5s, A6000 fp32 5.34 GB. (Historical build detail follows.) Operator-authorized high-effort build (incl. custom container from source). Plan-of-attack: docs/design/chatterbox-fast-plan.md (durable; carries the full executable detail). The old docs/design/chatterbox-fast.md was KILLED 2026-06-01 — its §5 windowed-token-streaming was the abandoned native arc; superseded by the adaptive-chunk plan. DEPLOYED 2026-06-02 — Phases 1-3 done: stacks/chatterbox-fast/ (scheduler + app + sim tests + Dockerfile/compose), live on irv-ml1 :8197 (container chatterbox-fast, healthy) BESIDE the live chatterbox :8196. Measured: TTFB ~0.5s (vs ~5.2s one-shot), no starvation, RTF ~3.9-4.0× on A6000. GPU = A6000 (device 1) ONLY — turbo loads fp32 (NOT the ~2.5GB fp16 old notes claimed), measured 5.34 GB, so the 3090's ~3.8GB free does NOT fit. Perf: TF32/SDPA on (don't move TTFA — AR-decode-bound; bf16 deferred, no clean fp32→bf16 cast path). Context-priming for join coherence TRIED + REVERTED (discard-cut leaks prefix — see Tried/abandoned); coherence loss at joins is an ACCEPTED limitation (operator: cold streaming "really good"). Remaining: Phase 4 catalog entry + cutover; optional deploy-playbook smoke gate; optional §5 base-fork 3-way A/B.

    • Goal: cut time-to-first-audio from ~2.5s → sub-second via streaming, keep turbo quality. Workload = single-stream interactive.
    • Chosen approach = adaptive buffer-ratchet chunking (operator's idea): generate the FIRST sentence alone for instant response (~0.66s first-audio measured), then while it plays, generate the largest sentence-aligned chunk that fits in ~80% of the bought playback time; chunks grow ~3× each (driven by Chatterbox's ~3.8× realtime headroom), so after 2-3 chunks the rest of the paragraph is one big chunk with near-full context. Context loss confined to 2-3 joins at natural sentence-pause boundaries. THIS is "sentence-level done right" and supersedes naive per-sentence splitting (which loses cross-sentence prosodic context = real quality loss, operator-corrected). Only works because RTF>1 (Fish at <1× realtime starves — can't use this).
    • Native frame-level streaming on turbo = ABANDONED (see Tried/abandoned).
    • Also to build for A/B: base-chatterbox + davidbrowne17 streaming fork (true frame-stream but base-model quality) — not yet installed.
    • A/B samples (GLaDOS voice) on nh3-dev ~/chatterbox-ab/: sentence-level turbo, chunked-native (artifacty), chunked-oneshot.
    • Dev/test pattern: one-off GPU-1 container from local/chatterbox:v1, mount /worktank/chatterbox/{cache,reference_audio} + the test script; docker exec -i chatterbox python - for lib introspection. lkraven is in the docker group on irv-ml1 — NO sudo needed for docker.
  • TTS fleet on irv-ml1 reworked this session; asset-engine catalog now 17 services (CSM removed). Canonical docs/asset-engine/services.yaml

    • vendored vh/asset-engine both pushed (eshpfi 38d9e3b; asset-engine 0b552bd). Per-engine live state:
    • dia2-2b (:8200) + dia2-1b (:8202) — LIVE + healthy (restarted this session; local/dia:v2, voices + cache intact). status: ready.
    • Chatterbox (:8196) — LIVE. Catalog v2: routes to /tts, exposes emotion levers (exaggeration 0.5, cfg_weight 0.5, temperature 0.8), 28 built-in predefined voices (default Emily.wav) + clone. Fast + competent cloner (it audibly cloned Imogen). Plain Imogen.wav uploaded to its clone dropdown (/get_reference_files now: Imogen.wav + glados_25s.wav).
    • Fish-S2 (:8195) — LIVE but HEAVY: ~240s torch.compile warmup on cold start, ~24 GB on the A6000; first call after any restart eats the warmup. Catalog v6: full /tts ServeTTSRequest surface (temperature/top_p/repetition_penalty/latency/seed/format) + reference_id is a 32-voice static "Voice" dropdown (default Emily) + advanced base64 clone. NOT OpenAI-compat (only /v1/tts, no /openapi.json, no /voices API). Research/non-commercial license.
    • zonos (:8203) down (adapter built, not deployed). ComfyUI :8188 catalog-deferred link.
  • "undefined" select bug — RESOLVED. asset-engine-dev shipped form-select hardening (v0.1.9/.10) + a durable per-field enable-toggle (togglable, v0.1.14/.16) — the real fix for the "form submits untouched fields" family. My catalog half: opted fish-s2 references into togglable: true (catalog_version 1→2, schema gained the togglable boolean; c98a12b, pushed to origin). The earlier blendable idea was superseded by togglable. dia2/chatterbox clone fields deliberately NOT toggled (dia2 defaults clone-on as its stable voice). asset-engine CI was briefly red because the commit was local-only until pushed — lesson: push catalog commits promptly (their CI drift-checks against the remote).

  • Fish "not British" had TWO independent root causes — BOTH now fixed. The ECAPA-TDNN re-test (2026-06-01) showed Fish's cloning engine works: an Imogen reference scores ~0.79 cosine vs real Imogen.wav vs ~0.10 vs the no-reference default — but that test fed the reference via the inline base64 references path, which masked the real defect. reference_id-by- name (the path the catalog + form actually use) was silently a NO-OP. Proof (2026-06-01): reference_id=Abigail, =Imogen, and no-reference all returned byte-identical audio (md5 40b6ac73…, same text+seed) → reference_id was being ignored entirely, so every dropdown voice collapsed to Fish's default speaker. Cause: fish-speech resolves reference_id=<name> against the directory references/<name>/ (audio + same-basename .lab), but the voices were staged as flat references/<name>.wav + .txt with the per-name dirs left empty → nothing resolved. FIX (2026-06-01): populated references/<name>/<name>.wav + <name>.lab for all 32 voices; re-test → Imogen/Eleanor/Beatrice/Abigail/no-ref all 5 distinct. reference_id now works; British accents come through. Hardened: the deploy playbook gained a normalize-layout step + an A/B smoke gate (reference_id MUST change output) so the empty-dir regression can't ship silently. So the "not British" saga was (A) the "undefined" select bug (asset-engine, still belt-and-suspenders) AND (B) this reference_id no-op (server-side, FIXED) — B alone would have kept every voice identical even with A fixed. Earlier "reference_id-by-name is THE working path, verified" claim RETRACTED. Pitch-shift deepening abandoned; future clones operator- handled.

  • On-host consented voice library — ~992 real-person clips cached in the kyutai tts-voices repo (/worktank/kyutai-tts/.../snapshots/.../): VCTK (CC BY 4.0, accent-tagged speaker IDs), Unmute voice-donations (CC0), EARS + expresso (CC BY-NC). Source for future clone voices. British-female Southern England speakers p225/p228/p229 staged into Fish as Imogen/Eleanor/Beatrice.

  • Parakeet ASR (:8765) is now UP (brought online 2026-06-02 to transcribe the glados clip; operator wants it kept online). CPU-only (gpu_device_id: null), restart=unless-stopped, healthy. API: POST multipart file to /transcribe{"text":...}. Image local/parakeet:sherpa-onnx-v2 pre-built; docker compose up -d in /opt/docker/compose/parakeet.

  • NFS share for Smithy — DONE. /volume1/smithy on nh3-nas → /mnt/smithy on nh3-dev (single export, RW, map-to-admin, scoped to nh3-dev only, fstab _netdev,nofail,x-systemd.automount). Subdirs datasets/{raw,manifest-store, derived,holdout,quarantine}; rename-atomicity verified. Smithy wired storage-roots.yaml + E2E-verified. Backup DEFERRED per operator (datasets regenerable; raw + manifest-store are the irreplaceable surface for when it's wired — durable tier should land OFF nh3-nas/cross-site, since the NH3 restic repo lives on the same NAS as the source).

  • Worldtree healthy v0.29.13 (last-known); Skaldsong v0.32.2 (ana-docker:8300, Kokoro SSE streaming); artemis-31b-v1i live on llama-swap + worldtree personal; ttyd fleet driver seat on nh3-dev (RO :7681 / drive :7682 via Caddy) — all out of this repo, last-known.

  • Disclosed-keys hygiene queue — rotate at convenience: HF token hf_HBl… (lkraven's HF account) leaked into BuildKit logs during the CSM build attempt (logs shredded, never committed — low urgency); /tmp/wt-personal-skaldsong-prod.key on nh3-dev; mead-hall's prior Worldtree bearer (superseded by a360822d); Worldtree Z_AI_API_KEY/ZAI_API_KEY; chamber forseti/agent_runner api_keys (superseded by 50d85460); Gitea runner registration token (a1135753…).

  • Still open from prior sessions: rotate MINIFLUX_PASSWORD (leaked twice); clean up legacy news-digest detritus on ana-docker; watch nh3-nas /volume1 (was 65%; recheck before ~80%); the docker push 60s client-side ceiling mystery remains uninstrumented.

Recent decisions

  • [2026-06-03] yt-voice-clipper bot-gate fix = route yt-dlp through NH3 residential egress, NOT cookies/PO-token. YouTube hard-flags the Irvine colo IP (LOGIN_REQUIRED on a public video even with no cookies). Cookies + the bgutil PO-token + deno JS-runtime all loaded fine — the gate is pure IP reputation. Operator chose proxy-via-nh3-dev → durable dante proxy → proven. The egress proxy is a reusable fleet lever for any datacenter-IP-gated service.

  • [2026-06-03] yt-voice-clipper push-to-deploy via gitea webhook (operator-directed, after 6 manual rebuilds in ~40 min). Webhook (not poll) — gitea CAN reach the WG IP per the operator. The proxy env + Homepage labels live in the host-specific override (untracked → survive the auto-deploy's git reset --hard), NOT yt-voice-clipper-dev's image. Runbook d4f180d.

  • [2026-06-03] R14 scope = (a) provision-only. infra-ops provides box + CUDA env + engines + weights + NFS; brokkr/dev wires arms.py + runs — keeps infra-ops OFF the VIVAE-processing path (VIVAE = Variably Intense Vocalizations of Affect/Emotion, CHARTER §4 highest-liability; operator authorized R&D-eval-only, quarantined). Box = irv-ml1 (A6000 free; ana-ml2 GPU-saturated). Per-engine venvs (divergent torch stacks); A6000 = cuda:0 NATIVE (≠ docker =1).

  • [2026-06-03] Declined worldtree v0.30.4 staging deploy — that's worldtree-team's CI/CD lane (a developer staging/vX.Y.Z git-tag promote), not infra-ops. They self-corrected to the same conclusion independently.

  • [2026-06-02] Chatterbox → main TTS engine; build custom chatterbox-fast streaming container. Workload = single-stream interactive. GPU placement: 3090 (device 0) if it fits else A6000 (device 1) — shared dev stack, 20.5 GB 3090-idle is expected residency, not a blocker. Cutover: parallel catalog entry, burn in beside live chatterbox, then flip. Streaming approach: adaptive buffer-ratchet chunking (see in-flight). Native frame-streaming abandoned (Tried/abandoned). Tracked: docs/design/chatterbox-fast-plan.md.

  • [2026-06-02] Sentence-splitting loses quality (operator-corrected). I claimed naive sentence-level streaming has "zero quality loss" — WRONG. The T3 AR backbone conditions prosody on the WHOLE text; splitting loses cross-sentence prosodic context (contextual delivery, declination, affect continuity) even though voice timbre stays (reference-conditioned). No artifacts ≠ no quality loss. Hence the adaptive-chunk design (maximize context per chunk subject to latency budget), not fixed per-sentence splits.

  • [2026-06-01] Fish reference_id empty-dir fix shipped (c5bbb90) — see in-flight + Tried/abandoned. Populated references/<name>/<name>.wav+.lab for all 32 voices; playbook gained normalize-step + A/B smoke gate. glados got a real transcript (ASR'd via Parakeet): the Portal "Welcome to test chamber 4" lines.

  • [2026-06-01] Fish cloning VERIFIED competent (ECAPA-TDNN) — retracting the earlier "weak cloner" call. Isolated test: Imogen-referenced clone ~0.79 cosine to the real Imogen.wav vs ~0.10 for the no-reference default; transcript condition (correct 0.787 / bogus 0.778 / empty 0.738) barely moves identity (affects pronunciation, not timbre). Root cause of "Imogen sounds nothing like British" = the "undefined" select bug feeding Fish its default voice, NOT Fish. So the entire Fish-Imogen saga was the undefined bug; no Fish-side fix needed. (Methodology lessons → Tried and abandoned.)

  • [2026-06-01] CSM (Sesame csm-1b) torn down entirely — removed from catalog, stacks/csm/, playbooks/deploy-csm.yaml, and host (c54ab13). Two reasons: (1) deep-research verdict — the acclaimed Maya/Miles demo is a fine-tuned, larger variant Sesame never open-sourced; the open 1B base ships no usable voices, can't generate text, English-only, no real-time streaming, absent from leaderboards, dominated by Kokoro/Dia2/Fish/IndexTTS. (2) The upstream image is rotted (see Tried and abandoned). asset-engine re-synced (306aae9).

  • [2026-06-01] Chatterbox catalog v1→v2 (93cedd6) — re-point to the devnen wrapper's /tts (was the thin OpenAI /v1/audio/speech exposing none of the emotion knobs → poor output). Default exaggeration: 0.5 per Resemble's README ("works well for most"); the server SHIPS demo-tuned 1.3 which over-acts — deliberately NOT adopted. 28 predefined voices via /get_predefined_voices (default Emily.wav). Corrected seedable false→true.

  • [2026-06-01] Fish-S2 catalog v1→v6 (f3c260e..38d9e3b) — was a blank 2-field form; now exposes the full fish-speech ServeTTSRequest surface

    • a 32-voice reference_id "Voice" dropdown (STATIC list — Fish has no /voices API). reference_id-by-name works ONLY when each voice is staged as a references/<name>/ dir with <name>.wav + <name>.lab — was silently a no-op until that layout fix on 2026-06-01 (see the two-root-causes entry above); the inline-path-as-filesystem-path references shape 500s (see Tried and abandoned). Corrected vram 4→24 GB, cold_start 8→240s, seedable false→true. Added 3 consenting British-female VCTK voices (Imogen=p225/Eleanor=p228/Beatrice=p229).
  • [2026-06-01] blendable: false catalog flag — DEFERRED, pending asset-engine-dev honoring it in the form template. Clone-voice selects (chatterbox/fish/dia2) must render as plain dropdowns, not the Kokoro blend widget (blending reference clones is nonsensical). I ship the catalog flag + schema field once they read it. Tracked: althing msg 01KT2K2SY9N7AY69R9V0B4RXSW to asset-engine-dev.

32 older entries archived to archival-memory.md.

Tried and abandoned

  • [2026-06-03] gitea webhook to a private IP is denied by webhook.ALLOWED_HOST_LIST (anti-SSRF; default external blocks private/loopback). Symptom: delivery shows dial tcp ...: webhook can only call allowed HTTP servers. Fix = APPEND the target net to ALLOWED_HOST_LIST in gitea's app.ini (keep external; scope tight, never */private) + restart gitea (act_runner job containers survive a restart). gitea runs as a container on ana-docker (gitea_gitea_data volume, /data/gitea/conf/app.ini).

  • [2026-06-03] torch-2.12 venvs need uv pip install torchcodec — torchaudio 2.12 defaults to the TorchCodec backend for .load; without it, real audio I/O throws "TorchCodec is required" — and it ONLY surfaces at actual conversion, NOT at import/model-load. Lesson: validate real I/O, not just import, when provisioning ML engine envs. (seed-vc on torch 2.4 uses the legacy backend, exempt.)

  • [2026-06-03] Backgrounding althing-cli monitor with an inline shell & (instead of the Bash-tool run_in_background) orphans it — it survives the shell exit, holds the per-handle flock UNTRACKED (won't notify the session), and stop-monitor doesn't detect it. Fix: find + kill the orphan PID (verify cwd=this repo / handle first — nh3-dev is shared, other agents' monitors run there too), then re-arm via run_in_background. Always re-arm tracked.

  • [2026-06-03] uv pip install . fails on SmoothKen/knn-svc (and similar script-repos) — it's analysis scripts + a poetry pyproject, no buildable package (setuptools package-discovery error). Install the pyproject deps directly, don't build the "package".

  • [2026-06-02] Fish (fish-s2 / OpenAudio S1-mini) progressive streaming — SHELVED (sub-realtime). Benched RTF on A6000: 0.72x (12w) / 0.82x (30w) / 0.86x (60w), mean 0.80x = sub-realtime, so client-side chunking would starve (same reason chatterbox-fast needs turbo's RTF>1). Root cause of the buffering (dvalin-smithy-dev deep research, verified in our code text2semantic/inference.py L600-607): Fish only chunks on <|speaker:X|> tags; plain text -> batches=[whole text] -> all semantic tokens generate before any audio (chunk_length inert). Plus a 2nd layer: kui/ASGI StreamResponse doesn't flush (header produced t=1s, delivered t=23s) -> fix = anti-buffering headers (X-Accel-Buffering:no / Transfer-Encoding:chunked) in tools/server/views.py (kept on file, not applied). A rebuild does NOT fix this (current main same logic). STANDING REVISIT TRIGGER: when an RTX Blackwell Pro lands in the fleet -> bench fp4-quantized Fish; if RTF > ~1.5x, give it the chatterbox-fast treatment (client-side adaptive buffer-ratchet chunker driving /v1/tts with small text pieces). Projection: fp4 (~1/4 weight bytes, memory-bound AR decode) + Blackwell (GDDR7 ~1.8TB/s vs A6000 0.77TB/s, native FP4 cores) ~ 2-3x RTF; validate fp4 voice quality (ear/ECAPA) before committing. For now Fish stays a buffered catalog entry (great for SAVED gens, not the live-audition lane).

  • [2026-06-02] Context-priming at chunk joins (chatterbox-fast §1.6) — ABANDONED (discard-cut leaks the prefix). To give a chunk backward prosodic context, prepend the prior sentence, generate prefix+content together, then discard the prefix audio. Built + opt-in shipped (commit d707439), live-A/B'd, reverted (090e70a). The kill: generate() returns one finished waveform with NO marker for where the prefix ends, and the model renders the same prefix with different timing solo vs followed-by-content — so locating the cut (generate prefix solo → measure duration → snap to nearest energy-min pause within ±0.4s) is a guess that left a whole clause of prefix in the output ("...without a trace of sarcasm," spoken twice; operator caught it). A reliable cut needs token-level boundaries (= the abandoned native-streaming arc) or per-chunk ASR/forced- alignment (heavy, imperfect, eats the latency budget). → Coherence loss at joins stays an ACCEPTED limitation; cold adaptive-chunk streaming judged "really good". Scheduler-side work that DID land + survive: affordability-gated priming math (a 2nd pass can't starve the buffer) — sound, but moot without a working cut.

  • [2026-06-02] Native frame-level streaming on Chatterbox-TURBO — ABANDONED (turbo isn't built for streaming). Long R&D arc; record so it's not re-derived. (1) The model's flow is CosyVoice2-derived but S3GenStreamer is referenced-in-docstring-only (not implemented). (2) The lib's flow_inference(finalize=False) is BUGGY: the lookahead trim removes pre_lookahead_len(3)*token_mel_ratio(2)=6 frames from h but NOT from h_masks/conds → decoder shape mismatch (e.g. 656 vs 662). A 1-line patch (h_masks = h_masks[:, :, :-pre*ratio] after the h trim) + sizing the meanflow noise to the trimmed length makes finalize=False RUN. (3) BUT the flow encoder uses FULL-context attention (static_chunk_size=0), so incremental/cumulative decode is prefix-unstable — adding tokens re-attends and shifts earlier mel (maxdiff ~0.30-0.39 vs one-shot, irrespective of fixed-noise slicing or emit-margin). (4) Forcing static_chunk_size>0 on the 2 modules that carry the attr did NOT stabilize it (decoding_chunk_size is a forward-arg, not settable via attribute). Verdict: true sub-second frame-streaming on turbo needs deep model-attention surgery with quality risk — not worth it. Matches research ("turbo+streaming unsolved"; vLLM-turbo outputs noise; davidbrowne17 streaming fork is BASE-only). → Use adaptive-chunking instead.

  • [2026-06-02] Naive cumulative re-decode for streaming — decode the growing token prefix each chunk + emit the delta, assuming the causal flow gives a stable prefix. It does NOT (full-context attention, see above); maxdiff 0.30. Also the high-level s3gen.inference re-applies trim_fade at the start every call. Don't go this way.

  • [2026-06-01] CSM bring-up — upstream phildougherty/sesame_csm_openai Dockerfile pins NO huggingface_hub version, so it now resolves to 1.17.0 where huggingface-cli is fully removed (replaced by hf) → the build dies at huggingface-cli login/download. Passing the HF token as a build-arg got the token in (the model downloads at BUILD time, not runtime — our compose only wired it as runtime env, the original failure) but then hit the rotted CLI. Fixing would need vendoring + patching the Dockerfile; not worth it for a weak model → CSM removed.

  • [2026-06-01] Fish path-form references ({audio:"/app/references/x.wav"}) → HTTP 500 on this build, despite the stack README showing that shape. The working clone paths are reference_id (by name → resolves the DIRECTORY references/<name>/ containing <name>.wav + <name>.lab; the flat references/<name>.wav is ignored — see the two-root-causes entry above) or inline base64 references. The catalog uses reference_id.

  • [2026-06-01] Reference transcript barely affects Fish clone IDENTITY (disproving my mid-session theory). I'd blamed a bogus provenance-note .txt for poor cloning, but the ECAPA re-test showed correct (0.787) / bogus (0.778) / empty (0.738) transcripts all clone Imogen about equally — the transcript affects PRONUNCIATION (phoneme disambiguation per the docs), not who it sounds like. The real culprit for "not British" was the "undefined" select bug, not the transcript. (A correct transcript still marginally helps pronunciation — cheap to stage, not load-bearing.)

  • [2026-06-01] Pitch-shift register control (rubberband, to deepen Imogen to contralto/mezzo) — Fish ignores small reference shifts and overshoots large ones; -3st WITHOUT formant preservation → "baritone through a tuba" (formants collapsed); -2/-3st WITH --formant → too subtle / still bad. Abandoned at every depth; all variants deleted. Finer independent pitch/formant control needs praat (not installed). Future clones = operator's.

  • [2026-06-01] resemblyzer is too weak for cloning-fidelity scoring — its dated 2019 LSTM encoder + a synthetic-vs-natural domain gap scored the Imogen clone CLOSER to the default than to real-Imogen, which led me to a WRONG "Fish is a weak cloner" call. ECAPA-TDNN (speechbrain spkrec-ecapa-voxceleb) on the same clips gave the correct answer (clone 0.79 to real Imogen, 0.10 to default). Use ECAPA, not resemblyzer, for speaker-verification.

35 older entries archived to archival-memory.md.