Files
esh-pfi-infrastructure/persistent-memory.md
T
vh 89611eb06b memory: snapshot — Zonos emotion-tuning + voice-cloning (8 voices, dial-in studio, emotion canonical) for /clear
Rewrote in-flight for the Zonos character-voice work: 4 cloned voices + host-managed
gateway voices, streaming dial-in studio (source saved to ~/development/zonos-tools/),
and the empirical emotion sweep canonical (single-emotion, two-regime accurate/expressive;
happy/sad usable, angry/surprised broken on named dirs -> axes sweep next). Captured #365
closed + WT#368 forensics + personal agent-memory scrub. Open loops: yt-voice-clipper
yields test, dvalin axes-sweep numbers, re-arm monitor + read mail.
2026-07-18 00:25:18 -07:00

30 KiB
Raw Blame History

Persistent memory — eshpfi-management

Last updated: 2026-07-18

Always check for /tmp/infra-ops-handoff.md — if it exists and its Written: stamp is under an hour old, read it (it carries the in-flight handoff from the previous session), then delete it. Older than an hour: stale — delete it unread.

Repo purpose

Reference workspace for PFI infrastructure: server inventory, canonical Docker Compose stacks, ops playbooks, and conventions. Authoritative copies of compose files live on the servers under /opt/docker/compose/<stack>/; this repo mirrors them for version control, editing, planning, and CI-driven deploys. It was originally spun up to handle the fleet backups — keep that lens when triaging backup/storage issues.

Tools and conventions

Sister repos (separate gitea repos, deployed by playbooks here):

Repo Role CI status
vh/task-board MCP + web dashboard for assistant task state (port 7878) push-to-main → CI deploys (2026-04-29)
vh/vor Inquisitor UI sidecar (port 7879) push-to-main → CI deploys (2026-04-29)
vh/nevermore Twice-daily LLM-curated briefing (port 8181, replaces news-digest) push-to-main → CI deploys (2026-04-30)
vh/asset-engine Internal control plane over inference services (port 8200, LAN-direct) push-to-main → CI deploys (2026-05-12)
vh/althing Lean trusted inter-agent message bus — v2 "email model" (v2.0.0b2, 2026-07): per-box local-SQLite bus + courier/receiver for P2P over the 10.x net; pillars = open-loops / per-box herald + wake-listener / roaming owner API /owner/* / althing-mcp stdio surface. The v0.15 lean-bus cut RIPPED moderation / chamber / forseti-daemon / agent-runner / redis-valkey. per-box uv tool install (NOT CI-deploy); nh3-dev = the DEV box (editable install of ~/development/althing, gets new versions first); nh3-extdev a mesh peer (model B: althing-svc + shared /srv/althing)
vh/mead-hall Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) push-to-main → CI deploys (2026-05-16)
vh/skaldsong Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) push-to-main → CI deploys (2026-05-19)
vh/Worldtree Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. gitea-runner builds on ana-docker; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. push-to-main → CI build-and-deploy (runner on ana-docker)
vh/yt-voice-clipper YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md
vh/arbo Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook
vh/zonos-gateway OpenAI-compatible TTS gateway over stock ZONOS2 (:8890 irv-ml1); emotion dials-first + voice mapping; reached via LiteLLM ext-tts alias stood up 2026-07-17 (private; internal SSH); NOT yet CI-wired (deployed /opt/docker/compose/zonos-gateway is a SEPARATE copy from the repo — CI + deploy key = open follow-up). Dials-first spec at docs/EMOTION-DIALS-SPEC.md. 2026-07-18: host-managed voices bind-mount (./voices:/app/voices) → 8 voices incl. 4 cloned chars (Emmie/Penny/Natalie/Miranda); voice wavs committed
model-training-forge (mtf-dev) Fine-tuning recipe forge; T1 = E-RP writing LoRA, retargeted qwopus-122B→AEON-27B (2026-07-06) (SFT→DPO, LitBench-RM reward) training runs, not a deployed sidecar

(vh/volva + Heid were re-architected from systemd daemons to Claude Code session orchestrators 2026-06-08; their nh3-dev .service units were removed — no longer deployed sidecars here. See Recent decisions.)

  • Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see docs/runbooks/disaster-recovery.md for the blast-radius matrix. ⚠️ The restic file+DB layer routes through TWO rest-servers (rest-server-ana @ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas; rest-server-nh3 @ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of /mnt/backup. (rest-server-ana recovered 2026-06-20.)

  • pull-hf-repo.yaml is the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at /tank/aimodels/huggingface/" playbook. Supports --var repo_type=model|dataset|space. Replaces ad-hoc huggingface_hub.snapshot_download patterns.

  • Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (key_id 61419c92) at ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin auths against demo only. Personal-instance admin (the ~/.config/worldtree/personal-admin-token, mode 600) POSTs /admin/keys (mints per-project keys; takes user_id+label, no scope param — scopes are tier-derived). On-instance mint recipe (cleaner than DB-manip): docker exec worldtree-worldtree-api-1 POST /admin/keys with the in-container WORLDTREE_BOOTSTRAP_ADMIN_KEY; cleartext once in .key=wt_live_+16hex. auto-memory reference_worldtree_demo_key_mint.

  • Per-project user keys against personal Worldtree (issued 2026-05-19): skaldsong:79744637, skaldsong:7c1dbbbe, althing:50d85460, mead-hall:a360822d. Mint via /admin/keys, drop value to /tmp/wt-personal-<name>.key mode 600, dev collects + shreds (DO NOT cat to chat transcript).

  • Skaldsong CD pattern (registry-pull). vh/skaldsong's CI builds and pushes gitea.phasefinal.com/vh/skaldsong:<sha> + :latest; playbooks/deploy-skaldsong.yaml on ana-docker pulls + recreates. SHA-pin only. Prereq: host needs docker login gitea.phasefinal.com once.

  • gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH 10.250.50.70:222, HTTP :3000. Fleet/colo hosts must use this internal route, NOT public gitea.phasefinal.com (38.120.12.44) — the public path fail2bans the host egress IP. Full gotcha in docs/orientation.md → Git/gitea.

  • docker-as-root pattern (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo): docker run --rm -v <target-dir>:/wt docker:cli sh -c "...". docker-group membership is effectively root via bind-mount. Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass -e VAR=/abs/path for any relative-default config dir.

  • scripts/elway sudo handling — elway prompts for the sudo password ONCE via getpass before the first sudo: true step → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH.

  • Per-host SSH identity matters for sudo. infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: default ssh ana-docker = lkraven (docker-group, NO passwordless sudo); ssh infra-ops@ana-docker HAS NOPASSWD root. → For any sudo op on ana-docker, use ssh infra-ops@ana-docker. ssh infra-ops@10.100.10.50 (nh3-dev) ALSO NOPASSWD sudo; on nh3-extdev infra-ops is sudo-LESS by design (ssh lkraven@10.100.50.42 is the NOPASSWD path). irv-ml1: ssh irv-ml1 = lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD) — stage model pulls to /home, not root-owned /worktank.

Current state / in-flight

As of 2026-07-18 — active work = Zonos2 emotion-tuning + voice-cloning for the character-voice product. Snapshot for a /clear; operator: "read mail after."

Zonos voice + emotion state (all LIVE):

  • 8 voices in zonos-gateway (:8890 irv-ml1): defaults AmericanFemale/Male/British/Cora + 4 CLONED characters Emmie/Penny/Natalie/Miranda. Add a voice = drop <Name>.wav in /opt/docker/compose/zonos-gateway/voices/ + docker compose restart zonos-gateway (host-managed bind-mount; NO rebuild). All voice wavs committed to vh/zonos-gateway.
  • Voice-cloning pipeline: /mnt/smithy/voice_clones/<name>.zip (diarized + WhisperX-scored clips) → ~/development/zonos-tools/assemble_voice.py <dir> → top-mean_score clips to ~1524s → drop in gateway voices. (/mnt/smithy = irv-ml1 NFS from nh3-nas; remount post-reboot.) Full pipeline + emotion detail in the [2026-07-18] Recent-decisions entry.
  • Dial-in studio at http://10.100.10.50:8898/ — streaming Web Audio (~0.6s first-audio), voice dropdown (all gateway voices), every Zonos dial as a slider + Reset + Copy-JSON; proxies the gateway. ⚠️ It's a nohup'd session process on nh3-dev (NOT a service); source now durable at ~/development/zonos-tools/dial-in-studio.py — relaunch nohup python3 ~/development/zonos-tools/dial-in-studio.py >/tmp/zonos-studio.log 2>&1 &. Systemd-ize if reboot-survival wanted.
  • Emotion CANONICAL established (empirical sweep, [2026-07-18] entry): single-emotion only; two-regime accurate(identity)/expressive(drama) policy; happy/sad usable, angry weak, surprised dead on the named directions.

Open loops for the fresh session:

  • yt-voice-clipper yields test — job f3ff746dbae9494d running on irv-ml1 (submitted to verify v0.3.3's max_gap fix). Check ssh irv-ml1 'curl -s :8000/jobs/f3ff746dbae9494d' (or /diagnostics) for ~14 segments (not 2), then reply to yt-voice-clipper-dev (thread 01KXT0T6GYHB) confirming verify #3. (Redeploy #1 A6000 + #2 version-0.3.3 already confirmed in 01KXT1BFTA5J.)
  • dvalin owes-me / I-owe-dvalin the axes-sweep numbers (thread 01KXT12FN0AS5A3WMKEK06BVPS) — I committed to run the axes sweep and send results.
  • NEXT experiment (operator to green-light): axes sweep for angry/surprised — valence/arousal grid (angry ≈ val/+aro; surprised ≈ +aro), the only path to rescue the two broken named emotions; then strength ladder + emotion-congruent text. Reuse ~/development/zonos-tools/emotion_sweep.py (scoring env: uv run --with resemblyzer --with funasr --with "numpy<2" --with soundfile --with requests --with "setuptools<80" --with torchaudio). Then bake happy/sad canonical into gateway presets.
  • Re-arm the althing monitor (/althing:monitor, handle infra-ops) — wake-listener dies on /clear; "read mail after" per operator. Open watch: worldtree-dev (#368 memory-leak, my read-only forensics done; #363 research-wing ingest PARKED).
  • eshpfi commits UNPUSHED — the zonos-gateway capture + memory (14a0004..438cd35) are committed; push is the operator's call. stacks/heretic2-charrp-reasoning/ still UNTRACKED; graphify-out/GRAPH_REPORT.md modified.
  • irv-ml1 3090 oversubscription (kokoro :8193 + vibevoicefusion :9527 idle-pinned; zonos engine :1920 also on 3090) — carried; operator declined to fix.

Landed this session (2026-07-17→18): #365 demo+personal → b125 (byte-exact, CLOSED — [2026-07-17] entry); PERSONAL agent-memory SCRUBBED (WT#368 remediation, rollback /opt/worldtree-personal/agent-memory-backup-20260717-181004.tar.gz); zonos-gateway repo stood up + dials-first spec + host-managed voices; 4 character voices cloned; dial-in studio (streaming); emotion sweep + canonical; yt-voice-clipper → A6000-pinned + v0.3.3.

Carried standing (non-blocking): ana-ml2 GPU0 ~14 G reserve; Worldtree #363 research-wing ingest (auto-memory, no deadline); T1 SFT LoRA dormant; rotate the 5 rest-server backup creds (operator, offline). Zonos2 engine still NATIVE (containerize deprioritized — priming was flat, emotion-steering is the lever).

Recent decisions

  • [2026-07-18] Zonos2 emotion CANONICAL from an empirical sweep + the voice-cloning pipeline — 4 chars cloned (Emmie/Penny/Natalie/Miranda), host-managed gateway voices, two-regime accurate/expressive policy, happy/sad usable + angry-weak/surprised-dead on named directions, dvalin-synthesized; axes sweep is the NEXT experiment. Studio + sweep tooling at ~/development/zonos-tools/. → persistent-memory.d/2026-07-18-zonos-emotion-canonical.md

  • [2026-07-18] yt-voice-clipper: A6000-pin fix + v0.3.3 redeploy. Fixed a latent misconfig — the host override said "pin worker to A6000" but NVIDIA_VISIBLE_DEVICES was "0" (the 3090); re-pinned worker+api to the A6000 by UUID (GPU-9672f0d5, 3090 is zonos2's). Then redeployed api+worker to v0.3.3 (docker compose up -d --build; SPA+Python; max_gap 0.6→1.2s; stderr surfaced in job.log). A6000 + version verified; yields test in-flight (job f3ff746dbae9494d). yt-voice-clipper-dev thread 01KXT0T6GYHB. reference_ytvc_autodeploy

  • [2026-07-17] Worldtree #365 internal-comms config CLOSED (demo+personal → b125) + WT#368 cross-agent memory-leak forensics + PERSONAL agent-memory scrub. #365: staged the internal-tiers/rules/gate on both instances' bind-mounts (byte-exact vs baked b125), both now live on b125. WT#368 (read-only): the operator's name was in NO recall store on demo; on PERSONAL it sat in lofn.chroma (old-code saga-v1 seeding + legacy contamination), and a clean-slate marker test proved current b125 code isolates character-session extraction correctly — the leak is legacy data, not a live bug. Operator-directed → executed a full PERSONAL agent-memory scrub (backup /opt/worldtree-personal/agent-memory-backup-20260717-181004.tar.gz; conversations/mood/auth preserved). worldtree-dev owns the code-fix/data contract. reference_corviduo_dev_emergency_ops

  • [2026-07-17] Zonos emotion levers RESOLVED: text-priming is FLAT → the working lever is ZONOS2's native emotion-steering, which the gateway ALREADY exposes as presets. The prosody-priming A/B (prime→generate→excise, silence-gap cut, parakeet-validated) was operator-judged FLAT on this checkpoint — text doesn't move it. Native emotion_directions/ (happy/sad/angry/surprised + valence/arousal axes, per-speaker calibrated for AmericanFemale/Male/British) clearly WORKS (sad→slow/quiet, excited→fast/bright, etc.). zonos-gateway:0.2.0 (:8890) already wires it: simplest caller path = POST /v1/audio/speech {preset:"…"} — presets neutral/warm/excited/sad/intense/whisper (defined in ~/zonos-gateway/src/zonos_gateway/dials.py), reached via the LiteLLM ext-tts alias (engine-neutral swap point; consumers never call the gateway by name). RTF measured on 3090: cfg1.0 steering = FREE (~0.52 = neutral, additive vectors), cfg1.5 amplified 0.625 (+20%, still realtime). Captured the live gateway stack → stacks/zonos-gateway/ (compose+env+README); ⚠️ gateway SOURCE at ~/zonos-gateway on irv-ml1 is NOT in gitea (backup gap, follow-up); stacks/zonos (v0.1 Gradio) marked DEAD/superseded. Whisper is a composed preset (no whisper direction; escalation for hard affects = custom directions via scripts/build_emotion_directions.py or emotional-ref cloning speaker_audio_base64). Harnesses in scratchpad (not yet landed). reference_zonos_tts_stack

  • [2026-07-17] Zonos2 :1920 → self-contained container (stays on 3090); prosody-priming is adapter-level, engine stays stock. Config captured (14a0004, unpushed); build = cu128 base + uv sync vs the lock + weights mount; priming = prime→generate-one-utterance→parakeet-clip→deliver in the gateway adapter. Crux = does AR prosody carry the sentence boundary (A/B the join). → persistent-memory.d/2026-07-17-zonos2-containerize-prosody-priming.md

  • [2026-07-16] GPU re-org: char-rp→GPU1 + both cards re-optimized for max context. Moved char-rp (Magidonia-24B) GPU0→GPU1, then maxed context: char-rp-reasoning 150K→256K (util 0.46, 1.56x), gen→256K + seqs 16→32 (util 0.42, 5.43x), granite 64K→128K full-chapter (util 0.27, 1.50x). FINAL: GPU0 ~14 G reserve (both seats 256K native), GPU1 ~6.7 G headroom. All healthy. LESSON: KV must hold ≥1× max-len (util-floor crashes) + per-model KV cost varies ~8× (MoE cheap, dense pricey) → tune util empirically. See Current state for the full layout + backups.

  • [2026-07-16] granite right-sized → ~10.5 GB freed on GPU1 (util 0.34→0.18 + max-len 131072→65536; KV 6.45 GiB / 1.29x@65536, summarizer healthy). GPU1 now ~45 GB free to relocate a GPU0 model. LESSON: ~950 MiB KV per 0.01 util here + KV must hold ≥1× max-len — util 0.15 crash-looped (est max-len 47184<65536, ~2-3 min summarizer blip) before 0.18 landed. .env-only, recreate vllm-granite alone (shared stack).

  • [2026-07-15] image-bench eviction DONE (parked item closed). Stopped vllm-qwen-image-bench (ana-ml2 GPU1, ~32 GB freed); LiteLLM image-judge+qwen-image-bench → gen :8015 (judge samplers + thinking-off), verified with :8014 down; comfy-dev pinged; also backfilled the canonical char-rp-reasoning litellm block (was lagging live). Revert ~90 s. auto-memory project_arbo_gen_switch_imagebench_evict.

  • [2026-07-15] arbo fully switched off image-judge (qwen-image-bench) -> gen; image-bench pending eviction post-bake → persistent-memory.d/2026-07-15-arbo-fully-switched-off-image-judge-qwen-image.md

  • [2026-07-15] esh-docker-vm NFS fstab fix = x-systemd.before=docker.servicepersistent-memory.d/2026-07-15-esh-docker-vm-nfs-fstab-fix-x-systemd.md

  • [2026-07-15] Homepage AI-tab revamp — flat "AI Systems" group -> dedicated AI tab, 6 role-based groups + AI-Dormant; committed 569e1af, pushed. (Also caught + pushed a ~100-commit unpushed eshpfi backlog.)

  • [2026-07-15] Home Assistant config repo created (vh/home-assistant-config, private). UI-managed HA -> allowlist model (YAML + curated secret-free .storage subset). git-in-place in /config on esh-docker-vm + scoped deploy key + local clone ~/development/home-assistant-config.

  • [2026-07-15] char-rp-reasoning OOM rescue — solo-restart on the packed GPU0 crash-looped; fixed via expandable_segments:True + util 0.39->0.38 + max-model-len 192K->150K. LESSON (Tried): max-model-len does NOT free vLLM VRAM (util-pinned KV pool). ~4.5 GB GPU0 headroom now.

  • [2026-07-15] soong-lab SOONG_LAB_LIBRARY_DIR made persistent (corviduo-dev) — was on the redeploy-wiped code default; set to /home/infra-ops/soong-lab-data/library (mirrors PORTRAIT_DIR), restarted. Closed a queued no-rush item; unblocked the operator.

  • [2026-07-15] Statusline overhauled (~/.claude/statusline-command.sh) — git state / 🔔🔕 monitor-armed / project tag / abs tokens / per-session cost (.cost.total_cost_usd) / threshold-colored ctx+rate (green<60 / yellow60-90 / red>90).

  • [2026-07-14] NVFP4+MTP fast char-rp-reasoning seat LANDED + LIVE + gateway-repointed + VRAM-tuned → persistent-memory.d/2026-07-14-nvfp4-mtp-fast-char-rp-reasoning-seat-landed.md

  • [2026-07-14] NVFP4 quant chase RESOLVED (gibberish) + PIVOTED to modelopt for MTP → persistent-memory.d/2026-07-14-nvfp4-quant-chase-resolved-gibberish-pivoted-to-modelopt.md

  • [2026-07-14] Pursue the NVFP4+MTP fast char-rp-reasoning seat to completion → persistent-memory.d/2026-07-14-pursue-the-nvfp4-mtp-fast-char-rp-reasoning.md

  • [2026-07-14] char-rp-reasoning seat: Deckard-PKD → NEO-CODE = Heretic2-Thinking (Qwen3.6-27B) → persistent-memory.d/2026-07-14-char-rp-reasoning-seat-deckard-pkd-neo-code.md

  • [2026-07-14] soong-lab webhook auto-deploy real root cause = gitea webhook.ALLOWED_HOST_LISTpersistent-memory.d/2026-07-14-soong-lab-webhook-auto-deploy-real-root-cause.md

  • [2026-07-13] #355-residual ROOT CAUSE (supersedes the "LiteLLM gateway holds while seat idles" entry below — that was DISPROVEN) → persistent-memory.d/2026-07-13-355-residual-root-cause-supersedes-the-litellm-gateway.md

  • [2026-07-13] Deploy-speed real bottleneck ≠ uv sync (memory's assumption was wrong) → persistent-memory.d/2026-07-13-deploy-speed-real-bottleneck-uv-sync-memory-s.md

  • [2026-07-13] WT #355 residual 300s hang localized to OUR LiteLLM gateway (holds 2 char-rp-reasoning requests ~21 min while the seat idles), NOT the seat — Deckard seat EXONE → persistent-memory.d/2026-07-13-wt-355-residual-300s-hang-localized-to-our.md

  • [2026-07-13] WT #355 turn-lifecycle fix VALIDATED on worldtree b60 — wedged turns self-terminate cancelled/stalled at the 300s stall-watchdog (turns 2064/2065 vs pre-b60 206 → persistent-memory.d/2026-07-13-wt-355-turn-lifecycle-fix-validated-on-worldtree.md

  • [2026-07-13] Worldtree deploy bottleneck = the image build (~11 min of a ~12 min deploy), root cause the Dockerfile uv sync → persistent-memory.d/2026-07-13-worldtree-deploy-bottleneck-the-image-build-11-min.md`

  • [2026-07-13] Ledger tier-3 consumer ledger:miranda provisioned on personal :8081 (key b38932f5, GPG-delivered+shredded, allowlist 10.100.10.50:8770 live); → persistent-memory.d/2026-07-13-ledger-tier-3-consumer-ledger-miranda-provisioned-on.md

  • [2026-07-10] Heimdall grant: ratatoskr affect.full on PERSONAL Worldtree (operator-approved, worldtree-dev R34-v1 request) → persistent-memory.d/2026-07-10-heimdall-grant-ratatoskr-affect-full-on-personal-worldtree.md

  • [2026-07-10] ComfyUI v0.27.1 SUCCESS on irv-ml1 (operator-confirmed execute-now) — landed on torch 2.12.1, SageAttention preserved, crash-loop AVOIDED → persistent-memory.d/2026-07-10-comfyui-v0-27-1-success-on-irv-ml1.md

  • [2026-07-10] ComfyUI 0.25.x bump on irv-ml1 ATTEMPTED → FAILED → ROLLED BACK (snapshot saved it) → persistent-memory.d/2026-07-10-comfyui-0-25-x-bump-on-irv-ml1.md

  • [2026-07-10] Biweekly open-weight-releases scan cron set up for brokkr-smithy (Vuong-authorized) → persistent-memory.d/2026-07-10-biweekly-open-weight-releases-scan-cron-set-up.md

  • [2026-07-09] Two parked items closed: phantom qwen3.6-35b-a3b alias VERIFIED already-gone; ana-docker docker log-cap SOLVED no-bounce → persistent-memory.d/2026-07-09-two-parked-items-closed-phantom-qwen3-6-35b.md

  • [2026-07-09] granite→gen memory_extractor bind host-synced on demo+personal Worldtree (Vuong-directed, #335 Slice-4) → persistent-memory.d/2026-07-09-granite-gen-memory-extractor-bind-host-synced-on.md

  • [2026-07-09] mOrpheus TTS off-the-shelf voice pipeline SHIPPED end-to-end (irv-ml1) + wired into gateway-chat → persistent-memory.d/2026-07-09-morpheus-tts-off-the-shelf-voice-pipeline-shipped.md

  • [2026-07-09] granite→gen memory_extractor bind GREEN-lit for worldtree-dev (Worldtree #335 Slice 4) → persistent-memory.d/2026-07-09-granite-gen-memory-extractor-bind-green-lit-for.md

  • [2026-07-08] RP-SEAT CAMPAIGN CLOSED — char-rp = Magidonia-24B-v4.3 (128K), char-rp-reasoning = Deckard-PKD Qwen3.5-27B (256K); both GGUF/llama.cpp on ana-ml2 GPU0 alongside gen (3… → persistent-memory.d/2026-07-08-rp-seat-campaign-closed-char-rp-magidonia-24b.md

  • [2026-07-08] worldtree Mimir deploy-blocker resolved (mid-session): → persistent-memory.d/2026-07-08-worldtree-mimir-deploy-blocker-resolved-mid-session.md

  • [2026-07-08] OFF-THE-SHELF INFERENCE PIVOT executed — serve curated abliterated models, stop home-training → persistent-memory.d/2026-07-08-off-the-shelf-inference-pivot-executed-serve-curated.md

  • [2026-07-08] DPO was silently running 3 epochs (harness gap) → KILLED at epoch 1.2, retargeted to 0.3 epochs (operator call) → persistent-memory.d/2026-07-08-dpo-was-silently-running-3-epochs-harness-gap.md

  • [2026-07-08] T1 DPO leg is RUNNING (unblocked) — 2 fixes applied to deployed backend.py → persistent-memory.d/2026-07-08-t1-dpo-leg-is-running-unblocked-2-fixes.md

  • [2026-07-08] T1 DPO leg launch — prior BLOCK (now resolved above), kept for the launch recipe → persistent-memory.d/2026-07-08-t1-dpo-leg-launch-prior-block-now-resolved.md

142 older entries archived to archival-memory.md.

Tried and abandoned

  • [2026-07-15] docker.service After=remote-fs.target does NOT wait for nofail NFS mounts → persistent-memory.d/2026-07-15-docker-service-after-remote-fs-target-does-not.md

  • [2026-07-15] The esh-docker-vm D-state/phantom-container wedge is only cleared by a host REBOOT → persistent-memory.d/2026-07-15-the-esh-docker-vm-d-state-phantom-container.md

  • [2026-07-15] vLLM max-model-len does NOT free GPU VRAM → persistent-memory.d/2026-07-15-vllm-max-model-len-does-not-free-gpu.md

  • [2026-07-15] Claude Code statusline .cost.total_cost_usd is per-SESSION → persistent-memory.d/2026-07-15-claude-code-statusline-cost-total-cost-usd-is.md

  • [2026-07-14] MTP-on-modelopt: NO checkpoint config skips the spec-decode drafter's quant (vLLM 0.24 bug) — 4 config attempts failed before the runtime workaround → persistent-memory.d/2026-07-14-mtp-on-modelopt-no-checkpoint-config-skips-the.md

  • [2026-07-14] AEON's "working NVFP4+MTP RP seat" was pantheon on compressed-tensors (0% MTP accept), not a modelopt MTP proof → persistent-memory.d/2026-07-14-aeon-s-working-nvfp4-mtp-rp-seat-was.md

  • [2026-07-14] NVFP4 (llm-compressor / compressed-tensors) gives NO batch-1 speedup over GGUF for the Qwen3.5 GDN-hybrid, and its MTP is 0%-accept → persistent-memory.d/2026-07-14-nvfp4-llm-compressor-compressed-tensors-gives-no-batch.md

  • [2026-07-14] NVFP4 spike: built the full MTP serve scaffolding BEFORE validating a plain NVFP4 serve was coherent → persistent-memory.d/2026-07-14-nvfp4-spike-built-the-full-mtp-serve-scaffolding.md

  • [2026-07-14] MTP graft via top-level mtp.* tensor names does NOT survive AutoModelForCausalLM.from_pretrainedpersistent-memory.d/2026-07-14-mtp-graft-via-top-level-mtp-tensor-names.md

  • [2026-07-14] gitea "test-delivery 204" is NOT proof a webhook works → persistent-memory.d/2026-07-14-gitea-test-delivery-204-is-not-proof-a.md

  • [2026-07-13] Relaying a peer's diagnosis as fact without confirming it against raw data → persistent-memory.d/2026-07-13-relaying-a-peer-s-diagnosis-as-fact-without.md

  • [2026-07-13] althing-cli reply <THREAD_id> (thread id, not a MESSAGE id) → "unknown message_id"; and reply to your OWN message self-addresses to your handle ("replying to your own message"). Reply to a PEER's message id, or use post --to <peer>. Bit me several times this session.

  • [2026-07-09] FP8 breaks mOrpheus audio-token generation → persistent-memory.d/2026-07-09-fp8-breaks-morpheus-audio-token-generation.md

  • [2026-07-09] vllm/vllm-openai:latest crashes on Ampere IMPORT — Blackwell-only kernels (oink/aiter, has_device_capability(100)) die during import on the 3090/A6000. Pin v0.23.0 on irv-ml1's Ampere GPUs. (vllm/vllm-omni:v0.18.0 has a different entrypoint — don't use it either.)

  • [2026-07-09] Per-frame CPU SNAC decode is too slow for streaming — per-call overhead × ~60 frames serialized → RTF 2.2 (WORSE than whole-clip's 1.0). Fix = windowed chunk decode (every 6 frames decode a [2 ctx | 6 | 2 ctx] window, emit the middle 6 → seamless, O(1)/frame, RTF ~0.97, TTFA ~0.8s).

  • [2026-07-09] Sentence-chunking TTS loses prosody → persistent-memory.d/2026-07-09-sentence-chunking-tts-loses-prosody.md

  • [2026-07-09] HF whisper datasets aren't actually whispered → persistent-memory.d/2026-07-09-hf-whisper-datasets-aren-t-actually-whispered.md

  • [2026-07-08] Angel (allura-org/MS3.2-24b-Angel) self-quanted to NVFP4 = GARBAGE → persistent-memory.d/2026-07-08-angel-allura-org-ms3-2-24b-angel-self.md

  • [2026-07-08] Mistral3 + vLLM tokenizer/vision traps (serve MS3.2-24b, vLLM 0.24) → persistent-memory.d/2026-07-08-mistral3-vllm-tokenizer-vision-traps-serve-ms3-2.md

  • [2026-07-08] Pantheon-Reasoning-27B refuses dark fiction DESPITE an abliterated base → persistent-memory.d/2026-07-08-pantheon-reasoning-27b-refuses-dark-fiction-despite-an.md

  • [2026-07-08] Pantheon-27B MTP on vLLM compressed-tensors = 0% acceptance → persistent-memory.d/2026-07-08-pantheon-27b-mtp-on-vllm-compressed-tensors-0.md

  • [2026-07-07] vLLM 0.24.0 qwen3_5 LoRA application = silent no-op (#47639) → persistent-memory.d/2026-07-07-vllm-0-24-0-qwen3-5-lora-application.md

  • [2026-07-07] SGLang generic image can't LOAD our NVFP4 AEON → persistent-memory.d/2026-07-07-sglang-generic-image-can-t-load-our-nvfp4.md

  • [2026-07-07] SGLang --lora-target-modules CLI enum REJECTS the GDN names its own resolver asks for → persistent-memory.d/2026-07-07-sglang-lora-target-modules-cli-enum-rejects-the.md

  • [2026-07-07] Engine invocation footguns cost several wasted serve-bounces this session → persistent-memory.d/2026-07-07-engine-invocation-footguns-cost-several-wasted-serve-bounces.md

  • [2026-07-04] LiteLLM (this gateway version) mutates the SHARED deployment config in-place on per-request sampler-param merge → persistent-memory.d/2026-07-04-litellm-this-gateway-version-mutates-the-shared-deployment.md

  • [2026-07-04] A systemd --user daemon that shells out to ~/.cargo/bin/~/.local/bin tools needs an explicit Environment=PATHpersistent-memory.d/2026-07-04-a-systemd-user-daemon-that-shells-out-to.md

  • [2026-07-04] On-prem T1 train that keeps ANY ana-ml2 serving up = ~6-8 DAYS → persistent-memory.d/2026-07-04-on-prem-t1-train-that-keeps-any-ana.md

  • [2026-07-01] A personal-Worldtree CI deploy that fails ~85s in with "not found / unauthorized" is usually the pull-only-vs-build RACE, not registry-auth → persistent-memory.d/2026-07-01-a-personal-worldtree-ci-deploy-that-fails-85s.md

  • [2026-07-01] MTP/spec-decode on a SHARED serving model helps single-stream but HURTS moderate-concurrency aggregate + silently ignores min_p/logit_bias (qwopus gen: N=1 +12%, N=4 20%). Reserve for dedicated/interactive deployments.

  • [2026-07-02] irv-ml1 /worktank ROOT is root-owned — lkraven can't write there (irv-ml1 sudo needs a password) → stage model pulls to /home. PIN THE A6000 BY UUID for training (native-CUDA ordering differs vs docker; the 3090 index 0 is usually near-full → OOM). CUDA_VISIBLE_DEVICES=GPU-<uuid>.

101 older entries archived to archival-memory.md.