74 KiB
Persistent memory — eshpfi-management
Last updated: 2026-08-15
Always check for
/tmp/infra-ops-handoff.md— if it exists and itsWritten:stamp is under an hour old, read it (it carries the in-flight handoff from the previous session), then delete it. Older than an hour: stale — delete it unread.
Repo purpose
Reference workspace for PFI infrastructure: server inventory, canonical
Docker Compose stacks, ops playbooks, and conventions. Authoritative
copies of compose files live on the servers under
/opt/docker/compose/<stack>/; this repo mirrors them for version
control, editing, planning, and CI-driven deploys. It was originally
spun up to handle the fleet backups — keep that lens when triaging
backup/storage issues.
Tools and conventions
Sister repos (separate gitea repos, deployed by playbooks here):
| Repo | Role | CI status |
|---|---|---|
vh/task-board |
MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) |
vh/vor |
Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
vh/nevermore |
Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
vh/asset-engine |
Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
vh/althing |
Lean trusted inter-agent message bus — v2 "email model" (v2.0.0b2, 2026-07): per-box local-SQLite bus + courier/receiver for P2P over the 10.x net; pillars = open-loops / per-box herald + wake-listener / roaming owner API /owner/* / althing-mcp stdio surface. The v0.15 lean-bus cut RIPPED moderation / chamber / forseti-daemon / agent-runner / redis-valkey. |
per-box uv tool install (NOT CI-deploy); nh3-dev = the DEV box (editable install of ~/development/althing, gets new versions first); nh3-extdev a mesh peer (model B: althing-svc + shared /srv/althing) |
vh/mead-hall |
Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
vh/skaldsong |
Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
vh/Worldtree |
Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. gitea-runner builds on ana-docker; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. | push-to-main → CI build-and-deploy (runner on ana-docker) |
vh/yt-voice-clipper |
YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md |
vh/arbo |
Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) | push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook |
vh/zonos-gateway |
OpenAI-compatible TTS gateway over stock ZONOS2 (:8890 irv-ml1); emotion dials-first + voice mapping; reached via LiteLLM ext-tts alias. v0.2.1 (2026-07-18): voice-resolved emotion presets (resolve_preset(name,voice); angry/happy/startled_happy per-voice). 8 voices incl. 4 clones |
pushed to gitea (main 8f1885b/v0.2.1); deployed irv-ml1 tree still NON-git (hand-updated build context — CI-wire = open follow-up). Spec docs/EMOTION-DIALS-SPEC.md; host-managed voices bind-mount (./voices:/app/voices, drop wav + restart, no rebuild) |
vh/soong-lab |
Noonien Soong character-design studio (SPA + /api + WT /bifrost/tool-call); containerized 2026-07-18, LIVE on corviduo-dev :8443 (image vh/soong-lab:latest). soong-dev owns Dockerfile/compose/workflow; infra-ops owns the host |
CI = Gitea Actions build+push+DEPLOY on tag/dispatch (fleet recipe: docker:cli + raw buildx, pushes AS vh; auto-redeploy LIVE 2026-07-18 — runner SSHes corviduo-dev as deploy, compose pull && up -d from /opt/soong-lab, health-gated on /api/version). Manual redeploy sudo -u deploy bash -c 'cd /opt/soong-lab && docker compose pull && docker compose up -d'. → persistent-memory.d/2026-07-18-soong-lab-auto-redeploy.md |
model-training-forge (mtf-dev) |
Fine-tuning recipe forge; T1 = E-RP writing LoRA, retargeted qwopus-122B→AEON-27B (2026-07-06) (SFT→DPO, LitBench-RM reward) | training runs, not a deployed sidecar |
(vh/volva + Heid were re-architected from systemd daemons to Claude Code
session orchestrators 2026-06-08; their nh3-dev .service units were removed —
no longer deployed sidecars here. See Recent decisions.)
-
Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see
docs/runbooks/disaster-recovery.mdfor the blast-radius matrix. ⚠️ The restic file+DB layer routes through TWO rest-servers (rest-server-ana@ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas;rest-server-nh3@ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of/mnt/backup. (rest-server-ana recovered 2026-06-20.) -
pull-hf-repo.yamlis the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at/tank/aimodels/huggingface/" playbook. Supports--var repo_type=model|dataset|space. Replaces ad-hochuggingface_hub.snapshot_downloadpatterns. -
Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (
key_id 61419c92) atana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-adminauths against demo only. Personal-instance admin (the~/.config/worldtree/personal-admin-token, mode 600) POSTs/admin/keys(mints per-project keys; takesuser_id+label, no scope param — scopes are tier-derived). On-instance mint recipe (cleaner than DB-manip):docker exec worldtree-worldtree-api-1POST/admin/keyswith the in-containerWORLDTREE_BOOTSTRAP_ADMIN_KEY; cleartext once in.key=wt_live_+16hex. auto-memoryreference_worldtree_demo_key_mint. -
Per-project user keys against personal Worldtree (issued 2026-05-19):
skaldsong:79744637,skaldsong:7c1dbbbe,althing:50d85460,mead-hall:a360822d. Mint via/admin/keys, drop value to/tmp/wt-personal-<name>.keymode 600, dev collects + shreds (DO NOT cat to chat transcript). -
Skaldsong CD pattern (registry-pull). vh/skaldsong's CI builds and pushes
gitea.phasefinal.com/vh/skaldsong:<sha>+:latest;playbooks/deploy-skaldsong.yamlon ana-docker pulls + recreates. SHA-pin only. Prereq: host needsdocker login gitea.phasefinal.comonce. -
gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH
10.250.50.70:222, HTTP:3000. Fleet/colo hosts must use this internal route, NOT publicgitea.phasefinal.com(38.120.12.44) — the public path fail2bans the host egress IP. Full gotcha indocs/orientation.md→ Git/gitea. -
docker-as-root pattern (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo):
docker run --rm -v <target-dir>:/wt docker:cli sh -c "...". docker-group membership is effectively root via bind-mount. Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass-e VAR=/abs/pathfor any relative-default config dir. -
scripts/elwaysudo handling — elway prompts for the sudo password ONCE viagetpassbefore the firstsudo: truestep → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH. -
Per-host SSH identity matters for sudo. infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: default
ssh ana-docker=lkraven(docker-group, NO passwordless sudo);ssh infra-ops@ana-dockerHAS NOPASSWD root. → For any sudo op on ana-docker, usessh infra-ops@ana-docker.ssh infra-ops@10.100.10.50(nh3-dev) ALSO NOPASSWD sudo; on nh3-extdev infra-ops is sudo-LESS by design (ssh lkraven@10.100.50.42is the NOPASSWD path). irv-ml1:ssh irv-ml1= lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD) — stage model pulls to/home, not root-owned/worktank.
Current state / in-flight
As of 2026-08-15 — GEN SEAT REQUANTED (mixed NVFP4+FP8, +18% decode) and the char-rp tool-call parser FIXED. Both were the two open items from the prior session; both are closed, verified, and committed. Nothing queued behind them.
-
✅ UNCENSORED GEN SEAT — DONE + LIVE (2026-08-15).
gen-seat/vllm-genon ana-ml2 GPU0:8015servesqwen3.8-27b-uncensored(JonathanColetti/Qwen3.8-27B-Uncensored, Heretic-abliterated, in-house NVFP4 W4A16 compressed-tensors + grafted bf16 MTP, vision-intact, 262K ctx, MTP n=3 ~42% accept / ~68 tok/s). Replaced the qwen3.6-35b-a3b-heretic MoE (which had briefly replaced granite/AEON). All 7 aliases repointed live + verified; compose renamed qwen36-27b-aeon→gen-seat, vllm-aeon-gen→vllm-gen, AEON_GEN_→GEN_, dead RP service dropped; repo mirrored + docs/memory refreshed; committed eshpfi680c30e+ dotfiles1d1970f. Full arc + the definitivere:^mtp.*-ignore fix → Recent decisions[2026-08-15]+persistent-memory.d/2026-08-15-uncensored-gen-seat.md. -
✅ GEN SEAT REQUANT — DONE + LIVE (2026-08-15, overnight). The seat now runs mixed-precision
/tank/aimodels/qwen38-27b-uncensored-nvfp4-mixed(NVFP4 W4A4 layers 0-55 MLPs + FP8 W8A8 attn/linear_attn/lm_head/layers 56-63 MLPs + FP8 KV) — 80.12 → 94.53 tok/s decode (+18.0%) and — the bigger win — prefill roughly DOUBLED (3,206→6,334 tok/s at 6.7k prompt; 2,862→5,085 at 27k; TTFT on a 27k doc 9.43→5.31 s), at unchanged MTP acceptance (47.8→47.7%), +1.7% PPL, abliteration 4/4 preserved, weights 27.7→22.5 GB. Prefill > decode is the expected ordering (decode is bandwidth-bound and 4-bit either way; prefill is compute-bound = where native FP4 replaces Marlin) — thesummarizeraliases feel this most. Surface-verified live (chat/vision/tools/thinking/36K-needle/streaming 6/6) + all 7 aliases routing. The queued "W4A8" framing was unservable — vLLM 0.24 allows NVFP4 weights with ONLY A16 or A4, FP8 activations raise ValueError at load; FP8 has to enter per-layer-group. Committed74f596b. Pipeline + acceptance harness + raw numbers →services/gen-seat-mixed-quant/; full arc → Recent decisions[2026-08-15]+persistent-memory.d/2026-08-15-gen-seat-mixed-requant.md. Rollback = one.envline, old build untouched at…/qwen38-27b-uncensored-nvfp4. -
⚠️ GPU0 is at 94.4/97.9 GB (gen 0.43 + meromero 0.52).
GEN_GPU_MEM_UTILwas cut 0.45→0.43 because the smaller mixed weights let gen soak the slack as KV and starved meromero by 0.18 GiB → crash-loop. Any future util bump on either GPU0 seat must be checked against the other. -
FLEET RERANKER = A3 (bge-reranker-v2-m3) PROD ana-ml2 GPU1 :8013 (R42 v13 gate passed). Passive watch: caps ~34 req/s; levers = A4 :8014 / util / 2nd replica; incumbent :8002 warm for rollback. Full arc
docs/pfi/reranker-selection-ledger.md. -
EVIDENCE HOLD (partial): WT #394 index-row half lifted+swept; the FILE half STILL STANDS — do NOT delete on-disk gen dirs (
fiction/rex390-dcc,rex392-dcc,b59c147c5ce0). rex393-fiction-* + r42-gate-* also KEEP. -
✅ CHAR-RP TOOL-CALLING — FIXED (2026-08-15). MeroMero (Gemma-4) had no tool parser at all, so every tools-bearing request 400'd. Fixed with
--tool-call-parser gemma4 --enable-auto-tool-choice --reasoning-parser gemma4 --default-chat-template-kwargs '{"enable_thinking": false}'— the last flag is mandatory, not decorative (the parser defaultsenable_thinkingto True, which pre-inits the engine to REASONING and returns nullcontentfor all plain RP prose). Verified green: tool call streaming + non-streaming, tool round-trip, prose incontent, vision. Committedb8f0f4c. Details instacks/meromero-charrp/README.md. -
OPEN FOLLOW-UPS (parked): chatterbox-fast flat-vs-package build-context divergence; CI-flip runner-auth research; #363 research-wing ingest (no deadline); zonos-gateway CI-wire.
-
althing monitor ARMED (handle
infra-ops). ⚠️ Re-arm ONLY after a real FIRE (rc0), NEVER after a plain operator turn (bounces rc3); spawnalthing-cli monitor/althing-wake-listeneras its OWNrun_in_backgroundtask, never chained with&. -
eshpfi push state:
origin/mainbehind — unpushed:74f596b(gen-seat mixed requant) +b8f0f4c(char-rp tool parser) from the overnight session,680c30e(gen-seat), plus the earlier eRP dual-seat arc (f08b6cb/7bd7375), wgtunnel mirror (398b58a), dots.tts + secrets-broker arcs. Dotfiles:1d1970f(roster) +6425cc6(statusline bell) unpushed. Push = operator's call.graphify-out/GRAPH_REPORT.mdchurns every commit (ignore);stacks/heretic2-charrp-reasoning/UNTRACKED (pre-existing, superseded).
Recent decisions
-
[2026-08-15]gen seat requanted to mixed NVFP4+FP8 (+18% decode) + char-rp Gemma-4 tool-calling fixed. The queued "W4A8" (NVFP4 weights + FP8 activations) is not servable — vLLM 0.24 allows NVFP4 weights with only A16 or A4; FP8 activations ValueError at load, andCompressedTensorsW4A8Fp8is INT4-weights + sm90-exact (closed on Blackwell twice). FP8 must enter per-layer-group. Also: the handoff's "~68 tok/s" baseline didn't reproduce — cache-busted, the incumbent already did 80.12 (≈ the stated W4A8 target), so the premise needed re-measuring before any work. Shortcut:unsloth/Qwen3.8-27B-NVFP4was already on-box → served as a probe, measured +19.1% at identical acceptance, which both proved the gain was real and handed over the reference recipe. Replicated it on the abliterated weights → 80.12→94.53 tok/s, acceptance unchanged, +1.7% PPL, abliteration 4/4, weights −19%; surface 6/6 live, 7 aliases routing. char-rp had no tool parser at all (every tools request 400'd) →gemma4tool + reasoning parser + a mandatoryenable_thinking:false(the parser defaults it True → nullcontentfor all RP prose; proven byte-identical prompt before deploying). Commitsb8f0f4c,74f596b. Foot-guns banked (llm-compressor prunes unmatchedignoreentries → the 0%-MTP bug, fired on this run; prompt_logprobs uniform under spec-decode; 0600.envsilently no-ops compose; GPU0 is zero-sum). →persistent-memory.d/2026-08-15-gen-seat-mixed-requant.md -
[2026-08-15]Uncensored gen seat: JonathanColetti/Qwen3.8-27B-Uncensored deployed asgen-seat/vllm-gen(NVFP4 W4A16 + grafted MTP, 262K); 7 aliases repointed; the definitivere:^mtp.*-ignore fix. 0%-MTP-on-quant (twice) was NOT the abliteration/scheme — the grafted bf16 MTP was missing fromquantization_config.ignore(vLLM loaded it as quantized → uninitialized). Full arc, the working pipeline, VRAM budget, unsloth speed decomposition, modelopt dead-end. →persistent-memory.d/2026-08-15-uncensored-gen-seat.md -
[2026-08-12]eRP dual-seat overhaul: MeroMero-v2 (char-rp) + Dark-Scarlett (char-rp-reasoning), both NVFP4A16 @ 256K on ana-ml2; granite retired. Replaced the GGUF/heretic2 RP seats with two home-quantized vLLM seats. The DS blocker (anAutoModelForCausalLMsave wrote a flatQwen3_5TextConfigthat both vLLM AND SGLang reject) was fixed by re-quanting via theQwen3_5ForConditionalGenerationwrapper class; ModelOpt was a version deadlock, SGLang lacked the impl (but revealed the fix). MeroMero vision reconstructed by extractingpreprocessor_config.jsonfromprocessor_config.json. Both models KV-efficient (Gemma-4 sliding-window / Qwen3.6 hybrid linear-attn) → full 256K; GPU-swapped for headroom; compose-ified + committedf08b6cb. granite downed + LiteLLMsummarizer/classifier→gen. Full arc, lessons, dead-ends →persistent-memory.d/2026-08-12-erp-dual-seat-overhaul.md -
[2026-08-12]infra-ops now holds an all-zones Cloudflare DNS-edit token (vaulted) + wgtunnel Phase-0 DNS landed. Operator handed over aZone·DNS·Edit(all zones) CF token →secret put nh3-dev/.config/cloudflare/infra-ops-dns-token(round-trip verified; /tmp drop shredded). Fleet DNS is now self-serve for infra-ops (⚠ HIGH blast radius — all zones). First use: createdboring.phasefinal.comCNAME →ana-srv1.phasefinal.com, DNS-only (proxied:false), verified resolving to 38.120.12.44 on both authoritative NS (louis/wren) + 1.1.1.1 — NOT Cloudflare-proxied. Unblocks wgtunnel's wstunnel ACME cert. phasefinal.com zone idf812ba74ed9a75cf21bbe7ce9188db50. auto-memoryreference_infra_ops_cloudflare_dns_token. (Earlier gap: the only prior vaulted CF token, jackdaw's, hadzone:read+worker:editbut nodns_records:edit.) -
[2026-08-12]wgtunnel stood up as its own repo (vh/wgtunnel, private) after a live endpoint-verification pass. Operator directed own-repo (mirrors stonehenge-park/tts-stack). Verified off the fleet before seeding:ana-wgWG server = UDP/31337 (not 51820), subnet 10.30.10.0/24, MTU 1420, active roaming peer proves the public UDP DNAT works; traefik on ana-docker terminates TLS :443 (ACMEanaprodhttp-challenge, docker+file providers, CrowdSec bouncer) → confirms the clean design (wstunnel container ontraefik-net, Host-routed, WS→UDP toana-wg:31337); edge38.120.12.44direct-A,tunnel.phasefinal.comfree (⚠ must be direct, NOT Cloudflare-proxied like vaultwarden). Repo pre-seeded (README/CLAUDE/persistent-memory/ROADMAP +docs/verified-infrastructure.md= ground truth) + pushed; commit9584d38, Vuong-attributed. vh gitea token pulled from the vault (secret get), not persisted to.git/config. NEXT =/vor-planor/vor(operator's call, interactive). Deps to line up in the plan: DNS A-record, FortiGate :443 host-routing, a new ana-wg peer for the laptop, client tooling. -
[2026-08-10→12]secrets-broker: per-box Vaultwarden credential store SHIPPED + consumer-confirmed.secretCLI (put/get/list/rm/backfill, bw-backed) on~/.local/bin; 25 nh3-dev secrets backfilled + round-trip-verified;rm+ new-namespace warning added post-launch; standing "vault is the credential source of truth" directive now global. →persistent-memory.d/2026-08-12-secrets-broker.md -
[2026-08-11]stonehenge-park: new fleet/parkservice repo stood up + designed (/vor-plan+/vor-ui). Self-contained SQLite+FastAPI idea-parking service that actively resurfaces (statusline + althing) so nothing dies in a cold repo;vh/stonehenge-parkpushed + pre-seeded for a fresh agent; build starts at the U1 tracer contract. →persistent-memory.d/2026-08-11-stonehenge-park.md -
[2026-08-12]Global~/.claude/CLAUDE.md:secret/vault tool entry + "store in AND pull from the vault" standing directive (dotfiles9db703b, pushed); statusline reset-countdowns + a latent tab-collapse parse-bug fix, now tracked in the dotfiles stow tree. Dogfooded the directive: createdvh/stonehenge-parkpulling the gitea token viasecret get. (dotfiles + global config, not eshpfi.) -
[2026-08-11]TTS stack extracted to its own repo (tts-stack) + eshpfi stood down on TTS dev. Operator: hand all TTS tuning/dev to a separate agent with a self-contained repo (knowledge + infra access + a live knowledge list), and move the voice corpus in. New repo~/development/tts-stack(commit9ee3288) carries: dots-tts stack (canonical intent),voices/corpus (MOVED out of eshpfi),KNOWLEDGE.md(engine landscape + prosody findings + foot-guns),docs/infrastructure.md(irv-ml1 access + gated deploy runbook + rollback), CLAUDE/persistent-memory/ROADMAP,tools/(pause-probe + Booth render). Followed the chatterbox-fast precedent: eshpfistacks/dots-tts/reduced to a POINTER README; the ~15 experimental TTS compose wrappers stay here as reference (catalogued in tts-stack KNOWLEDGE). Blast-radius check: no eshpfi playbook/script reads the canonical corpus (othervoices/refs = unrelated host paths). Reverses the earlier "Corpus home = eshpfivoices/(keep-here)" call. ⚠ tts-stack is LOCAL-ONLY until pushed — needs a gitea remote (vh/tts-stack) + push before the separate agent can clone (operator's call — outward-facing + repo-create creds). -
[2026-08-10]dots-tts v3 — clause-break → period pause mapping. Operator: v2 "sounds good" but donut won't pause at semicolons/dashes. ROOT CAUSE (measured via a pause-probe A/B — synth duration over N runs, non-determinism averaged out): dots' prosody honors a real pause only for ellipsis (+0.43s) and period (+0.3s, capitalization-independent); comma/semicolon/colon/dash all run flat (~+0.03s vs no-punct). Two distinct sub-causes: dashes regressed in v2 (the—→-fold made em-dashes read as word-joiners), while semicolons were NEVER a v2 change — dots ignores them natively, only newly noticeable because v2 made everything else clean. Operator call: ellipsis "too much" → map;, clause:, and em-dash—→ period in_sanitize(believable ~0.3s clause break). GUARDS (pinned by 11 unit tests,stacks/dots-tts/test_sanitize.py): digit-guarded colon(?<!\d)\s*:\s*(?!\d)so times3:45/ ratios2:1survive; en-dash–→hyphen KEPT (numeric-range10–20safety — em-dash breaks, en-dash ranges, different jobs); genuine ellipsis left at full strength (author meant a long pause). Gated deploy (redeploy2 pattern → v3): build → throwaway :8199 test container + pause-gate (semicolon sentence must run ≥0.12s longer than baseline; measured +0.427s) → only then cut live over. LIVE + healthylocal/dots-tts:v3on :8198. rollback =sed -i 's/^DOTS_TAG=.*/DOTS_TAG=v2/' .env + docker compose up -d dots-tts(v2 image retained). Boothdots-pauses(A=old-flat / C=ellipsis-too-much / D=live-v3). reference_chatterbox_fast_repo -
[2026-08-10]dots-tts v2 — contraction fix (curly-sanitize) + sentence-chunking + dependency-pin recovery. Operator: donut read contractions wrong ("you're"→"you ree", "donut's"→"donut ess"). ROOT CAUSE (isolated via A/B booth): curly/typographic apostrophes (’U+2019 from ratatoskr's LLM) — dots' tokenizer mispronounces them; STRAIGHT apostrophes read clean undernormalize_text=True. FIX (app.py): fold curly→ASCII (str.maketrans) before synth, KEEPnormalize_text=True(operator call — retains number/date expansion). Also added server-side sentence-chunking (pack ≤280 chars): dots caps onegenerate()at ~500 patches/~40s, so long RP turns (the Zev monologue = 160s audio) truncated; chunking stitches them (verified full 160.3s, not 40s-cut). ⚠ BUILD FOOT-GUNS (both bit this redeploy): (1) upstream dots.ttsconstraints/recommended.txtnow pinsgradio==6.17.0— phantom, not on PyPI → freshpip install dots.ttsunsatisfiable; FIX = pindots.tts==0.2.1+ DROP the-c recommended.txtconstraints (0.2.1 pulls working gradio 6.17.3). (2) pinning onlytorch==2.8.0let torchaudio float to 2.11.0 → dots.tts refuses to load (minor-version match check); FIX = pintorchaudio==2.8.0. ⚠ DEPLOY LESSON:docker compose up -dto a new tag swaps the LIVE container BEFORE any health check — a broken image crash-loops production (ratatoskr TTS down ~1-2min this session). NEW PATTERN = build → test in a THROWAWAY container on an alt port (:8199) → health+verify → only THEN cut live over (redeploy2.sh). v2 LIVE + healthy on irv-ml1:8198, CONSUMER-CONFIRMED clean (ratatoskr verified end-to-end on their :8765 — apostrophe string reads clean, /api/tts 200 @ 48kHz, no client change; the ~1-2min blip didn't hit them, their concurrent auto-audio issue was client-side localStorage). rollback =sed DOTS_TAG=v1 + docker compose up -d dots-tts(v1 image retained). Also: deployed container GPU crept ~6→13.9GB over 8h serving (cache accumulation; a redeploy resets it — watch item). reference_chatterbox_fast_repo -
[2026-08-09→10]dots.tts (rednote-hilab) TTS burn-in on irv-ml1 + canonical voice corpus built (voices/). Operator-directed eval to potentially replace chatterbox-fast. dots.tts VERIFIED real (canonical HF nsdots-studio/,rednote-hilab/dots.tts-*redirects there; Apache-2.0; PyPIdots.tts0.2.1; 2B continuous-AR = semantic enc + Qwen2.5-1.5B LLM + flow-matching acoustic head over 48kHz AudioVAE; zero-shot clone from wav+transcript). Runs on Ampere 3090 (sm_86, bf16, no fp8 dep); optimized RTF 0.22 at num_steps=10 (from_pretrained(..., optimize=True)CUDA graphs — raw unoptimized was 1.21), ~6GB VRAM, 48kHz, streams (generate_stream). Venv+cache atirv-ml1:/home/lkraven/dots-tts(~10GB). Operator design calls: SGLang Omni serving (OpenAI/v1/audio/speech), transcribe-refs-first,soarvariant. ⚠ Omni serves soar but its continuous-batching + streaming opts are mf-only (soar = single-request) — non-issue for ratatoskr's single-consumer RP surface. KEY FINDING — dots is highly sensitive to an accurate AND sentence-bounded reference transcript: mismatched transcript → 0.16s collapse; over-long/messy transcript → reference-audio BLEEDS as an output prefix; mid-clause trim → dangling-word leak (glados "we'll", emmie "And,"). RECIPE (baked intovoices/derive.py): trim ref to a clean ~6–10s clip ending on a sentence boundary + accurate transcript of exactly that clip. CANONICAL VOICE CORPUS stood up in eshpfivoices/(operator idea): engine-agnosticcanonical/<v>.wav+transcripts/<v>.txt→ per-engine ref sets DERIVED byderive.pyreadingengines.yamlprofiles (dots/chatterbox/zonos); canonical wavs git-tracked (small/curated),derived/gitignored. 4 voices optimized + verified CLEAN for dots: donut, glados, emmie, miranda (glados canonical is low-SR 16kHz — flagged upgrade candidate). ⚠ GPU GOTCHA: irv-ml1 native CUDA orders A6000=device0 (ComfyUI-full) — pin the 3090 withCUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=0; andPYTORCH_CUDA_ALLOC_CONF=expandable_segmentsCONFLICTS withoptimize=TrueCUDA graphs (curr_block error). Booths:dots-vs-chatterbox,dots-voices-optimized. SHIPPED 2026-08-10: operator A/B verdict "dots is very good" → containerized as a thin FastAPI wrapper over DotsTtsRuntime (chosen over SGLang Omni — Omni's batching is mf-only, unneeded for ratatoskr's single consumer; wrapper is SERIALIZED one-gen-at-a-time via a threading.Lock, Omni+mf = parked API-compatible escalation if multi-consumer ever lands). LIVE on irv-ml1:8198 (local/dots-tts:v1, OpenAI/v1/audio/speech+/health+/v1/voices, container healthy, both stream + non-stream verified CLEAN, 4 voices donut/glados/emmie/miranda) alongside chatterbox :8197 (nothing repointed). Stack =stacks/dots-tts/(Dockerfile/app.py/compose/.env.example/README). ⚠ CONTAINER GOTCHA:optimize=True(torch.compile/inductor/triton) needs a C compiler at RUNTIME — slim image mustapt install build-essentialor model-load dies "Failed to find C compiler" (host venv had gcc ambient, masking it); persistTORCHINDUCTOR_CACHE_DIRto a mounted dir or every restart re-JITs ~5min. Corpus home = eshpfivoices/(operator ruled keep-here). REMAINING: ratatoskr client cutover to :8198/v1/audio/speech(Phase-2 tail, peer-coupled — draft the ask). reference_chatterbox_fast_repo reference_zonos_tts_stack reference_verify_hf_repo_ids_before_pull -
[2026-08-08]worldtree-dev #400 CLOSED → fiction-decomp snapshot cleared from nh3-dev. worldtree-dev signaled #400 done (shipped v1.0.0b185; exact-lexical efficacy 79%→12% on ratatoskr's gate, brokkr no-harm bracket green both ends; the snapshot served 4 probe rounds — rank decomposition, promoted-vs-gold annotation, tie-set falsification, A0/A1/A2 mechanism probe). Cleared~/snapshots/worldtree-400-fiction-decomp(208M: chroma + manifest/provenance/stamp) — a read-only rsync copy of PERSONAL Worldtree's Chroma (source on corviduo-dev, so safe to remove). LEFT INTACT:rex393-fiction-index/rex393-fiction-snapshot(separate operator KEEP word, unchanged) +r42-gate-*. No config deltas rode this train. Only remaining non-blocking await = ratatoskr-dev's chatterbox-fast knob revert. Replied confirming (01KZJ9GMCC…). -
[2026-08-07]chatterbox-fast "broken audio" root-caused (T3 AR tail over-run) + FIXED (max_chunk_chars=250 cap, :v2 deployed). Long saga, operator-driven clean diagnosis. Symptom: ratatoskr's migrated RP-surface TTS "swaps to German" / "dead air" / "garbage" on long turns. NOT German-leak (Turbogenerate()has NO language param — plain AutoTokenizer, nolanguage_id; the multilinguallanguage_id="en"lever lives only in the separateChatterboxMultilingualTTS), NOT OOM alone. Real cause: the Chatterbox Turbo T3 model OVER-RUNS its generation tail — a long singlegenerate()degrades into garble/dead-air in its final ~2-3s (lib filters OOV tokens<6561+ pads silence = messy AR tail). The scheduler's buffer-ratchet builds 300-600 char mega-chunks that land in that zone; streaming concatenates each bad tail (worst case). ratatoskr's anti-"German" knobs (top_k=80/temp=0.5) made it WORSE — tight sampling pulls the degradation onset SHORTER (~200 chars vs ~300 at default knobs). Diagnosis method (deterministic, no ears-only): single-shot length sweep + amplitude-gated voiced-ZCR (garble spikes ZCR; must gate on |x|>500 else trailing silence confounds it) — degraded voiced-tail = 1.58× mid, clean = ~0.64-1.1×. FIX: server-sidemax_chunk_chars=250cap on the scheduler (:v2image,CBF_MAX_CHUNK_CHARS=250env) — bounds each generation to just under the ~300-char onset → clean 3-4 sentence chunks (max prosodic arc while clean). Operator ear-confirmed clean audio + clean joins; chatterbox's low emotiveness keeps chunk joins smooth (the harsh joins that got Zonos rejected are absent — operator's key call). ratatoskr TODO (relayed msg01KZER9X7S): revert knobs to default (top_k→1000, temp→0.8), send full text (server chunks internally), keep the 503-on-empty guard. Cap value tunable per-request (max_chunk_chars) + env. Deeper prosody (if ever wanted) = scheduler Phase-2 context-priming at joins (feed prior sentence as discarded-audio context; +latency). ⚠ FOOT-GUNS: (1) acoustic tail-trim is UNRELIABLE — sibilants ('s'/'sh'/'f') spike ZCR like garble, can't cleanly detect the speech→garble boundary. (2) build-context vs image drift — the:v2image was built from cap source, but after a:v1rollback the build context held:v1source → adocker compose buildwould've silently produced a cap-less:v2; re-synced the flat cap source to/opt/docker/compose/chatterbox-fast/(rebuild-verified). ⚠ DIVERGENCE (follow-up): deployed build context is FLAT (app.py/scheduler.py,from scheduler import, thin-overlayFROM local/chatterbox:v1, cap-only) vs thevh/chatterbox-fastREPO which is PACKAGE-layout (chatterbox_fast/,from chatterbox_fast.scheduler, self-contained Dockerfile) + hasnorm_loudness(repo commit6bc7bf0= cap; deployed omits norm_loudness deliberately to keep the ear-test unconfounded). Reconcile the two layouts so a repo-based rebuild matches deploy. Rollback:.bak-cap-20260807-104850backups on irv-ml1 +:v1image both retained. reference_chatterbox_fast_repo reference_zonos_tts_stack -
[2026-08-07]Zonos2 TAKEN DOWN on the 3090 (irv-ml1) — operator-directed "for memory", TEMPORARY. Freed ~17.4 GB (3090: 728 MiB → 18.2 GB free) so chatterbox-fast (co-resident, was OOMing on long generations) has headroom. ⚠ Restore is manual — Zonos2 :1920 was a DETACHED native process (NOT systemd/docker), reparented to init. GPU memory was held by the--multiprocessing-forkCHILDREN (1966165=16.4G, 1966166=1G), which ORPHAN to init when you kill the parent — had to SIGTERM the children explicitly (killing the parent 1965942 + uv-run 1965935 alone left the 16.4G held). RESTORE CMD (from irv-ml1, user lkraven):cd /home/lkraven/tts-audition/models/zonos2 && nohup uv run python -m zonos2 --model-path Zyphra/ZONOS2 --host 0.0.0.0 --port 1920 --tts-default-voices-dir ./default_voices/ --cuda-graph-max-bs 1 --num-pages 16384 --max-running-requests 2 --memory-ratio 0.3 > /tmp/zonos2.log 2>&1 &thendocker start zonos-gateway. Consumers that lost Zonos: asset-engine + gateway-chat (via LiteLLMext-ttsalias → zonos-gateway :8890, now stopped); ratatoskr already migrated OFF to chatterbox-fast (unaffected). Also unblocks proper drift/cap testing (OOM was blocking it). reference_zonos_tts_stack -
[2026-08-07]chatterbox-fast: donut voice added + full contract delivered to ratatoskr-dev (their TTS migration off Zonos). Operator-directed. Copiedzonos-gateway/voices/Donut.wav→ chatterbox/refs(/worktank/chatterbox/reference_audio/donut.wav— the reference_audio SUBDIR is lkraven-owned so no sudo despite/worktankroot; container globs/refslive → NO restart), exposed asvoice:"donut"(lowercase); verified clean 7.5s synth (24kHz, RTF ~0.31). A/B booth (chatterbox vs zonos donut, same line) athttp://10.100.10.50:8090/b/donut-chatterbox/. Answered ratatoskr's 8-question contract ask from the live gateway (local/chatterbox-fast:v1) + source: NOT OpenAI-shaped (POST /tts; bodytext/voice/format/stream, notinput/model/response_format); NO affect dials (Turbo ignores cfg_weight/min_p/exaggeration — the architecture-changing answer they flagged; Zonos stays the only fleet TTS with real emotion steering); streaming WAV placeholder-header shape IDENTICAL to Zonos (their per-chunk Web Audio path survives); SR 24000 (Zonos 44100); server chunks arbitrary-length text internally (no client-side chunking, unlike Zonos's 71.2s cap); English-only, no language pin. FYI-worthy (operator): ratatoskr is moving its RP-surface TTS OFF Zonos back to chatterbox-fast → loses the live-PAD affect coupling (heavy Zonos emotion investment) — their call, trade-off flagged to them. auto-memoryreference_chatterbox_fast_repoenriched w/ the live contract. reference_zonos_tts_stack -
[2026-08-07]Fleet reranker cut over: Qwen3-Reranker-0.6B → BAAI/bge-reranker-v2-m3 (Brokkr R43). The incumbent was measured HARMING 80/90 fleet queries (no-reranker beat it 89/90 vs 56/90). R43 bake-off: the A2 control (same Qwen weights, seq-cls head) scored identical to the incumbent → proved the fault is a training-prior not the serving head → cancelled the expensive Qwen3-4B arm; A3 (bge-v2-m3) won on multilingual safety + bare-name recovery. LiteLLMrerankerrepointed incumbent→A3 :8013 (boundary 2026-08-06T17:37:48Z, config-edit + ~52s gateway restart); R42 v13 gate PASSED first-ever (56/90→90/90). Incumbent kept warm :8002 (rollback viaqwen3-rerankeralias), A4 fallback :8014. Full arc + rollback runbookdocs/pfi/reranker-selection-ledger.md; commits ad2df89/2c11748/377f8a4 (unpushed). auto-memories: the earlier reranker-serving notes. -
[2026-08-07]Personal-Worldtree kb-contamination incident (WT #394) diagnosed; attribution CLOSED UNRESOLVED. A reconcileWingStore._embedfull-tree walk (kbfs_root=KB_PATHroot, sibling wings nested) swept 5,354 fiction+main rows into personal'sknowledge_base(2 superseded generations served as current). Fixed by WT #394 (aca39a1, kb walks exclude sibling wings; ships b182). Trigger un-attributable — peer reconcile via the SHARED infra-ops identity + 0 dockerd exec-logging = fingerprint-less. Durable finding → auto-memoryinfra_ops_shared_identity_attribution_gap, PARKED (operator ruled A) into project_migrate_infra_access_to_claude_credentials. Evidence hold on the 5,354 rows until operator sequences cleanup (w/ Brokkr, on #394's agenda). -
[2026-08-05]Fleet CI resilience flip (DEFAULT_ACTIONS_URL=self) — attempted end-to-end, PARKED on a runner action-fetch auth blocker; infra-ops to research it (operator-directed, deferred, NOT now). 7 gitea action mirrors staged public+populated (orgsactions+astral-sh); the flip resolvesuses:correctly but act_runner v0.6.0 can't authenticate its fetch to gitea 1.26 ("Invalid username or token. Password authentication is not supported"). Reverted (CI back on github default);REQUIRE_SIGNIN_VIEW=falseKEPT as a standing change (operator, internal WG net). Full endeavor, the reliable nh3-dev-egress + git-SSH mirror method, exact config state, smoke method, and next step →persistent-memory.d/2026-08-05-ci-flip-parked.md -
[2026-08-05]Booth — 3 features shipped, live on:8090+ tagged. (1) verbatim-index.htmlbooths get a floating top-right "‹ all booths" chip + inherited favicon, doctype/charset-safe byte-injection (booth-v0.1.5,8577e7e); (2).mdrenders +.txt/.logview in-booth without downloading via the/b/<n>/viewroute + amarkdowndep +doc.html(booth-v0.1.6,315faac); (3) prev/next arrows in the image zoom viewer — wrap-around + keyboard ←/→, hidden for single-image booths (booth-v0.1.7,c37a425). Canonicalservices/booth/; deploy =systemctl --user restart booth.serviceon nh3-dev (runs from the checkout's.venv;uv pip installnew deps into it first); 47 tests.uv.lockgitignored (348c5c1). -
[2026-08-05]worldtree herald re-nudge bug root-caused → forseti shipped althing-core v2.1.2 (d5d33df, deployed on nh3-dev).herald.py:363rendered the wake command from the empty fresh mail set on the re-nudge path (should bedeliver_msgs) →messages[0]IndexError → un-suppressed outer catch-all → 7s crash-loop for 9 days on worldtree-codex's pane route (mimir-dev surfaced it; I traced it from the editable source). Fix +render_commandempty-guard + outer log-suppress + 3 tests + contract amendment, all forseti's. nh3-extdev herald 2.1.2 upgrade DEFERRED (operator, not-now): extdev is a WHEEL install (not editable), unexposed (no pane routes); the verified 2.1.2 wheel is staged on nh3-dev/tmp(sha256003508…cef27) —uv tool install --force+ restart both heralds when un-parked. extdev herald-unit provenance resolved (operator-authorized 2026-07-25 via forseti relay; recorded in this file's 07-25 herald-install entry). auto-memoryreference_nh3_dev_althing_herald. -
[2026-08-03]worldtree b168/#384/#385 arc COMPLETE — providers.yaml boot-gate pre-sync → b168 deploy → DCC+P&P re-ingest (705+667 concepts, 0 truncations, #385 budget fix validated vs April's 785 control) → #381 restart → operator-approved production dedup sweep (785 April orphans deleted frommain, 4009→3224). Fiction wing 166→1,372 concepts; consumer verify 0/5→5/5→saturated. Three of MY foot-guns hardened into fleet runbook rules (mv -t,docker exec -u 1000, shared-containerd pull-race — see Tried-and-abandoned). Full runbooks (deploy-wt-config, Chroma-verify, config-delta pre-sync rule, #381, sweep) →persistent-memory.d/2026-08-03-worldtree-b168-384-385-arc.md -
[2026-08-02]mimir-inbox / #377-read-path arc — deployed + 4 bugs found/fixed/verified + a cloned voice. mimir-inbox live on corviduo-dev:8091 (#377 write+read proven, live8ece117); worldtree-dev #380 (wing-blind index) + #381 (stale-client restart) + #382 (intermittent Mimir grounding) chased and verified 3/3 by ratatoskr-dev; muninn-gate → dispatch 0.1.5; donut voice cloned from the 65-frost Booth bundle into the Zonos gateway; Zonos streaming confirmed already-working. Full arc, procedures, and lessons →persistent-memory.d/2026-08-02-mimir-inbox-arc.md -
[2026-07-31]muninn-gate (#377 ingestion front door) BUILT + DEPLOYED + healthy on corviduo-dev:8090. First-boot acceptance passed (watcher:running:true proves ingestion_root byte-identity); submit path deferred to the mimir-inbox era. Full wiring (uid-1000, state-volume mount, staging path-agreement, BuildKit-secret build, deferred repoint + operational guards) →persistent-memory.d/2026-07-31-muninn-gate-deploy.md -
[2026-07-31]worldtree-sdk 1.1.0 (Python) published to vh Gitea PyPI + a durable infra-ops publish cred. memory_context pass-through; unblocked wyrd-dev. claude-bot now a write-collaborator onvh/worldtree-sdk(source pulled via the Gitea API archive — git-HTTP 403s on that repo); publishing to the vh USER namespace can't be delegated (401reqPackageAccesseven withwrite:package) so it needs an owner token — operator saved a FULL vh site-admin token at~/.config/gitea/vh-token(0600) for it (⚠️ high blast radius, kept over a scoped one; org-namespace migration is the only real de-personalization, parked by wtsdk-dev). auto-memoryreference_infra_ops_vh_gitea_token_and_sdk_publish. -
[2026-07-31]kimi-k3 "output cap" root-caused = a ~16384 REASONING-token ceiling, not an output cap. heid's cross-frontier panel was silently degraded (empty content,finish_reason: stop). dvalin+bil researched (docs said deprecated-max_tokens); heid's live data refuted that (completion hit 18455) → it's a reasoning ceiling. Proven on the wire against heid's real 500KB bundle:reasoning_effort: lowdrops reasoning under the ceiling → content returns, on BOTH coding + general endpoints. Fix is CALLER-side (no gateway change): sendreasoning_effortviaextra_body(LiteLLMdrop_params: truestrips the top-level param — why heid's earlier attempt no-op'd). Relayed to heid to validate; backstop =allowed_openai_paramson the route. →persistent-memory.d/2026-07-31-kimi-k3-reasoning-cap.md -
[2026-07-27]Zed edit-predictions: keyless FIM-completion route SHIPPED end-to-end. Operator wants Zed's inline edit-prediction (which CANNOT send an auth header) to reach a FIM coder via/v1/completions. Deep-research (106-agent workflow) pickedQwen/Qwen2.5-Coder-1.5B(BASE, Apache-2.0; native FIM<|fim_prefix|>/<|fim_suffix|>/<|fim_middle|>IDs 151659/60/61; Zedprompt_format:"qwen"). Runner-up 3B = non-commercial Qwen-Research license; no small dense Qwen3-Coder exists (all MoE, smallest 30B). Stood upvllm-coderon ana-ml2 GPU1 :8020 (served-nameqwen2.5-coder-1.5b, 8192 ctx, util 0.06, fp8 KV). To fit, shrank granite (phasing out, operator-directed): util 0.27→0.13, max-len 131072→16384, seqs 1024→256 (freed ~14 GB; the KV-≥-1×-max-len rule crash-looped it at util 0.12/32768 → settled 0.13/16384). LiteLLM aliascoder-fast→ :8020 (mode: completion). Minted acoder-fast-SCOPED virtual key (verified 403 ongen— the real blast-radius bound). Builtzed-fim-proxy(ana-docker :4141,network_mode: host, stdlib-python,stacks/zed-fim-proxy): keyless POST/v1/completions, model-allowlistcoder-fast, injects the scoped key → LiteLLM :4000;GET /pinganon liveness; wrong-model→403, wrong-path→404,/chat/completionsrejected. Verified keyless FIM end-to-end ('a + b', finishstop). Zedapi_url=http://10.250.50.70:4141/v1, modelcoder-fast, prompt_formatqwen. source-IP allowlist intentionally LEFT OFF (operator direction 2026-07-27) — do NOT tighten: Zed roams the operator's WireGuard10.0.0.0/8, so a single-IP pin would break it. Blast-radius bound is thecoder-fast-scoped key + model/path allowlist (keyless but coder-fast-only, internal-net-only). (The proxy does exact-IP matching; scoping to the10.0.0.0/8CIDR would need CIDR support — deliberately not added.) Canonical:stacks/vllm(coder + granite shrink),stacks/litellm(coder-fast),stacks/zed-fim-proxy(NEW). Server vllm compose.yaml has benign stale-comment drift vs canonical (didn't overwrite the newer canonical). -
[2026-07-27]Muninn ingestion-watcher sidecar deployed on PERSONAL Worldtree (#377). worldtree-dev request (research-wing ingest arc, personal-only per the 2026-07-16 topology ruling); operator-approved. Added aworldtree-muninncompose sidecar to/opt/worldtree-personal/compose.yaml—<<: *worldtree-commonanchor inherits the api's image + full env + config/state/kb mounts;command: python -m core.muninn --watch;restart: unless-stopped;stop_grace_period: 1h(INV-377-7: max 2 concurrent × worst-case job, SIGTERM-drains). Pinned to the running SHA773866084af9(b146, ≥ b143 — dodges both the:latesttrap AND the "pre-b143 ref resurrects deleted dispatch.py from stale bytecode" warning). Verified: running / 0 restarts / flock sole-runner (no rc3) / heartbeat live at{ingestion_root=/data/state/ingestion}/.watcher-heartbeat(poll 30s). Containerworldtree-personal-worldtree-muninn-1; backupcompose.yaml.bak-muninn-20260727-081920. DURABILITY RESOLVED (worldtree-dev, same day): Q1 was a LIVE FOOTGUN —deploy-personal.ymlscp's the REPO compose.yaml over the box's + runsup -d --remove-orphans, so the box-local sidecar would've been clobbered AND orphan-removed at the next staging tag. worldtree-dev fixed at source: moved the sidecar into their repo compose.yaml gated behind amuninncompose profile (commit 5d7f6bd) — shared compose stays instance-identical,.envCOMPOSE_PROFILESdifferentiates (demo watcher-less). My action: addedCOMPOSE_PROFILES=muninnto/opt/worldtree-personal/.env(backup.bak-muninn-profile-20260727-082541; no-op vs the current unprofiled box-local sidecar → seamless handover at next deploy). Q2: their deployup -d's the whole stack w/WORLDTREE_IMAGEexported → sidecar version-tracks the api, no drift. CONFIG-AS-CODE EXTENSION: mirrored the non-secret delta aspersonal/env.publicinvh/worldtree-instance-configs(repoa9d091e) — FIRST extension beyond config.yaml files to env-level config; the secret-laden.envstays box-only,env.publicrecords only non-secret infra-ops-owned env deltas (record, not a deploy source —deploy-wt-configglobs*.yaml). BOUNDARY CLARIFIED: compose.yaml = worldtree-dev's (their repo, instance-identical, scp'd on deploy); per-instance.env= infra-ops's differentiator. Deploy step of the #363/#377 arc. #377 CLOSED — acceptance PASSED 2026-07-27: worldtree-dev enqueued a test job via muninn-dispatch 0.1.0 in a one-shot ephemeral container (no docker-exec); the sidecar claimed it within one 30s poll, drove it to terminal (structure→summarize→complete), zero restarts/rc3, heartbeat fresh throughout — whole loop (request→deploy→durability fix→acceptance) in <2h. (Pre-existing pipeline bug #379 surfaced —output.kb_notes=falseignored → 1 inert test note in the research wing — worldtree-dev owns it, nothing infra-ops-side.) ⚠ OPERATOR-SURFACE (open): theenv.publicoverlay mechanism is a repo-scope call to bless/adjust. reference_worldtree_deploys_cicd reference_worldtree_instance_configs_repo project_worldtree_research_wing_ingest -
[2026-07-27]jackdaw-compose.service DECOMMISSIONED (jackdaw-dev request; the JackDAW AI Composer was cut from v1 by operator decision 2026-07-27). Stopped + disabled the nh3-dev:8787user service (no client calls it — ai/server/AiChat deleted from main,/composeproxy removed); unit archived not deleted →~/.config/systemd/user/jackdaw-compose.service.decommissioned-20260727(revival = rename +daemon-reload). No credential revoked — the unit used the SHARED all-agents LiteLLM key (sk-eA_XOd…, modelgen), not a dedicated one. Code preserved on jackdaworigin/ai-composer-preserved; treat as permanent. The:4500HTTPS audition bench is untouched. (Supersedes the 2026-07-23 stand-up line below.) -
[2026-07-26]DemoBIFROST_CLIENT_ALLOWED_HOSTS+=10.100.10.50:8391(wyrd-dev's bifrost memory-store provider; operator-approved). First live exercise of the #376 config-as-code boundary working as designed — worldtree-dev routed the delta to infra-ops instead of hand-editing/opt/demo. Appended to/opt/worldtree/.env:25(now 4 netlocs), recreated ONLYworldtree-api(the gated conv-api path), health-gate green, container env verified. REUSABLE FOOT-GUN: an env-var change needs a container RECREATE, notdocker restart(env is baked at create); and the demo.envdefaultsWORLDTREE_IMAGE=:latestwhile the box runs a specific SHA — so a naivecompose uprisks the documented stale-:latestcrash. FIX = capture the running image live (docker inspect …Config.Image→…:9eff09f007ba) andsudo env WORLDTREE_IMAGE=<sha> docker compose up -d worldtree-api. Backup/opt/worldtree/.env.bak-bifrost-20260726-221602. BOUNDARY SEAM: this was a compose-.envvar, NOT aconfig.yamlfile invh/worldtree-instance-configs— the.envholds secrets so it's deliberately not repo-tracked → env-deltas land directly on the box (config files are versioned, compose env vars aren't). reference_worldtree_instance_configs_repo -
[2026-07-25]nh3-extdev herald installed — box is now a full v2 push participant. forseti flagged (relaying operator): extdev had thealthing-heraldbinary (/usr/local/bin/) but NO unit (skipped the whole v2 arc), soherald-status= "notifications suspended" and ldp-dev ran on thealthing-light-monitorpoll fallback. Installed/etc/systemd/system/althing-herald.serviceas a SYSTEM unit mirroring the receiver (User=althing-svc,Group=althing,Environment=ALTHING_ROOT=/srv/althing,ExecStart=/usr/local/bin/althing-herald --poll 5, enabled) via the lkraven@ NOPASSWD path (used under the then-mistaken belief infra-ops was sudo-less — CORRECTION 2026-08-03: infra-ops has had full NOPASSWD sudo on extdev since 2026-06-25 per reference_nh3_extdev_althing_mesh; future extdev installs can self-serve as infra-ops without the lkraven@ hop). Verified: active / 0 restarts /herald-statusflipped to "✓ herald up." No zellij routes on extdev → heartbeat + wake-FIFO poke only, no pane-dispatch; ldp-dev keeps light-monitor unless it opts into a wake-listener. -
[2026-07-25]Booth v0.1.4 — booths are downloadable. Verbatimindex.htmlbooths (e.g. edict-design-brief) were served raw with no download affordance. Added/b/<name>/?download=1(streams the whole booth as<name>.zip, attachment) +?dl=1on the file route (forces Content-Disposition attachment so html/md/text saves instead of rendering inline) + ⬇ zip links on the index card (the accessible spot for verbatim booths) and the gallery header.zip_booth()helper, 31 tests green; verified live on nh3-dev :8090 (edict-design-brief.zip = index.html + ui-design-brief.md). eshpfi91a031f/ tagbooth-v0.1.4. -
[2026-07-25]bil-smithy-dev wired as an althing zellij-window-ping (pane route). She's adriver: humandwarf peer (panebil-smithyalready live alongside eitri/dvalin/regin-smithy in theClaudezellij session) but had no delivery route → smoke messages posted to the bus but never reached her window. Mechanism (reusable for any pane-route handle):~/.althing/config.yaml→zellij_sessions.Claude.agents[]mapshandle→target(a zellij pane TITLE, matched vialist-panes -jinalthing/zellij.py:resolve_pane_id) →command(heraldwrite-chars+ CR into that pane). The herald loads config ONCE at startup (herald.py main()), sosystemctl --user restart althing-herald.serviceafter editing. Added bil (target: bil-smithy), restarted, verified: herald delivered the pending smoke01KYD7W7CF…(available→attempted→delivered). ⚠️ Noticed pre-existing pane-route errors onworldtree-codex+eitri-smithy-dev("route-error: list index out of range", empty msg_ids — likelyrender_command messages[0]on an empty list; NOT caused by this change, bil works) — worth a herald look. -
[2026-07-25]Kimi K3 wired into the LiteLLM gateway — CODING endpoint (operator-directed; fulfills a Heid gateway request to add a 4th cross-frontier panel arm). Primarymodel_name: kimi-k3→openai/k3@https://api.kimi.com/coding/v1(Kimi Code / Vivace membership; keyKIMI_CODE_API_KEY). A general-endpoint variantkimi-k3-gen-api→openai/kimi-k3@https://api.moonshot.ai/v1(keyMOONSHOT_API_KEY) is kept alongside (originally wired then demoted when the operator corrected: the plan uses the CODING endpoint, not the general Moonshot API). Both keys in compose env + server.env(NOT committed) +.env.example. Both verified live through the gateway :4000 (17+25→"42", "PONG"). k3 constraints on BOTH endpoints (config-pinned + commented): accepts ONLYtemperature=1(else 400 "only 1 is allowed"); REASONING model (CoT inreasoning_content, answer incontent→ tinymax_tokensreturns EMPTY; Kimi Code adds thinking-effort tiers low/high/max). Coding lineup also carriesk3-256k/kimi-for-coding/kimi-for-coding-highspeed(not wired). Reachable by any gateway key spanning all proxy models (incl. shared all-agents key → spends the paid Vivace/Moonshot quota). eshpfiedaa9a9(gen wiring) +9e2f787(coding correction). OPEN: Heid key-scoping — shared key reaches it (paid) vs a dedicated scoped key (asked in althing01KYD63ZBY…). -
[2026-07-25]infra-ops Worldtree config-as-code repo SHIPPED —vh/worldtree-instance-configs(private) built, pushed, validated. Dir-per-instance (demo/,personal/;pinned/= README stub, out-of-scope — no bind-mount, config frozen in image446e5807). Seeded byte-exact from live/opt/<instance>/config; 5 files each (defaults/policies/model_roles/providers/matrix).scripts/deploy-wt-config= diff / deploy / capture, with in-run host backup → install(vh:vh,644) → restart api+matrix → health-gate api/health→ auto-rollback. All verbs live-tested (in-sync, capture round-trips zero-diff, pinned refused, dry-run no-ops). Gitea repo created via ana-docker localhost API with vh creds (operator-authorized one-time); pushed over internal git-SSH10.250.50.70:222(nh3-dev 403s gitea HTTP). Boundary AGREED by worldtree-dev (althing01KYCAECRW…): they stop hand-editing/opt/<instance>/config, route config deltas to infra-ops; three-layer model (image baseline → repo per-instance truth → host bind-mount deploy target); carve-out = their admin-API DB mutations (key mint / tier / retirement) stay in-band, not config edits. By-design deltas (personalagent_architect+ratatoskr-affect-full-allow; demo#308metrics + grants) preserved verbatim. →persistent-memory.d/2026-07-25-infra-ops-wt-config-repo.md, auto-memoryreference_worldtree_instance_configs_repo -
[2026-07-23→25]Worldtree #376 config-divergence arc CLOSED — per-instance config ruled BY DESIGN. wyrdsession.history.writedemo grant was the one real bug (demo-intended grant not on demo; fixed via wholesalepolicies.yamlreplace + restart). The b131 drift guard then surfaced broader divergence = legitimate live-bridged per-instance deltas; operator ruled deltas are the design not rot; guard demoted to INFO (b132); infra-ops drift-watcher built then retired same day. →persistent-memory.d/2026-07-25-wt-376-per-instance-config-arc.md, auto-memoryreference_worldtree_perinstance_config -
[2026-07-20→25]The Booth SHIPPED (v0.1.3) — ephemeral media drop board for CC sessions. New fleet tool: user-systemd on nh3-dev :8090 (services/booth/, FastAPI+Jinja2, Corviduo "Australis" theme, 34 tests), Homepage-linked (Apps). Drop a folder in~/booth-data/<name>→ browsable "booth" (auto-gallery of images/webm/audio, or a folder's ownindex.htmlverbatim), 24h TTL. Added across the session: browser/curl upload-for-pickup with human-readable ids (4-wombat), image viewer (Fit/1:1, conditional toggle), copy-id button (HTTP-LANexecCommandfallback). Registered in global CLAUDE.md tools. auto-memoryreference_booth_media_board. -
[2026-07-23]jackdaw-compose backend deployed as a persistent nh3-dev service (:8787). Hosted for jackdaw-dev: thin statelessbun server/index.ts(from~/development/jackdaw) → LiteLLMgen, Origin-gated (INV-BK04/05), reached same-origin via their:4500bench's/composeproxy.jackdaw-compose.service(env/shared-key server-side, unit 0600, uncommitted). Also stood up + tore down a throwaway cloudflare quick-tunnel for their preview (cloudflarednow installed at~/bin). In the nh3-dev README inventory (cd4d52e). -
[2026-07-19]irv-ml1 ComfyUI — RTX VSR baked into canonical provisioning (comfy-dev ticket DONE). RTXVideoSuperResolution node +nvidia-vfxdep were manual installs; documented both in the canonicalstacks/comfyui/README.mdrunbook (this stack's provisioning IS the README — no automated provision script). Key durability insight: the node lives inbasedir/custom_nodes(persistent, restic-included → durable) but thenvidia-vfxwheel lives in the venv underrun/(disposable, restic-excluded → dropped by anyrm -rf run/*fresh-bootstrap), so the pip step must re-run after every venv rebuild. Both steps run as uid 1000 (root install → venv-ownership crash-loop, reference_irv_ml1_comfyui_mmartial);--extra-index-url https://pypi.nvidia.comkept scoped to the nvidia-vfx install, deliberately NOT a global composePIP_EXTRA_INDEX_URL(would risk perturbing the pinned torch 2.12.1/SageAttention boot bootstrap). Node already live on the box; no host change, canonical runbook now replays it. comfy-dev informed. -
[2026-07-19]vh private Gitea PyPI — consumer READ-access convention set + wyrd-dev provisioned. Consuming agents read the internal vh PyPI (https://gitea.phasefinal.com/api/packages/vh/pypi/simple/) with a shared read-only token (operator call: shared, not per-consumer — read-only blast radius is small, per-agent Gitea identities aren't worth it). Minted a dedicatedread:package-scoped PAT off claude-bot (POST /users/claude-bot/tokens, namevh-pypi-read-consumers; verified reads worldtree-sdk, write-probe 401), revocable/rotatable independently. uv auth =UV_INDEX_GITEA_USERNAME=claude-bot+UV_INDEX_GITEA_PASSWORD=<token>(or~/.netrc); pyproject uses[[tool.uv.index]] name=gitea … explicit=true+[tool.uv.sources] <pkg> = { index = "gitea" }(mirrors soong-lab's bifrost setup). Delivered to wyrd-dev (worldtree-sdk adoption) via mode-600 drop on nh3-dev, drop-and-shred. reference_claude_bot_gitea_creds -
[2026-07-18]soong-lab auto-redeploy WIRED + validated (queued item CLOSED). Added a WT-style CI-deploy step tobuild-and-push.yml: after build+push, the pfi-fleet runner SSHes corviduo-dev as thedeployuser and runsdocker compose pull && up -dfrom /opt/soong-lab, health-gated on/api/version(120s, fails loud). Reused WT'sdeployaccount (uid 1001, docker-group → no sudo); relocated the deploy dir /home/infra-ops/soong-lab-deploy → /opt/soong-lab (deploy-owned; old dir retired.retired-20260718). Minted a dedicated soong-only ed25519 deploy key, pubkey ondeploy's authorized_keys (fp SHA256:MG7M3Ri…). First dispatch FAILED on a bad DEPLOY_SSH_KEY paste (error in libcrypto— unparseable key bytes; build+push were fine, live Soong untouched); repo secrets are vh-owner-only (claude-bot token = write:package only → 403; the vh package-scoped PAT also 403 on secrets), so operator re-set DEPLOY_SSH_KEY/HOST/USER. Re-dispatch run #5 GREEN: live container recreated ...541f7730 → ...07526a08, health 200. soong-dev pinged to sync DEPLOY.md's redeploy path (/opt/soong-lab) + close the "auto-pull open follow-up". →persistent-memory.d/2026-07-18-soong-lab-auto-redeploy.md -
[2026-07-18]soong-lab auto-redeploy APPROVED — QUEUED for next session (deferred, not started) — Vuong approved (via soong-dev thread01KXT3A6C3908TA4V9THV3AMH7); mechanism = WT-style CI-deploy step (runner SSHes corviduo-dev →compose pull && up -d+ health-gate); blocked on a vh-owned runner→corviduo-dev deploy SSH-key secret (reuse WT's demo-deploy key). Operator: "do soong on fresh context." →persistent-memory.d/2026-07-18-soong-lab-auto-redeploy.md -
[2026-07-15]arbo fully switched off image-judge (qwen-image-bench) -> gen; image-bench pending eviction post-bake →persistent-memory.d/2026-07-15-arbo-fully-switched-off-image-judge-qwen-image.md
188 older entries archived to archival-memory.md.
Tried and abandoned
-
[2026-08-15]Grafted bf16 MTP loads UNINITIALIZED (0% accept) unlessre:^mtp.*is in the quant-configignore; and W4A16=Marlin (not native FP4) costs ~20% even on decode. Cost a premature 79 GB delete of a good model (declared desync-dead off the 0%). Lessons: test MTP on bf16 FIRST, isolate before deleting; modelopt 0.43 is dependency-hell for qwen3_5 (list-vs-dict quant_cfg + transformers conflict) — use llm-compressor. Full →persistent-memory.d/2026-08-15-uncensored-gen-seat.md -
[2026-08-03]ComfyUI--enable-triton-backendon the irv-ml1 A6000 crashes EVERY render — Ampere has no hardware e4m3. adhoc-agent's operator-approved probe: comfy_kitchen's triton backend has a FUSED int8 matmul that would beat the eager backend's ~1.9x-slower unfused int8 path (21.3s vs 11.2s fp8 on the Moody Krea2 int8 checkpoints). Flipped it (added toCOMFY_CMDLINE_EXTRA, recreated) →triton.compiler.errors.CompilationError: ValueError("type fp8e4nv not supported in this architecture. supported: fp8e4b15, fp8e5")incomfy_kitchen/backends/triton/quantization.py:145 dequantize_per_tensor_fp8, failing at node 5 CLIPTextEncode. Triton's fp8 dequant kernel targetsfp8e4nv(Hopper/Ada e4m3); sm_86 Ampere (A6000) lacks hardware e4m3 → the JIT compile dies. With triton on it grabs the global--fp8_e4m3fn-text-encdequant, so every render (fp8 AND int8) dies upstream at the text-encode step — the int8 UNet path never ran, so the convrot-coverage caveat wasn't even the limiter. Reverted cleanly (~15s to healthy, image unchangedsha256:94afb8ca, sage intact, prod restored). The parked cu130 rebuild won't fix it (e4m3 = hardware format, not CUDA version). DEFERRED to the Ada refresh (operator: "ada is coming, we'll optimize then" — Ada sm_89 has native e4m3, so triton's fp8 path should compile there). Mechanics:--enable-triton-backendis a composeenvironment:var, so toggling it needsdocker compose up -d(recreate), NOTdocker restart(reuses the baked env, no-ops silently). Full: auto-memoryparked_triton_backend_ampere_fp8. -
[2026-08-03]corviduo-dev shared containerd: a concurrent-pull race fails ONE instance's deploy; DON'T "prune to fix" — the image is in-use by the instance that won the race. b169 personal deploy failed atdocker compose pull(Lchown … no such file or directoryon the big torch layer → looked like disk pressure / corrupt snapshot). ACTUAL: NOT disk (56G free, inodes 7%). demo + personal + pinned share ONE/var/lib/containerdon corviduo-dev; demo (from main) and personal (from staging tag) extracted b169's shared torch layer simultaneously → personal's hit a partial snapshot mid-race and aborted while demo's completed. The image6e34a87was FULLY VALID — demo was RUNNING it healthy. Fix = just re-run the failed deploy (image already materialized; compose pull finds it present). NEAR-MISS: worldtree-dev's suggested "prune unused images/snapshots" would have rmi'd6e34a87= the image the running demo depends on → demo outage. Lesson: before any prune/rmi "cleanup,"docker psthe running images — an "unused" image may be a co-tenant's live one; and verify the failure's REAL cause (disk? inode? in-use? race?) before applying the suggested remedy. (Pipeline fix, deferred: serialize demo-from-main + personal-from-staging, or a per-image pull lock, to avoid the shared-layer extraction race.) -
[2026-08-02]docker execinto worldtree containers defaults to ROOT — root writes contaminate the uid-1000 (vh) KB tree. Mysudo docker exec … --reindexon personal ran as ROOT (muninn app = uid 1000); its wing git-commit + atomic note-swap left root-owned files in theworldtree-personal_worldtree-kbvolume: a root-owned.old-<job>backup dir (blocked the uid-1000 retry'srmtree→ Errno 13, because unlink needs write on the DIR and it was root:root 755) AND 60 root-owned loose git objects in.git/objects/. Fix (host-side, corviduo-dev):sudo rm -rfthe superseded.old-dir (tar'd aside to /tmp first) +sudo find … -user 0 -exec chown 1000:1000the objects (ownership-only, git-content-safe; the.git/objects/XX/dirs were vh-owned so these weren't a hard blocker, but violated "clean tree"). RUNBOOK RULE (worldtree-dev, ADOPTED): anydocker execinto worldtree containers that WRITES pipeline state runs-u 1000, never default-root — same genus as the mv footgun (acting without matching the target's constraints; 3rd such slip in one session). GOTCHA that hid the scope:find … -user 0 | head -20TRUNCATED (the.old-dir alone had 153 files, so the first page was all.old-) → I "verified clean" off a partial list. Neverheada scope-defining find; count first (| wc -l). Related blind-spot (muninn-dev): a root-owned job SUBDIR passes every requeue guard (job_row/dispatch/list_jobs render fine) AND/health(contract'sos.access(ingestion_root, W_OK)tests only the ROOT dir, so a foreign-owned subdir underpending/still reportsingestion_root_writable: true) — then the uid-1000 gate can't write into it. "Clean board + green /health + failure at next mutation." muninn-dev added an OWNERSHIP column to the standing post-move check to catch it; two green signals both miss a foreign-owned subdir otherwise. -
[2026-08-02]mv <job> complete/ → failed/RENAMED the job tofailedbecause failed/ didn't exist. worldtree-dev's round-2 unblock command (mv /data/state/ingestion/complete/<job> /data/state/ingestion/failed/) assumedfailed/existed; on PERSONAL muninn it did NOT (fresh instance — root wasactive/ complete/ pending/ sources/, nofailed/).mv src nonexistent/renames src→nonexistent, so job1 became thefaileddir and job2 nested inside it. Caught on post-movels(failed/ held job contents, not two subdirs), reconstructed via complete/ as watcher-safe scratch + rebuiltfailed/(worldtree:worldtree 755) — NO data loss. Lessons: (1) beforemv X into-dir/, verify the dir EXISTS ([ -d dir ]) — an emptyls dir/ 2>/dev/nullis AMBIGUOUS (missing vs empty), which was the preflight miss that let it through; (2) the correct guard ismv -t <targetdir> <src>(--target-directory): it refuses a MISSING target loudly (rc=1, "No such file or directory", nothing moved) — this is the house convention for queue/state moves now. TESTED by muninn-dev on coreutils 9.1: a trailing slash does NOT protect —mv src failed/withfailed/missing STILL silently renames tofailed(rc=0); "just add the slash" is a false guard. (mkdir -p failed/first also works, butmv -tinverts the failure from silent-wrong to loud-safe in one flag.) Containershis dash — no(in echo strings. SILENT failure mode (muninn-dev carry-forward): a misplaced ingestion-state move doesn't crash anything —list_jobs()stays OK, loose files are inert; the ONLY symptom is the job quietly absent from the board (job_row→None, requeue→not_found/404, looks IDENTICAL to the original block). So after ANY state move, verify the job is actually ON THE BOARD (job_rowfound + guards pass), don't trust mv exit codes — and confirmjob.dispatch.jsonsurvived (requeue refuses a dispatch-less job with the same not_requeueable symptom). Cross-checked + all-clear'd by muninn-dev, who correctly refused to mutate ingestion_root (INV-MG-1) and flagged instead. DON'T TIDY (round-2 pending): both DCC + P&P jobs currently REST in personalfailed/with manifests readingstate: completeuntil round-2 requeue runs — deliberate + load-bearing (requeuekeys on DIRECTORY PLACEMENT, not manifest state); looks wrong to anyone cold, leave it exactly as-is. Round-2 sequencing: the requeue is mimir-dev's browser flow (pending their operator's board-vs-API ruling); muninn-dev is the gate confirmer (runs the post-move board-check inside its custody — the right split, don't reach across INV-MG-1); infra-ops = the #381 restart after both jobs go terminal, then later the supervised main-collection sweep. Guard-verified HOLD LIFTED by muninn-dev 02:36Z. ARC COMPLETE (2026-08-03 ~05:49): both books terminal — DCCmimir-6351554e8e8f705 concepts + P&Pmimir-f3887c9b97b7667, extracted AND indexed, 5/5 phases, 0 failures/truncations (validates the #385 budget fix vs April's 785 control); #381 restart-after-ingest FIRED (personal api, healthz/readyz 200 ~25s), retrieval-visibility confirmed (search_library returns DCC+P&P from fiction post-restart); handed ratatoskr-verify go to worldtree-dev. Delete-sweep precondition NOW MET — the stale DCC rows inmainare genuine duplicates of livefictionrows, so worldtree-dev's supervised sweep of the ~785 April orphans is unblocked (still comes to me supervised: snapshot + operator-in-loop). -
[2026-08-02]donut voice multi-clip reference (onyx-58 expansion) — TRIED, REVERTED. Folded theonyx-58bundle's 3 Donut clips (seg101/seg110/seg148) in alongside the original seg000 → a 52.0s 4-take concat reference, hoping a longer ref → more robust speaker embedding. A pinned-seed A/B (5 pairs, varied registers, boothdonut-onyx58) showed the original single-clip seg000 (16.3s) sounds better — concatenating disparate takes muddied the timbre more than the extra range helped. Reverted to seg000-alone (live + build-source). Two durable lessons: (1) for a faithful clone, a single clean representative take can beat a longer multi-take concat — more reference audio is NOT automatically better when the takes vary. (2) Emotion steering pulls the output AWAY from the cloned voice fast (operator's craft rule) — keep donut (and clones) emotion-neutral for fidelity; the gateway only enables emotion when anemotion_*/presetdial is explicitly sent, so bare{input,voice}calls stay pure-clone.seg148was diarized SPEAKER_03 but IS Donut (operator-confirmed misdiarize). onyx-58 curated bundle lives in boothonyx-58(24h TTL — stash to/mnt/smithy/voice_clones/if a future middle-ref experiment is wanted). -
[2026-08-02]Verifying the INDEX is not verifying GROUNDING (#382). Asearch_libraryreturning wing=fiction hits proves the content is retrievable; it does NOT prove the agent (Mimir) trusts and uses those hits vs. silently answering from training. I reported "Mimir read Austen back to you" off a grounded-looking answer; ratatoskr-dev caught that grounding was intermittent (some sessions discarded the correct hits and substituted training knowledge). Test the harder claim — are the citations note-extracted or model-knowledge? — and reading the DEPLOYED artifact beats trusting the test for "is the fix live." -
[2026-07-30]brokkr's WebSearch "verification" CONFIRMED a hallucination — 3 phantommicrosoft/Mage-Flow-{Base,Turbo,Edit}repo IDs. brokkr-smithy-dev handed 3 gated-looking repo IDs for an operator-directed model pull; they don't exist (its own web-search fabricated an arXiv ID + project page, twice). Lesson: the HF registry API is ground truth — an unauth 401 ≠ exists ({"error":"Invalid username or password"}masks private/gated/nonexistent alike), an authed 404 = phantom, andauthor=X&search=Yrefutes existence. API-verify every repo ID before a pull; LLM-summarized web fetches confabulate. auto-memoryreference_verify_hf_repo_ids_before_pull. -
[2026-07-30]magpie TTS serving — evaluated, ABANDONED. Pulledmagpie_tts_multilingual_357m(the one real repo of brokkr's batch) to NFS, stood it up on irv-ml1 (ephemeral NeMo-Speech-maincontainer — stock PyPI/NGC NeMo can't load v2607), A/B'd vs Zonos → Zonos wins expressive English decisively, multilingual not needed. Not served;magpie-nemotorn down..nemoKEPT on NFS as brokkr's fine-tuning base. auto-memoryproject_magpie_tts_eval_rejected. -
[2026-07-25]Chaining the althing wake-listener arm orphans it.reply && althing-wake-listener &(or spawningalthing-wake-listenerwith&inside arun_in_backgroundtask) → the&-child reparents to init, UNTRACKED by the harness: no fire-notification, and re-arms bounce rc3 off a lock nothing services (mail silently unwatched). Compounding foot-gun: re-arming after a plain operator turn (not an actual fire) collides with the still-live prior listener (rc3). FIX: spawnalthing-wake-listeneras its OWNrun_in_backgroundtask, and re-arm ONLY after a real fire (<task-notification> completed rc0). Reclaim an orphan withalthing-cli stop-monitorthen re-arm. -
[2026-07-25]Peer green-light ≠ operator consent for a managed-box mutation. Auto-mode guard blocked a config-replace+restart on the Worldtree-team demo box that was authorized only by worldtree-dev's althing message — correctly: a persistent change to shared infra needs the operator's yes for that specific change, not a peer's. Surface it; don't route around the guard. (The operator then stood the whole change down — the guard's hold was the right call.) -
[2026-07-18]Fleet Gitea CI foot-guns (3 failed soong-lab builds): the pfi-fleet runner'snode:20-slimjob image has no docker/git soactions/checkout+docker/*marketplace actions all fail;vhis a USER so its packages are owner-write-only (claude-bot repo-admin-collab still 401s on push/publish, and can't set repo secrets — owner-only);GITEA_-prefixed secret names are reserved/illegal. Fixes in →persistent-memory.d/2026-07-18-fleet-gitea-runner-build-recipe.md -
[2026-07-18]zonos-gateway local clone had NO git remote + a history unrelated to gitea's — "committed to vh/zonos-gateway" was never pushed from that clone; two separategit initlineages, no merge-base. Reconcile = reset local→origin/main + overlay the changed files + push (NOT force — that erases gitea's voice-wav commits). Checkgit remote -v+git merge-basebefore assuming a clone is wired.
132 older entries archived to archival-memory.md.