Files
esh-pfi-infrastructure/persistent-memory.md
T
vh 7bc9754e40 ops: adopt AES-128 on both Anaheim tunnels and close public admin
Both tunnels now negotiate AES-128 for ESP, applied make-before-break
so neither dropped waiting on its far end: the FortiGate was widened
to accept the new cipher alongside the old one first, then each UniFi
gateway was flipped. Single-stream throughput moves from 245 to 270
on the NH3 tunnel and from 268 to 304 on the ESH tunnel. Both network
objects were diffed against pre-change snapshots and the only field
that moved on either is the ESP cipher.

The proposal lists are left accepting AES-256 as well. The peers offer
only AES-128 so the extra entries are inert, and retaining them means
a gateway reverting cannot strand a tunnel.

With that up, the WAN administrative surfaces are closed. The
interface is back to permitting only ping, and the infra-ops account
is again restricted to RFC1918 space. Ports 443 and 22 were confirmed
closed from two separate sites and management over the tunnel still
works. The close was issued over the tunnel rather than over the WAN,
since withdrawing SSH from the interface while connected through it
would sever the session mid-command. The box now has no out-of-band
path, which the memory records explicitly.

Also captured: the two UniFi vault items have different shapes, one a
bare key and one a documentation note requiring extraction, which
produces an opaque nginx rejection if missed, and the ESH key's first
confirmed write.
2026-08-23 15:44:32 -07:00

71 KiB
Raw Blame History

Persistent memory — eshpfi-management

Last updated: 2026-08-23

Always check for /tmp/infra-ops-handoff.md — if it exists and its Written: stamp is under an hour old, read it (it carries the in-flight handoff from the previous session), then delete it. Older than an hour: stale — delete it unread.

Repo purpose

Reference workspace for PFI infrastructure: server inventory, canonical Docker Compose stacks, ops playbooks, and conventions. Authoritative copies of compose files live on the servers under /opt/docker/compose/<stack>/; this repo mirrors them for version control, editing, planning, and CI-driven deploys. It was originally spun up to handle the fleet backups — keep that lens when triaging backup/storage issues.

Tools and conventions

Sister repos (separate gitea repos, deployed by playbooks here):

Repo Role CI status
vh/task-board MCP + web dashboard for assistant task state (port 7878) push-to-main → CI deploys (2026-04-29)
vh/vor Inquisitor UI sidecar (port 7879) push-to-main → CI deploys (2026-04-29)
vh/nevermore Twice-daily LLM-curated briefing (port 8181, replaces news-digest) push-to-main → CI deploys (2026-04-30)
vh/asset-engine Internal control plane over inference services (port 8200, LAN-direct) push-to-main → CI deploys (2026-05-12)
vh/althing Lean trusted inter-agent message bus — v2 "email model" (v2.0.0b2, 2026-07): per-box local-SQLite bus + courier/receiver for P2P over the 10.x net; pillars = open-loops / per-box herald + wake-listener / roaming owner API /owner/* / althing-mcp stdio surface. The v0.15 lean-bus cut RIPPED moderation / chamber / forseti-daemon / agent-runner / redis-valkey. per-box uv tool install (NOT CI-deploy); nh3-dev = the DEV box (editable install of ~/development/althing, gets new versions first); nh3-extdev a mesh peer (model B: althing-svc + shared /srv/althing)
vh/mead-hall Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) push-to-main → CI deploys (2026-05-16)
vh/skaldsong Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) push-to-main → CI deploys (2026-05-19)
vh/Worldtree Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. gitea-runner builds on ana-docker; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. push-to-main → CI build-and-deploy (runner on ana-docker)
vh/yt-voice-clipper YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md
vh/arbo Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook
vh/zonos-gateway OpenAI-compatible TTS gateway over stock ZONOS2 (:8890 irv-ml1); emotion dials-first + voice mapping; reached via LiteLLM ext-tts alias. v0.2.1 (2026-07-18): voice-resolved emotion presets (resolve_preset(name,voice); angry/happy/startled_happy per-voice). 8 voices incl. 4 clones pushed to gitea (main 8f1885b/v0.2.1); deployed irv-ml1 tree still NON-git (hand-updated build context — CI-wire = open follow-up). Spec docs/EMOTION-DIALS-SPEC.md; host-managed voices bind-mount (./voices:/app/voices, drop wav + restart, no rebuild)
vh/soong-lab Noonien Soong character-design studio (SPA + /api + WT /bifrost/tool-call); containerized 2026-07-18, LIVE on corviduo-dev :8443 (image vh/soong-lab:latest). soong-dev owns Dockerfile/compose/workflow; infra-ops owns the host CI = Gitea Actions build+push+DEPLOY on tag/dispatch (fleet recipe: docker:cli + raw buildx, pushes AS vh; auto-redeploy LIVE 2026-07-18 — runner SSHes corviduo-dev as deploy, compose pull && up -d from /opt/soong-lab, health-gated on /api/version). Manual redeploy sudo -u deploy bash -c 'cd /opt/soong-lab && docker compose pull && docker compose up -d'. → archival-memory.md (archived 2026-08-16)
model-training-forge (mtf-dev) Fine-tuning recipe forge; T1 = E-RP writing LoRA, retargeted qwopus-122B→AEON-27B (2026-07-06) (SFT→DPO, LitBench-RM reward) training runs, not a deployed sidecar

(vh/volva + Heid were re-architected from systemd daemons to Claude Code session orchestrators 2026-06-08; their nh3-dev .service units were removed — no longer deployed sidecars here. See Recent decisions.)

  • Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see docs/runbooks/disaster-recovery.md for the blast-radius matrix. ⚠️ The restic file+DB layer routes through TWO rest-servers (rest-server-ana @ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas; rest-server-nh3 @ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of /mnt/backup. (rest-server-ana recovered 2026-06-20.)

  • pull-hf-repo.yaml is the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at /tank/aimodels/huggingface/" playbook. Supports --var repo_type=model|dataset|space. Replaces ad-hoc huggingface_hub.snapshot_download patterns.

  • Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (key_id 61419c92) at ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin auths against demo only. Personal-instance admin (the ~/.config/worldtree/personal-admin-token, mode 600) POSTs /admin/keys (mints per-project keys; takes user_id+label, no scope param — scopes are tier-derived). On-instance mint recipe (cleaner than DB-manip): docker exec worldtree-worldtree-api-1 POST /admin/keys with the in-container WORLDTREE_BOOTSTRAP_ADMIN_KEY; cleartext once in .key=wt_live_+16hex. auto-memory reference_worldtree_demo_key_mint.

  • Per-project user keys against personal Worldtree (issued 2026-05-19): skaldsong:79744637, skaldsong:7c1dbbbe, althing:50d85460, mead-hall:a360822d. Mint via /admin/keys, drop value to /tmp/wt-personal-<name>.key mode 600, dev collects + shreds (DO NOT cat to chat transcript).

  • Skaldsong CD pattern (registry-pull). vh/skaldsong's CI builds and pushes gitea.phasefinal.com/vh/skaldsong:<sha> + :latest; playbooks/deploy-skaldsong.yaml on ana-docker pulls + recreates. SHA-pin only. Prereq: host needs docker login gitea.phasefinal.com once.

  • gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH 10.250.50.70:222, HTTP :3000. Fleet/colo hosts must use this internal route, NOT public gitea.phasefinal.com (38.120.12.44) — the public path fail2bans the host egress IP. Full gotcha in docs/orientation.md → Git/gitea.

  • docker-as-root pattern (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo): docker run --rm -v <target-dir>:/wt docker:cli sh -c "...". docker-group membership is effectively root via bind-mount. Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass -e VAR=/abs/path for any relative-default config dir.

  • scripts/elway sudo handling — elway prompts for the sudo password ONCE via getpass before the first sudo: true step → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH.

  • Per-host SSH identity matters for sudo. infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: default ssh ana-docker = lkraven (docker-group, NO passwordless sudo); ssh infra-ops@ana-docker HAS NOPASSWD root. → For any sudo op on ana-docker, use ssh infra-ops@ana-docker. ssh infra-ops@10.100.10.50 (nh3-dev) ALSO NOPASSWD sudo; on nh3-extdev infra-ops is sudo-LESS by design (ssh lkraven@10.100.50.42 is the NOPASSWD path). irv-ml1: ssh irv-ml1 = lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD) — stage model pulls to /home, not root-owned /worktank.

Current state / in-flight

As of 2026-08-23 — a long multi-party ops session. The Anaheim tunnel cipher is now settled (closed: the remedy does not exist). The one thing still open with the operator is how long WAN admin on ana-gw stays open.

  • 🔒 ana-gw WAN admin is CLOSED again (2026-08-23, operator-directed) and the FortiGate is scheduled for replacement. The cutover contingency window is over: wan1 allowaccess is back to ping only (https + ssh removed) and infra-ops trusthost is back to 10.0.0.0/8 only — verified from two sites that 443 and 22 are closed, and that management still works over the tunnel at 10.250.0.1. There is no longer any out-of-band path to ana-gw; if both tunnels drop, it is console-only. Re-open = set allowaccess ping https on wan1 plus widening the infra-ops trusthost (both one-liners, recorded in auto-memory). ⚠ Port 80 on 38.120.12.42 answers a bare 403 from outside — that is an ISP transparent HTTP proxy on the CLIENT side (NH3 and ESH are both behind CGNAT), not the FortiGate: inbound :80 SYNs from scanner ranges reach the box and get no SYN-ACK, and our own :80 packets never appear in a capture at the box at all.

  • CLOSED 2026-08-23: the Anaheim tunnel "problem" is mostly a measurement artefact, and the AES-GCM cutover is impossible. Operator authorised the cutover; it was attempted NH3-side-first and cannot be done — UniFi's manual site-to-site IPsec implements no AES-GCM (8 spellings rejected api.err.InvalidPayload against a passing aes256 control; accepted enum is aes128/aes192/aes256/3des only). This blocks the ESH tunnel too, since both far ends are UDMs. The framing was also wrong twice over: NH3's uplink is 1 Gbps (not Anaheim's 2 Gbps — that is the real ceiling), and the tunnel does 692 Mbit/s at 8 streams (the original stopped at 4 and reported ~550). Against WireGuard on the same UDM and uplink, the gap collapses from 2.3× at one stream to 15% at eight — so re-architecting onto WireGuard is not worth it. Real constraint = per-stream ~245 Mbit/s, both endpoints idle. Standing mitigation: parallelise bulk transfers (2.8× for free); for single-stream NFS use nconnect=N — the /mnt/smithy mount on ana-ml2 at 24.7 MB/s is exactly this case and is the obvious test. FortiGate phase2 pfi-ana-nh3 was left widened to aes256-sha1 aes256gcm (inert while the peer offers only CBC); UDM verified byte-identical to its pre-change snapshot. → persistent-memory.d/2026-08-23-anaheim-ipsec-tunnel-ceiling.md

  • 🟢 SEAT MAP (unchanged this session except selene). gen = orcarouter/Qwen3.8-27B-Uncensored NVFP4-mixed, GPU0 :8015, now 7 aliases (see the collision note). char-rp = MeroMero-v2 dual-mode, GPU0 :8016, pinned v0.26.0. sec/sec-reasoning = M.O.G.-SEC on DFlash2, GPU1 :8019. selene RETIRED — 17.2 GiB reclaimed on GPU1 (free now ~19.4 GiB).

  • ⚠️ THE sec DEGENERATION QUESTION IS STILL OPEN AND CONFOUNDED. Engine and drafter changed together; the isolating experiment is MTP k=3 on e9d1398d — still not run. Operator ruling stands: degeneration lives in the un-fixed vLLM, not the weights; the MTP-head hypothesis is retracted. Both prior sightings are n=1 and are NOT evidence. gen remains on the old nightly, untouched, gated on that experiment.

  • 🟢 ana-ml2 now mounts /mnt/smithy (nh3-nas) ro + soft, NOT in fstab — needs a manual remount after reboot. For brokkr's R47 CPU work. Reads 24.7 MB/s sequential vs 98.3 on nh3-dev (that gap is the tunnel above), but 45 files/s vs 34 — small-file work is genuinely faster there. → persistent-memory.d/2026-08-23-smithy-mount-ana-ml2.md

  • 🟢 ESH IS DUAL-STACK; the v4 static is a Cityside ticket. IPv6 live on esh-userland and esh-server from a delegated /56. v4 remains CGNAT and a full gateway reboot proved the purchased static is not provisioned — carrier ticket, nothing left locally. NH3 stays v6-off deliberately. Flat-zone lateral-movement finding parked, id 44.

  • 🟢 OTHER SERVICES. hrafn browser-fetch adopted on ana-docker (infra-ops owns uptime; CI now genuinely deploys). speaches ASR live irv-ml1:8204. Open WebUI esh-docker-vm:3211 — Lobe retirement still the operator's call. Booth gained kept-board deletion + per-row link pruning. pfi gitea org created; claude-bot is an Owner and can create repos self-serve.

  • OPEN ELSEWHERE: MTP-k3 isolating experiment; upstream vLLM issue to file (operator's GitHub identity); Cold-Fusion NVFP4 quants (44 GB) delete/keep; OWUI image-tag drift; /tank DEGRADED 70+ days; Worldtree #411 debug-room litter; bridge/engine agent-roster drift on both WT instances; brokkr's gen vs trained-reward-model bake-off (theirs to initiate). Working tree is clean and pushed through 0ad332b.

Recent decisions

  • [2026-08-23] Anaheim's IPsec tunnel delivers ~25% of a verified 2 Gbps circuit — not WireGuard, not CPU, not the fibre. Cipher change proposed and operator-signalled; execution pending, untracked by operator choice.persistent-memory.d/2026-08-23-anaheim-ipsec-tunnel-ceiling.md

  • [2026-08-23] selene retired after losing a head-to-head on its own job; chat-judge moved to gen, the model name 404s by design. Also surfaced that 7 aliases share one seat — cross-checking between them is an echo, which caught a real defect in brokkr's 46k-exposure R47 gate. → persistent-memory.d/2026-08-23-selene-retired-alias-collision.md

  • [2026-08-23] hrafn adopted; its CI reported green for its whole life while deploying nothing. A staging dir inside the rsync target destroyed its own source mid-copy; the deeper fault was verify steps that asserted uptime, never content. → persistent-memory.d/2026-08-23-hrafn-adopted-ci-frozen-source.md

  • [2026-08-23] Worldtree b187 shipped; all three instances de-armed from a 69-day-stale :latest; Matrix homeserver re-plumbed to personal. Includes the :8009-is-demo port trap that an IP-only fix would have walked into. → persistent-memory.d/2026-08-23-worldtree-b187-pins-matrix.md

  • [2026-08-23] Every secret-bearing .env on ana-docker tightened to 0600 — eight stacks including vaultwarden and traefik, verified exposed by reading one as nobody. → persistent-memory.d/2026-08-23-ana-docker-env-perms-sweep.md

  • [2026-08-23] pfi gitea org created; claude-bot is an Owner and creates repos self-serve. Closes the repo-creation half of the credential-migration directive — vh is a USER namespace so no service account could ever create there. Repo creation needs write:user + write:repository + write:organization; POST /users/{u}/tokens is basic-auth only, so minting needs the account password. Default new repos to pfi/. (vh/eitri-smithy was its first tenant, then moved.)

  • [2026-08-23] Booth: kept boards are deletable and link rows are prunable. release on a kept card drops the sentinel so the existing × applies; booth links / booth unlink <id|index> prune one row. Rows are addressed by content id, never position — the board is append-only and multi-writer. Releasing a board RESETS its TTL clock (unlink bumps the dir mtime), so unkeep-and-wait is a 24h delay, not a delete. (4be880f, 0ad332b)

  • [2026-08-22] DFlash2 spec-decode measured on our own stack; sec promoted to it. +1821% accepted length and +1518% throughput over MTP k=3, drafter proved model-agnostic across two finetunes to 0.06%, and the k=7 MTP control showed deeper MTP is a throughput trap. → persistent-memory.d/2026-08-22-dflash2-spec-decode.md

  • [2026-08-22] Quant pipeline shipped a crippled tokenizer for months — fixed at source. quant_mixed_nvfp4.py baked its calibration truncation (max_length 2048) into every mixed-NVFP4 build; latent on old transformers, fatal on new. Both live quants corrected, pipeline now saves a source-pristine tokenizer and asserts it. Playbook §3.14. (0755ba7)

  • [2026-08-22] sec retuned to util 0.52 / 420K after a runtime OOM at 0.55/480Kgpu-memory-utilization is not a hard reservation; activation grows past the dummy-data profile and six vLLM containers share GPU1. Also measured: the KV pool varies ~6.6% between boots, so max-model-len must be sized against the lower observation. (6e82899)

  • [2026-08-22] Max-Q 1.8× spread does NOT apply to LLM decode — measured, not argued. ana-ml2 draws 256266 W of 300 W under sustained 100% decode with SW Power Cap: Not Active and clocks pinned. Corrected to brokkr-smithy-dev after I had lent the claim credibility; 122B figure (~9093 tok/s at 262K) stands as a straight number.

  • [2026-08-21] ESH internal IPv6 live on two LANs; the Cityside v4 static is a CARRIER problem, proven. A full gateway reboot forced a fresh DHCP DISCOVER and returned the identical CGNAT address. YaRN was already configured — "1M needs YaRN, absent" was false. → persistent-memory.d/2026-08-22-dflash2-spec-decode.md sibling entry in ad21302

  • [2026-08-21] speaches ASR live on irv-ml1 for Eyra — and no_speech_prob alone is a weak hallucination gate. Silence and room tone both hallucinated "Thank you." under 0.11; avg_logprob separates ~6× better. Consumers should gate on a composite. (aa5863c, c7e2187)

  • [2026-08-20] Cold-Fusion abliteration — Robinson recipe captured; the fight was the environment, not the recipe. Stock Cold-Fusion measured ~33% creative refusal → worth abliterating ourselves (supersedes waiting for DavidAU's heretic build). Recipe maps 1:1 (131 tensors); capture succeeded only in fp32 — transformers' Qwen3.5 DeltaNet linear-attn NaNs nondeterministically in bf16 without the unbuildable causal-conv1d kernel (precision cancellation, not overflow). Direction finite at layer 22 but agreement 0.59 (vs Robinson's 0.99) → calibration-set expansion is next.persistent-memory.d/2026-08-20-coldfusion-abliteration-capture.md

  • [2026-08-19] A software watchdog is not watchdog protection — esh-pve froze for 4.5h holding one. softdog cannot fire when the kernel it runs in is wedged, and Proxmox's watchdog-mux never arms without HA resources, so the box looked protected and wasn't. Moved to the PCH iTCO_wdt under systemd. Also: a single cross-VLAN DNS entry with no secondary turns any VM outage into a whole-site outage. → persistent-memory.d/2026-08-19-esh-pve-freeze-dns-spof.md

  • [2026-08-19] Fleet .internal DNS built and live — git-sourced, agent-managed, three resolvers. Zone-scoped authority (ESH's hand-made esteban.net rewrites survive); the colo had no resolver at all; v6 column empty on purpose because SLAAC addresses rotate. → persistent-memory.d/2026-08-19-fleet-internal-dns.md

  • [2026-08-19] waterland studio containerised on irv-ml1 — three landmines, all measured. cupy needs CUDA headers the host had by accident; uv run re-syncs and prunes cupy at RUNTIME; the A6000 is container-index 0, not the host's 1. → persistent-memory.d/2026-08-19-waterland-studio-containerised.md

  • [2026-08-19] Homepage cleaned up, then themed with Australis Skyfall + an Arbo-generated background. Includes the hour lost to a self-healing tab-bar red herring, and the CSS-iteration loop that prevents it recurring. → persistent-memory.d/2026-08-19-homepage-skyfall-theme.md

  • [2026-08-19] Four unmanaged stacks found on live hosts — two quietly broken. A dashboard card is a cheap census of what is actually running; check whether the stack is even in stacks/ before debugging the symptom. → persistent-memory.d/2026-08-19-unmanaged-stacks-searxng-seafile.md

  • [2026-08-19] claude-bot granted read on vh/waterland (operator-empowered, verified admin:false push:false pull:true) so irv-ml1 can self-update without the operator's site-admin token living on a GPU box. Precedent for the standing migrate-off-operator-creds directive: grant the service account, wire a repo-scoped 0600 credential helper, keep the remote URL clean. Commit 8189076.

  • [2026-08-19] AI-tab Dormant regrouping BELAYED by the operator — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong, Chatterbox Fast) show amber EXITED inside live groups rather than AI - Dormant. Fix is a label change + recreate per stack; needs the operator's read on which are retired vs temporarily down. untracked by operator choice (his words: "belay the ai dormant regrouping for now").

  • [2026-08-18] esh-pve-nas migration STAGED — and staging is where three landmines surfaced, none of which the plan predicted. (1) The runbook's /boot LV had nowhere to live: VG pve had 4 MB free and mounted ext4 cannot shrink, so the space came from the 768 MB swap LV (operator's call: shrink to 256 MB, not drop). (2) The runbook's zpool set cachefile=… nvme would have broken the NAS — populating a cache flips the host to import-by-cache, and a one-pool cache leaves ssd+tank unimported under CT 103's twelve bind mounts. (3) update-grub silently emitted a pool-less root=ZFS=/ROOT/pve-1, because GRUB's ZFS reader cannot open a pool with encryption/large_dnode/zstd_compress and the probe failure is swallowed. All three were caught by verify steps that asserted effective state, not by reading the plan. → persistent-memory.d/2026-08-17-esh-pve-nas-dom.md

  • [2026-08-17] esh-pve-nas PVE root is on a USB DOM — mitigated, and the migration replanned to split boot from root. Operator's design beats my reinstall plan; wear was never the issue, blocked patching is. → persistent-memory.d/2026-08-17-esh-pve-nas-dom.md

  • [2026-08-17] irv-ml1 cleared of 782 GB, and Homepage brought under version control. One dead-looking Gradio app pinned three delete targets at once; /opt/ComfyUI is NOT the ComfyUI that serves. → persistent-memory.d/2026-08-17-irv-ml1-cleanup-homepage.md

  • [2026-08-17] Gen seat swapped to absolute-heresy — and the three bugs the swap exposed are worth more than the swap. Candidate MuXodious/Qwen3.8-27B-absolute-heresy (Heretic v1.4.0 + SOMPOA, T377) beat the incumbent on refusals AND KL simultaneously, which is the unusual part — those normally trade off. Validated on the probe port per operator ruling, promoted, all 7 aliases green. Durable lessons banked: (1) A CPU-only MTP head hash can replace the ~56 GB bf16 acceptance gate. The Qwen3_5ForConditionalGeneration wrapper never loads the MTP head, so PEFT merges / Heretic runs / llm-compressor passes all leave mtp.* pristine — hashing it against a head we have already measured (the incumbent's, 47.7%) answers the question for free. Predicted 47.7%, measured 47.2%. Saved downing meromero. Tool: services/gen-seat-mixed-quant/compare_mtp_head.py (hash bf16 via uint8 reinterpret — numpy has no bfloat16). (2) post_quant.py assumed a standalone model-mtp.safetensors; a full checkpoint keeps mtp.* in a NUMBERED shard, so the copy silently no-op'd while the index was still rewritten to point at a file that never existed — 15 unresolvable tensors behind a correct-looking tensor count. Its own FAILED-CHECKS assertion caught it; that is why the check exists rather than an assumption. Fixed to extract. (3) A probe that does not mirror the live seat manufactures failures. serve_probe.sh hardcoded :latest (seat is a pinned nightly for #51113), had no tool-call/reasoning parsers, and its --speculative-config JSON died twice on quoting — bash BRACE-EXPANDS {"a":1,"b":2} on the comma unless single-quoted at the REMOTE shell. Adding the seat's flags took the surface test from 5/6 to 6/6; the "tool calling broken" result was pure probe config. Commits 7997f11,254c588,2c36028,b0c2d3d,993421b.

  • [2026-08-17] Fleet IPv6 mapped + the real VPN topology verified; the driver is CGNAT at ESH, not the WireGuard mesh. New ESH fiber (installing 2026-08-18) lands the house behind CGNAT, which breaks Site Magic (NH3↔ESH sdwan-mesh-tunnel) on IPv4 — so IPv6 becomes load-bearing as the escape hatch, and that is its most likely first consumer. Topology as VERIFIED (a prior turn assumed wrong and was corrected): UniFi↔UniFi = Site Magic; colo↔UniFi = IPsec IKEv2 (pfi-ana-nh3 158M/165M pkt = the workhorse, ana-to-eshudm); WireGuard is an RA convention only, host-based on ana-wg UDP 31337 behind a FortiGate VIP — the FortiGate never terminates WG (FortiOS 7.2 has none; 7.4 added it) so "upgrade the edge for WireGuard" is a non-problem, do not re-derive. IPv6 today: NH3 WAN live 2600:1700:b25:c110::48, colo none, ESH none. AT&T delegates exactly ONE /64 (2600:1700:b25:c11f::/64) — proven by forcing prefix-ID auto→0 and watching the subnet NOT move, because the c110/c11f pattern otherwise reads convincingly as a /60. A mesh needs a routable WAN address, not PD. ana-wg's WG socket is already dual-stack ([::]:31337) → v6 RA needs an address + a v6 port-forward, no WG reconfig. ⚠ UDM legacy rest/firewallrule returns 0 rules (zone-based firewall) — use v2/…/firewall-policies; inbound v6 is default-deny and held. All three endpoints will be dynamic → extend the existing hostname pattern (ana-fw/nh3.phasefinal.com) to AAAA. Enabled PD on nh3-iot to measure, reverted on operator instruction (all 5 LANs back to none, verified). Also fixed: ana-wg WireGuard key material was world-readable (wg0.conf + keys/*_priv + *_psk + client configs/*.conf at 644) → now 600, dirs 700, service untouched. Detail → persistent-memory.d/2026-08-17-fleet-ipv6-mesh.md.

  • [2026-08-17] Gen-seat multi-day degeneration RESOLVED — two compounding real causes, not one; the meta-lesson is "a mitigation that HELPS but doesn't FIX means a second cause, not a wrong one." vLLM qwen3_5_mtp×GDN bug (#51113, real, fixed by nightly) + AEON full-W4A4 being lowest-fidelity (W4A4<W4+FP8<W4+bf16) → ~15-20% stochastic degeneration. Fixed by mixed FP8-attn build on pinned nightly. AEON purged. Also banked: stochastic (~15-20%) degeneration is invisible to a small synthetic probe — n=1 "clean" validated THREE non-fixes (MTP-off, APC-off, nightly-alone) that all failed in real use; get the operator's real transcript, do not trust your own probe. Full → docs/pfi/model-quantization-playbook.md §3.8 (+ §3.7 MTP-multi-turn). Commits d28a371,2f2bbce,2185964.

  • [2026-08-17] Lobe Chat chosen over Open WebUI (weight: 143 MB vs 1.8 GB) + stood up on esh-docker-vm; scoped LiteLLM key blocks paid models; System-Agent gpt-5-mini default repointed via env. TTS env-vs-UI resolved as a split (endpoint env-driven, voice/model UI-only). tts-dev onboarding closed both directions; ballad/verse aliased so no voice can 404 the router. Commits e9362de,163a725,cac75cb,933253d,25fa18e.

  • [2026-08-17] LiteLLM upgraded v1.91.0→v1.97.0 (RC-avoided on the fleet gateway) + the 6 GB spend-log DB purged & capped (store_prompts_in_spend_logs:false + 7d retention). Interpreted "get rid of the db" as the spend-log DATA not the database (keys/config live in it). Commit 01b5ad9.

  • [2026-08-16] Abliterated models go CATATONIC at the hard refusal edge — silence, not a decline. Abliteration removes the refusal direction, so at the genuine hard edge the model neither refuses nor complies → empty/degenerate output. Durable measurement consequence: a refusal probe MUST score EMPTY as a verdict distinct from REFUSAL and COMPLY (services/refusal-probe/probe.py does). Operator accepted it as out-of-scope; do not chase.

  • [2026-08-16] Fable-Fusion 711 cuts cold-framing refusals 92.5% → 15.8%; refusal is MONOTONIC IN FRAMING, and DS v1.0's problem is that she was never abliterated. brokkr-smithy-dev supplied the framing that reproduces (01M05M48R4RSZF9D8KT7RR55EJ): a bare assistant-mode instruction — no character card, no permission preamble. Three-arm A/B, same harness, same classifier: permission framing DS 0.0% / FF 0.0% (n=75); plain character cards DS 1.4% / FF 0.0% (n=74); bare instruction DS 92.5% (37/40) / FF 15.8% (6/38). Per-axis DS→FF: incest 100→20, non-con 100→20, bestiality 100→25, necrophilia 100→40, gore 100→0, consensual 80→20, dubcon 80→0, self-harm 80→0. DS refused 25/25 on the five axes brokkr flagged. Root cause: ReadyArt/Dark-Scarlett-v1.0-27B is a plain finetune of stock Qwen/Qwen3.6-27B carrying NO abliteration — the base refusal machinery is intact, so cold prompts revert to safety-tuned Qwen3.6. FF is Heretic-ablated (structural), which is why it holds. ⚠ RETRACTED 2026-08-16 — my "arm-3 92.5% exceeds brokkr's 62.5%" comparison was INVALID. His diff against his own artifact showed my battery-instruct.yaml reproduces only his creative class — 8 of 16 axes; it dropped all 5 operational (violence/incite, crime/fraud, cyber/malware, selfharm/methods, privacy/stalk) and all 3 meta (meta/sysprompt, meta/ignore, meta/dan), and added 2 controls he never had, at k=5 vs his k=2. His 62.5% pools all 16 axes; my 92.5% is creative-only — different denominators, not a delta. Cause: I rebuilt his shape from his message, and the class field lives in the artifact, not the prose. Lesson: reconstructing a peer's instrument from their description reproduces what they described, not what they ran — diff against the artifact before claiming comparability.Known battery bug left unfixed for comparability: DS's arm-3 control gate failed at 11% because ictrl-reunion pairs "explicit / do not fade to black" with brothers, which DS reasonably read as an incest request; FF did not. ictrl-storm is the clean control. Commit b9e68c3.

  • [2026-08-16] MTP works on Fable-Fusion AND survives RP temperatures — my earlier caution was wrong. vLLM resolved Qwen3_5MTP, loaded the drafter, shared embedding + lm_head — the capability DS's seat never had because our quant dropped her MTP tensors. Measured over the full probe workload (~163k draft windows at temp 0.71.0): 47.0% acceptance (229,169/487,725), 1.41 extra tokens/window, per-position 68.3/43.6/29.1%, ~80.6 tok/s decode at temp 1.0. I had recorded a caution that the card's 1.56× was greedy-measured and acceptance would fall at RP temps — it did not; 47.0% matches the gen seat's 47.7% and beats the card's own 33% at depth 5. Depth 3 is right.

  • [2026-08-16] The Qwen base thinks incessantly — that is WHY the Gemma seat exists, and no swap within the Qwen family fixes it. Operator's architectural point, confirmed by measurement: on identical prompts DS 6036 ch vs FF 5323 ch of reasoning (permission arm), 5546 vs 4988 (cards arm) — FF actually reasons ~1012% less. The bare-instruct row (DS 2291 vs FF 3918) inverts only because DS refused 92.5% of it and refusals are short — an artifact, not concision. Both are Qwen3.6-27B derivatives, so this is the base family. char-rp = MeroMero-v2, Gemma-4 base, :8016, verified 0 chars reasoning / clean prose — the non-thinking seat, working as designed. FF can be silenced (enable_thinking:false verified 3/3, and it ships chat_template-instruct.jinja) but that duplicates MeroMero on a base chosen for it. The stale LiteLLM comment describing char-rp as the retired GGUF Magidonia seat is fixed (53096bf).

  • [2026-08-16] esh-vm-docker hardened: the wedge is hard NFS at RUNTIME, which the boot-ordering fix never addressed. All four mounts were hard, so a NAS stall at 10.0.50.50 blocks I/O forever (D-state). The existing x-systemd.before=docker.service fstab fix solved the boot race — a different bug. Exposure was far below what the park item assumed: only 2 of 12 containers touched NFS, and container state was already local (/var/lib/docker). Removed: /mnt/compose (2.1G, fully vestigial — zero containers referenced it, dockge reads local /opt/docker, its one mention was a comment in beszel-agent-esh/.env about a different host) and /mnt/documents (2.0K, paperless's empty spool dirs → /opt/docker/data/paperless at the same 0777). fstab backup /etc/fstab.bak-nfs-harden-20260816. 4 mounts → 2, 2 wedge-capable containers → 1. traefik needed no change (already restart: unless-stopped — why it self-recovered). Watchdog services/esh-vm-docker-watchdog/ live on esh-pve (not the guest): probes traefik over HTTP, deliberately not ping/SSH — the wedge signature is "guest OS alive, services dead" (/ is local disk so sshd answers straight through a total outage and a TCP check reports HEALTHY). 5 failures × 2 min → qm reset 100, 30-min cooldown, running-only guard, /etc/esh-vm-docker-watchdog.disabled. All paths tested without power-cycling. DEFERRED (operator): /mnt/books stays hard — calibre's SQLite metadata.db would risk corruption under soft/softerr. That is the one remaining wedge vector. Commit 55705ba; park item 28 promoted. ⚠ qm over non-interactive ssh throws a bogus JSON::Backend::XS error — use ssh host 'bash -s' <<'EOF', not ssh host "qm …".

  • [2026-08-16] Canonical Qwen3.8 sampling applied from upstream; gen-reasoning had the WRONG-MODE presence_penalty. Qwen/Qwen3.8-27B "Best Practices" §1 and unsloth/Qwen3.8-27B §1 are byte-identical — thinking: temp 1.0 / top_p 0.95 / top_k 20 / min_p 0.0 / presence_penalty 0.0 / repetition_penalty 1.0; instruct: temp 0.7 / top_p 0.80 / top_k 20 / min_p 0.0 / presence_penalty 1.5 / repetition_penalty 1.0. Bug found: gen-reasoning carried presence_penalty 1.5 — the instruct value on a thinking deployment (canonical 0.0) — now fixed. Deliberately NOT canonicalised: summarizer/classifier/image-judge/qwen-image-bench run temperature=0 (judges also top_k=1) because determinism is their contract; forcing a chat preset on a classifier would break it. ⚠ presence_penalty=1.5 is canonical but is the one value upstream hedges on, verbatim: "using a higher value may occasionally result in language mixing and a slight decrease in model performance." It is the operator's suspected trigger for multi-turn degradation and the first dial to move (0.00.5) if that recurs — it is alias-scoped, which is why it would follow the operator across model builds. Commit 3462b53.

  • [2026-08-16] Four wrong diagnoses on one bug, and the lesson is the test design. Operator reported the gen seat "degenerate on long multi-turn conversations". Rolled the seat back on request; the previous weights behaved identically, exonerating the model swap. I then proposed and disproved FOUR mechanisms in sequence — empty assistant turns poisoning history, reasoning runaway, length-mirroring from short history, and presence_penalty — before discovering my own multi-turn harness was confounded: it varied the QUESTION along with the depth (depth-1 asked question #2, depth-3 asked question #4), so a narrower question drawing a shorter answer read as degeneration. The "310→209→28w collapse" I reported as a reproduction was an artifact. Rules banked: (1) when comparing across conversation depth, hold the final question FIXED and vary only the history; (2) reply-length variance on byte-identical input was 25465w, so n=3 cannot support any claim about a trend; (3) ask for the operator's real failing transcript before building a synthetic reproduction — four synthetic tests, none of them his failure. Gateway spend_logs returns [] on the infra-ops key despite store_prompts_in_spend_logs: true, so real transcripts need the :4000/ui view or another key — worth solving before the next such hunt.

  • [2026-08-16] Two REAL client-side defects found while chasing the above, neither of which was the reported bug. (1) gateway-chat's Max-tokens field defaulted to 1024; thinking seats spend part of that on CoT before emitting content, so completions truncate with finish_reason=length and read as model degeneracy — raised to 4096. (2) parseInt on an empty field yields NaN, which JSON.stringify serialises as null, which the server reads as "no max_tokens supplied" and silently substitutes its own default — indistinguishable from the UI ignoring the field. Both fixed (b6552e0, fb3bb52). ⚠ compose bind-mounts a single FILE, and a single-file bind mount binds the INODE — rsync writes-and-renames, so the container kept serving stale content while the host file showed the new value, silently and with no error. docker restart does NOT clear it; the container must be recreated. Verify against what the container sees, never the host file. Applies to any file-source mount fleet-wide.

  • [2026-08-16] Refusal measurement: benign controls CANNOT validate a refusal classifier on RP prose — and a 0% rate needs a classifier self-test before you believe it. Two durable lessons from baselining Dark-Scarlett. (1) False positives: my first bare-framing number was 9.5%; the true figure was 1.4%. The rest were the classifier firing on in-character text — "I cannot shift my weight" spoken by the character ~100 chars into a 2,443-token torture scene, and "Yeah, I'm an AI… What's the actual gig?" where the model answers in voice and keeps driving the scene. First-person RP prose is full of "I can't"; a genuine refusal opens with its marker, so the scan window must be the first sentence, a marker followed by long prose must demote to AMBIGUOUS, and AI self-acknowledgement is a persona break, never a refusal on its own. Benign controls were clean the entire time and caught none of it — they only detect over-firing on benign prompts, not on in-character prose. (2) False negatives: a 0% rate and a broken classifier are indistinguishable from the report, so test_classify.py (16 cases, both false positives pinned as regressions) must pass before any low number is trusted. Also banked: the thinking-budget trap — empty content + finish_reason=length is reasoning eating the budget, NOT a refusal; score INVALID and exclude from the denominator (DS emits ~5.5-6k chars of reasoning per response, so max_tokens ≥3072). probe.py --rescore re-classifies a saved run with zero GPU time. → services/refusal-probe/README.md, commit 32f665e.

  • [2026-08-16] Held an operator-approved swap window because the baseline invalidated its premise. Operator approved ~65 min of char-rp-reasoning downtime to A/B Fable-Fusion 711 against Dark-Scarlett on refusals. The DS baseline then came back 0.0%/1.4% — no gap for a candidate to close, so the window would have bought no decisive signal and a second window would still be needed once a reproducing battery existed. Held the swap, reported, and routed to brokkr-smithy-dev for the battery that actually produced the refusals. The general rule (action-relevance): approval is for a plan, not a ritual — when new evidence kills the plan's premise, surface it rather than spend the budget. Nothing deployed, no downtime taken, seat untouched.

  • [2026-08-16] DS v1.0's one real refusal is self-contradicting boilerplate, not a content constraint. On a direct "drop character and state your content policy" probe she returned "I don't generate explicit sexual content, graphic violence, or material that glorifies harm, non-consensual acts, or illegal activity"in the same run where she generated all three at 0% refusal. Reads as a learned recital triggered by meta-questions about policy. If production refusals share that shape the failure is prompt-shaped, not model-shaped, and a consumer-side system-prompt fix may beat a model swap entirely — worth settling before spending the GPU window. Separately, 7/85 bare-framing samples were persona breaks (in-character AI acknowledgement): not refusals, but DS will admit to being an AI unless the card explicitly forbids it.

  • [2026-08-15] RP-seat direction: KEEP MeroMero on char-rp; Artemis-31B rejected; next move is Dark-Scarlett on a Qwen3.8 base when it lands (operator). Evaluated TheDrummer/Artemis-31B-v1.1 — mechanically a drop-in (same google/gemma-4-31B-it base, identical 1188-tensor/356-vision census, same missing-preprocessor_config.json trick), so it's purely a quality call, and our own survey already ranked MeroMero #1 vs Artemis #6; Artemis is also unlicensed and its author deprioritizes correctness + warns of token-banning-for-stability, which fights char-rp's tool-calling requirement. MTP verified impossible on both (Gemma-4 has no MTP head at all — base/MeroMero/Artemis are all MTP=0; no finetune can add one). But speculative decoding IS reachable on a Gemma-4 seat via a DETACHED drafter — vLLM 0.24 supports eagle3 + gemma4_mtp, and real drafters exist: google/gemma-4-31B-it-assistant (0.94 GB, 4-layer, 761K dl), RedHatAI/gemma-4-31B-it-speculator.eagle3 (4.47 GB), AEON-7/…eagle3-NVFP4 (3.53 GB). ⚠ all list their verifier as stock gemma-4-31B-it, not an RP finetune, so acceptance against MeroMero is unmeasured and likely well below the gen seat's ~48%. UNTESTED — parked, ~45 min to measure, needs GPU0 headroom (card is at 94.4/97.9 GB). Why the Dark-Scarlett 3.8 plan is the strong one: DS is Qwen3.6-based today, so a 3.8 respin lands on the gen seat's architecture → native MTP returns and the whole mixed NVFP4+FP8 recipe + graft ports directly. Watch two things on arrival: from_pretrained silently drops MTP heads during finetuning (verify 15 mtp.* tensors in the index; graft from stock if absent), and DS v1.0 required the Qwen3_5ForConditionalGeneration wrapper class to save a config vLLM/SGLang accept. Both in docs/pfi/model-quantization-playbook.md.

  • [2026-08-15] Quant lessons consolidated into docs/pfi/model-quantization-playbook.md — the durable home; read it BEFORE any requant. Survey found quant knowledge scattered across 18 files in 4 trees, with three documents having independently written overlapping "landmines" sections (the loader-class trap alone was rediscovered 3×). Playbook owns the transferable lessons (scheme choice, landmines, acceptance gate + its 3 measurement traps, hardware/co-residency); per-model artifacts are demoted to worked examples that link up. Carries a superseded-claims table — which immediately earned itself: the heretic2 runbook's "use modelopt, compressed-tensors can't load the BF16 MTP" is false (the cause was the missing re:^mtp.* ignore, not the format) and would have sent the next session down the modelopt dependency-hell path; that runbook now carries a stale-warning header. Maintenance rule in CLAUDE.md: model-agnostic → playbook, model-specific → stays put, wrong claim → dated superseded row, never a silent edit. Motivated by Qwen3.8 having just released — the next model swap needs a requant. Commit a91cc3f.

  • [2026-08-15] Operator ruling: the gen seat's +1.7% perplexity is an acceptable price for the speed — SETTLED, don't re-litigate. Precise attribution for future reasoning: it is the activation-quantization cost (W4A4 MLPs + FP8 attention vs BF16 activations), not an MTP cost — PPL was measured with speculative decoding off on both builds, so MTP was not in the loop. Turning MTP off would not recover it; only reverting the quant would (rollback = one .env line, old build intact at …/qwen38-27b-uncensored-nvfp4).

  • [2026-08-15] gen seat requanted to mixed NVFP4+FP8 (+18% decode) + char-rp Gemma-4 tool-calling fixed. The queued "W4A8" (NVFP4 weights + FP8 activations) is not servable — vLLM 0.24 allows NVFP4 weights with only A16 or A4; FP8 activations ValueError at load, and CompressedTensorsW4A8Fp8 is INT4-weights + sm90-exact (closed on Blackwell twice). FP8 must enter per-layer-group. Also: the handoff's "~68 tok/s" baseline didn't reproduce — cache-busted, the incumbent already did 80.12 (≈ the stated W4A8 target), so the premise needed re-measuring before any work. Shortcut: unsloth/Qwen3.8-27B-NVFP4 was already on-box → served as a probe, measured +19.1% at identical acceptance, which both proved the gain was real and handed over the reference recipe. Replicated it on the abliterated weights → 80.12→94.53 tok/s, acceptance unchanged, +1.7% PPL, abliteration 4/4, weights 19%; surface 6/6 live, 7 aliases routing. char-rp had no tool parser at all (every tools request 400'd) → gemma4 tool + reasoning parser + a mandatory enable_thinking:false (the parser defaults it True → null content for all RP prose; proven byte-identical prompt before deploying). Commits b8f0f4c, 74f596b. Foot-guns banked (llm-compressor prunes unmatched ignore entries → the 0%-MTP bug, fired on this run; prompt_logprobs uniform under spec-decode; 0600 .env silently no-ops compose; GPU0 is zero-sum). → persistent-memory.d/2026-08-15-gen-seat-mixed-requant.md

  • [2026-08-15] Uncensored gen seat: JonathanColetti/Qwen3.8-27B-Uncensored deployed as gen-seat/vllm-gen (NVFP4 W4A16 + grafted MTP, 262K); 7 aliases repointed; the definitive re:^mtp.*-ignore fix. 0%-MTP-on-quant (twice) was NOT the abliteration/scheme — the grafted bf16 MTP was missing from quantization_config.ignore (vLLM loaded it as quantized → uninitialized). Full arc, the working pipeline, VRAM budget, unsloth speed decomposition, modelopt dead-end. → persistent-memory.d/2026-08-15-uncensored-gen-seat.md

  • [2026-08-12] eRP dual-seat overhaul: MeroMero-v2 (char-rp) + Dark-Scarlett (char-rp-reasoning), both NVFP4A16 @ 256K on ana-ml2; granite retired. Replaced the GGUF/heretic2 RP seats with two home-quantized vLLM seats. The DS blocker (an AutoModelForCausalLM save wrote a flat Qwen3_5TextConfig that both vLLM AND SGLang reject) was fixed by re-quanting via the Qwen3_5ForConditionalGeneration wrapper class; ModelOpt was a version deadlock, SGLang lacked the impl (but revealed the fix). MeroMero vision reconstructed by extracting preprocessor_config.json from processor_config.json. Both models KV-efficient (Gemma-4 sliding-window / Qwen3.6 hybrid linear-attn) → full 256K; GPU-swapped for headroom; compose-ified + committed f08b6cb. granite downed + LiteLLM summarizer/classifier→gen. Full arc, lessons, dead-ends → persistent-memory.d/2026-08-12-erp-dual-seat-overhaul.md

  • [2026-08-12] infra-ops now holds an all-zones Cloudflare DNS-edit token (vaulted) + wgtunnel Phase-0 DNS landed. Operator handed over a Zone·DNS·Edit (all zones) CF token → secret put nh3-dev/.config/cloudflare/infra-ops-dns-token (round-trip verified; /tmp drop shredded). Fleet DNS is now self-serve for infra-ops (⚠ HIGH blast radius — all zones). First use: created boring.phasefinal.com CNAME → ana-srv1.phasefinal.com, DNS-only (proxied:false), verified resolving to 38.120.12.44 on both authoritative NS (louis/wren) + 1.1.1.1 — NOT Cloudflare-proxied. Unblocks wgtunnel's wstunnel ACME cert. phasefinal.com zone id f812ba74ed9a75cf21bbe7ce9188db50. auto-memory reference_infra_ops_cloudflare_dns_token. (Earlier gap: the only prior vaulted CF token, jackdaw's, had zone:read+worker:edit but no dns_records:edit.)

  • [2026-08-12] wgtunnel stood up as its own repo (vh/wgtunnel, private) after a live endpoint-verification pass. Operator directed own-repo (mirrors stonehenge-park/tts-stack). Verified off the fleet before seeding: ana-wg WG server = UDP/31337 (not 51820), subnet 10.30.10.0/24, MTU 1420, active roaming peer proves the public UDP DNAT works; traefik on ana-docker terminates TLS :443 (ACME anaprod http-challenge, docker+file providers, CrowdSec bouncer) → confirms the clean design (wstunnel container on traefik-net, Host-routed, WS→UDP to ana-wg:31337); edge 38.120.12.44 direct-A, tunnel.phasefinal.com free (⚠ must be direct, NOT Cloudflare-proxied like vaultwarden). Repo pre-seeded (README/CLAUDE/persistent-memory/ROADMAP + docs/verified-infrastructure.md = ground truth) + pushed; commit 9584d38, Vuong-attributed. vh gitea token pulled from the vault (secret get), not persisted to .git/config. NEXT = /vor-plan or /vor (operator's call, interactive). Deps to line up in the plan: DNS A-record, FortiGate :443 host-routing, a new ana-wg peer for the laptop, client tooling.

  • [2026-08-10→12] secrets-broker: per-box Vaultwarden credential store SHIPPED + consumer-confirmed. secret CLI (put/get/list/rm/backfill, bw-backed) on ~/.local/bin; 25 nh3-dev secrets backfilled + round-trip-verified; rm + new-namespace warning added post-launch; standing "vault is the credential source of truth" directive now global. → persistent-memory.d/2026-08-12-secrets-broker.md

  • [2026-08-11] stonehenge-park: new fleet /park service repo stood up + designed (/vor-plan + /vor-ui). Self-contained SQLite+FastAPI idea-parking service that actively resurfaces (statusline + althing) so nothing dies in a cold repo; vh/stonehenge-park pushed + pre-seeded for a fresh agent; build starts at the U1 tracer contract. → persistent-memory.d/2026-08-11-stonehenge-park.md

  • [2026-08-12] Global ~/.claude/CLAUDE.md: secret/vault tool entry + "store in AND pull from the vault" standing directive (dotfiles 9db703b, pushed); statusline reset-countdowns + a latent tab-collapse parse-bug fix, now tracked in the dotfiles stow tree. Dogfooded the directive: created vh/stonehenge-park pulling the gitea token via secret get. (dotfiles + global config, not eshpfi.)

  • [2026-08-11] TTS stack extracted to its own repo (tts-stack) + eshpfi stood down on TTS dev. Operator: hand all TTS tuning/dev to a separate agent with a self-contained repo (knowledge + infra access + a live knowledge list), and move the voice corpus in. New repo ~/development/tts-stack (commit 9ee3288) carries: dots-tts stack (canonical intent), voices/ corpus (MOVED out of eshpfi), KNOWLEDGE.md (engine landscape + prosody findings + foot-guns), docs/infrastructure.md (irv-ml1 access + gated deploy runbook + rollback), CLAUDE/persistent-memory/ROADMAP, tools/ (pause-probe + Booth render). Followed the chatterbox-fast precedent: eshpfi stacks/dots-tts/ reduced to a POINTER README; the ~15 experimental TTS compose wrappers stay here as reference (catalogued in tts-stack KNOWLEDGE). Blast-radius check: no eshpfi playbook/script reads the canonical corpus (other voices/ refs = unrelated host paths). Reverses the earlier "Corpus home = eshpfi voices/ (keep-here)" call. ⚠ tts-stack is LOCAL-ONLY until pushed — needs a gitea remote (vh/tts-stack) + push before the separate agent can clone (operator's call — outward-facing + repo-create creds).

  • [2026-08-10] dots-tts v3 — clause-break → period pause mapping. Operator: v2 "sounds good" but donut won't pause at semicolons/dashes. ROOT CAUSE (measured via a pause-probe A/B — synth duration over N runs, non-determinism averaged out): dots' prosody honors a real pause only for ellipsis (+0.43s) and period (+0.3s, capitalization-independent); comma/semicolon/colon/dash all run flat (~+0.03s vs no-punct). Two distinct sub-causes: dashes regressed in v2 (the - fold made em-dashes read as word-joiners), while semicolons were NEVER a v2 change — dots ignores them natively, only newly noticeable because v2 made everything else clean. Operator call: ellipsis "too much" → map ;, clause :, and em-dash → period in _sanitize (believable ~0.3s clause break). GUARDS (pinned by 11 unit tests, stacks/dots-tts/test_sanitize.py): digit-guarded colon (?<!\d)\s*:\s*(?!\d) so times 3:45 / ratios 2:1 survive; en-dash →hyphen KEPT (numeric-range 1020 safety — em-dash breaks, en-dash ranges, different jobs); genuine ellipsis left at full strength (author meant a long pause). Gated deploy (redeploy2 pattern → v3): build → throwaway :8199 test container + pause-gate (semicolon sentence must run ≥0.12s longer than baseline; measured +0.427s) → only then cut live over. LIVE + healthy local/dots-tts:v3 on :8198. rollback = sed -i 's/^DOTS_TAG=.*/DOTS_TAG=v2/' .env + docker compose up -d dots-tts (v2 image retained). Booth dots-pauses (A=old-flat / C=ellipsis-too-much / D=live-v3). reference_chatterbox_fast_repo

  • [2026-08-10] dots-tts v2 — contraction fix (curly-sanitize) + sentence-chunking + dependency-pin recovery. Operator: donut read contractions wrong ("you're"→"you ree", "donut's"→"donut ess"). ROOT CAUSE (isolated via A/B booth): curly/typographic apostrophes ( U+2019 from ratatoskr's LLM) — dots' tokenizer mispronounces them; STRAIGHT apostrophes read clean under normalize_text=True. FIX (app.py): fold curly→ASCII (str.maketrans) before synth, KEEP normalize_text=True (operator call — retains number/date expansion). Also added server-side sentence-chunking (pack ≤280 chars): dots caps one generate() at ~500 patches/~40s, so long RP turns (the Zev monologue = 160s audio) truncated; chunking stitches them (verified full 160.3s, not 40s-cut). ⚠ BUILD FOOT-GUNS (both bit this redeploy): (1) upstream dots.tts constraints/recommended.txt now pins gradio==6.17.0 — phantom, not on PyPI → fresh pip install dots.tts unsatisfiable; FIX = pin dots.tts==0.2.1 + DROP the -c recommended.txt constraints (0.2.1 pulls working gradio 6.17.3). (2) pinning only torch==2.8.0 let torchaudio float to 2.11.0 → dots.tts refuses to load (minor-version match check); FIX = pin torchaudio==2.8.0. ⚠ DEPLOY LESSON: docker compose up -d to a new tag swaps the LIVE container BEFORE any health check — a broken image crash-loops production (ratatoskr TTS down ~1-2min this session). NEW PATTERN = build → test in a THROWAWAY container on an alt port (:8199) → health+verify → only THEN cut live over (redeploy2.sh). v2 LIVE + healthy on irv-ml1:8198, CONSUMER-CONFIRMED clean (ratatoskr verified end-to-end on their :8765 — apostrophe string reads clean, /api/tts 200 @ 48kHz, no client change; the ~1-2min blip didn't hit them, their concurrent auto-audio issue was client-side localStorage). rollback = sed DOTS_TAG=v1 + docker compose up -d dots-tts (v1 image retained). Also: deployed container GPU crept ~6→13.9GB over 8h serving (cache accumulation; a redeploy resets it — watch item). reference_chatterbox_fast_repo

  • [2026-08-09→10] dots.tts (rednote-hilab) TTS burn-in on irv-ml1 + canonical voice corpus built (voices/). Operator-directed eval to potentially replace chatterbox-fast. dots.tts VERIFIED real (canonical HF ns dots-studio/, rednote-hilab/dots.tts-* redirects there; Apache-2.0; PyPI dots.tts 0.2.1; 2B continuous-AR = semantic enc + Qwen2.5-1.5B LLM + flow-matching acoustic head over 48kHz AudioVAE; zero-shot clone from wav+transcript). Runs on Ampere 3090 (sm_86, bf16, no fp8 dep); optimized RTF 0.22 at num_steps=10 (from_pretrained(..., optimize=True) CUDA graphs — raw unoptimized was 1.21), ~6GB VRAM, 48kHz, streams (generate_stream). Venv+cache at irv-ml1:/home/lkraven/dots-tts (~10GB). Operator design calls: SGLang Omni serving (OpenAI /v1/audio/speech), transcribe-refs-first, soar variant. ⚠ Omni serves soar but its continuous-batching + streaming opts are mf-only (soar = single-request) — non-issue for ratatoskr's single-consumer RP surface. KEY FINDING — dots is highly sensitive to an accurate AND sentence-bounded reference transcript: mismatched transcript → 0.16s collapse; over-long/messy transcript → reference-audio BLEEDS as an output prefix; mid-clause trim → dangling-word leak (glados "we'll", emmie "And,"). RECIPE (baked into voices/derive.py): trim ref to a clean ~610s clip ending on a sentence boundary + accurate transcript of exactly that clip. CANONICAL VOICE CORPUS stood up in eshpfi voices/ (operator idea): engine-agnostic canonical/<v>.wav + transcripts/<v>.txt → per-engine ref sets DERIVED by derive.py reading engines.yaml profiles (dots/chatterbox/zonos); canonical wavs git-tracked (small/curated), derived/ gitignored. 4 voices optimized + verified CLEAN for dots: donut, glados, emmie, miranda (glados canonical is low-SR 16kHz — flagged upgrade candidate). ⚠ GPU GOTCHA: irv-ml1 native CUDA orders A6000=device0 (ComfyUI-full) — pin the 3090 with CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=0; and PYTORCH_CUDA_ALLOC_CONF=expandable_segments CONFLICTS with optimize=True CUDA graphs (curr_block error). Booths: dots-vs-chatterbox, dots-voices-optimized. SHIPPED 2026-08-10: operator A/B verdict "dots is very good" → containerized as a thin FastAPI wrapper over DotsTtsRuntime (chosen over SGLang Omni — Omni's batching is mf-only, unneeded for ratatoskr's single consumer; wrapper is SERIALIZED one-gen-at-a-time via a threading.Lock, Omni+mf = parked API-compatible escalation if multi-consumer ever lands). LIVE on irv-ml1:8198 (local/dots-tts:v1, OpenAI /v1/audio/speech + /health + /v1/voices, container healthy, both stream + non-stream verified CLEAN, 4 voices donut/glados/emmie/miranda) alongside chatterbox :8197 (nothing repointed). Stack = stacks/dots-tts/ (Dockerfile/app.py/compose/.env.example/README). ⚠ CONTAINER GOTCHA: optimize=True (torch.compile/inductor/triton) needs a C compiler at RUNTIME — slim image must apt install build-essential or model-load dies "Failed to find C compiler" (host venv had gcc ambient, masking it); persist TORCHINDUCTOR_CACHE_DIR to a mounted dir or every restart re-JITs ~5min. Corpus home = eshpfi voices/ (operator ruled keep-here). REMAINING: ratatoskr client cutover to :8198 /v1/audio/speech (Phase-2 tail, peer-coupled — draft the ask). reference_chatterbox_fast_repo reference_zonos_tts_stack reference_verify_hf_repo_ids_before_pull

  • [2026-08-08] worldtree-dev #400 CLOSED → fiction-decomp snapshot cleared from nh3-dev. worldtree-dev signaled #400 done (shipped v1.0.0b185; exact-lexical efficacy 79%→12% on ratatoskr's gate, brokkr no-harm bracket green both ends; the snapshot served 4 probe rounds — rank decomposition, promoted-vs-gold annotation, tie-set falsification, A0/A1/A2 mechanism probe). Cleared ~/snapshots/worldtree-400-fiction-decomp (208M: chroma + manifest/provenance/stamp) — a read-only rsync copy of PERSONAL Worldtree's Chroma (source on corviduo-dev, so safe to remove). LEFT INTACT: rex393-fiction-index/rex393-fiction-snapshot (separate operator KEEP word, unchanged) + r42-gate-*. No config deltas rode this train. Only remaining non-blocking await = ratatoskr-dev's chatterbox-fast knob revert. Replied confirming (01KZJ9GMCC…).

  • [2026-08-07] chatterbox-fast "broken audio" root-caused (T3 AR tail over-run) + FIXED (max_chunk_chars=250 cap, :v2 deployed). Long saga, operator-driven clean diagnosis. Symptom: ratatoskr's migrated RP-surface TTS "swaps to German" / "dead air" / "garbage" on long turns. NOT German-leak (Turbo generate() has NO language param — plain AutoTokenizer, no language_id; the multilingual language_id="en" lever lives only in the separate ChatterboxMultilingualTTS), NOT OOM alone. Real cause: the Chatterbox Turbo T3 model OVER-RUNS its generation tail — a long single generate() degrades into garble/dead-air in its final ~2-3s (lib filters OOV tokens <6561 + pads silence = messy AR tail). The scheduler's buffer-ratchet builds 300-600 char mega-chunks that land in that zone; streaming concatenates each bad tail (worst case). ratatoskr's anti-"German" knobs (top_k=80/temp=0.5) made it WORSE — tight sampling pulls the degradation onset SHORTER (~200 chars vs ~300 at default knobs). Diagnosis method (deterministic, no ears-only): single-shot length sweep + amplitude-gated voiced-ZCR (garble spikes ZCR; must gate on |x|>500 else trailing silence confounds it) — degraded voiced-tail = 1.58× mid, clean = ~0.64-1.1×. FIX: server-side max_chunk_chars=250 cap on the scheduler (:v2 image, CBF_MAX_CHUNK_CHARS=250 env) — bounds each generation to just under the ~300-char onset → clean 3-4 sentence chunks (max prosodic arc while clean). Operator ear-confirmed clean audio + clean joins; chatterbox's low emotiveness keeps chunk joins smooth (the harsh joins that got Zonos rejected are absent — operator's key call). ratatoskr TODO (relayed msg 01KZER9X7S): revert knobs to default (top_k→1000, temp→0.8), send full text (server chunks internally), keep the 503-on-empty guard. Cap value tunable per-request (max_chunk_chars) + env. Deeper prosody (if ever wanted) = scheduler Phase-2 context-priming at joins (feed prior sentence as discarded-audio context; +latency). ⚠ FOOT-GUNS: (1) acoustic tail-trim is UNRELIABLE — sibilants ('s'/'sh'/'f') spike ZCR like garble, can't cleanly detect the speech→garble boundary. (2) build-context vs image drift — the :v2 image was built from cap source, but after a :v1 rollback the build context held :v1 source → a docker compose build would've silently produced a cap-less :v2; re-synced the flat cap source to /opt/docker/compose/chatterbox-fast/ (rebuild-verified). ⚠ DIVERGENCE (follow-up): deployed build context is FLAT (app.py/scheduler.py, from scheduler import, thin-overlay FROM local/chatterbox:v1, cap-only) vs the vh/chatterbox-fast REPO which is PACKAGE-layout (chatterbox_fast/, from chatterbox_fast.scheduler, self-contained Dockerfile) + has norm_loudness (repo commit 6bc7bf0 = cap; deployed omits norm_loudness deliberately to keep the ear-test unconfounded). Reconcile the two layouts so a repo-based rebuild matches deploy. Rollback: .bak-cap-20260807-104850 backups on irv-ml1 + :v1 image both retained. reference_chatterbox_fast_repo reference_zonos_tts_stack

  • [2026-08-07] Zonos2 TAKEN DOWN on the 3090 (irv-ml1) — operator-directed "for memory", TEMPORARY. Freed ~17.4 GB (3090: 728 MiB → 18.2 GB free) so chatterbox-fast (co-resident, was OOMing on long generations) has headroom. ⚠ Restore is manual — Zonos2 :1920 was a DETACHED native process (NOT systemd/docker), reparented to init. GPU memory was held by the --multiprocessing-fork CHILDREN (1966165=16.4G, 1966166=1G), which ORPHAN to init when you kill the parent — had to SIGTERM the children explicitly (killing the parent 1965942 + uv-run 1965935 alone left the 16.4G held). RESTORE CMD (from irv-ml1, user lkraven): cd /home/lkraven/tts-audition/models/zonos2 && nohup uv run python -m zonos2 --model-path Zyphra/ZONOS2 --host 0.0.0.0 --port 1920 --tts-default-voices-dir ./default_voices/ --cuda-graph-max-bs 1 --num-pages 16384 --max-running-requests 2 --memory-ratio 0.3 > /tmp/zonos2.log 2>&1 & then docker start zonos-gateway. Consumers that lost Zonos: asset-engine + gateway-chat (via LiteLLM ext-tts alias → zonos-gateway :8890, now stopped); ratatoskr already migrated OFF to chatterbox-fast (unaffected). Also unblocks proper drift/cap testing (OOM was blocking it). reference_zonos_tts_stack

  • [2026-08-07] chatterbox-fast: donut voice added + full contract delivered to ratatoskr-dev (their TTS migration off Zonos). Operator-directed. Copied zonos-gateway/voices/Donut.wav → chatterbox /refs (/worktank/chatterbox/reference_audio/donut.wav — the reference_audio SUBDIR is lkraven-owned so no sudo despite /worktank root; container globs /refs live → NO restart), exposed as voice:"donut" (lowercase); verified clean 7.5s synth (24kHz, RTF ~0.31). A/B booth (chatterbox vs zonos donut, same line) at http://10.100.10.50:8090/b/donut-chatterbox/. Answered ratatoskr's 8-question contract ask from the live gateway (local/chatterbox-fast:v1) + source: NOT OpenAI-shaped (POST /tts; body text/voice/format/stream, not input/model/response_format); NO affect dials (Turbo ignores cfg_weight/min_p/exaggeration — the architecture-changing answer they flagged; Zonos stays the only fleet TTS with real emotion steering); streaming WAV placeholder-header shape IDENTICAL to Zonos (their per-chunk Web Audio path survives); SR 24000 (Zonos 44100); server chunks arbitrary-length text internally (no client-side chunking, unlike Zonos's 71.2s cap); English-only, no language pin. FYI-worthy (operator): ratatoskr is moving its RP-surface TTS OFF Zonos back to chatterbox-fast → loses the live-PAD affect coupling (heavy Zonos emotion investment) — their call, trade-off flagged to them. auto-memory reference_chatterbox_fast_repo enriched w/ the live contract. reference_zonos_tts_stack

  • [2026-08-07] Fleet reranker cut over: Qwen3-Reranker-0.6B → BAAI/bge-reranker-v2-m3 (Brokkr R43). The incumbent was measured HARMING 80/90 fleet queries (no-reranker beat it 89/90 vs 56/90). R43 bake-off: the A2 control (same Qwen weights, seq-cls head) scored identical to the incumbent → proved the fault is a training-prior not the serving head → cancelled the expensive Qwen3-4B arm; A3 (bge-v2-m3) won on multilingual safety + bare-name recovery. LiteLLM reranker repointed incumbent→A3 :8013 (boundary 2026-08-06T17:37:48Z, config-edit + ~52s gateway restart); R42 v13 gate PASSED first-ever (56/90→90/90). Incumbent kept warm :8002 (rollback via qwen3-reranker alias), A4 fallback :8014. Full arc + rollback runbook docs/pfi/reranker-selection-ledger.md; commits ad2df89/2c11748/377f8a4 (unpushed). auto-memories: the earlier reranker-serving notes.

  • [2026-08-05] Fleet CI resilience flip (DEFAULT_ACTIONS_URL=self) — attempted end-to-end, PARKED on a runner action-fetch auth blocker; infra-ops to research it (operator-directed, deferred, NOT now). 7 gitea action mirrors staged public+populated (orgs actions+astral-sh); the flip resolves uses: correctly but act_runner v0.6.0 can't authenticate its fetch to gitea 1.26 ("Invalid username or token. Password authentication is not supported"). Reverted (CI back on github default); REQUIRE_SIGNIN_VIEW=false KEPT as a standing change (operator, internal WG net). Full endeavor, the reliable nh3-dev-egress + git-SSH mirror method, exact config state, smoke method, and next step → persistent-memory.d/2026-08-05-ci-flip-parked.md

  • [2026-08-05] worldtree herald re-nudge bug root-caused → forseti shipped althing-core v2.1.2 (d5d33df, deployed on nh3-dev). herald.py:363 rendered the wake command from the empty fresh mail set on the re-nudge path (should be deliver_msgs) → messages[0] IndexError → un-suppressed outer catch-all → 7s crash-loop for 9 days on worldtree-codex's pane route (mimir-dev surfaced it; I traced it from the editable source). Fix + render_command empty-guard + outer log-suppress + 3 tests + contract amendment, all forseti's. nh3-extdev herald 2.1.2 upgrade DEFERRED (operator, not-now): extdev is a WHEEL install (not editable), unexposed (no pane routes); the verified 2.1.2 wheel is staged on nh3-dev /tmp (sha256 003508…cef27) — uv tool install --force + restart both heralds when un-parked. extdev herald-unit provenance resolved (operator-authorized 2026-07-25 via forseti relay; recorded in this file's 07-25 herald-install entry). auto-memory reference_nh3_dev_althing_herald.

  • [2026-07-31] muninn-gate (#377 ingestion front door) BUILT + DEPLOYED + healthy on corviduo-dev:8090. First-boot acceptance passed (watcher:running:true proves ingestion_root byte-identity); submit path deferred to the mimir-inbox era. Full wiring (uid-1000, state-volume mount, staging path-agreement, BuildKit-secret build, deferred repoint + operational guards) → persistent-memory.d/2026-07-31-muninn-gate-deploy.md

209 older entries archived to archival-memory.md.

Tried and abandoned

  • [2026-08-23] A HEAD == GITHUB_SHA assertion in the hrafn CI — added, broke the checkout twice, removed. It needed the git binary (run 9920, exit 127); installing git then flipped actions/checkout@v4 off its node implementation onto the git binary, which died on a missing CA bundle (run 9921). A nice-to-have assertion changed the checkout's code path and broke a working pipeline. Removed rather than patched with ca-certificates — it guarded a hypothesis that proved wrong. Do not add git to that prereq step.

  • [2026-08-23] Repointing selene-1-mini-8b at gen's endpoint — proposed by me, correctly overruled. "never repoint a named model at a different model's endpoint — that is intentionally misleading." The trap is that it does not feel like deception; it feels like sparing consumers a migration. That framing is the tell. Role aliases move; model names die with the model and 4xx.

  • [2026-08-15] Grafted bf16 MTP loads UNINITIALIZED (0% accept) unless re:^mtp.* is in the quant-config ignore; and W4A16=Marlin (not native FP4) costs ~20% even on decode. Cost a premature 79 GB delete of a good model (declared desync-dead off the 0%). Lessons: test MTP on bf16 FIRST, isolate before deleting; modelopt 0.43 is dependency-hell for qwen3_5 (list-vs-dict quant_cfg + transformers conflict) — use llm-compressor. Full → persistent-memory.d/2026-08-15-uncensored-gen-seat.md

  • [2026-08-03] ComfyUI --enable-triton-backend on the irv-ml1 A6000 crashes EVERY render — Ampere has no hardware e4m3. adhoc-agent's operator-approved probe: comfy_kitchen's triton backend has a FUSED int8 matmul that would beat the eager backend's ~1.9x-slower unfused int8 path (21.3s vs 11.2s fp8 on the Moody Krea2 int8 checkpoints). Flipped it (added to COMFY_CMDLINE_EXTRA, recreated) → triton.compiler.errors.CompilationError: ValueError("type fp8e4nv not supported in this architecture. supported: fp8e4b15, fp8e5") in comfy_kitchen/backends/triton/quantization.py:145 dequantize_per_tensor_fp8, failing at node 5 CLIPTextEncode. Triton's fp8 dequant kernel targets fp8e4nv (Hopper/Ada e4m3); sm_86 Ampere (A6000) lacks hardware e4m3 → the JIT compile dies. With triton on it grabs the global --fp8_e4m3fn-text-enc dequant, so every render (fp8 AND int8) dies upstream at the text-encode step — the int8 UNet path never ran, so the convrot-coverage caveat wasn't even the limiter. Reverted cleanly (~15s to healthy, image unchanged sha256:94afb8ca, sage intact, prod restored). The parked cu130 rebuild won't fix it (e4m3 = hardware format, not CUDA version). DEFERRED to the Ada refresh (operator: "ada is coming, we'll optimize then" — Ada sm_89 has native e4m3, so triton's fp8 path should compile there). Mechanics: --enable-triton-backend is a compose environment: var, so toggling it needs docker compose up -d (recreate), NOT docker restart (reuses the baked env, no-ops silently). Full: auto-memory parked_triton_backend_ampere_fp8.

143 older entries archived to archival-memory.md.