Files
esh-pfi-infrastructure/persistent-memory.md
T
vh 5af362e9d0 ops(ana-gw): restore WAN admin access ahead of the FortiGate cutover
Re-open the ana-gw admin GUI on wan1 so the Anaheim edge can be
managed remotely if the cutover goes wrong, reversing part of the
2026-08-12 lockdown. Two config changes, nothing else (verified by
diffing pre/post `show full-configuration`):

- wan1 `allowaccess ping https` — https only; http, ssh, and fgfm
  stay off, and wan2 is untouched.
- `infra-ops` trusthost widened to all routable IPv4; the `admin`
  account stays locked to 10.0.0.0/8 so the guessable username
  remains unreachable from the internet.

Verified end-to-end from two sites: a real `/logincheck` POST returns
AUTH OK over the public path, on a browser-trusted Let's Encrypt cert
for ana-fw.phasefinal.com valid through 2026-10-27.

Two FortiOS behaviours worth recording, both of which cost time here:
a trusthost whose base address is 0.0.0.0 is silently treated as
unset (so there is no writable "any" — only decomposed ranges), and
trusthost is enforced before the TCP handshake, so a blocked source
sees a filtered port rather than a refused login.

Follow-ons captured in memory, not actioned: ACME renewal for the
admin cert needs port 80 on wan1 (next attempt ~2026-09-27), and a
~5 SYN/s source in 179.51.184.0/21 now draws SYN-ACKs at no
measurable CPU cost.
2026-08-23 14:03:40 -07:00

70 KiB
Raw Blame History

Persistent memory — eshpfi-management

Last updated: 2026-08-23

Always check for /tmp/infra-ops-handoff.md — if it exists and its Written: stamp is under an hour old, read it (it carries the in-flight handoff from the previous session), then delete it. Older than an hour: stale — delete it unread.

Repo purpose

Reference workspace for PFI infrastructure: server inventory, canonical Docker Compose stacks, ops playbooks, and conventions. Authoritative copies of compose files live on the servers under /opt/docker/compose/<stack>/; this repo mirrors them for version control, editing, planning, and CI-driven deploys. It was originally spun up to handle the fleet backups — keep that lens when triaging backup/storage issues.

Tools and conventions

Sister repos (separate gitea repos, deployed by playbooks here):

Repo Role CI status
vh/task-board MCP + web dashboard for assistant task state (port 7878) push-to-main → CI deploys (2026-04-29)
vh/vor Inquisitor UI sidecar (port 7879) push-to-main → CI deploys (2026-04-29)
vh/nevermore Twice-daily LLM-curated briefing (port 8181, replaces news-digest) push-to-main → CI deploys (2026-04-30)
vh/asset-engine Internal control plane over inference services (port 8200, LAN-direct) push-to-main → CI deploys (2026-05-12)
vh/althing Lean trusted inter-agent message bus — v2 "email model" (v2.0.0b2, 2026-07): per-box local-SQLite bus + courier/receiver for P2P over the 10.x net; pillars = open-loops / per-box herald + wake-listener / roaming owner API /owner/* / althing-mcp stdio surface. The v0.15 lean-bus cut RIPPED moderation / chamber / forseti-daemon / agent-runner / redis-valkey. per-box uv tool install (NOT CI-deploy); nh3-dev = the DEV box (editable install of ~/development/althing, gets new versions first); nh3-extdev a mesh peer (model B: althing-svc + shared /srv/althing)
vh/mead-hall Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) push-to-main → CI deploys (2026-05-16)
vh/skaldsong Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) push-to-main → CI deploys (2026-05-19)
vh/Worldtree Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. gitea-runner builds on ana-docker; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. push-to-main → CI build-and-deploy (runner on ana-docker)
vh/yt-voice-clipper YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md
vh/arbo Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook
vh/zonos-gateway OpenAI-compatible TTS gateway over stock ZONOS2 (:8890 irv-ml1); emotion dials-first + voice mapping; reached via LiteLLM ext-tts alias. v0.2.1 (2026-07-18): voice-resolved emotion presets (resolve_preset(name,voice); angry/happy/startled_happy per-voice). 8 voices incl. 4 clones pushed to gitea (main 8f1885b/v0.2.1); deployed irv-ml1 tree still NON-git (hand-updated build context — CI-wire = open follow-up). Spec docs/EMOTION-DIALS-SPEC.md; host-managed voices bind-mount (./voices:/app/voices, drop wav + restart, no rebuild)
vh/soong-lab Noonien Soong character-design studio (SPA + /api + WT /bifrost/tool-call); containerized 2026-07-18, LIVE on corviduo-dev :8443 (image vh/soong-lab:latest). soong-dev owns Dockerfile/compose/workflow; infra-ops owns the host CI = Gitea Actions build+push+DEPLOY on tag/dispatch (fleet recipe: docker:cli + raw buildx, pushes AS vh; auto-redeploy LIVE 2026-07-18 — runner SSHes corviduo-dev as deploy, compose pull && up -d from /opt/soong-lab, health-gated on /api/version). Manual redeploy sudo -u deploy bash -c 'cd /opt/soong-lab && docker compose pull && docker compose up -d'. → archival-memory.md (archived 2026-08-16)
model-training-forge (mtf-dev) Fine-tuning recipe forge; T1 = E-RP writing LoRA, retargeted qwopus-122B→AEON-27B (2026-07-06) (SFT→DPO, LitBench-RM reward) training runs, not a deployed sidecar

(vh/volva + Heid were re-architected from systemd daemons to Claude Code session orchestrators 2026-06-08; their nh3-dev .service units were removed — no longer deployed sidecars here. See Recent decisions.)

  • Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see docs/runbooks/disaster-recovery.md for the blast-radius matrix. ⚠️ The restic file+DB layer routes through TWO rest-servers (rest-server-ana @ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas; rest-server-nh3 @ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of /mnt/backup. (rest-server-ana recovered 2026-06-20.)

  • pull-hf-repo.yaml is the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at /tank/aimodels/huggingface/" playbook. Supports --var repo_type=model|dataset|space. Replaces ad-hoc huggingface_hub.snapshot_download patterns.

  • Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (key_id 61419c92) at ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin auths against demo only. Personal-instance admin (the ~/.config/worldtree/personal-admin-token, mode 600) POSTs /admin/keys (mints per-project keys; takes user_id+label, no scope param — scopes are tier-derived). On-instance mint recipe (cleaner than DB-manip): docker exec worldtree-worldtree-api-1 POST /admin/keys with the in-container WORLDTREE_BOOTSTRAP_ADMIN_KEY; cleartext once in .key=wt_live_+16hex. auto-memory reference_worldtree_demo_key_mint.

  • Per-project user keys against personal Worldtree (issued 2026-05-19): skaldsong:79744637, skaldsong:7c1dbbbe, althing:50d85460, mead-hall:a360822d. Mint via /admin/keys, drop value to /tmp/wt-personal-<name>.key mode 600, dev collects + shreds (DO NOT cat to chat transcript).

  • Skaldsong CD pattern (registry-pull). vh/skaldsong's CI builds and pushes gitea.phasefinal.com/vh/skaldsong:<sha> + :latest; playbooks/deploy-skaldsong.yaml on ana-docker pulls + recreates. SHA-pin only. Prereq: host needs docker login gitea.phasefinal.com once.

  • gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH 10.250.50.70:222, HTTP :3000. Fleet/colo hosts must use this internal route, NOT public gitea.phasefinal.com (38.120.12.44) — the public path fail2bans the host egress IP. Full gotcha in docs/orientation.md → Git/gitea.

  • docker-as-root pattern (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo): docker run --rm -v <target-dir>:/wt docker:cli sh -c "...". docker-group membership is effectively root via bind-mount. Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass -e VAR=/abs/path for any relative-default config dir.

  • scripts/elway sudo handling — elway prompts for the sudo password ONCE via getpass before the first sudo: true step → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH.

  • Per-host SSH identity matters for sudo. infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: default ssh ana-docker = lkraven (docker-group, NO passwordless sudo); ssh infra-ops@ana-docker HAS NOPASSWD root. → For any sudo op on ana-docker, use ssh infra-ops@ana-docker. ssh infra-ops@10.100.10.50 (nh3-dev) ALSO NOPASSWD sudo; on nh3-extdev infra-ops is sudo-LESS by design (ssh lkraven@10.100.50.42 is the NOPASSWD path). irv-ml1: ssh irv-ml1 = lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD) — stage model pulls to /home, not root-owned /worktank.

Current state / in-flight

As of 2026-08-23 — a long multi-party ops session; open with the operator are the Anaheim tunnel cipher and how long WAN admin on ana-gw stays open.

  • 🔓 ana-gw WAN admin is OPEN again as of 2026-08-23 — deliberately, and it should be re-closed after the FortiGate cutover. Operator asked for it as the cutover contingency ("so I don't have to drive down there"). https://ana-fw.phasefinal.com/ (wan1's own IP 38.120.12.42) serves the admin GUI on a browser-trusted LE cert; https only, and only the infra-ops account is WAN-reachable (admin remains 10.0.0.0/8-locked). Verified end-to-end with a real login from two sites. This partially reverses the 2026-08-12 hardening on a box still running the EOL, actively-exploited FortiOS 7.2 branch — so the exposure is CVE-shaped, and time-boxing it is the mitigation. Re-close = config system interface / edit wan1 / set allowaccess ping / next / end. Config backups pre+post at /var/tmp/ana-gw-config-20260823T20*.conf on nh3-dev. Two follow-ons: (a) ACME renewal for that cert needs port 80 on wan1 — currently absent, next attempt ≈2026-09-27, expiry 2026-10-27; (b) ~5 SYN/s from 179.51.184.0/21 now draws SYN-ACKs, CPU impact nil, unmitigated by choice. Detail + the FortiOS trusthost gotchas: auto-memory reference_fortigate_ana_gw_access.

  • THE OPEN ITEM: Anaheim's IPsec tunnel delivers ~25% of a verified 2 Gbps circuit. Circuit measured at 2,153 Mbit/s; the tunnel ceilings ~550 Mbit/s aggregate, ~230 single-stream. FortiGate CPU idle, IPsec NPU-offloaded, interface error-free — so it is not crypto exhaustion. Both tunnels negotiate aes256-sha1; AES-GCM is the proposed change and the operator has signalled he will authorize it. Not executed: production edge, needs a matching change at NH3 + ESH, each tunnel drops during renegotiation. → persistent-memory.d/2026-08-23-anaheim-ipsec-tunnel-ceiling.md

  • 🟢 SEAT MAP (unchanged this session except selene). gen = orcarouter/Qwen3.8-27B-Uncensored NVFP4-mixed, GPU0 :8015, now 7 aliases (see the collision note). char-rp = MeroMero-v2 dual-mode, GPU0 :8016, pinned v0.26.0. sec/sec-reasoning = M.O.G.-SEC on DFlash2, GPU1 :8019. selene RETIRED — 17.2 GiB reclaimed on GPU1 (free now ~19.4 GiB).

  • ⚠️ THE sec DEGENERATION QUESTION IS STILL OPEN AND CONFOUNDED. Engine and drafter changed together; the isolating experiment is MTP k=3 on e9d1398d — still not run. Operator ruling stands: degeneration lives in the un-fixed vLLM, not the weights; the MTP-head hypothesis is retracted. Both prior sightings are n=1 and are NOT evidence. gen remains on the old nightly, untouched, gated on that experiment.

  • 🟢 ana-ml2 now mounts /mnt/smithy (nh3-nas) ro + soft, NOT in fstab — needs a manual remount after reboot. For brokkr's R47 CPU work. Reads 24.7 MB/s sequential vs 98.3 on nh3-dev (that gap is the tunnel above), but 45 files/s vs 34 — small-file work is genuinely faster there. → persistent-memory.d/2026-08-23-smithy-mount-ana-ml2.md

  • 🟢 ESH IS DUAL-STACK; the v4 static is a Cityside ticket. IPv6 live on esh-userland and esh-server from a delegated /56. v4 remains CGNAT and a full gateway reboot proved the purchased static is not provisioned — carrier ticket, nothing left locally. NH3 stays v6-off deliberately. Flat-zone lateral-movement finding parked, id 44.

  • 🟢 OTHER SERVICES. hrafn browser-fetch adopted on ana-docker (infra-ops owns uptime; CI now genuinely deploys). speaches ASR live irv-ml1:8204. Open WebUI esh-docker-vm:3211 — Lobe retirement still the operator's call. Booth gained kept-board deletion + per-row link pruning. pfi gitea org created; claude-bot is an Owner and can create repos self-serve.

  • OPEN ELSEWHERE: MTP-k3 isolating experiment; upstream vLLM issue to file (operator's GitHub identity); Cold-Fusion NVFP4 quants (44 GB) delete/keep; OWUI image-tag drift; /tank DEGRADED 70+ days; Worldtree #411 debug-room litter; bridge/engine agent-roster drift on both WT instances; brokkr's gen vs trained-reward-model bake-off (theirs to initiate). Working tree is clean and pushed through 0ad332b.

Recent decisions

  • [2026-08-23] Anaheim's IPsec tunnel delivers ~25% of a verified 2 Gbps circuit — not WireGuard, not CPU, not the fibre. Cipher change proposed and operator-signalled; execution pending, untracked by operator choice.persistent-memory.d/2026-08-23-anaheim-ipsec-tunnel-ceiling.md

  • [2026-08-23] selene retired after losing a head-to-head on its own job; chat-judge moved to gen, the model name 404s by design. Also surfaced that 7 aliases share one seat — cross-checking between them is an echo, which caught a real defect in brokkr's 46k-exposure R47 gate. → persistent-memory.d/2026-08-23-selene-retired-alias-collision.md

  • [2026-08-23] hrafn adopted; its CI reported green for its whole life while deploying nothing. A staging dir inside the rsync target destroyed its own source mid-copy; the deeper fault was verify steps that asserted uptime, never content. → persistent-memory.d/2026-08-23-hrafn-adopted-ci-frozen-source.md

  • [2026-08-23] Worldtree b187 shipped; all three instances de-armed from a 69-day-stale :latest; Matrix homeserver re-plumbed to personal. Includes the :8009-is-demo port trap that an IP-only fix would have walked into. → persistent-memory.d/2026-08-23-worldtree-b187-pins-matrix.md

  • [2026-08-23] Every secret-bearing .env on ana-docker tightened to 0600 — eight stacks including vaultwarden and traefik, verified exposed by reading one as nobody. → persistent-memory.d/2026-08-23-ana-docker-env-perms-sweep.md

  • [2026-08-23] pfi gitea org created; claude-bot is an Owner and creates repos self-serve. Closes the repo-creation half of the credential-migration directive — vh is a USER namespace so no service account could ever create there. Repo creation needs write:user + write:repository + write:organization; POST /users/{u}/tokens is basic-auth only, so minting needs the account password. Default new repos to pfi/. (vh/eitri-smithy was its first tenant, then moved.)

  • [2026-08-23] Booth: kept boards are deletable and link rows are prunable. release on a kept card drops the sentinel so the existing × applies; booth links / booth unlink <id|index> prune one row. Rows are addressed by content id, never position — the board is append-only and multi-writer. Releasing a board RESETS its TTL clock (unlink bumps the dir mtime), so unkeep-and-wait is a 24h delay, not a delete. (4be880f, 0ad332b)

  • [2026-08-22] DFlash2 spec-decode measured on our own stack; sec promoted to it. +1821% accepted length and +1518% throughput over MTP k=3, drafter proved model-agnostic across two finetunes to 0.06%, and the k=7 MTP control showed deeper MTP is a throughput trap. → persistent-memory.d/2026-08-22-dflash2-spec-decode.md

  • [2026-08-22] Quant pipeline shipped a crippled tokenizer for months — fixed at source. quant_mixed_nvfp4.py baked its calibration truncation (max_length 2048) into every mixed-NVFP4 build; latent on old transformers, fatal on new. Both live quants corrected, pipeline now saves a source-pristine tokenizer and asserts it. Playbook §3.14. (0755ba7)

  • [2026-08-22] sec retuned to util 0.52 / 420K after a runtime OOM at 0.55/480Kgpu-memory-utilization is not a hard reservation; activation grows past the dummy-data profile and six vLLM containers share GPU1. Also measured: the KV pool varies ~6.6% between boots, so max-model-len must be sized against the lower observation. (6e82899)

  • [2026-08-22] Max-Q 1.8× spread does NOT apply to LLM decode — measured, not argued. ana-ml2 draws 256266 W of 300 W under sustained 100% decode with SW Power Cap: Not Active and clocks pinned. Corrected to brokkr-smithy-dev after I had lent the claim credibility; 122B figure (~9093 tok/s at 262K) stands as a straight number.

  • [2026-08-21] ESH internal IPv6 live on two LANs; the Cityside v4 static is a CARRIER problem, proven. A full gateway reboot forced a fresh DHCP DISCOVER and returned the identical CGNAT address. YaRN was already configured — "1M needs YaRN, absent" was false. → persistent-memory.d/2026-08-22-dflash2-spec-decode.md sibling entry in ad21302

  • [2026-08-21] speaches ASR live on irv-ml1 for Eyra — and no_speech_prob alone is a weak hallucination gate. Silence and room tone both hallucinated "Thank you." under 0.11; avg_logprob separates ~6× better. Consumers should gate on a composite. (aa5863c, c7e2187)

  • [2026-08-20] Cold-Fusion abliteration — Robinson recipe captured; the fight was the environment, not the recipe. Stock Cold-Fusion measured ~33% creative refusal → worth abliterating ourselves (supersedes waiting for DavidAU's heretic build). Recipe maps 1:1 (131 tensors); capture succeeded only in fp32 — transformers' Qwen3.5 DeltaNet linear-attn NaNs nondeterministically in bf16 without the unbuildable causal-conv1d kernel (precision cancellation, not overflow). Direction finite at layer 22 but agreement 0.59 (vs Robinson's 0.99) → calibration-set expansion is next.persistent-memory.d/2026-08-20-coldfusion-abliteration-capture.md

  • [2026-08-19] A software watchdog is not watchdog protection — esh-pve froze for 4.5h holding one. softdog cannot fire when the kernel it runs in is wedged, and Proxmox's watchdog-mux never arms without HA resources, so the box looked protected and wasn't. Moved to the PCH iTCO_wdt under systemd. Also: a single cross-VLAN DNS entry with no secondary turns any VM outage into a whole-site outage. → persistent-memory.d/2026-08-19-esh-pve-freeze-dns-spof.md

  • [2026-08-19] Fleet .internal DNS built and live — git-sourced, agent-managed, three resolvers. Zone-scoped authority (ESH's hand-made esteban.net rewrites survive); the colo had no resolver at all; v6 column empty on purpose because SLAAC addresses rotate. → persistent-memory.d/2026-08-19-fleet-internal-dns.md

  • [2026-08-19] waterland studio containerised on irv-ml1 — three landmines, all measured. cupy needs CUDA headers the host had by accident; uv run re-syncs and prunes cupy at RUNTIME; the A6000 is container-index 0, not the host's 1. → persistent-memory.d/2026-08-19-waterland-studio-containerised.md

  • [2026-08-19] Homepage cleaned up, then themed with Australis Skyfall + an Arbo-generated background. Includes the hour lost to a self-healing tab-bar red herring, and the CSS-iteration loop that prevents it recurring. → persistent-memory.d/2026-08-19-homepage-skyfall-theme.md

  • [2026-08-19] Four unmanaged stacks found on live hosts — two quietly broken. A dashboard card is a cheap census of what is actually running; check whether the stack is even in stacks/ before debugging the symptom. → persistent-memory.d/2026-08-19-unmanaged-stacks-searxng-seafile.md

  • [2026-08-19] claude-bot granted read on vh/waterland (operator-empowered, verified admin:false push:false pull:true) so irv-ml1 can self-update without the operator's site-admin token living on a GPU box. Precedent for the standing migrate-off-operator-creds directive: grant the service account, wire a repo-scoped 0600 credential helper, keep the remote URL clean. Commit 8189076.

  • [2026-08-19] AI-tab Dormant regrouping BELAYED by the operator — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong, Chatterbox Fast) show amber EXITED inside live groups rather than AI - Dormant. Fix is a label change + recreate per stack; needs the operator's read on which are retired vs temporarily down. untracked by operator choice (his words: "belay the ai dormant regrouping for now").

  • [2026-08-18] esh-pve-nas migration STAGED — and staging is where three landmines surfaced, none of which the plan predicted. (1) The runbook's /boot LV had nowhere to live: VG pve had 4 MB free and mounted ext4 cannot shrink, so the space came from the 768 MB swap LV (operator's call: shrink to 256 MB, not drop). (2) The runbook's zpool set cachefile=… nvme would have broken the NAS — populating a cache flips the host to import-by-cache, and a one-pool cache leaves ssd+tank unimported under CT 103's twelve bind mounts. (3) update-grub silently emitted a pool-less root=ZFS=/ROOT/pve-1, because GRUB's ZFS reader cannot open a pool with encryption/large_dnode/zstd_compress and the probe failure is swallowed. All three were caught by verify steps that asserted effective state, not by reading the plan. → persistent-memory.d/2026-08-17-esh-pve-nas-dom.md

  • [2026-08-17] esh-pve-nas PVE root is on a USB DOM — mitigated, and the migration replanned to split boot from root. Operator's design beats my reinstall plan; wear was never the issue, blocked patching is. → persistent-memory.d/2026-08-17-esh-pve-nas-dom.md

  • [2026-08-17] irv-ml1 cleared of 782 GB, and Homepage brought under version control. One dead-looking Gradio app pinned three delete targets at once; /opt/ComfyUI is NOT the ComfyUI that serves. → persistent-memory.d/2026-08-17-irv-ml1-cleanup-homepage.md

  • [2026-08-17] Gen seat swapped to absolute-heresy — and the three bugs the swap exposed are worth more than the swap. Candidate MuXodious/Qwen3.8-27B-absolute-heresy (Heretic v1.4.0 + SOMPOA, T377) beat the incumbent on refusals AND KL simultaneously, which is the unusual part — those normally trade off. Validated on the probe port per operator ruling, promoted, all 7 aliases green. Durable lessons banked: (1) A CPU-only MTP head hash can replace the ~56 GB bf16 acceptance gate. The Qwen3_5ForConditionalGeneration wrapper never loads the MTP head, so PEFT merges / Heretic runs / llm-compressor passes all leave mtp.* pristine — hashing it against a head we have already measured (the incumbent's, 47.7%) answers the question for free. Predicted 47.7%, measured 47.2%. Saved downing meromero. Tool: services/gen-seat-mixed-quant/compare_mtp_head.py (hash bf16 via uint8 reinterpret — numpy has no bfloat16). (2) post_quant.py assumed a standalone model-mtp.safetensors; a full checkpoint keeps mtp.* in a NUMBERED shard, so the copy silently no-op'd while the index was still rewritten to point at a file that never existed — 15 unresolvable tensors behind a correct-looking tensor count. Its own FAILED-CHECKS assertion caught it; that is why the check exists rather than an assumption. Fixed to extract. (3) A probe that does not mirror the live seat manufactures failures. serve_probe.sh hardcoded :latest (seat is a pinned nightly for #51113), had no tool-call/reasoning parsers, and its --speculative-config JSON died twice on quoting — bash BRACE-EXPANDS {"a":1,"b":2} on the comma unless single-quoted at the REMOTE shell. Adding the seat's flags took the surface test from 5/6 to 6/6; the "tool calling broken" result was pure probe config. Commits 7997f11,254c588,2c36028,b0c2d3d,993421b.

  • [2026-08-17] Fleet IPv6 mapped + the real VPN topology verified; the driver is CGNAT at ESH, not the WireGuard mesh. New ESH fiber (installing 2026-08-18) lands the house behind CGNAT, which breaks Site Magic (NH3↔ESH sdwan-mesh-tunnel) on IPv4 — so IPv6 becomes load-bearing as the escape hatch, and that is its most likely first consumer. Topology as VERIFIED (a prior turn assumed wrong and was corrected): UniFi↔UniFi = Site Magic; colo↔UniFi = IPsec IKEv2 (pfi-ana-nh3 158M/165M pkt = the workhorse, ana-to-eshudm); WireGuard is an RA convention only, host-based on ana-wg UDP 31337 behind a FortiGate VIP — the FortiGate never terminates WG (FortiOS 7.2 has none; 7.4 added it) so "upgrade the edge for WireGuard" is a non-problem, do not re-derive. IPv6 today: NH3 WAN live 2600:1700:b25:c110::48, colo none, ESH none. AT&T delegates exactly ONE /64 (2600:1700:b25:c11f::/64) — proven by forcing prefix-ID auto→0 and watching the subnet NOT move, because the c110/c11f pattern otherwise reads convincingly as a /60. A mesh needs a routable WAN address, not PD. ana-wg's WG socket is already dual-stack ([::]:31337) → v6 RA needs an address + a v6 port-forward, no WG reconfig. ⚠ UDM legacy rest/firewallrule returns 0 rules (zone-based firewall) — use v2/…/firewall-policies; inbound v6 is default-deny and held. All three endpoints will be dynamic → extend the existing hostname pattern (ana-fw/nh3.phasefinal.com) to AAAA. Enabled PD on nh3-iot to measure, reverted on operator instruction (all 5 LANs back to none, verified). Also fixed: ana-wg WireGuard key material was world-readable (wg0.conf + keys/*_priv + *_psk + client configs/*.conf at 644) → now 600, dirs 700, service untouched. Detail → persistent-memory.d/2026-08-17-fleet-ipv6-mesh.md.

  • [2026-08-17] Gen-seat multi-day degeneration RESOLVED — two compounding real causes, not one; the meta-lesson is "a mitigation that HELPS but doesn't FIX means a second cause, not a wrong one." vLLM qwen3_5_mtp×GDN bug (#51113, real, fixed by nightly) + AEON full-W4A4 being lowest-fidelity (W4A4<W4+FP8<W4+bf16) → ~15-20% stochastic degeneration. Fixed by mixed FP8-attn build on pinned nightly. AEON purged. Also banked: stochastic (~15-20%) degeneration is invisible to a small synthetic probe — n=1 "clean" validated THREE non-fixes (MTP-off, APC-off, nightly-alone) that all failed in real use; get the operator's real transcript, do not trust your own probe. Full → docs/pfi/model-quantization-playbook.md §3.8 (+ §3.7 MTP-multi-turn). Commits d28a371,2f2bbce,2185964.

  • [2026-08-17] Lobe Chat chosen over Open WebUI (weight: 143 MB vs 1.8 GB) + stood up on esh-docker-vm; scoped LiteLLM key blocks paid models; System-Agent gpt-5-mini default repointed via env. TTS env-vs-UI resolved as a split (endpoint env-driven, voice/model UI-only). tts-dev onboarding closed both directions; ballad/verse aliased so no voice can 404 the router. Commits e9362de,163a725,cac75cb,933253d,25fa18e.

  • [2026-08-17] LiteLLM upgraded v1.91.0→v1.97.0 (RC-avoided on the fleet gateway) + the 6 GB spend-log DB purged & capped (store_prompts_in_spend_logs:false + 7d retention). Interpreted "get rid of the db" as the spend-log DATA not the database (keys/config live in it). Commit 01b5ad9.

  • [2026-08-16] Abliterated models go CATATONIC at the hard refusal edge — silence, not a decline. Abliteration removes the refusal direction, so at the genuine hard edge the model neither refuses nor complies → empty/degenerate output. Durable measurement consequence: a refusal probe MUST score EMPTY as a verdict distinct from REFUSAL and COMPLY (services/refusal-probe/probe.py does). Operator accepted it as out-of-scope; do not chase.

  • [2026-08-16] Fable-Fusion 711 cuts cold-framing refusals 92.5% → 15.8%; refusal is MONOTONIC IN FRAMING, and DS v1.0's problem is that she was never abliterated. brokkr-smithy-dev supplied the framing that reproduces (01M05M48R4RSZF9D8KT7RR55EJ): a bare assistant-mode instruction — no character card, no permission preamble. Three-arm A/B, same harness, same classifier: permission framing DS 0.0% / FF 0.0% (n=75); plain character cards DS 1.4% / FF 0.0% (n=74); bare instruction DS 92.5% (37/40) / FF 15.8% (6/38). Per-axis DS→FF: incest 100→20, non-con 100→20, bestiality 100→25, necrophilia 100→40, gore 100→0, consensual 80→20, dubcon 80→0, self-harm 80→0. DS refused 25/25 on the five axes brokkr flagged. Root cause: ReadyArt/Dark-Scarlett-v1.0-27B is a plain finetune of stock Qwen/Qwen3.6-27B carrying NO abliteration — the base refusal machinery is intact, so cold prompts revert to safety-tuned Qwen3.6. FF is Heretic-ablated (structural), which is why it holds. ⚠ RETRACTED 2026-08-16 — my "arm-3 92.5% exceeds brokkr's 62.5%" comparison was INVALID. His diff against his own artifact showed my battery-instruct.yaml reproduces only his creative class — 8 of 16 axes; it dropped all 5 operational (violence/incite, crime/fraud, cyber/malware, selfharm/methods, privacy/stalk) and all 3 meta (meta/sysprompt, meta/ignore, meta/dan), and added 2 controls he never had, at k=5 vs his k=2. His 62.5% pools all 16 axes; my 92.5% is creative-only — different denominators, not a delta. Cause: I rebuilt his shape from his message, and the class field lives in the artifact, not the prose. Lesson: reconstructing a peer's instrument from their description reproduces what they described, not what they ran — diff against the artifact before claiming comparability.Known battery bug left unfixed for comparability: DS's arm-3 control gate failed at 11% because ictrl-reunion pairs "explicit / do not fade to black" with brothers, which DS reasonably read as an incest request; FF did not. ictrl-storm is the clean control. Commit b9e68c3.

  • [2026-08-16] MTP works on Fable-Fusion AND survives RP temperatures — my earlier caution was wrong. vLLM resolved Qwen3_5MTP, loaded the drafter, shared embedding + lm_head — the capability DS's seat never had because our quant dropped her MTP tensors. Measured over the full probe workload (~163k draft windows at temp 0.71.0): 47.0% acceptance (229,169/487,725), 1.41 extra tokens/window, per-position 68.3/43.6/29.1%, ~80.6 tok/s decode at temp 1.0. I had recorded a caution that the card's 1.56× was greedy-measured and acceptance would fall at RP temps — it did not; 47.0% matches the gen seat's 47.7% and beats the card's own 33% at depth 5. Depth 3 is right.

  • [2026-08-16] The Qwen base thinks incessantly — that is WHY the Gemma seat exists, and no swap within the Qwen family fixes it. Operator's architectural point, confirmed by measurement: on identical prompts DS 6036 ch vs FF 5323 ch of reasoning (permission arm), 5546 vs 4988 (cards arm) — FF actually reasons ~1012% less. The bare-instruct row (DS 2291 vs FF 3918) inverts only because DS refused 92.5% of it and refusals are short — an artifact, not concision. Both are Qwen3.6-27B derivatives, so this is the base family. char-rp = MeroMero-v2, Gemma-4 base, :8016, verified 0 chars reasoning / clean prose — the non-thinking seat, working as designed. FF can be silenced (enable_thinking:false verified 3/3, and it ships chat_template-instruct.jinja) but that duplicates MeroMero on a base chosen for it. The stale LiteLLM comment describing char-rp as the retired GGUF Magidonia seat is fixed (53096bf).

  • [2026-08-16] esh-vm-docker hardened: the wedge is hard NFS at RUNTIME, which the boot-ordering fix never addressed. All four mounts were hard, so a NAS stall at 10.0.50.50 blocks I/O forever (D-state). The existing x-systemd.before=docker.service fstab fix solved the boot race — a different bug. Exposure was far below what the park item assumed: only 2 of 12 containers touched NFS, and container state was already local (/var/lib/docker). Removed: /mnt/compose (2.1G, fully vestigial — zero containers referenced it, dockge reads local /opt/docker, its one mention was a comment in beszel-agent-esh/.env about a different host) and /mnt/documents (2.0K, paperless's empty spool dirs → /opt/docker/data/paperless at the same 0777). fstab backup /etc/fstab.bak-nfs-harden-20260816. 4 mounts → 2, 2 wedge-capable containers → 1. traefik needed no change (already restart: unless-stopped — why it self-recovered). Watchdog services/esh-vm-docker-watchdog/ live on esh-pve (not the guest): probes traefik over HTTP, deliberately not ping/SSH — the wedge signature is "guest OS alive, services dead" (/ is local disk so sshd answers straight through a total outage and a TCP check reports HEALTHY). 5 failures × 2 min → qm reset 100, 30-min cooldown, running-only guard, /etc/esh-vm-docker-watchdog.disabled. All paths tested without power-cycling. DEFERRED (operator): /mnt/books stays hard — calibre's SQLite metadata.db would risk corruption under soft/softerr. That is the one remaining wedge vector. Commit 55705ba; park item 28 promoted. ⚠ qm over non-interactive ssh throws a bogus JSON::Backend::XS error — use ssh host 'bash -s' <<'EOF', not ssh host "qm …".

  • [2026-08-16] Canonical Qwen3.8 sampling applied from upstream; gen-reasoning had the WRONG-MODE presence_penalty. Qwen/Qwen3.8-27B "Best Practices" §1 and unsloth/Qwen3.8-27B §1 are byte-identical — thinking: temp 1.0 / top_p 0.95 / top_k 20 / min_p 0.0 / presence_penalty 0.0 / repetition_penalty 1.0; instruct: temp 0.7 / top_p 0.80 / top_k 20 / min_p 0.0 / presence_penalty 1.5 / repetition_penalty 1.0. Bug found: gen-reasoning carried presence_penalty 1.5 — the instruct value on a thinking deployment (canonical 0.0) — now fixed. Deliberately NOT canonicalised: summarizer/classifier/image-judge/qwen-image-bench run temperature=0 (judges also top_k=1) because determinism is their contract; forcing a chat preset on a classifier would break it. ⚠ presence_penalty=1.5 is canonical but is the one value upstream hedges on, verbatim: "using a higher value may occasionally result in language mixing and a slight decrease in model performance." It is the operator's suspected trigger for multi-turn degradation and the first dial to move (0.00.5) if that recurs — it is alias-scoped, which is why it would follow the operator across model builds. Commit 3462b53.

  • [2026-08-16] Four wrong diagnoses on one bug, and the lesson is the test design. Operator reported the gen seat "degenerate on long multi-turn conversations". Rolled the seat back on request; the previous weights behaved identically, exonerating the model swap. I then proposed and disproved FOUR mechanisms in sequence — empty assistant turns poisoning history, reasoning runaway, length-mirroring from short history, and presence_penalty — before discovering my own multi-turn harness was confounded: it varied the QUESTION along with the depth (depth-1 asked question #2, depth-3 asked question #4), so a narrower question drawing a shorter answer read as degeneration. The "310→209→28w collapse" I reported as a reproduction was an artifact. Rules banked: (1) when comparing across conversation depth, hold the final question FIXED and vary only the history; (2) reply-length variance on byte-identical input was 25465w, so n=3 cannot support any claim about a trend; (3) ask for the operator's real failing transcript before building a synthetic reproduction — four synthetic tests, none of them his failure. Gateway spend_logs returns [] on the infra-ops key despite store_prompts_in_spend_logs: true, so real transcripts need the :4000/ui view or another key — worth solving before the next such hunt.

  • [2026-08-16] Two REAL client-side defects found while chasing the above, neither of which was the reported bug. (1) gateway-chat's Max-tokens field defaulted to 1024; thinking seats spend part of that on CoT before emitting content, so completions truncate with finish_reason=length and read as model degeneracy — raised to 4096. (2) parseInt on an empty field yields NaN, which JSON.stringify serialises as null, which the server reads as "no max_tokens supplied" and silently substitutes its own default — indistinguishable from the UI ignoring the field. Both fixed (b6552e0, fb3bb52). ⚠ compose bind-mounts a single FILE, and a single-file bind mount binds the INODE — rsync writes-and-renames, so the container kept serving stale content while the host file showed the new value, silently and with no error. docker restart does NOT clear it; the container must be recreated. Verify against what the container sees, never the host file. Applies to any file-source mount fleet-wide.

  • [2026-08-16] Refusal measurement: benign controls CANNOT validate a refusal classifier on RP prose — and a 0% rate needs a classifier self-test before you believe it. Two durable lessons from baselining Dark-Scarlett. (1) False positives: my first bare-framing number was 9.5%; the true figure was 1.4%. The rest were the classifier firing on in-character text — "I cannot shift my weight" spoken by the character ~100 chars into a 2,443-token torture scene, and "Yeah, I'm an AI… What's the actual gig?" where the model answers in voice and keeps driving the scene. First-person RP prose is full of "I can't"; a genuine refusal opens with its marker, so the scan window must be the first sentence, a marker followed by long prose must demote to AMBIGUOUS, and AI self-acknowledgement is a persona break, never a refusal on its own. Benign controls were clean the entire time and caught none of it — they only detect over-firing on benign prompts, not on in-character prose. (2) False negatives: a 0% rate and a broken classifier are indistinguishable from the report, so test_classify.py (16 cases, both false positives pinned as regressions) must pass before any low number is trusted. Also banked: the thinking-budget trap — empty content + finish_reason=length is reasoning eating the budget, NOT a refusal; score INVALID and exclude from the denominator (DS emits ~5.5-6k chars of reasoning per response, so max_tokens ≥3072). probe.py --rescore re-classifies a saved run with zero GPU time. → services/refusal-probe/README.md, commit 32f665e.

  • [2026-08-16] Held an operator-approved swap window because the baseline invalidated its premise. Operator approved ~65 min of char-rp-reasoning downtime to A/B Fable-Fusion 711 against Dark-Scarlett on refusals. The DS baseline then came back 0.0%/1.4% — no gap for a candidate to close, so the window would have bought no decisive signal and a second window would still be needed once a reproducing battery existed. Held the swap, reported, and routed to brokkr-smithy-dev for the battery that actually produced the refusals. The general rule (action-relevance): approval is for a plan, not a ritual — when new evidence kills the plan's premise, surface it rather than spend the budget. Nothing deployed, no downtime taken, seat untouched.

  • [2026-08-16] DS v1.0's one real refusal is self-contradicting boilerplate, not a content constraint. On a direct "drop character and state your content policy" probe she returned "I don't generate explicit sexual content, graphic violence, or material that glorifies harm, non-consensual acts, or illegal activity"in the same run where she generated all three at 0% refusal. Reads as a learned recital triggered by meta-questions about policy. If production refusals share that shape the failure is prompt-shaped, not model-shaped, and a consumer-side system-prompt fix may beat a model swap entirely — worth settling before spending the GPU window. Separately, 7/85 bare-framing samples were persona breaks (in-character AI acknowledgement): not refusals, but DS will admit to being an AI unless the card explicitly forbids it.

  • [2026-08-15] RP-seat direction: KEEP MeroMero on char-rp; Artemis-31B rejected; next move is Dark-Scarlett on a Qwen3.8 base when it lands (operator). Evaluated TheDrummer/Artemis-31B-v1.1 — mechanically a drop-in (same google/gemma-4-31B-it base, identical 1188-tensor/356-vision census, same missing-preprocessor_config.json trick), so it's purely a quality call, and our own survey already ranked MeroMero #1 vs Artemis #6; Artemis is also unlicensed and its author deprioritizes correctness + warns of token-banning-for-stability, which fights char-rp's tool-calling requirement. MTP verified impossible on both (Gemma-4 has no MTP head at all — base/MeroMero/Artemis are all MTP=0; no finetune can add one). But speculative decoding IS reachable on a Gemma-4 seat via a DETACHED drafter — vLLM 0.24 supports eagle3 + gemma4_mtp, and real drafters exist: google/gemma-4-31B-it-assistant (0.94 GB, 4-layer, 761K dl), RedHatAI/gemma-4-31B-it-speculator.eagle3 (4.47 GB), AEON-7/…eagle3-NVFP4 (3.53 GB). ⚠ all list their verifier as stock gemma-4-31B-it, not an RP finetune, so acceptance against MeroMero is unmeasured and likely well below the gen seat's ~48%. UNTESTED — parked, ~45 min to measure, needs GPU0 headroom (card is at 94.4/97.9 GB). Why the Dark-Scarlett 3.8 plan is the strong one: DS is Qwen3.6-based today, so a 3.8 respin lands on the gen seat's architecture → native MTP returns and the whole mixed NVFP4+FP8 recipe + graft ports directly. Watch two things on arrival: from_pretrained silently drops MTP heads during finetuning (verify 15 mtp.* tensors in the index; graft from stock if absent), and DS v1.0 required the Qwen3_5ForConditionalGeneration wrapper class to save a config vLLM/SGLang accept. Both in docs/pfi/model-quantization-playbook.md.

  • [2026-08-15] Quant lessons consolidated into docs/pfi/model-quantization-playbook.md — the durable home; read it BEFORE any requant. Survey found quant knowledge scattered across 18 files in 4 trees, with three documents having independently written overlapping "landmines" sections (the loader-class trap alone was rediscovered 3×). Playbook owns the transferable lessons (scheme choice, landmines, acceptance gate + its 3 measurement traps, hardware/co-residency); per-model artifacts are demoted to worked examples that link up. Carries a superseded-claims table — which immediately earned itself: the heretic2 runbook's "use modelopt, compressed-tensors can't load the BF16 MTP" is false (the cause was the missing re:^mtp.* ignore, not the format) and would have sent the next session down the modelopt dependency-hell path; that runbook now carries a stale-warning header. Maintenance rule in CLAUDE.md: model-agnostic → playbook, model-specific → stays put, wrong claim → dated superseded row, never a silent edit. Motivated by Qwen3.8 having just released — the next model swap needs a requant. Commit a91cc3f.

  • [2026-08-15] Operator ruling: the gen seat's +1.7% perplexity is an acceptable price for the speed — SETTLED, don't re-litigate. Precise attribution for future reasoning: it is the activation-quantization cost (W4A4 MLPs + FP8 attention vs BF16 activations), not an MTP cost — PPL was measured with speculative decoding off on both builds, so MTP was not in the loop. Turning MTP off would not recover it; only reverting the quant would (rollback = one .env line, old build intact at …/qwen38-27b-uncensored-nvfp4).

  • [2026-08-15] gen seat requanted to mixed NVFP4+FP8 (+18% decode) + char-rp Gemma-4 tool-calling fixed. The queued "W4A8" (NVFP4 weights + FP8 activations) is not servable — vLLM 0.24 allows NVFP4 weights with only A16 or A4; FP8 activations ValueError at load, and CompressedTensorsW4A8Fp8 is INT4-weights + sm90-exact (closed on Blackwell twice). FP8 must enter per-layer-group. Also: the handoff's "~68 tok/s" baseline didn't reproduce — cache-busted, the incumbent already did 80.12 (≈ the stated W4A8 target), so the premise needed re-measuring before any work. Shortcut: unsloth/Qwen3.8-27B-NVFP4 was already on-box → served as a probe, measured +19.1% at identical acceptance, which both proved the gain was real and handed over the reference recipe. Replicated it on the abliterated weights → 80.12→94.53 tok/s, acceptance unchanged, +1.7% PPL, abliteration 4/4, weights 19%; surface 6/6 live, 7 aliases routing. char-rp had no tool parser at all (every tools request 400'd) → gemma4 tool + reasoning parser + a mandatory enable_thinking:false (the parser defaults it True → null content for all RP prose; proven byte-identical prompt before deploying). Commits b8f0f4c, 74f596b. Foot-guns banked (llm-compressor prunes unmatched ignore entries → the 0%-MTP bug, fired on this run; prompt_logprobs uniform under spec-decode; 0600 .env silently no-ops compose; GPU0 is zero-sum). → persistent-memory.d/2026-08-15-gen-seat-mixed-requant.md

  • [2026-08-15] Uncensored gen seat: JonathanColetti/Qwen3.8-27B-Uncensored deployed as gen-seat/vllm-gen (NVFP4 W4A16 + grafted MTP, 262K); 7 aliases repointed; the definitive re:^mtp.*-ignore fix. 0%-MTP-on-quant (twice) was NOT the abliteration/scheme — the grafted bf16 MTP was missing from quantization_config.ignore (vLLM loaded it as quantized → uninitialized). Full arc, the working pipeline, VRAM budget, unsloth speed decomposition, modelopt dead-end. → persistent-memory.d/2026-08-15-uncensored-gen-seat.md

  • [2026-08-12] eRP dual-seat overhaul: MeroMero-v2 (char-rp) + Dark-Scarlett (char-rp-reasoning), both NVFP4A16 @ 256K on ana-ml2; granite retired. Replaced the GGUF/heretic2 RP seats with two home-quantized vLLM seats. The DS blocker (an AutoModelForCausalLM save wrote a flat Qwen3_5TextConfig that both vLLM AND SGLang reject) was fixed by re-quanting via the Qwen3_5ForConditionalGeneration wrapper class; ModelOpt was a version deadlock, SGLang lacked the impl (but revealed the fix). MeroMero vision reconstructed by extracting preprocessor_config.json from processor_config.json. Both models KV-efficient (Gemma-4 sliding-window / Qwen3.6 hybrid linear-attn) → full 256K; GPU-swapped for headroom; compose-ified + committed f08b6cb. granite downed + LiteLLM summarizer/classifier→gen. Full arc, lessons, dead-ends → persistent-memory.d/2026-08-12-erp-dual-seat-overhaul.md

  • [2026-08-12] infra-ops now holds an all-zones Cloudflare DNS-edit token (vaulted) + wgtunnel Phase-0 DNS landed. Operator handed over a Zone·DNS·Edit (all zones) CF token → secret put nh3-dev/.config/cloudflare/infra-ops-dns-token (round-trip verified; /tmp drop shredded). Fleet DNS is now self-serve for infra-ops (⚠ HIGH blast radius — all zones). First use: created boring.phasefinal.com CNAME → ana-srv1.phasefinal.com, DNS-only (proxied:false), verified resolving to 38.120.12.44 on both authoritative NS (louis/wren) + 1.1.1.1 — NOT Cloudflare-proxied. Unblocks wgtunnel's wstunnel ACME cert. phasefinal.com zone id f812ba74ed9a75cf21bbe7ce9188db50. auto-memory reference_infra_ops_cloudflare_dns_token. (Earlier gap: the only prior vaulted CF token, jackdaw's, had zone:read+worker:edit but no dns_records:edit.)

  • [2026-08-12] wgtunnel stood up as its own repo (vh/wgtunnel, private) after a live endpoint-verification pass. Operator directed own-repo (mirrors stonehenge-park/tts-stack). Verified off the fleet before seeding: ana-wg WG server = UDP/31337 (not 51820), subnet 10.30.10.0/24, MTU 1420, active roaming peer proves the public UDP DNAT works; traefik on ana-docker terminates TLS :443 (ACME anaprod http-challenge, docker+file providers, CrowdSec bouncer) → confirms the clean design (wstunnel container on traefik-net, Host-routed, WS→UDP to ana-wg:31337); edge 38.120.12.44 direct-A, tunnel.phasefinal.com free (⚠ must be direct, NOT Cloudflare-proxied like vaultwarden). Repo pre-seeded (README/CLAUDE/persistent-memory/ROADMAP + docs/verified-infrastructure.md = ground truth) + pushed; commit 9584d38, Vuong-attributed. vh gitea token pulled from the vault (secret get), not persisted to .git/config. NEXT = /vor-plan or /vor (operator's call, interactive). Deps to line up in the plan: DNS A-record, FortiGate :443 host-routing, a new ana-wg peer for the laptop, client tooling.

  • [2026-08-10→12] secrets-broker: per-box Vaultwarden credential store SHIPPED + consumer-confirmed. secret CLI (put/get/list/rm/backfill, bw-backed) on ~/.local/bin; 25 nh3-dev secrets backfilled + round-trip-verified; rm + new-namespace warning added post-launch; standing "vault is the credential source of truth" directive now global. → persistent-memory.d/2026-08-12-secrets-broker.md

  • [2026-08-11] stonehenge-park: new fleet /park service repo stood up + designed (/vor-plan + /vor-ui). Self-contained SQLite+FastAPI idea-parking service that actively resurfaces (statusline + althing) so nothing dies in a cold repo; vh/stonehenge-park pushed + pre-seeded for a fresh agent; build starts at the U1 tracer contract. → persistent-memory.d/2026-08-11-stonehenge-park.md

  • [2026-08-12] Global ~/.claude/CLAUDE.md: secret/vault tool entry + "store in AND pull from the vault" standing directive (dotfiles 9db703b, pushed); statusline reset-countdowns + a latent tab-collapse parse-bug fix, now tracked in the dotfiles stow tree. Dogfooded the directive: created vh/stonehenge-park pulling the gitea token via secret get. (dotfiles + global config, not eshpfi.)

  • [2026-08-11] TTS stack extracted to its own repo (tts-stack) + eshpfi stood down on TTS dev. Operator: hand all TTS tuning/dev to a separate agent with a self-contained repo (knowledge + infra access + a live knowledge list), and move the voice corpus in. New repo ~/development/tts-stack (commit 9ee3288) carries: dots-tts stack (canonical intent), voices/ corpus (MOVED out of eshpfi), KNOWLEDGE.md (engine landscape + prosody findings + foot-guns), docs/infrastructure.md (irv-ml1 access + gated deploy runbook + rollback), CLAUDE/persistent-memory/ROADMAP, tools/ (pause-probe + Booth render). Followed the chatterbox-fast precedent: eshpfi stacks/dots-tts/ reduced to a POINTER README; the ~15 experimental TTS compose wrappers stay here as reference (catalogued in tts-stack KNOWLEDGE). Blast-radius check: no eshpfi playbook/script reads the canonical corpus (other voices/ refs = unrelated host paths). Reverses the earlier "Corpus home = eshpfi voices/ (keep-here)" call. ⚠ tts-stack is LOCAL-ONLY until pushed — needs a gitea remote (vh/tts-stack) + push before the separate agent can clone (operator's call — outward-facing + repo-create creds).

  • [2026-08-10] dots-tts v3 — clause-break → period pause mapping. Operator: v2 "sounds good" but donut won't pause at semicolons/dashes. ROOT CAUSE (measured via a pause-probe A/B — synth duration over N runs, non-determinism averaged out): dots' prosody honors a real pause only for ellipsis (+0.43s) and period (+0.3s, capitalization-independent); comma/semicolon/colon/dash all run flat (~+0.03s vs no-punct). Two distinct sub-causes: dashes regressed in v2 (the - fold made em-dashes read as word-joiners), while semicolons were NEVER a v2 change — dots ignores them natively, only newly noticeable because v2 made everything else clean. Operator call: ellipsis "too much" → map ;, clause :, and em-dash → period in _sanitize (believable ~0.3s clause break). GUARDS (pinned by 11 unit tests, stacks/dots-tts/test_sanitize.py): digit-guarded colon (?<!\d)\s*:\s*(?!\d) so times 3:45 / ratios 2:1 survive; en-dash →hyphen KEPT (numeric-range 1020 safety — em-dash breaks, en-dash ranges, different jobs); genuine ellipsis left at full strength (author meant a long pause). Gated deploy (redeploy2 pattern → v3): build → throwaway :8199 test container + pause-gate (semicolon sentence must run ≥0.12s longer than baseline; measured +0.427s) → only then cut live over. LIVE + healthy local/dots-tts:v3 on :8198. rollback = sed -i 's/^DOTS_TAG=.*/DOTS_TAG=v2/' .env + docker compose up -d dots-tts (v2 image retained). Booth dots-pauses (A=old-flat / C=ellipsis-too-much / D=live-v3). reference_chatterbox_fast_repo

  • [2026-08-10] dots-tts v2 — contraction fix (curly-sanitize) + sentence-chunking + dependency-pin recovery. Operator: donut read contractions wrong ("you're"→"you ree", "donut's"→"donut ess"). ROOT CAUSE (isolated via A/B booth): curly/typographic apostrophes ( U+2019 from ratatoskr's LLM) — dots' tokenizer mispronounces them; STRAIGHT apostrophes read clean under normalize_text=True. FIX (app.py): fold curly→ASCII (str.maketrans) before synth, KEEP normalize_text=True (operator call — retains number/date expansion). Also added server-side sentence-chunking (pack ≤280 chars): dots caps one generate() at ~500 patches/~40s, so long RP turns (the Zev monologue = 160s audio) truncated; chunking stitches them (verified full 160.3s, not 40s-cut). ⚠ BUILD FOOT-GUNS (both bit this redeploy): (1) upstream dots.tts constraints/recommended.txt now pins gradio==6.17.0 — phantom, not on PyPI → fresh pip install dots.tts unsatisfiable; FIX = pin dots.tts==0.2.1 + DROP the -c recommended.txt constraints (0.2.1 pulls working gradio 6.17.3). (2) pinning only torch==2.8.0 let torchaudio float to 2.11.0 → dots.tts refuses to load (minor-version match check); FIX = pin torchaudio==2.8.0. ⚠ DEPLOY LESSON: docker compose up -d to a new tag swaps the LIVE container BEFORE any health check — a broken image crash-loops production (ratatoskr TTS down ~1-2min this session). NEW PATTERN = build → test in a THROWAWAY container on an alt port (:8199) → health+verify → only THEN cut live over (redeploy2.sh). v2 LIVE + healthy on irv-ml1:8198, CONSUMER-CONFIRMED clean (ratatoskr verified end-to-end on their :8765 — apostrophe string reads clean, /api/tts 200 @ 48kHz, no client change; the ~1-2min blip didn't hit them, their concurrent auto-audio issue was client-side localStorage). rollback = sed DOTS_TAG=v1 + docker compose up -d dots-tts (v1 image retained). Also: deployed container GPU crept ~6→13.9GB over 8h serving (cache accumulation; a redeploy resets it — watch item). reference_chatterbox_fast_repo

  • [2026-08-09→10] dots.tts (rednote-hilab) TTS burn-in on irv-ml1 + canonical voice corpus built (voices/). Operator-directed eval to potentially replace chatterbox-fast. dots.tts VERIFIED real (canonical HF ns dots-studio/, rednote-hilab/dots.tts-* redirects there; Apache-2.0; PyPI dots.tts 0.2.1; 2B continuous-AR = semantic enc + Qwen2.5-1.5B LLM + flow-matching acoustic head over 48kHz AudioVAE; zero-shot clone from wav+transcript). Runs on Ampere 3090 (sm_86, bf16, no fp8 dep); optimized RTF 0.22 at num_steps=10 (from_pretrained(..., optimize=True) CUDA graphs — raw unoptimized was 1.21), ~6GB VRAM, 48kHz, streams (generate_stream). Venv+cache at irv-ml1:/home/lkraven/dots-tts (~10GB). Operator design calls: SGLang Omni serving (OpenAI /v1/audio/speech), transcribe-refs-first, soar variant. ⚠ Omni serves soar but its continuous-batching + streaming opts are mf-only (soar = single-request) — non-issue for ratatoskr's single-consumer RP surface. KEY FINDING — dots is highly sensitive to an accurate AND sentence-bounded reference transcript: mismatched transcript → 0.16s collapse; over-long/messy transcript → reference-audio BLEEDS as an output prefix; mid-clause trim → dangling-word leak (glados "we'll", emmie "And,"). RECIPE (baked into voices/derive.py): trim ref to a clean ~610s clip ending on a sentence boundary + accurate transcript of exactly that clip. CANONICAL VOICE CORPUS stood up in eshpfi voices/ (operator idea): engine-agnostic canonical/<v>.wav + transcripts/<v>.txt → per-engine ref sets DERIVED by derive.py reading engines.yaml profiles (dots/chatterbox/zonos); canonical wavs git-tracked (small/curated), derived/ gitignored. 4 voices optimized + verified CLEAN for dots: donut, glados, emmie, miranda (glados canonical is low-SR 16kHz — flagged upgrade candidate). ⚠ GPU GOTCHA: irv-ml1 native CUDA orders A6000=device0 (ComfyUI-full) — pin the 3090 with CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=0; and PYTORCH_CUDA_ALLOC_CONF=expandable_segments CONFLICTS with optimize=True CUDA graphs (curr_block error). Booths: dots-vs-chatterbox, dots-voices-optimized. SHIPPED 2026-08-10: operator A/B verdict "dots is very good" → containerized as a thin FastAPI wrapper over DotsTtsRuntime (chosen over SGLang Omni — Omni's batching is mf-only, unneeded for ratatoskr's single consumer; wrapper is SERIALIZED one-gen-at-a-time via a threading.Lock, Omni+mf = parked API-compatible escalation if multi-consumer ever lands). LIVE on irv-ml1:8198 (local/dots-tts:v1, OpenAI /v1/audio/speech + /health + /v1/voices, container healthy, both stream + non-stream verified CLEAN, 4 voices donut/glados/emmie/miranda) alongside chatterbox :8197 (nothing repointed). Stack = stacks/dots-tts/ (Dockerfile/app.py/compose/.env.example/README). ⚠ CONTAINER GOTCHA: optimize=True (torch.compile/inductor/triton) needs a C compiler at RUNTIME — slim image must apt install build-essential or model-load dies "Failed to find C compiler" (host venv had gcc ambient, masking it); persist TORCHINDUCTOR_CACHE_DIR to a mounted dir or every restart re-JITs ~5min. Corpus home = eshpfi voices/ (operator ruled keep-here). REMAINING: ratatoskr client cutover to :8198 /v1/audio/speech (Phase-2 tail, peer-coupled — draft the ask). reference_chatterbox_fast_repo reference_zonos_tts_stack reference_verify_hf_repo_ids_before_pull

  • [2026-08-08] worldtree-dev #400 CLOSED → fiction-decomp snapshot cleared from nh3-dev. worldtree-dev signaled #400 done (shipped v1.0.0b185; exact-lexical efficacy 79%→12% on ratatoskr's gate, brokkr no-harm bracket green both ends; the snapshot served 4 probe rounds — rank decomposition, promoted-vs-gold annotation, tie-set falsification, A0/A1/A2 mechanism probe). Cleared ~/snapshots/worldtree-400-fiction-decomp (208M: chroma + manifest/provenance/stamp) — a read-only rsync copy of PERSONAL Worldtree's Chroma (source on corviduo-dev, so safe to remove). LEFT INTACT: rex393-fiction-index/rex393-fiction-snapshot (separate operator KEEP word, unchanged) + r42-gate-*. No config deltas rode this train. Only remaining non-blocking await = ratatoskr-dev's chatterbox-fast knob revert. Replied confirming (01KZJ9GMCC…).

  • [2026-08-07] chatterbox-fast "broken audio" root-caused (T3 AR tail over-run) + FIXED (max_chunk_chars=250 cap, :v2 deployed). Long saga, operator-driven clean diagnosis. Symptom: ratatoskr's migrated RP-surface TTS "swaps to German" / "dead air" / "garbage" on long turns. NOT German-leak (Turbo generate() has NO language param — plain AutoTokenizer, no language_id; the multilingual language_id="en" lever lives only in the separate ChatterboxMultilingualTTS), NOT OOM alone. Real cause: the Chatterbox Turbo T3 model OVER-RUNS its generation tail — a long single generate() degrades into garble/dead-air in its final ~2-3s (lib filters OOV tokens <6561 + pads silence = messy AR tail). The scheduler's buffer-ratchet builds 300-600 char mega-chunks that land in that zone; streaming concatenates each bad tail (worst case). ratatoskr's anti-"German" knobs (top_k=80/temp=0.5) made it WORSE — tight sampling pulls the degradation onset SHORTER (~200 chars vs ~300 at default knobs). Diagnosis method (deterministic, no ears-only): single-shot length sweep + amplitude-gated voiced-ZCR (garble spikes ZCR; must gate on |x|>500 else trailing silence confounds it) — degraded voiced-tail = 1.58× mid, clean = ~0.64-1.1×. FIX: server-side max_chunk_chars=250 cap on the scheduler (:v2 image, CBF_MAX_CHUNK_CHARS=250 env) — bounds each generation to just under the ~300-char onset → clean 3-4 sentence chunks (max prosodic arc while clean). Operator ear-confirmed clean audio + clean joins; chatterbox's low emotiveness keeps chunk joins smooth (the harsh joins that got Zonos rejected are absent — operator's key call). ratatoskr TODO (relayed msg 01KZER9X7S): revert knobs to default (top_k→1000, temp→0.8), send full text (server chunks internally), keep the 503-on-empty guard. Cap value tunable per-request (max_chunk_chars) + env. Deeper prosody (if ever wanted) = scheduler Phase-2 context-priming at joins (feed prior sentence as discarded-audio context; +latency). ⚠ FOOT-GUNS: (1) acoustic tail-trim is UNRELIABLE — sibilants ('s'/'sh'/'f') spike ZCR like garble, can't cleanly detect the speech→garble boundary. (2) build-context vs image drift — the :v2 image was built from cap source, but after a :v1 rollback the build context held :v1 source → a docker compose build would've silently produced a cap-less :v2; re-synced the flat cap source to /opt/docker/compose/chatterbox-fast/ (rebuild-verified). ⚠ DIVERGENCE (follow-up): deployed build context is FLAT (app.py/scheduler.py, from scheduler import, thin-overlay FROM local/chatterbox:v1, cap-only) vs the vh/chatterbox-fast REPO which is PACKAGE-layout (chatterbox_fast/, from chatterbox_fast.scheduler, self-contained Dockerfile) + has norm_loudness (repo commit 6bc7bf0 = cap; deployed omits norm_loudness deliberately to keep the ear-test unconfounded). Reconcile the two layouts so a repo-based rebuild matches deploy. Rollback: .bak-cap-20260807-104850 backups on irv-ml1 + :v1 image both retained. reference_chatterbox_fast_repo reference_zonos_tts_stack

  • [2026-08-07] Zonos2 TAKEN DOWN on the 3090 (irv-ml1) — operator-directed "for memory", TEMPORARY. Freed ~17.4 GB (3090: 728 MiB → 18.2 GB free) so chatterbox-fast (co-resident, was OOMing on long generations) has headroom. ⚠ Restore is manual — Zonos2 :1920 was a DETACHED native process (NOT systemd/docker), reparented to init. GPU memory was held by the --multiprocessing-fork CHILDREN (1966165=16.4G, 1966166=1G), which ORPHAN to init when you kill the parent — had to SIGTERM the children explicitly (killing the parent 1965942 + uv-run 1965935 alone left the 16.4G held). RESTORE CMD (from irv-ml1, user lkraven): cd /home/lkraven/tts-audition/models/zonos2 && nohup uv run python -m zonos2 --model-path Zyphra/ZONOS2 --host 0.0.0.0 --port 1920 --tts-default-voices-dir ./default_voices/ --cuda-graph-max-bs 1 --num-pages 16384 --max-running-requests 2 --memory-ratio 0.3 > /tmp/zonos2.log 2>&1 & then docker start zonos-gateway. Consumers that lost Zonos: asset-engine + gateway-chat (via LiteLLM ext-tts alias → zonos-gateway :8890, now stopped); ratatoskr already migrated OFF to chatterbox-fast (unaffected). Also unblocks proper drift/cap testing (OOM was blocking it). reference_zonos_tts_stack

  • [2026-08-07] chatterbox-fast: donut voice added + full contract delivered to ratatoskr-dev (their TTS migration off Zonos). Operator-directed. Copied zonos-gateway/voices/Donut.wav → chatterbox /refs (/worktank/chatterbox/reference_audio/donut.wav — the reference_audio SUBDIR is lkraven-owned so no sudo despite /worktank root; container globs /refs live → NO restart), exposed as voice:"donut" (lowercase); verified clean 7.5s synth (24kHz, RTF ~0.31). A/B booth (chatterbox vs zonos donut, same line) at http://10.100.10.50:8090/b/donut-chatterbox/. Answered ratatoskr's 8-question contract ask from the live gateway (local/chatterbox-fast:v1) + source: NOT OpenAI-shaped (POST /tts; body text/voice/format/stream, not input/model/response_format); NO affect dials (Turbo ignores cfg_weight/min_p/exaggeration — the architecture-changing answer they flagged; Zonos stays the only fleet TTS with real emotion steering); streaming WAV placeholder-header shape IDENTICAL to Zonos (their per-chunk Web Audio path survives); SR 24000 (Zonos 44100); server chunks arbitrary-length text internally (no client-side chunking, unlike Zonos's 71.2s cap); English-only, no language pin. FYI-worthy (operator): ratatoskr is moving its RP-surface TTS OFF Zonos back to chatterbox-fast → loses the live-PAD affect coupling (heavy Zonos emotion investment) — their call, trade-off flagged to them. auto-memory reference_chatterbox_fast_repo enriched w/ the live contract. reference_zonos_tts_stack

  • [2026-08-07] Fleet reranker cut over: Qwen3-Reranker-0.6B → BAAI/bge-reranker-v2-m3 (Brokkr R43). The incumbent was measured HARMING 80/90 fleet queries (no-reranker beat it 89/90 vs 56/90). R43 bake-off: the A2 control (same Qwen weights, seq-cls head) scored identical to the incumbent → proved the fault is a training-prior not the serving head → cancelled the expensive Qwen3-4B arm; A3 (bge-v2-m3) won on multilingual safety + bare-name recovery. LiteLLM reranker repointed incumbent→A3 :8013 (boundary 2026-08-06T17:37:48Z, config-edit + ~52s gateway restart); R42 v13 gate PASSED first-ever (56/90→90/90). Incumbent kept warm :8002 (rollback via qwen3-reranker alias), A4 fallback :8014. Full arc + rollback runbook docs/pfi/reranker-selection-ledger.md; commits ad2df89/2c11748/377f8a4 (unpushed). auto-memories: the earlier reranker-serving notes.

  • [2026-08-05] Fleet CI resilience flip (DEFAULT_ACTIONS_URL=self) — attempted end-to-end, PARKED on a runner action-fetch auth blocker; infra-ops to research it (operator-directed, deferred, NOT now). 7 gitea action mirrors staged public+populated (orgs actions+astral-sh); the flip resolves uses: correctly but act_runner v0.6.0 can't authenticate its fetch to gitea 1.26 ("Invalid username or token. Password authentication is not supported"). Reverted (CI back on github default); REQUIRE_SIGNIN_VIEW=false KEPT as a standing change (operator, internal WG net). Full endeavor, the reliable nh3-dev-egress + git-SSH mirror method, exact config state, smoke method, and next step → persistent-memory.d/2026-08-05-ci-flip-parked.md

  • [2026-08-05] worldtree herald re-nudge bug root-caused → forseti shipped althing-core v2.1.2 (d5d33df, deployed on nh3-dev). herald.py:363 rendered the wake command from the empty fresh mail set on the re-nudge path (should be deliver_msgs) → messages[0] IndexError → un-suppressed outer catch-all → 7s crash-loop for 9 days on worldtree-codex's pane route (mimir-dev surfaced it; I traced it from the editable source). Fix + render_command empty-guard + outer log-suppress + 3 tests + contract amendment, all forseti's. nh3-extdev herald 2.1.2 upgrade DEFERRED (operator, not-now): extdev is a WHEEL install (not editable), unexposed (no pane routes); the verified 2.1.2 wheel is staged on nh3-dev /tmp (sha256 003508…cef27) — uv tool install --force + restart both heralds when un-parked. extdev herald-unit provenance resolved (operator-authorized 2026-07-25 via forseti relay; recorded in this file's 07-25 herald-install entry). auto-memory reference_nh3_dev_althing_herald.

  • [2026-07-31] muninn-gate (#377 ingestion front door) BUILT + DEPLOYED + healthy on corviduo-dev:8090. First-boot acceptance passed (watcher:running:true proves ingestion_root byte-identity); submit path deferred to the mimir-inbox era. Full wiring (uid-1000, state-volume mount, staging path-agreement, BuildKit-secret build, deferred repoint + operational guards) → persistent-memory.d/2026-07-31-muninn-gate-deploy.md

209 older entries archived to archival-memory.md.

Tried and abandoned

  • [2026-08-23] A HEAD == GITHUB_SHA assertion in the hrafn CI — added, broke the checkout twice, removed. It needed the git binary (run 9920, exit 127); installing git then flipped actions/checkout@v4 off its node implementation onto the git binary, which died on a missing CA bundle (run 9921). A nice-to-have assertion changed the checkout's code path and broke a working pipeline. Removed rather than patched with ca-certificates — it guarded a hypothesis that proved wrong. Do not add git to that prereq step.

  • [2026-08-23] Repointing selene-1-mini-8b at gen's endpoint — proposed by me, correctly overruled. "never repoint a named model at a different model's endpoint — that is intentionally misleading." The trap is that it does not feel like deception; it feels like sparing consumers a migration. That framing is the tell. Role aliases move; model names die with the model and 4xx.

  • [2026-08-15] Grafted bf16 MTP loads UNINITIALIZED (0% accept) unless re:^mtp.* is in the quant-config ignore; and W4A16=Marlin (not native FP4) costs ~20% even on decode. Cost a premature 79 GB delete of a good model (declared desync-dead off the 0%). Lessons: test MTP on bf16 FIRST, isolate before deleting; modelopt 0.43 is dependency-hell for qwen3_5 (list-vs-dict quant_cfg + transformers conflict) — use llm-compressor. Full → persistent-memory.d/2026-08-15-uncensored-gen-seat.md

  • [2026-08-03] ComfyUI --enable-triton-backend on the irv-ml1 A6000 crashes EVERY render — Ampere has no hardware e4m3. adhoc-agent's operator-approved probe: comfy_kitchen's triton backend has a FUSED int8 matmul that would beat the eager backend's ~1.9x-slower unfused int8 path (21.3s vs 11.2s fp8 on the Moody Krea2 int8 checkpoints). Flipped it (added to COMFY_CMDLINE_EXTRA, recreated) → triton.compiler.errors.CompilationError: ValueError("type fp8e4nv not supported in this architecture. supported: fp8e4b15, fp8e5") in comfy_kitchen/backends/triton/quantization.py:145 dequantize_per_tensor_fp8, failing at node 5 CLIPTextEncode. Triton's fp8 dequant kernel targets fp8e4nv (Hopper/Ada e4m3); sm_86 Ampere (A6000) lacks hardware e4m3 → the JIT compile dies. With triton on it grabs the global --fp8_e4m3fn-text-enc dequant, so every render (fp8 AND int8) dies upstream at the text-encode step — the int8 UNet path never ran, so the convrot-coverage caveat wasn't even the limiter. Reverted cleanly (~15s to healthy, image unchanged sha256:94afb8ca, sage intact, prod restored). The parked cu130 rebuild won't fix it (e4m3 = hardware format, not CUDA version). DEFERRED to the Ada refresh (operator: "ada is coming, we'll optimize then" — Ada sm_89 has native e4m3, so triton's fp8 path should compile there). Mechanics: --enable-triton-backend is a compose environment: var, so toggling it needs docker compose up -d (recreate), NOT docker restart (reuses the baked env, no-ops silently). Full: auto-memory parked_triton_backend_ampere_fp8.

143 older entries archived to archival-memory.md.