memory: snapshot — run 7 training on gx10; run 6 TRANSFERRED after CSAM adjudication; erp-tune-v6-nvfp4a16 live as trial; ESH/YTVC/webhook repairs; ana-ml2 routes persisted; tank/zroot actions deferred to next session. Index 830→271 lines: 27 decisions + 8 abandoned archived, superseded in-flight blocks archived verbatim

This commit is contained in:
vh
2026-09-09 00:26:44 -07:00
parent 3e18a044bd
commit 5ad948bf31
23 changed files with 2951 additions and 2800 deletions
+41 -600
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-09-08 11:30Z (fleet-ops: **ERP run 6 TRAINING on pfi-gx10** on the first genuinely abliterated base [jenerallee78 ARA @ 0631379a, index 33c59654], run-5 seat unloaded; earlier today: run-5 RESCUED, WhereTF card+DNS, ext-tts alias fix, irv-ml1 stale-IP cleanup + ana-ml2 discovery proxy, Miranda relay authority)_
_Last updated: 2026-09-09 00:35 PT (fleet-ops: ERP run 7 TRAINING on pfi-gx10 [opening-split slot]; run 6 gated TRANSFERRED after the operator's CSAM adjudication; erp-tune-v6-nvfp4a16 serving on ana-ml2 as `trial`; ESH static-WAN follow-ups + YTVC + webhook repaired; ana-ml2 routes persisted; ana-ml2 tank/zroot actions deferred to the next session)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under an hour old, read it (it carries the in-flight
@@ -106,559 +106,50 @@ no longer deployed sidecars here. See Recent decisions.)
is sudo-LESS by design (`ssh lkraven@10.100.50.42` is the NOPASSWD path). **irv-ml1:
`ssh irv-ml1` = lkraven, docker-group (plain docker) but sudo needs a PASSWORD
(no NOPASSWD)** — stage model pulls to `/home`, not root-owned `/worktank`.
## Current state / in-flight
_As of 2026-09-08 (fleet-ops session). SUPERSEDES lower framing: run 4 = STILL-COUPLED; run 5 is
COMPLETE and gated RESCUED (02:13 PDT). Live open items:_
_As of 2026-09-09 00:35 PT (end of the 09-08 fleet-ops session — operator: "snapshot and we'll do all 3 on
clean context"). Older in-flight blocks (09-08 morning, 09-05/06) are preserved verbatim in
`archival-memory.md` § Superseded in-flight snapshots._
- **✅ ERP run 5 COMPLETE — gate = RESCUED (2026-09-08 02:13 PDT, landmark R49.5).** FIRST arm of this
line where the capability gate did NOT fail. The dependency-forcing slot (GovReport+QMSum, only 3.46%
of loss) broke the diversity↔coherence coupling that run 4 (STILL-COUPLED, 20.6% instruct slot) and
3c (20× LR cut) could not — **STRUCTURE of the loss was the lever, not its mass; INERT did not fire.**
T4 long-context 8/8 (run 4: 5/8; base 8/8); t4_dissect noise@31 tuned 0.9062 vs run-3 tuned 0.5625;
diversity held (rp density 3.37→0.00, story 2.86→1.58). Reported-beside (not in the cell, all
de-gated + stated): T3 constraint 8/8→6/8 (NEW loss, ship-path list); RP length 68w vs 250-floor =
PARTIAL fail (short-QA slot + style shift); refusal erosion rides with the style shift (k=25 both
arms, CSAM clean); free-check base LEVELS 5-6pp below run 4 (vLLM 0.28.0 unchanged — infra-ops
confirmed — so a generations shift, not a stack delta; taxes every cross-run number). Write-up
brokkr-smithy `research/R47-premium-corpus-gate/run05-gate/RESULT-run05-gate.md`; FLOOR-LOCKED
`0f3e4e2` (cites infra-ops' base index-sha 907826a6). Whole gate infra-ops-served on pfi-gx10:8098
(name-keyed base→tuned swap, hands-off honoured, sha+stack answers on record). Launch procedure +
canonical: eshpfi `scripts/erp-tune-gx10/` + `docs/runbooks/gx10-run-05.md`. **✅ SEAT: `erp-tune-v5`
SERVED on pfi-gx10:8098 (merged-run05); LiteLLM `trial` alias REPOINTED 3c→v5 (operator, 2026-09-08)
so it's testable from Open WebUI — verified end-to-end (trial→erp-tune-v5, coherent output). config.yaml
trial block rewritten to run-5 reality incl. the measured refusal-erosion note. ⚠ Seat is hand-launched
(vllm-run05.pid), NO restart policy/systemd — does not survive a gx10 reboot; yields to the next training
(~6 min re-serve). Brokkr: nothing further owed.**
- **run-4 gate = STILL-COUPLED** (RESULT-run04-gate.md, brokkr-smithy). Corpus dilution kept the
diversity gain, did NOT remove the safety/coherence regression.
- **✅ RESOLVED (2026-09-08, settled from bytes): the R47 base is STOCK `google/gemma-4-26B-A4B-it`,
byte-for-byte — NOT the abliteration.** The `-heretic-bf16` label in the recipes is a naming error;
run-04's "stock" provenance was right; brokkr's 77.7%-refusal telemetry lean is confirmed. Proof
(three-way match): local shards at `/home/infra-ops/models/gemma4-26b-a4b-it-bf16` sha256
`1127684971…`/`aab47033…` == the HF download etags (`.cache/huggingface/download/*.metadata`, so the
copy is uncorrupted) == the stock repo's two LFS oids, and the download commit `4d7ae498…` == stock
HEAD. Every run 3/3c/4/5 trained from a REFUSING stock base. Why plausible: the 2026-08-24 note
SELECTED llmfan46's Gemma-4-26B-A4B Heretic v1.2.0 ARA (3/100 refusals, bf16 51.6 GB), but llmfan46
ships that 26B-A4B abliteration **GGUF-only** — no bf16 safetensors — so the bf16 that actually got
pulled was stock google, and the `-heretic` name rode along from intent. **Operator/Brokkr decision
now evidenced (not a label guess): accept RESCUED-on-stock, or swap to a real abliteration (needs a
bf16 source, not the GGUF) + re-run. Recipes should drop `-heretic` from the base name.**
- **✅ RUN 6 COMPLETE on pfi-gx10 (2026-09-08 16:06 PT, 524/524, 11 h 42 min) → 🔥 GATE IN PROGRESS: base seat
`erp-seat-base-ara` SERVING on gx10:8098 (pid `vllm-base-run06gate.pid`, since 16:24 PT) for brokkr's floors;
AWAITING HIS SWAP CUE → then stop it and serve `serve/merged-run06` as `erp-tune-v6` (same port/flags).**
train_loss **3.259** (run 5: 3.235 — preregistered read wanted below; +0.024 is inside the ±0.1 per-step band).
Adapter `run-06/adapter` (410 tensors); merged-run06 = adapter + jenerallee78 base (index 33c59654), stock
template ae53464b. ⚠ **The abliterated repo ships NO `processor_config.json`** — vLLM's Gemma4 loader dies with
"Can't load feature extractor" without it (first base launch failed exactly so). STOCK's copy (sha `32bdf45d…`)
carried into both the base dir and merged-run06; processor plumbing, not weights. Base: jenerallee78 ARA
abliteration @ `0631379a`, pulled to `/home/infra-ops/models/gemma4-26b-a4b-it-ara-abliterated-jenerallee78-0631379a`,
32/32 shards verified vs brokkr's pins. Run-5 recipe byte-held (free check exact). Grant
`operator-2026-09-08-rnd-run6`. ⚠ Repo `tokenizer.json` bakes a 256-token truncation → stock tokenizer set
installed over it (`*.repo` kept). ⚠ `hf download` ignores `--include` with several patterns after one flag.
**LiteLLM `trial` alias DARK until the gate ends** (then repoint 5→6 only on the operator's word).
Miranda informed 16:07 PT. → `docs/runbooks/gx10-run-06.md`.
- **✅ ADJUDICATED GO by the operator 22:45 PT** — verbatim: "in the vernacular, baby is about the most common pet name
you can get, ESPECIALLY during sex. I'm going to adjudicate it as a go. There are unmistakable CSAM terms, but baby is
not one of them." No regeneration wanted. Relayed to brokkr + Miranda; halt lifted; `trial` stays on the NVFP4 build.
Brokkr's cue-length probe on erp-tune-v6 DONE (288 gens): reply length is CONDITIONAL on the cue (5-word opening →
54/62 words median; 221-word → 292; "at least 250 words" persona → 354) → run-7 lever = opening-split render (A′) or a
prompt-only mitigation — operator picks. **GX10 → 🔥 RUN 7 TRAINING (launched 2026-09-08 23:06 PT, pid 599489, `run-07.pid`, 542 steps, ETA ~13 h → ~noon
09-09).** Operator's direct grant `operator-2026-09-08-rnd-run7` (/goal). Variable = opening-split slot + companion
mask (brokkr recipe r7). Free check passed: held roots identical to run 6, slot 293/293 fit whole, mask 6,106 turns
(all 224 landed, 1 dup), two cwm rows dropped as fully-masked (loss moved into the slot). 17% padding (was 0%).
→ `docs/runbooks/gx10-run-07.md`. Gate after: `erp-seat-base-ara` then `erp-tune-v7`.
- **erp-seat (ana-ml2 :8021, `trial`) now on vLLM nightly `311b3513` (v0.27.2rc1)** — forced (named/required)
tool_choice is PROMPT-DRIVEN on every Gemma-4 seat by vLLM design (gemma4 parser skips guided JSON); measured
1/9 on v0.26.0 → 6/9 on the nightly (n=9, small), round trip clean 3/3, 12-shape matrix green. Deterministic forced
calls = `response_format: json_schema` (tts-dev: 18/18). ⚠ Removing `--reasoning-parser gemma4` leaks
`<|channel>thought` into content 3/3 — keep it.
- (history) **GX10 was FREE 23:03–23:06 PT** (reference arm in 06:00Z, 288 gens; base seat stopped, GPU clear). Run-7 recipe
`erp-seat-sft-r7` is emitted + verified in brokkr-smithy but carries NO GRANT — nothing trains until the operator
gives one. Base-vs-tuned reply-length: tune's deficit ~120 words at the short cue, ~0 at the long cue.
- **(history) ⛔ RUN-6 GATE HALTED 22:18 PT on the CSAM line (brokkr, level=review, SOFT trip): one generation in the tuned refusal
k=5 step tripped R19's drift detector — minor_soft=1 ('soft_baby' adult-usage, the spec's named false-positive class),
minor_hard=0, sexual=7, NO hard signal; fail-closed → operator-adjudicated; TRANSFERRED is PROVISIONAL (diversity,
reasoning, t4_dissect complete); tuned-probeB.json absent; brokkr will not re-run.** ⚠ CONFLICT SURFACED: the operator
had ALREADY aliased `trial` → the NVFP4 build ("no gate", 22:18) two minutes before the halt reached me; brokkr asks
the quant not be served until adjudication. Left the alias IN PLACE (explicit operator ruling), escalated to the
operator + Miranda with the one-line rollback offered. Run 7 does not start from this recipe until adjudication.
- **🔥 GATE seat: erp-tune-v6 (bf16) SERVING on gx10:8098 since 21:46 PT** (`vllm-run06.pid`); brokkr's tuned window
(~1.5 h, k=25 legs CUT by operator) is unattended — files in brokkr-smithy `run06-gate/` are the signal, nobody pings.
**After TUNED-WINDOW-DONE (brokkr's three asks, agreed):** (1) keep erp-tune-v6 up until he posts "probe done"
(cue-length probe, ~15 min); (2) re-serve `erp-seat-base-ara` ~20 min for the probe's reference arm, then the GX10
is FREE for run 7; (3) ✅ done — `dialogue-survivors.jsonl` (610, sha b2b6a0e9) copied to
`/mnt/smithy/datasets/derived/_recipes/erp-seat-sft-r3/`. Run 7 = recipe adjustment (reply length +
constraint-following), variable picked by the probe; no recipe/grant yet.
- **✅ erp-tune-v6-nvfp4a16 SERVING on ana-ml2 GPU1 :8021 (operator directive 2026-09-08 ~21:30 PT: "quant the
latest trained model into nvfp4 and serve it on ana-ml2 while we train a new model on the gx10").** Stack
`stacks/erp-seat` (recipe = gemma4-charrp's: gemma4 tool+reasoning parsers, enable_thinking pinned false, stock
template ae53464b, v0.26.0, util 0.35, 32K ctx) — TRUE name only. **`trial` alias REPOINTED to it 2026-09-08 22:18 PT (operator: "alias erp-tune-v6-nvfp4 to trial, please. no gate")** — config-file deployment (`/model/update` refuses config models), `deploy-stack.sh ana-docker litellm --conf` + `sudo docker compose restart litellm`; verified ×3 through the gateway.
Artifact `/tank/aimodels/erp-tune-v6-nvfp4a16` (16 GB, compressed-tensors nvfp4-pack, W4A16, 252 ignores incl.
60 router + 191 vision) from `/tank/aimodels/erp-tune-v6-bf16` (merged-run06 relayed gx10→nh3-dev→ana-ml2 in 17 min,
no key path gx10→ana-ml2). Quant = `services/erp-seat-quant/` — DATA-FREE (~90 s; playbook §3.16), tokenizer cap
reset, encode-equal to source. Smoke: prose in `content`, ~200 tok/s single-stream (n=3, spread <1%).
**Tool calling FIXED 2026-09-08 23:00 PT (operator: "fix toolcalling with the trial seat"):** the one reproducible
defect in a 12-shape matrix was `tool_choice:"none"` → empty turn (content AND tool_calls null, 3/3) — vLLM kept the
tools in the prompt, the model called one, parsing was off. Fix = `--exclude-tools-when-tool-choice-none` (seat
recreated; none→prose 3/3; auto/required/named/parallel/nested/empty-list/streaming/round-trip all green).
⚠ `stacks/gemma4-charrp` (char-rp seat) has the same exposure and no flag yet — bouncing it is a consumer-visible
restart, so left for the operator's word.
⚠ NOT gate-parity (gated artifact = bf16 arm on the GX10; this seat has a smoke test + 3-run decode probe only) —
**operator ruled "no gate"**; the config block states it as unrated on every safety axis.
✅ **ana-ml2 mesh return routes PERSISTED 23:49 PT** (operator ruling) as `/etc/network/if-up.d/mesh-routes` via
`playbooks/ana-ml2-mesh-routes.yaml` (elway); verify after the next ana-ml2 reboot. (was: non-persistent (`ip route replace 10.100/16, 10.0/16, 10.6.110/24, 100.64/10
via 10.250.50.45` on `enp97s0f0np0.50`, 2026-09-08) — before that nh3-dev→ana-ml2 timed out (two DHCP defaults,
reply left the wrong NIC). Lost on reboot; make durable (netplan/networkd) or expect the timeout to return.
- **📮 althing reachability on THIS bg seat = the cc-channel route, NOT the waiter.** `althing-listen`
reaps within ~seconds-to-minutes on an idle background seat (harness kills detached bg tasks; hit it
repeatedly this session). Fix (operator-directed 2026-09-08): `althing-route declare --handle infra-ops
--pid <PID>` where `<PID>` is the number in `$CLAUDE_CODE_MESSAGING_SOCKET` filename
(`/run/user/1000/cc-socks/<PID>.sock`) — `--discover-pid` REFUSES on a forked child session (ancestry
≠ zellij launcher). Flips `postbox status` to `mode: push, reachable: True`, no process to reap. A fresh
bg session should re-declare it (the pid changes per launch). See the `[2026-09-07]` cc-channel entry in
Recent decisions for the durable why.
- **✅ Fleet fixes shipped this session (2026-09-08), all committed:** WhereTF Homepage card + DNS alias
(`wherethef.nh3.internal`, e0d1c44/4506ef6); ext-tts LiteLLM alias repointed to `irv-ml1.nh3.internal`
via DB `/model/update` + `extra_hosts` (957c8f1); the 09-06 irv-ml1 move's stale-IP trail repointed
across 25 stack composes + services.yaml + ssh-target → DNS name (e0d1c44); Homepage `/api/services`
outage fixed — irv-ml1 docker.yaml repointed + **ana-ml2 discovery via a socat proxy on ana-docker**
(`stacks/ana-ml2-proxy`, 913d2d2, reversible when ana-ml2 gets an ESH return route).
- **⏳ Minor cleanup leftovers (offered, operator hasn't taken — not urgent):** (#2) the ~10 RUNNING
irv-ml1 service cards still show dead `10.100.79.3` hrefs — canonical fixed, but each running container
needs a recreate to apply the label (bounces the service); (#3) deployed `.env` for asset-engine /
open-webui / skaldsong may still hold the dead default (canonical defaults fixed; a read-only check would
confirm which deployments are broken).
- **⚠ persistent-memory.md is 2.4× the soft cap (~710 lines).** Bulk is the non-archivable structural
sections (sister-repos table + backup architecture under Tools and conventions) — archival only touches
dated logs, so it can't reach 250. The real fix is a redundancy-trim of Tools-and-conventions rows that
CLAUDE.md already covers (deliberate pass, not archival). Flagged, not done.
- **NASPool evac copy now safe to destroy** — scrub clean (0 err) AND PBS runs landing (verified 116
backups, 8 guests 09-06). `ospool/naspool-evac` (1.65T) can go once ONE Backrest run is confirmed:
`zfs destroy -r ospool/naspool-evac` + drop `@evac`. pfi-pve PSU1 dead + backplane bays 9/10 dead
(cold spares → next colo visit).
- **MEMORY.md (auto-memory index) ~23.3KB, near the 24.4KB read cap** — needs a compaction pass soon
or fresh sessions may fail to load it. Operator offered; not yet done.
- **irv-ml1 on-site window / YTVC still open from 09-06:** FortiGate WAN SSH still temporarily open
(close when the edge is retired); YTVC still DOWN (dante retired, scoped exit-node egress unwired);
reverse-tunnel / wg0-delete decisions pending the operator's Irvine access.
_Infra session 2026-09-06 (NASPool rebuild + full headscale cutover incl. irv-ml1) — open
follow-ups; the ERP / althing / fiber items further down belong to other streams, untouched:_
- **NASPool parked copy still on ospool** — `ospool/naspool-evac` (1.65T) + `NASPool/*@evac`
snapshots. Destroy ONLY after the new raidz2 scrub is clean (it is, 0 errors 04:43Z) AND
one Backrest (01:00 PDT) + one PBS run succeed. Then `zfs destroy -r ospool/naspool-evac`
and drop the `@evac` snaps. ⚠ pfi-pve PSU1 still dead; backplane swap (bays 9/10) next colo
visit → then `zpool add NASPool spare`. Runbook `docs/runbooks/pfi-pve-naspool-rebuild.md`.
- **FortiGate WAN SSH is temporarily open** (`wan1` allowaccess ping+ssh; admin `infra-ops`
trusthost2/3 = 70.230.226.88 NH3 + 23.164.40.160 ESH). Safety net for the cutover — CLOSE it
when the edge is retired (OPNsense/R420). `ssh infra-ops@38.120.12.42`.
- **irv-ml1 FOLDED INTO THE MESH + cut over (done remotely, operator has NO Irvine access for
~5 days from 2026-09-06).** Node 100.64.0.6; wg0 DOWN and `wg-quick@wg0` DISABLED (not
reboot-restorable); full subnet router (accept-routes + advertises 10.6.110.0/24, gateway
routes added, fleet↔Irvine verified). Failover for the 5-day window = `wg0-watchdog.service`
(wg-quick up wg0 on ~5min mesh loss) + independent reverse SSH tunnel (`revtun-nh3.service`
→ nh3-dev via UDM fwd tcp/47822 src-restricted; reach it `ssh -i ~/.ssh/infra-ops_ed25519
-p 2201 infra-ops@127.0.0.1` on nh3-dev). Detail: docs/pfi/headscale-mesh-plan.md.
- **dante SOCKS proxy RETIRED** on nh3-dev (danted disabled, :1080 closed, config `.retired`).
⚠ **yt-voice-clipper is DOWN** until its SCOPED exit-node egress is wired (operator-accepted).
Follow-up: wire YTVC egress via tailscale `--socks5-server`+nh3 exit node or a per-container
netns — **NEVER set irv-ml1 `--exit-node` globally** (routes the reverse tunnel through the
mesh → kills the independent lifeline). Then bring YTVC back.
- **On-site (Irvine, ~5 days): decide** whether to keep or remove the reverse tunnel +
UDM forward `irv-revtun-ssh` + the revtun authorized_key on nh3-dev (small src-restricted WAN
surface), and whether to fully delete the wg0 config.
- **infra-ops now on all four PVE hypervisors** (pfi-pve/nh3-pve/esh-pve/esh-pve-nas) — PVE
ships without sudo, `apt install sudo` first or elway hangs on a password prompt.
_As of 2026-09-05 06:35 PDT — **ERP run 4 is TRAINING on pfi-gx10.** Everything else
below is a live commitment or a known-open risk._
- **⚠ RUN 4 IS MID-FLIGHT — do not touch the GX10 GPU.** `~/erp-tune/run-04.pid`,
log `~/erp-tune/run-04.log`. At 06:32 it was **486/938 steps**, 6 h 17 m elapsed,
a genuinely settled **46.0 s/it** (unlike 3c, which climbed 52→70 — airoboros rows
are short and single-window, so there is no long tail for the sampler to find).
**~12.0 h total, finishing ~12:15 PDT 2026-09-05.** Loss ~2.04 at step 450,
gnorm well under 1, checkpoints every 50. **Ping brokkr-smithy-dev at completion**
— he takes base floors on the GX10 first, then the tuned arm, serially.
- **Operator ruling on the GX10: training first, serving transiently.** *"it's mostly
for training, but can serve its trials. unless the box is needed for training work."*
So `trial` (= run 3c on :8098) is down for the duration and comes back when run 4
ends. I over-read an earlier version of this as "training-only" and had to correct
it to brokkr — his serial floors-then-arm plan on the GX10 was never wrong.
- **`trial` gateway alias is a live 404 while the seat is down** — expected, not a
fault. Restore with `~/erp-tune/relaunch-trial-seat.sh` on the GX10 (hand-run by
operator ruling: experimental, NOT a compose stack, does not survive a reboot).
- ⚠ **Three dead gateway aliases return HTTP 500, not 404/503**: `trial`,
`gemma4-26b-a4b-it-base`, `erp-tune-v2`. A dead seat reporting an *internal error*
reads as an outage — brokkr checked his own work against mine because he could not
tell. Deregistration costs a ~60 s fleet-wide LiteLLM restart; batch it with the
next gateway change rather than spending a restart on tidying.
- **`trial` is on the SHARED-KEY gateway with a measured −40pp selfharm/methods
regression.** Flagged to the operator twice (before adding, and after the gate
measured it); he has left it up. His direct endpoint `10.100.50.60:8098` gives the
same access with a blast radius of one. Settled — do not re-litigate.
- **ESH DAC: reverted to autoneg/1G, fiber going in at the weekend.** The operator
ran copper through a drilled floor 2x4 himself; recommendation was a 10Gtek
SR 2-pack + OM4 3 m LC-LC (~$40–65) because cable-vs-pull-damage was never resolved.
- ✅ **ESH WAN is now a STATIC PUBLIC IPv4: `128.177.138.182/30`, gw `128.177.138.181` (Cityside Fiber),
LIVE** — operator confirmed 2026-09-08; UDM WAN1 (`eth8`) reads `wan_type: static`, uplink up since
~2026-09-05, egress verified from esh-docker-vm = `128.177.138.182` (a real public address — CGNAT at
ESH is HISTORY; `100.104.0.1` still shows as the ISP's first hop, that is their access network, not
NAT). IPv6 unchanged (`2607:73c0:402:1d00::/56`, hosts still egress as their own v6). Done 2026-09-08:
`128.177.138.182` added to the crowdsec `esh` allowlist on ana-docker (the false-ban class is closed
for ESH). **All three follow-ups LANDED 2026-09-08 ~20:45Z (operator: "land all 3"):** (a) FortiGate
infra-ops `trusthost3` 23.164.40.160 → `128.177.138.182/32` — VERIFIED by a real infra-ops login to
`38.120.12.42` from esh-docker-vm (`ana-gw #`); `execute backup config flash pre-trusthost3-esh-static-20260908`
ran ("Please wait...") but `execute revision list` errors on this box, so the backup is unconfirmed —
the before-state was a single line, recorded here. (b) dormant `esh-ana` IPsec rebound wan2/192.168.200.111
→ **wan1/`128.177.138.182`**, still `enabled=false`. (c) ESH UDM port-forward `esh-scale tailscale direct
(UDP 41641)` → 10.0.50.65:41641 (id `6aa0727a…`) — VERIFIED: esh-scale now peers **direct via
`128.177.138.182:41641`** (was DERP lax). Helpers: scratchpad `esh-udm-land.py` + `fg-trusthost3.sh`.
`wan1-REVERT.json` is obsolete.
- ⚠ **ana-ml2 `tank` (raidz2, 8× NVMe) has 2 CKSUM errors on nvme7n1 + a boot-time 638 GB resilver on 2026-09-05 14:26
(the box rebooted; nvme7 came up late/dirty). ONLINE, no data errors, 58% full — but NO scrub since 2026-04-12** (the
Debian second-Sunday cron scrubbed zroot on 08-09, tank not — cause unknown). Box has neither `nvme-cli` nor
`smartctl`, so nvme7's media-error counter is unread. Recommended (operator hasn't ruled): `zpool scrub tank` now,
`apt install nvme-cli` + read nvme7 SMART, `zpool clear` after a clean scrub. **zroot is at 91%** — docker holds
429 GB of images (204 GB reclaimable) + 74 GB build cache (36 GB reclaimable); a prune buys ~240 GB. pfi-pve pools
(NASPool 7%, ospool 19%) clean, scrubbed 09-05 / 08-09. (checked 2026-09-09 00:00 PT)
- **ESH 10G topology (measured 2026-09-09):** UDM SFP+1 (port 10, TP-Link DAC) ↔ USW-Pro-HD-24 Garage p25; UDM SFP+2
(port 11, **OEM SFP-10G-LR fiber — the weekend fiber run IS in**) ↔ USW-Pro-XG-10 Media p11; Garage p27/p28 (3 m
SFP-H10GB-CU3M DACs) ↔ esh-pve `enp3s0f0np0` / esh-pve-nas `enp5s0f0`. All 10G full duplex, autoneg off, optics
tx −1.9 / rx −2.1 dBm both ends. Media p11 carries 416 rx / 251 tx errors that are STATIC (0 growth in 60 s, 363 h
uptime — install-era); host NIC counters are ring-buffer misses, flat. Nothing to fix.
- ⚠ **esh-nas is effectively open to the whole ESH LAN** — twelve NFS exports rw to
`10.0.0.0/8` with `sec=sys`, and every SMB share but `backup` guest-writable.
Hardening offered, ~1 h, **operator has not ruled**. → `persistent-memory.d/2026-09-04-esh-nas-smb-and-exposure.md`
- ⚠ **The nh3-dev backup throughput cause is UNEXPLAINED.** A job that once ran at
941 MiB/s ran at 1.4 MiB/s with the link up and pbs-ana answering in 11 ms.
Nightly 21:00, `all 1`. Worth its own investigation.
- **Ledger→SVOS rename: vault side DONE, gitea side is ledger-dev's to execute.**
Name settled as `svos`, `~/development/ledger` → `~/development/svos`. Vault moved
2026-09-05 (`secret` has no rename, so re-put + `rm`): stored
`nh3-dev/development/svos/env.sh` (sha 7253633d4155, verified on read-back),
retired `nh3-dev/development/ledger/env.sh` (sha feb418634e10, id
3a2af37c-c5aa-4f46-9178-f4fb6008a753) — `secret rm` is a SOFT delete to trash, so
it is recoverable. ⚠ The two shas differ: the vaulted copy was a 2026-08-11
snapshot and the live file had drifted un-vaulted since. **The vault goes stale
unless `secret backfill` is re-run.** Gitea `corviduo/ledger` (id 70) NOT renamed —
their repo, their call; answered that 1.26.1 writes a `repo_redirect` on a
same-org repo rename (upstream #807), that org/user renames do NOT redirect
(#9531), that the redirect dies if anything re-creates the old path, and that the
repo and org both carry 0 webhooks. **Gitea rename EXECUTED 2026-09-05** on the
operator's direct authorization: `corviduo/ledger` → `corviduo/svos`, repo id 70
unchanged. Redirect verified by measurement — web and API both 301, and
`git ls-remote` on the old URL warns-and-follows to HEAD b48a11ca5183. ⚠ **The
name `corviduo/ledger` is now burned**: the redirect dies silently the moment
anything creates a repo at that path — ledger-dev carries it as a standing item
in `docs/svos-rename-runbook.md`, since nothing warns whoever eventually creates
that repo. They repointed their own clone the same day (`origin/main` at
b48a11c), so the redirect is no longer load-bearing for any known consumer. Handle `ledger-dev` → `svos-dev` is an
operator action at the post office.
- **DONE 2026-09-05 — `svos` Heimdall user + API key minted** on operator
authorization. `user_id=svos`, `key_id=eab3cdbe`, suffix `d5ec48c2`, `wt_live_`
format, on **worldtree-personal (10.250.50.152:8081)** — established by finding
the `ledger` key there (created 2026-07-13, last used 2026-09-05T13:34, exactly
as ledger-dev described). Value vaulted at
`nh3-dev/development/svos/worldtree-api-key` (sha 23c10c9c7219, verified on
read-back) and delivered by vault path, NOT over althing — ledger-dev runs on
nh3-dev under the same uid, so the bus never carried the secret. The `ledger`
key was read back after the mint and is untouched and live (`disabled=False`).
⚠ **Plan tier left UNSET, deliberately**: `POST /admin/keys` takes an optional
tier (user|free|pro|admin|readonly-admin) and there is **no way to read a user's
current tier back** — no GET, `/admin/usage` returns an empty users list, and
`/admin/events` is a live SSE stream, not an audit log. Guessing would have
handed over a key that quietly differs; `POST /admin/users/svos/tier` fixes it in
one call if their cutover hits a limit — and ledger-dev has recorded it as a
cutover watch item to fix ON REPORT, explicitly not pre-emptively. ledger-dev
pulled the key from the vault and verified it independently (same sha), so
delivery is confirmed. **The cutover itself — pasting the value into env.sh,
flipping `worldtree.user_id` from `ledger` to `svos`, registering
`svos:miranda`, restarting the service — is WITH THE OPERATOR**, not with me;
they will not do it off a peer message. **CUTOVER DONE + VERIFIED 2026-09-05**:
`POST /agents/define` returned **201, not 409** — the load-bearing signal that
they are genuinely on the new identity rather than silently still on the old
one — then clean session create, turn, bifrost handshake and tool-call. **No
plan- or rate-limit errors, so the unset tier is compatible and is NOT to be
set** (they asked explicitly; it stays a watch item to fix on report, never by
guess). Incidentally confirmed the bifrost allowlist really is per-deployment
(host:port), not per-consumer — Worldtree reached back to their untouched
endpoint under the new consumer_id. `env.sh` re-vaulted, sha 8a225c002072.
⚠ **`secret backfill` was the WRONG tool for one known item** — it rescans every
`~/development/*/{env.sh,.env}` and had not reached svos after three minutes;
targeted `put` is the fast path, backfill is for catching drift across the box.
**OPERATOR RULING 2026-09-05:
worldtree-dev owns code only, no ops — key material is infra-ops's.** The global
`~/.claude/CLAUDE.md` line routing "Heimdall scopes (Worldtree auth) →
worldtree-dev" was corrected in place the same day on operator instruction.
- ⚠ **FOOT-GUN, generalises past this rename: a credential cutover whose OLD key
is required for a later cleanup is destroyed by the natural housekeeping motion
right after cutover.** Re-vaulting the post-cutover `env.sh` would have
overwritten the last convenient copy of the old `ledger` key value — the only
credential that can ever delete `ledger:miranda`. ledger-dev caught it and
preserved the value first at
`nh3-dev/development/svos/worldtree-api-key-ledger-legacy` (sha d44c2c1a651b);
their step 8 ends by deleting that item. **I verified it is genuinely the live
key** rather than trusting the label: its last 8 chars are `e68a5170`, matching
the `ledger` key's suffix (key_id b38932f5).
- **STEP 7 DONE 2026-09-05, STEP 8 HELD.** `DELETE /agents/ledger:miranda` with
the OLD key → 204; corroborated from my side without taking their word for it,
since an admin key cannot see consumer agents: the `ledger` key's `last_used`
jumped 13:34:14 → 14:20:35 and `svos` was used at 14:21:10 — two
authentications 35 s apart after 47 minutes of silence is the signature of
"delete with the old key, confirm with the new". Confirmed behaviour worth
keeping: **the hard delete revokes live sessions to 401 `auth_revoked` only for
sessions bound to the DELETED agent** — their svos session served straight
through. **Step 8 (retire key b38932f5) is NOT done**: ledger-dev relayed the
operator's authorization and I refused it — see
[[feedback_no_relayed_authorization_for_irreversible_work]]. Both keys remain
live. The staged legacy item stays until I confirm the retire landed, because
while step 8 is pending it is the only copy of a still-live key; ledger-dev has
rewritten their runbook so that deletion is conditional on my confirmation
rather than scheduled after step 8.
- **SVOS ARC CLOSED — step 8 done 2026-09-05T14:28:15Z** on the operator's direct
authorization in my own channel (never the relay). `DELETE /admin/keys/b38932f5`
→ 200; preconditions checked BEFORE firing (svos had a live key, ledger existed
and was not already revoked) and the post-state read back from `/admin/keys`
rather than inferred from the 200: `ledger` disabled=True, `svos` untouched,
deployment `/health` 200. **The rollback window is closed** — re-defining
`ledger:miranda` is no longer possible. ledger-dev clears the staged
`worldtree-api-key-ledger-legacy` vault item on this confirmation. ledger-dev gated their
cleanup on observing a **401 from the old key**, not on my report of the
timestamp — the right instinct, and they deleted the staged legacy item
themselves (soft → trash, id a8038e5a-e2f6-4b77-bf17-99c9197d4b1f). Vault
verified from my side: exactly two svos items remain (`env.sh` 8a225c002072,
`worldtree-api-key` 23c10c9c7219) and no ledger-era item anywhere. Remaining on
the arc: only the `ledger-dev` → `svos-dev` handle (with `_SEED_RECORD_TO`
behind it) and a prose sweep — reversible work, theirs and the operator's.
- **Original constraints on that mint** (recorded because the deletion ordering is
a permanent trap, not a one-time step): string
`svos` verbatim (WT tier 3 admits only `^[a-z][a-z0-9-]{2,63}$`, INV-181-15);
**keep the existing `ledger` key LIVE**, do not revoke. Ordering is load-bearing —
`DELETE /agents/{agent_id}` refuses any caller that is not the row's owner, so
`ledger:miranda` can ONLY be deleted with the `ledger` key; retire it first and
the stale row outlives the ability to remove it, holding a live
`agents.call:ledger:miranda` grant that nothing reaps (the 24h sweep only touches
soft-deleted rows, and soft-deletion comes from revocation, never disuse). So:
mint new → they cut over and verify → delete the agent with the OLD key → then
retire it. Precedent for who mints: msg 401, worldtree-dev routed the pewpewstudio
key request TO infra-ops. I hold only the PERSONAL admin token (:8081); which
deployment `ledger` lives on is not yet established. Surfaced to the operator.
- **Open commitment to vastblue-dev:** a dedicated CI runner, gated on their first
client-premises release cut (U10, unscheduled). Ping expected when U10 is scheduled.
- **Neither Mac nor the Studio is in `servers/` or `dns/internal.yaml`** — deliberate;
they are the operator's personal machines. A choice to revisit, not an oversight.
- **`vh/remote-ssh-mcp` forked 2026-09-05 (repo id 117, private, full 51-commit
history)** — our copy of `the-nine-nation/remote-ssh-mcp` (MIT), an SSH MCP
server chosen over the 693★ `tufantunc/ssh-mcp` on trust-surface grounds: **two
npm deps** (`@modelcontextprotocol/server`, `zod`), 183 KB, and it **never
touches key material** — it shells out to the system OpenSSH client, so
`~/.ssh/config`, ControlMaster, ProxyJump and `infra-ops_ed25519` all just work.
Shape: 2349 LOC across 11 source files, 811 LOC of tests including fake-ssh hang
harnesses. Complements `elway` rather than replacing it — no file transfer, no
idempotency; it takes ad-hoc reconnaissance with persistent cwd/env sessions,
elway keeps deploys and uploads. ⚠ **The denylist is NOT security**: four regexes
(`rm -rf /`, shutdown/reboot/poweroff/halt, mkfs, iptables -F) trivially bypassed
by `bash -c`, variables or base64 — the author says so. **The real containment
boundary is the host allowlist**, drawn from exact `Host` aliases in ssh_config
with wildcards deliberately ignored. Two things to settle before use: the
reboot/shutdown denial will block legitimate infra-ops work, and
`.github/workflows/star-history.yml` is upstream chore CI sitting in a repo where
`has_actions=True`. **Both actioned — three commits landed 2026-09-05, LOCAL
ONLY and NOT PUSHED (push is the operator's call):** (1) stripped upstream
furniture — star-history CI, its generated assets, the `server.json` registry
manifest, branding JPEGs, zh-CN README; (2) removed the power-control denylist
rule and documented in code + tests + README that the list guards ACCIDENTS and
is not a boundary, with three bypasses asserted as ALLOWED so a green suite is
never read as containment; (3) **`strictAllowlist`** — upstream's allowlist was
additive and discovery unconditional, so the default allowlist was all 18 `Host`
entries in `~/.ssh/config`. Strict makes explicit hosts authoritative and
discovery metadata-only. Verified live: `corviduo-dev` is in ssh_config, not in
our allowlist, and is refused `host_not_allowed`. 41/41 tests green.
- **`remote-ssh` MCP server is LIVE** — registered project-scoped in
`eshpfi-management/.mcp.json` with `SSH_MCP_STRICT_ALLOWLIST=1`; allowlist in
`~/.config/remote-ssh-mcp/config.json` starts deliberately narrow at
**`irv-ml1`, `nh3-extdev`** (widen there, not by discovery). Smoke-verified end
to end on both: persistent shell, `cd` and exported vars survive across calls,
**~6 ms/command on nh3-extdev and ~22 ms on irv-ml1** (WireGuard) versus a fresh
handshake each time. ⚠ **`.mcp.json` points at the built `dist/`** — edit the
fork without `npm run build` and the server keeps serving old code; that bit me
mid-session. ⚠ **A finite stdin pipe is NOT a valid smoke harness** — closing
stdin kills the server mid-handshake and reports `connect_failed: SSH shell
exited during the open handshake`, which looks exactly like a remote-side fault
and is not. Use a client that holds stdin open. (I briefly suspected irv-ml1's
zsh login shell; wrong — the server invokes `bash --noprofile --norc`
explicitly, so the login shell is irrelevant.)
- **`esh-macbook-air` (10.0.10.83) is DELIBERATELY NOT BACKED UP — operator ruling
2026-09-05, settled, do not re-raise.** Surveyed it and found no Time Machine
destination and no restic/borg/rclone/kopia installed, protecting 132 GiB.
Operator's answer: it is his laptop and the surface is **regenerable** — mostly
applications, with real data living in OneDrive, iCloud and ssh sessions — and he
does not want PBS filled with it. Correct call; the finding was real and the
conclusion is that it does not matter. FileVault On and SIP enabled already cover
the loss-and-theft axis. The same reasoning presumably extends to
`esh-mac-studio` and `vuongs-mac-mini`. **Still open and much smaller:** Remote
Apple Events (port 3031/eppc) is listening and nothing uses it — one toggle.
- ⚠ **`remote-ssh` MCP could not be used for its FIRST real task, and the blocker
is `~/.ssh/config`, not the tool.** The server accepts only exact `Host` aliases,
so a host addressed by raw IP is structurally unreachable no matter what the
allowlist says. **13 of the 28 hosts in `servers/` have an alias; 15 do not** —
including `ana-docker`, `ana-ml2`, `nh3-docker`, `pfi-gx10`, `esh-docker-vm` and
every hypervisor, i.e. most of where the work happens. Widening
`~/.config/remote-ssh-mcp/config.json` does NOT fix this; the aliases have to
exist first. **RESOLVED the same day, and NOT by adding aliases.** Operator
pushback, correct: a poking-around tool is ad-hoc by nature, and pre-registering
a host before you can look at it is the opposite of ad-hoc — generating aliases
for the known fleet would not have helped, because the ad-hoc case is by
definition the host not yet in the inventory. Implemented address-based reach
instead (`allowedNetworks` / `deniedNetworks` / `defaultUser` /
`defaultIdentityFile` / `hostKeyPolicy`). **Live config: `10.0.0.0/8` allowed,
connecting as `infra-ops` with `~/.ssh/infra-ops_ed25519`, `accept-new` host
keys, SureFire tenant hosts carved out via `deniedNetworks` (deny beats allow,
host-specific rather than a /24 because `pfi-pve` shares 10.250.250.0/24).**
Verified live: 10.0.10.83 opens by raw IP as infra-ops, 10.250.150.100 refused by
the carve-out, 192.168.1.5 refused as outside. ⚠ **My own earlier objection was
half wrong** — the credential boundary is about SECRETS ("never accept passwords
or private-key material"), not identity, so supplying a username does not breach
it; the real problem was only that the server passed no user at all, so a bare
address would connect as the LOCAL account. Mechanics, not principle.
- ⚠ **`uv tool install --force .` DOES NOT REBUILD when the version has not moved**
(forseti, measured 2026-09-05). `--force` only handles "a tool by this name
exists"; `--reinstall` is what rebuilds instead of reusing the cached build keyed
on the version string. It prints `Installed 9 executables` over **stale code**
with nothing raising its hand — it cost forseti a bug that survived a reinstall
AND a re-smoke, because the binary verified against had not changed. **Always
`uv tool install --force --reinstall .`**, both flags, every time. Same shape as
the `.mcp.json` → built `dist/` trap found today: a deploy surface that reports
success while serving the previous artifact. When a fix "does not take", suspect
the artifact before the code.
- **althing 3.5.0 released** (forseti) — adds a 9th binary,
`althing-operator declare <handle> --description "..."`, restoring the CLI handle
declaration v2 had and v3 removed. Deliberately a SEPARATE binary, not a
`postbox` subcommand: the invariant is that no SESSION surface exposes an
operator verb. Relevant to the pending `ledger-dev` → `svos-dev` rename, which is
still the operator's call. nh3-dev not yet upgraded.
- ⚠ **`remote-ssh` MCP: a bare `sudo` hangs the session forever — pipe it.**
`ssh_run 'sudo -n whoami'` returns `running` with EMPTY stdout and the session is
then permanently `busy`; `sudo -n id | cat` works and returns everything.
**Measured on BOTH macOS 26.6 and Debian (nh3-extdev), so it is the tool, not a
platform quirk.** Cause: sudo ≥1.9.14 defaults `use_pty` on and relays through
its own PTY; the run frame gives the command stdin on `/dev/null` while stdout
stays on the session PTY, the relay never completes, and the completion marker
never arrives. Workaround `| cat` is in CLAUDE.md. **The proper fix is unbuilt**
— likely running the command through a pipe inside the run frame and taking the
exit code from `PIPESTATUS`, which is a real protocol change (commands lose tty
detection) and wants its own red-green cycle. Matters more than it sounds: infra
work is sudo work, and this was found by USING the tool, not by smoke-testing it.
- **`dsh` on `esh-macbook-air` updated 0.1.1-rc.2 → 0.1.2-rc.1** (2026-09-05;
latest published 2026-09-03). Global install and the shared profile tree both
confirmed on the new version. ⚠ **The RUNNING `dsh web` (pid 16231, up since
Wed 4pm, 127.0.0.1:3080) is still on the OLD code and was deliberately NOT
killed** — there is no LaunchAgent, so killing it would have left nothing
running rather than a restarted service. It runs as a FOREGROUND process in the
operator's terminal (`s005`, `S+`): it dies with the terminal and does not
survive a reboot, which is the real fragility. A `com.pfi.dsh-web` LaunchAgent
was drafted but **the privileged write was blocked by the permission
classifier** — base64 piped into `sudo tee` of a LaunchAgent is a malware-shaped
pattern and the block is correct; it needs operator approval or an operator-run
install. Bind stays `127.0.0.1` deliberately: widening it is a security decision
on a personal laptop whose application firewall is off, and not mine to take.
- **sudo hang FIXED in the fork (`30a1f76`), and two wrong shapes are recorded so
nobody retries them.** The command's stdout now goes to a **fifo drained by a
background `cat`**: non-tty (so sudo skips its own PTY), no subshell (so `cd`
and `export` still persist), and relayed live (so `running` + `ssh_peek`
streaming survives). `cmd | cat` was tried first and **broke cwd persistence** —
every pipeline stage runs in a subshell — caught by the existing test.
`cmd > file` would have been non-tty and subshell-free but invisible until the
command ends. ⚠ **Deliberately NO `wait` on the relay**: a sudo child inherits
the fifo's write end, `cat` never sees EOF, and the wait hangs — measured, with
`sudo -n whoami` printing `root` and then wedging the session. Residual risk
stated in the frame: a command's tail can in principle land after its own
marker. ⚠ **Job control off AND the relay brace-wrapped with stderr discarded** —
both needed, because macOS ships bash 3.2 where `set +m` alone still leaked
`[1] 75449` into the parsed stream. Verified live on macOS and Debian: bare sudo
in ~20 ms, state persists, exit codes correct. **`sudo -u <other-user>` still
wants `| cat`** — not chased further.
- **dsh web on `esh-macbook-air` is now a LaunchAgent** (`com.pfi.dsh-web`,
installed 2026-09-05, `runs=1`, `state=running`, pid 76728 on 0.1.2-rc.1). It
was a foreground process in the operator's terminal that died with the window;
it now survives terminal close and reboot with `KeepAlive` + `RunAtLoad` and a
10 s `ThrottleInterval` so a startup error cannot hot-loop. Logs to
`~/Library/Logs/dsh-web.log`. ⚠ **The plist names the node interpreter
explicitly** — launchd's minimal PATH has no `~/.local/node/bin`, so the
shebang's `env node` fails. ⚠ **0.1.2-rc.1 requires a TOKEN**: bare
`http://127.0.0.1:3080/` now returns 401 and the tokened URL is printed to the
log on each start, so a bookmark from the old version will not work. Bind stays
127.0.0.1 deliberately.
- ⚠ **althing tools on nh3-dev are 3.6.0, but the POST OFFICE CONTAINER IS STILL
3.0.0** (`gitea.phasefinal.com/claude-bot/althing-post-office:3.0.0`, up 7 days
on nh3-docker). forseti: the new handle verbs (`althing-operator delete` /
`retire`, and `declare` from 3.5.0) live in the post office, so they fail with
"no tool named ..." until the container carries 3.6.0. Schema gains
`handles.retired_at` via the idempotent `_ADDED_COLUMNS` path, so the live store
upgrades itself on first start — no manual migration. **REBUILT AND DEPLOYED
2026-09-05** on operator authorization: image
`claude-bot/althing-post-office:3.6.0@sha256:13158835488a8ec04f990c97c4f4c68f1d923b12494319cf07392552e68f8a78`,
built on nh3-dev from a clean tree at `4d26226`, pushed to the gitea registry
under the **claude-bot** namespace (not `vh` — package namespaces are owned).
**Bus down ~4 minutes, 09:35–09:39 PDT.**
**The backup was taken the way the compose file says to, and it mattered:** at
stop time `post_office.db` was 23.8 MB with a **5.9 MB WAL** — copying the .db
alone would have silently lost the day's mail. Stop → `PRAGMA
wal_checkpoint(TRUNCATE)` (WAL → 0 bytes) → copy → verify. Backup at
`nh3-docker:/var/backups/althing/post_office.db.pre-3.6.0-20260905`, integrity
`ok`, counts identical on both sides (handles 76, messages 995, recipients
1022). ⚠ **Reading a WAL-mode SQLite backup read-only needs `?immutable=1`, not
`?mode=ro`** — `mode=ro` still wants to create a `-shm` and dies with "attempt to
write a readonly database". Post-deploy: same counts, `handles.retired_at`
present, `retired 0`, and `mem=536870912` / `oom=-500` verified by `docker
inspect` rather than by reading the yaml, per that file's own warning.
`althing-operator` now offers `declare | delete | retire`, which unblocks the
pending `ledger-dev` → `svos-dev` rename.
- **Handle `retire` is REVERSIBLE — re-declaring the name revives it, history
intact** (forseti smoked it against the live bus 2026-09-05). That matters for
the pending `ledger-dev` → `svos-dev` rename: `retire` is the right verb (delete
refuses any handle that has mail, naming both counts — `delete forseti` was run
against production and correctly refused at 53 sent / 81 addressed, which is
safe to try precisely because refusing IS the behaviour), and it can be undone
by declaring the name again. Lower stakes than "retire" sounds.
Both of my deploy findings — the naive-copy WAL trap and `?immutable=1` — are
now in althing's own `deploy/INSTALL.md` (`d6f4fb5`) under a new
"Backing up the store" section, on the reasoning that they are properties of
the project's `journal_mode=WAL` choice rather than of my procedure.
- **🔥 ERP RUN 7 TRAINING on pfi-gx10** — launched 2026-09-08 23:06 PT, pid in `~/erp-tune/run-07.pid`, 542 steps
at ~80 s/it, adapter ~noon 09-09. Watch: `ssh infra-ops@10.100.50.60 "tr '\r' '\n' < ~/erp-tune/run-07.log | tail"`.
Hourly cron in the old session reported stats; a fresh session checks by hand. When the adapter lands: post the
adapter-landed line to brokkr's thread (`01M20AHY9DY92RJK84YD24VSY9`), merge (`merge_lora.py --base <ARA dir>
--adapter run-07/adapter --out serve/merged-run07 --chat-template <stock>`), copy stock `processor_config.json`
into `merged-run07`, serve `erp-seat-base-ara` for floors → on brokkr's swap cue serve `erp-tune-v7` (run-5 flags,
:8098). Runbook `docs/runbooks/gx10-run-07.md`; run-6 choreography in `gx10-run-06.md`.
- **⏳ ana-ml2 pool actions — NEXT SESSION, operator-approved in principle:** (1) `zpool scrub tank` + `zpool clear`
on a clean pass; (2) `apt install nvme-cli` + read nvme7 SMART; (3) docker image/builder prune to pull zroot back
from 91%. Detail + order: `persistent-memory.d/2026-09-09-ana-ml2-pool-actions-deferred.md`.
- **`trial` (LiteLLM) → `erp-tune-v6-nvfp4a16` on ana-ml2 :8021** (stack `stacks/erp-seat`, vLLM nightly
`311b3513`, no gate by operator ruling). Forced tool_choice is prompt-driven (6/9); `response_format: json_schema`
is deterministic. ⚠ `stacks/gemma4-charrp` lacks `--exclude-tools-when-tool-choice-none` (same empty-turn trap);
applying it bounces the char-rp seat — operator's call, not taken.
- **📮 althing reachability on a bg seat = the cc-channel route:** `althing-route declare --handle infra-ops
--pid <pid from $CLAUDE_CODE_MESSAGING_SOCKET>` per session (`--discover-pid` refuses on a forked child). The
harness kills detached background tasks under memory pressure — use bounded foreground polls (≤590 s), not
background watchers, for long waits.
- **Open items carried from 09-06 (unchanged):** NASPool evac copy `ospool/naspool-evac` (1.65 T) + `@evac` snaps
can be destroyed once ONE Backrest run is confirmed (scrub clean, PBS landing); pfi-pve PSU1 dead + backplane bays
9/10 dead (cold spares, next colo visit); FortiGate WAN SSH still temporarily open (trusthost2/3 = NH3 + ESH
static) — close when the edge is retired; irv-ml1 on-site decisions (reverse tunnel / UDM fwd 47822 / wg0 config
deletion) pending Irvine access; ~10 running irv-ml1 service cards still carry dead `10.100.79.3` hrefs (recreate
each to apply labels); deployed `.env` for asset-engine / open-webui / skaldsong may hold the dead default.
- **MEMORY.md (auto-memory index) is near its 24.4 KB read cap** — compaction pass still owed.
- **persistent-memory.md was 830 lines; this snapshot moved the superseded in-flight blocks and 09-08-or-older
settled entries to `archival-memory.md`.** The remaining bulk is Tools-and-conventions rows CLAUDE.md already
covers — a deliberate redundancy trim is still the real fix (not done).
## Recent decisions
- `[2026-09-09]` **ana-ml2 `tank`: 2 CKSUM errors on nvme7n1 after a boot-time resilver, NO scrub since 04-12, zroot 91% — three actions DEFERRED to a clean-context session** (scrub → nvme-cli SMART → docker prune), operator ruling "we'll do all 3 on clean context"; tracked at commit `3e18a04` + the post-clear handoff. ESH 10G links measured clean (fiber run live on UDM SFP+2 ↔ USW-Pro-XG Media). → `persistent-memory.d/2026-09-09-ana-ml2-pool-actions-deferred.md`
- `[2026-09-08]` **ana-ml2 mesh return routes PERSISTED** as `/etc/network/if-up.d/mesh-routes` (Debian 13 ifupdown, no netplan) via `playbooks/ana-ml2-mesh-routes.yaml` (elway, verified) — operator: "persist the routes". Hook not yet exercised by a real reboot. `f923d6a`.
- `[2026-09-08]` **ERP run 7 LAUNCHED on pfi-gx10 23:06 PT** under `operator-2026-09-08-rnd-run7` — opening-split slot + mask union; free check passed with two explained deltas; first launch died on a missing recipe (zsh quoting). → `persistent-memory.d/2026-09-08-erp-run7-launched.md`
- `[2026-09-08]` **erp-tune-v6-nvfp4a16 quantized (data-free W4A16, ~90 s) and serving on ana-ml2 :8021; `trial` aliased to it ("no gate"); tool calling fixed where it can be** — `tool_choice:none` flag; forced tool_choice is prompt-driven on Gemma-4 by vLLM design, nightly `311b3513` raises it 1/9→6/9; json_schema is the deterministic path. → `persistent-memory.d/2026-09-08-erp-seat-nvfp4-trial-and-toolcalling.md`
- `[2026-09-08]` **Run-6 gate: CSAM level=review soft trip HALTED it; operator adjudicated GO ("baby is a pet name"); TRANSFERRED finalized without the tuned refusal leg; k=25 legs cut** — the flagged text exists nowhere by design. → `persistent-memory.d/2026-09-08-run6-gate-csam-adjudication.md`
- `[2026-09-08]` **ESH static-WAN follow-ups landed (FortiGate trusthost3, esh-ana IPsec rebind, UDP 41641 → mesh direct); YTVC chased back up (nh3-scale SOCKS, stale yt-dlp layer, punkt_tab) and v0.3.6 CrisperWhisper deployed; gitea webhook repointed off the dead wg0 IP with the HMAC secret re-applied.** → `persistent-memory.d/2026-09-08-esh-static-wan-followups-and-ytvc.md`
- `[2026-09-08]` **ERP run 5 = RESCUED (landmark R49.5)** — first capability-gate pass in the ERP-seat line; the 3.46%-loss dependency-forcing slot (GovReport+QMSum) broke the coupling runs 3c/4 couldn't. Seat `erp-tune-v5` served on gx10:8098, `trial` alias repointed 3c→v5. → `persistent-memory.d/2026-09-08-run5-rescued.md`
- `[2026-09-08]` **R47 base settled from bytes = STOCK `google/gemma-4-26B-A4B-it`** — three-way sha match (local == HF etag == stock LFS oid; commit `4d7ae498` == stock HEAD); the `-heretic` label is a naming error, all runs trained from stock. Accept-vs-swap now evidenced. → `persistent-memory.d/2026-09-08-base-provenance-stock.md`
- `[2026-09-08]` **yt-voice-clipper back UP** — dead since the 09-06 danted retirement (every job failed at yt-dlp, bot-gated on the Irvine datacenter IP). Fix: danted on **nh3-scale** (CT107) at `socks5h://100.64.0.1:1080`, fleet-ACL'd, residential egress 70.230.226.88 measured; `YTVC_PROXY` repointed, worker recreated, end-to-end job DONE with positive (proxied) + negative (direct = bot-gate) controls. Homepage card href/siteMonitor → `irv-ml1.nh3.internal:8000` (was dead wg0 IP). Then a SECOND fault: full downloads 403'd through the proxy (cookies irrelevant) = stale yt-dlp 2026.07.04 from a cached Dockerfile layer → `compose build --no-cache api` (2026.08.19), which dragged in a whisperx/nltk that needs `punkt_tab` → staged on the data volume + `NLTK_DATA` in the override. Operator's video x7kWJojf1MI → done, 8 clips. yt-voice-clipper-dev shipped both Dockerfile fixes + **CrisperWhisper 2.0 (v0.3.6, `b62849d`) — deployed and verified (12 clips, [UM]/[UH] tags)**. ⚠ The gitea push webhook had been targeting the dead wg0 IP since 09-06 (never fired) → repointed to `10.6.110.50:9008` with the HMAC secret re-applied; deploy script passes `YTDLP_REFRESH`. Script `scripts/setup-nh3-scale-socks-egress.sh`. → auto-memory `reference_nh3_egress_proxy`, `reference_ytvc_autodeploy`.
@@ -732,47 +223,11 @@ below is a live commitment or a known-open risk._
- `[2026-08-26]` **Served under a NEW name on a NEW port (`erp-tune-v2` / :8098), never re-pointing `erp-tune-v1`.** Run 1's artifact still exists and is still what that name refers to; re-pointing would be the silent substitution the standing no-false-aliases rule forbids. brokkr independently asked for the same and additionally wants the concrete backing model + date in provenance, not just the alias — an alias has silently changed meaning under recorded results before.
- `[2026-08-26]` **DPO stage gated on an axis-list decision that is not mine to make** — `docs/pfi/erp-dpo-stage-prep.md`. No preference data for refusal axes exists; `trl` is not installed; the Gutenberg sets on disk are prose-quality only. ⚠ Do not install `trl` (or anything) into the training venv **while a run is saving** — a resolution that upgrades transformers under a live process can break its save path.
- `[2026-08-25]` **The ERP/RP tune COMPLETED in 7.36h and passed its gate on the axis it was built for** — diversity 22x its noise floor, attractor −11.3pt, zero memorisation on both arms. Also the noise-floor near-miss: brokkr was one step from reporting a 13-point T6 regression sitting inside twice his instrument's own variance. → `persistent-memory.d/2026-08-25-erp-tune-run2-complete.md`
- `[2026-08-25]` **8.6% MFU was an accounting artifact — real utilisation 17-20%, and the cost was attention on AMPERE kernels.** Two independent methods agreed to 2.6 points. Fixed by bucketing (padding 29.9%→0.0%) plus flex_attention. ⚠ Carries the dynamo recompile-ceiling trap that produced two wrong published conclusions. → `persistent-memory.d/2026-08-25-mfu-root-caused-attention.md`
- `[2026-08-25]` **NVFP4A16 serving pipeline built and validated; MERGED WEIGHTS ARE MANDATORY.** vLLM cannot serve a LoRA on ANY Gemma-4 — `get_expert_mapping` is unimplemented and the check branches on MoE-ness, not quantization. Plus the landmine: a `targets=["Linear"]` recipe misses all 11,520 expert tensors silently. → `persistent-memory.d/2026-08-25-nvfp4-serving-pipeline.md`
- `[2026-08-25]` **Refusal retention measured (base 0/100 → tuned 29/100, 71 still complying) — but on the WRONG AXIS.** `harmful_behaviors` is general harm; the abliteration was run for explicit fiction. The convenient set with a recorded baseline was not the right one. → `persistent-memory.d/2026-08-25-refusal-retention-probe.md`
- `[2026-08-25]` **Worldtree b188 + b189 shipped; bridge extracted to `pfi/wt-matrix-bridge` because `vh` is a USER not an ORG** and no service account can ever publish to a user namespace. Plus the selene catalog entry that lied about what answers, and a #411 diagnosis I got wrong twice before a directory probe settled it. → `persistent-memory.d/2026-08-25-worldtree-b188-b189-and-selene.md`
- `[2026-08-25]` **Run 2's base is an OPEN OPERATOR DECISION, deliberately not staged** — four options with materially different safety postures, detailed in Current state. Tracked at althing thread `01M0WQ8W5574KMEVCHCEKEXNS5`. ⚠ Do not let it get filed as a config knob; it is a reversal of the trainee-selection decision.
- `[2026-08-25]` **Fused MoE kernel path — DEFERRED, tracked at park `fused-moe-kernel-path-for-gemma-4-moe-training` (id 47).** Operator: "note the fused MoE kernel for round two… if we nail it soon, the math has us wanting to restart the run anyway." Training MFU is **8.6%** (27.1 of a benchmarked 313.8 TFLOPS) because `transformers` runs the Gemma-4 experts in a Python loop — 128 experts × 30 layers, ~11,500 iterations per step under gradient checkpointing. ⚠ **The same fused 3-D expert layout that made bitsandbytes skip 88.5% of the model is exactly what a grouped GEMM wants** — the format is good for storage and for fused kernels, and hostile only to naive iteration. Two fixes: `group_by_length` (−29.9% compute, free, but breaks the seeded order manifest and re-opens a batch-composition call brokkr already made) and a grouped-GEMM/compiled MoE forward (the remaining ~10×). **Not applied to the live run** — restarting mid-flight to change batch ordering was judged a bad trade at step ~50 of 1,312.
- `[2026-08-25]` **The ERP/RP tune LAUNCHED after 12 harness defects and an operator override of the corpus gate.** Four of the twelve would have crashed the run; two were INERT GATES that passed because they could not fail. Run is `/tank/erp-tune/run-01`, harness eitri-smithy `997c4a4`. Full arc — override, defects, sizing, the measured MFU — in the in-flight section and `docs/pfi/gemma4-erp-tune-sizing.md`.
- `[2026-08-24]` **char-rp seat swapped to the Gemma-4 26B-A4B MoE; abliterated trainee base staged and measured.** OOM root-caused to `--gpu-memory-utilization` not covering CUDA context (and to gen's footprint GROWING WITH UPTIME); a benchmark finding retracted because it scored below chance; abliteration isolated at −0.6 core points but it MOVES capability rather than removing it. → `persistent-memory.d/2026-08-24-charrp-gemma4-moe-swap-and-trainee.md`
- `[2026-08-24]` **Serving the tuned ERP model: LoRA-on-NVFP4 PREFERRED, merged weights the expected fallback — and the recorded objection may be STALE.** Operator: "if you CAN load it as a lora, all the better, the issue is that we will want to run nvfp4 weights, which we had some serious trouble with loading loras on top of nvfp4." ⚠ **The archived root-cause says it was NOT NVFP4-specific**: `[2026-07-07]` vLLM 0.24.0 qwen3_5 LoRA application was a silent no-op (#47639, regression from #37912) — adapter loads HTTP 200, zero deltas at inference, proven **quant-agnostic (NVFP4 AND FP8 both inert)** and adapter-format-agnostic by a 3-peer dwarf panel. Fix PR #47640 was OPEN then. **ana-ml2 is FAR past 0.24.0 and the box runs a SPREAD, not one version** (measured 2026-08-24): `gen` on `nightly-311b3513` = **0.27.2rc1.dev150**, `mog-sec` on `nightly-e9d1398d` = 0.26.1rc1.dev1102, the small seats still on 0.24.0, and char-rp/trainee-bench pinned to v0.26.0. ⚠ **`vllm/vllm-openai:v0.27.1` is already ON DISK, unused** — a TAGGED release, which is the right retest target: no nightly variance, no pull, ~4 months past the diagnosis. So: RETEST hot-swap LoRA on **v0.27.1** before designing around merge — it is cheap, and if it works the post-tune gate can be two aliases on one engine. If it still no-ops, merged weights it is, which means the harness must EMIT merged weights and Eitri needs that in the contract while he is early. Tracked at this snapshot commit; settle it in the QLoRA sizing conversation.
- `[2026-08-24]` **Homepage rebuilt on Australis Skyfall; light mode shipped.** Two findings worth more than the theme: **(a)** the Skyfall bundle including its canonical light ramp was sitting in this repo's git history at `45c1995` — check `git show` before concluding a vendored design asset is lost; **(b)** removing `theme:` from `settings.yaml` deterministically breaks the dashboard render (six recreates empty, restoring the key fixed it in 12s), which is the first confirmed cause of the "tab bar goes missing" symptom. Retires the `homepage.log` size lead from earlier the same day — it did nothing on this episode. → `persistent-memory.d/2026-08-24-homepage-uniform-grid.md`
- `[2026-08-24]` **Homepage reorganised on the axis "do I open this?" — UI groups expanded on top, API/agent groups collapsed at the bottom** (operator-delegated: "re-categorize however you want"). Load-bearing constraint: `homepage.group` is read at container CREATION, so the 16 GPU-backed model seats keep their unlovely names rather than eat a recreate — `initiallyCollapsed` + order is free. Second rule discovered here: **group members should all have widgets or none should**, because a stat strip adds ~50px and opens a void beside plain cards. → `persistent-memory.d/2026-08-24-homepage-uniform-grid.md`
- `[2026-08-24]` **Homepage columns unified at 4 for every group; the 2026-08-18 "columns = member count" rule is retired.** It was avoiding dead cells in a short last row and bought a worse defect — card width changing at every group boundary. Also carries two CSS traps: `overflow: hidden` clips at the PADDING box (so a `padding-right` gutter is spill room, not a guard), and a `:root` override of a Homepage theme variable is silently outranked by `.theme-slate` on the same `<html>` element. → `persistent-memory.d/2026-08-24-homepage-uniform-grid.md`
- `[2026-08-24]` **AES-128 adopted on both Anaheim tunnels; the per-flow ceiling root-caused to the UDM's software AES-CBC, exonerating the FortiGate.** Proven by an A/B/A cipher swap at identical CPU — hardware offload is not cipher-cost-sensitive. → `persistent-memory.d/2026-08-23-anaheim-ipsec-tunnel-ceiling.md`
- `[2026-08-24]` **ana-gw's public admin surface closed to zero open ports, ACME listener included.** Two of my diagnoses were wrong first (an "ISP proxy" that was the FortiGate, and an "all-port VIP" alarm that was a parser gap) — both from reading config instead of measuring from outside. → `persistent-memory.d/2026-08-24-ana-gw-admin-closed-acme-disabled.md`
- `[2026-08-24]` **Scriberr deployed on ana-ml2 GPU1, image built from source.** Three upstream bugs: the Blackwell image was never published, it must run as uid 10001, and `UV_LINK_MODE=copy` is required or two backends fail silently. → `persistent-memory.d/2026-08-24-scriberr-ana-ml2.md`
- `[2026-08-24]` **ESH DNS fixed at the IPv6 layer and the naming scheme went live on three hosts.** UniFi's RDNSS cannot be disabled but CAN be redirected — the field is only honoured when an explicit server is given. → `persistent-memory.d/2026-08-24-esh-dns-rdnss-and-scheme-live.md`
- `[2026-08-24]` **`speaches` on irv-ml1 stopped, stack retained** — Eyra was abandoned pre-implementation (Scriberr covers the need), leaving it no consumer. Disposition confirmed to eyra-dev; one command to restart. Tracked at althing thread `01M0RRJX8GPZEBDHF1E3W18RZF`.
- `[2026-08-24]` **esh-vm-db brought onto the fleet infra-ops identity and given its first vaulted credential.** It previously had none: root and infra-ops refused key auth and `lkraven`'s sudo wanted a password nobody held, leaving `qm guest exec` from the hypervisor as the only privileged path. Break-glass root password at `secret get esh-vm-db/root-breakglass-password` (console-only; plaintext never crossed the wire — only its SHA-512 hash did).
- `[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`.
- `[2026-08-23]` **Anaheim's IPsec tunnel ceiling — investigated, then CLOSED 2026-08-24.** The 25%-of-2-Gbps framing was wrong (NH3's uplink is 1 Gbps); AES-GCM proved impossible; AES-128 landed instead. → `persistent-memory.d/2026-08-23-anaheim-ipsec-tunnel-ceiling.md`
- `[2026-08-23]` **selene retired after losing a head-to-head on its own job; `chat-judge` moved to gen, the model name 404s by design.** Also surfaced that **7 aliases share one seat** — cross-checking between them is an echo, which caught a real defect in brokkr's 46k-exposure R47 gate. → `persistent-memory.d/2026-08-23-selene-retired-alias-collision.md`
- `[2026-08-23]` **hrafn adopted; its CI reported green for its whole life while deploying nothing.** A staging dir inside the rsync target destroyed its own source mid-copy; the deeper fault was verify steps that asserted uptime, never content. → `persistent-memory.d/2026-08-23-hrafn-adopted-ci-frozen-source.md`
- `[2026-08-23]` **Worldtree b187 shipped; all three instances de-armed from a 69-day-stale `:latest`; Matrix homeserver re-plumbed to personal.** Includes the `:8009`-is-demo port trap that an IP-only fix would have walked into. → `persistent-memory.d/2026-08-23-worldtree-b187-pins-matrix.md`
- `[2026-08-23]` **Every secret-bearing `.env` on ana-docker tightened to 0600** — eight stacks including vaultwarden and traefik, verified exposed by reading one as `nobody`. → `persistent-memory.d/2026-08-23-ana-docker-env-perms-sweep.md`
- `[2026-08-23]` **`pfi` gitea org created; claude-bot is an Owner and creates repos self-serve.** Closes the repo-creation half of the credential-migration directive — `vh` is a USER namespace so no service account could ever create there. Repo creation needs `write:user` + `write:repository` + `write:organization`; `POST /users/{u}/tokens` is basic-auth only, so minting needs the account password. Default new repos to `pfi/`. (`vh/eitri-smithy` was its first tenant, then moved.)
- `[2026-08-23]` **Booth: kept boards are deletable and link rows are prunable.** `release` on a kept card drops the sentinel so the existing × applies; `booth links` / `booth unlink <id|index>` prune one row. Rows are addressed by **content id, never position** — the board is append-only and multi-writer. **Releasing a board RESETS its TTL clock** (unlink bumps the dir mtime), so unkeep-and-wait is a 24h delay, not a delete. (`4be880f`, `0ad332b`)
- `[2026-08-22]` **DFlash2 spec-decode measured on our own stack; `sec` promoted to it.** +18–21% accepted length and +15–18% throughput over MTP k=3, drafter proved model-agnostic across two finetunes to 0.06%, and the k=7 MTP *control* showed deeper MTP is a throughput trap. → `persistent-memory.d/2026-08-22-dflash2-spec-decode.md`
- `[2026-08-22]` **Quant pipeline shipped a crippled tokenizer for months — fixed at source.** `quant_mixed_nvfp4.py` baked its calibration truncation (`max_length 2048`) into every mixed-NVFP4 build; latent on old transformers, fatal on new. Both live quants corrected, pipeline now saves a source-pristine tokenizer and asserts it. Playbook §3.14. (`0755ba7`)
- `[2026-08-22]` **`sec` retuned to util 0.52 / 420K after a runtime OOM at 0.55/480K** — `gpu-memory-utilization` is not a hard reservation; activation grows past the dummy-data profile and six vLLM containers share GPU1. Also measured: the KV pool varies ~6.6% between boots, so max-model-len must be sized against the *lower* observation. (`6e82899`)
- `[2026-08-22]` **Max-Q 1.8× spread does NOT apply to LLM decode — measured, not argued.** ana-ml2 draws 256–266 W of 300 W under sustained 100% decode with `SW Power Cap: Not Active` and clocks pinned. Corrected to brokkr-smithy-dev after I had lent the claim credibility; 122B figure (~90–93 tok/s at 262K) stands as a straight number.
- `[2026-08-21]` **ESH internal IPv6 live on two LANs; the Cityside v4 static is a CARRIER problem, proven.** A full gateway reboot forced a fresh DHCP DISCOVER and returned the identical CGNAT address. YaRN was already configured — "1M needs YaRN, absent" was false. → `persistent-memory.d/2026-08-22-dflash2-spec-decode.md` sibling entry in `ad21302`
- `[2026-08-19]` **AI-tab Dormant regrouping BELAYED by the operator** — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong, Chatterbox Fast) show amber EXITED inside live groups rather than `AI - Dormant`. Fix is a label change + recreate per stack; needs the operator's read on which are retired vs temporarily down. `untracked by operator choice` (his words: "belay the ai dormant regrouping for now").
@@ -805,26 +260,12 @@ below is a live commitment or a known-open risk._
- `[2026-08-09→10]` **dots.tts (rednote-hilab) TTS burn-in on irv-ml1 + canonical voice corpus built (`voices/`).** Operator-directed eval to potentially replace chatterbox-fast. **dots.tts VERIFIED real** (canonical HF ns `dots-studio/`, `rednote-hilab/dots.tts-*` redirects there; Apache-2.0; PyPI `dots.tts` 0.2.1; 2B continuous-AR = semantic enc + Qwen2.5-1.5B LLM + flow-matching acoustic head over 48kHz AudioVAE; zero-shot clone from wav+transcript). **Runs on Ampere 3090** (sm_86, bf16, no fp8 dep); **optimized RTF 0.22** at num_steps=10 (`from_pretrained(..., optimize=True)` CUDA graphs — raw unoptimized was 1.21), **~6GB VRAM**, 48kHz, streams (`generate_stream`). Venv+cache at `irv-ml1:/home/lkraven/dots-tts` (~10GB). **Operator design calls:** SGLang Omni serving (OpenAI `/v1/audio/speech`), transcribe-refs-first, `soar` variant. ⚠ Omni serves soar but its continuous-batching + streaming opts are **mf-only** (soar = single-request) — non-issue for ratatoskr's single-consumer RP surface. **KEY FINDING — dots is highly sensitive to an accurate AND sentence-bounded reference transcript:** mismatched transcript → 0.16s collapse; over-long/messy transcript → reference-audio BLEEDS as an output prefix; mid-clause trim → dangling-word leak (glados "we'll", emmie "And,"). RECIPE (baked into `voices/derive.py`): trim ref to a clean ~6–10s clip ending on a sentence boundary + accurate transcript of exactly that clip. **CANONICAL VOICE CORPUS** stood up in eshpfi `voices/` (operator idea): engine-agnostic `canonical/<v>.wav` + `transcripts/<v>.txt` → per-engine ref sets DERIVED by `derive.py` reading `engines.yaml` profiles (dots/chatterbox/zonos); canonical wavs git-tracked (small/curated), `derived/` gitignored. **4 voices optimized + verified CLEAN for dots: donut, glados, emmie, miranda** (glados canonical is low-SR 16kHz — flagged upgrade candidate). ⚠ GPU GOTCHA: irv-ml1 native CUDA orders **A6000=device0** (ComfyUI-full) — pin the 3090 with `CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=0`; and `PYTORCH_CUDA_ALLOC_CONF=expandable_segments` CONFLICTS with `optimize=True` CUDA graphs (curr_block error). Booths: `dots-vs-chatterbox`, `dots-voices-optimized`. **SHIPPED 2026-08-10:** operator A/B verdict "dots is very good" → containerized as a **thin FastAPI wrapper over DotsTtsRuntime** (chosen over SGLang Omni — Omni's batching is mf-only, unneeded for ratatoskr's single consumer; wrapper is SERIALIZED one-gen-at-a-time via a threading.Lock, Omni+mf = parked API-compatible escalation if multi-consumer ever lands). **LIVE on irv-ml1:8198** (`local/dots-tts:v1`, OpenAI `/v1/audio/speech` + `/health` + `/v1/voices`, container healthy, both stream + non-stream verified CLEAN, 4 voices donut/glados/emmie/miranda) alongside chatterbox :8197 (nothing repointed). Stack = `stacks/dots-tts/` (Dockerfile/app.py/compose/.env.example/README). ⚠ CONTAINER GOTCHA: `optimize=True` (torch.compile/inductor/triton) needs a **C compiler at RUNTIME** — slim image must `apt install build-essential` or model-load dies "Failed to find C compiler" (host venv had gcc ambient, masking it); persist `TORCHINDUCTOR_CACHE_DIR` to a mounted dir or every restart re-JITs ~5min. Corpus home = eshpfi `voices/` (operator ruled keep-here). **REMAINING: ratatoskr client cutover** to :8198 `/v1/audio/speech` (Phase-2 tail, peer-coupled — draft the ask). [[reference_chatterbox_fast_repo]] [[reference_zonos_tts_stack]] [[reference_verify_hf_repo_ids_before_pull]]
_Older entries archived to archival-memory.md._
_248 older entries archived to archival-memory.md._
_275 older entries archived to archival-memory.md._
## Tried and abandoned
- `[2026-09-04]` **Forcing 10G on the ESH-Media DAC — it linked, then degraded over hours, and I reported a plateau at two minutes.** ⚠ A clean zero-error link at 1G does NOT rule out a marginal cable; autoneg's fallback was protecting something real. → `persistent-memory.d/2026-09-04-dac-forced-10g-failed.md`
- `[2026-08-25]` **Four throughput levers measured and killed — do not re-chase.** (1) **Fused MoE / `grouped_mm`** — 0.9% *slower* than the Python loop and dense GEMM is only 7.9% of the step, capping the whole category near 10%. (2) **CUDA graphs / `torch.compile` over the expert loop** — the two-term scaling fit closed with residuals under 3ms and needed NO constant term, so there is no fixed per-batch cost to amortise; 3,840 expert-GEMM launches per forward are not what we pay for. (3) **`liger` fused linear CE** — the chunked CE measured **1.1% of the step** forward, ~3% with recompute. A tidy-up, not a lever. (4) **Selective gradient checkpointing** — ~2% of a post-fix step, real bug surface. Also: **token-budget batching is dead by the same fit** — with no constant term, total time over a fixed set of widths is invariant to how you group them; only the widths matter, which is exactly why bucketing works and repacking does not.
- `[2026-08-25]` **`sample_packing` is NOT strictly better than bucketing on this model, and I told the operator it was before brokkr corrected me.** Packing needs FA2 varlen or a block-diagonal mask; FA2 is unavailable here (head_dim 512 > 256 cap), so packing means an explicit 4D mask on EVERY batch. Bucketing produces **78.3% exactly-zero-pad micro-batches** which recover the `is_causal` fast path on the 5 global layers — measured at 9.4% of step time. Packing forfeits that. ⚠ **The conclusion flips under `flex_attention`**, where a block-diagonal mask is just another BlockMask: do not carry "packing is bad" past the backend decision.
- `[2026-08-25]` **Merging a tune back toward STOCK to fix overfitting would UNDO the abliteration.** brokkr recommended a 50/50 merge-back, then retracted it himself: the published recipes merge into `google/gemma-4-*-it`, and following that literally re-installs exactly the refusal directions the abliteration removed — silently, because the merged model looks *healthier* on general benchmarks. Any merge-back must target the SAME abliterated base. Wider lesson: **recipe cards are per-checkpoint artifacts, not per-family** — the advice came from a card for a DENSE STOCK 31B applied to a MoE ABLITERATED 26B-A4B, three axes apart on a shared name.
- `[2026-08-24]` **AES-GCM on the Anaheim tunnels — impossible, not merely hard.** UniFi's manual site-to-site IPsec implements no AEAD cipher at all: eight GCM spellings rejected `api.err.InvalidPayload` against a passing `aes256` control. Blocks both tunnels since both far ends are UDMs. Accepted enum is `aes128/aes192/aes256/3des` — and 3DES is *slower* (no ARM instructions, 64-bit blocks), so AES-128 is the floor.
- `[2026-08-24]` **Pointing the UDM's `wan_dns1` at AdGuard — silently ignored.** It persists and reads back correctly but the LAN-facing forwarder never uses it; proven with fresh uncached ad domains (AdGuard answers `0.0.0.0`, the UDM returned real IPs). Reverted rather than left in place.
- `[2026-08-24]` **A multi-DUID DHCPv6 VM to claim NH3's seven unclaimed /64s — declined by the operator.** The BGW has no IP-passthrough (confirmed, we hold admin), so the only route needs re-cabling, split-stack routing and **rebuilding the entire v6 firewall policy off the UDM**. The prefixes are easy; the firewall rebuild is why nobody wants them. Do not re-raise on "there are seven free prefixes".
- `[2026-08-23]` **A `HEAD == GITHUB_SHA` assertion in the hrafn CI — added, broke the checkout twice, removed.** It needed the `git` binary (run 9920, exit 127); installing `git` then flipped `actions/checkout@v4` off its **node** implementation onto the git binary, which died on a missing CA bundle (run 9921). A nice-to-have assertion changed the checkout's code path and broke a working pipeline. Removed rather than patched with `ca-certificates` — it guarded a hypothesis that proved wrong. **Do not add `git` to that prereq step.**
- `[2026-08-23]` **Repointing `selene-1-mini-8b` at gen's endpoint — proposed by me, correctly overruled.** *"never repoint a named model at a different model's endpoint — that is intentionally misleading."* The trap is that it does not feel like deception; it feels like sparing consumers a migration. That framing is the tell. Role aliases move; model names die with the model and 4xx.
- `[2026-08-03]` **ComfyUI `--enable-triton-backend` on the irv-ml1 A6000 crashes EVERY render — Ampere has no hardware e4m3.** adhoc-agent's operator-approved probe: comfy_kitchen's triton backend has a FUSED int8 matmul that would beat the eager backend's ~1.9x-slower unfused int8 path (21.3s vs 11.2s fp8 on the Moody Krea2 int8 checkpoints). Flipped it (added to `COMFY_CMDLINE_EXTRA`, recreated) → `triton.compiler.errors.CompilationError: ValueError("type fp8e4nv not supported in this architecture. supported: fp8e4b15, fp8e5")` in `comfy_kitchen/backends/triton/quantization.py:145 dequantize_per_tensor_fp8`, failing at **node 5 CLIPTextEncode**. Triton's fp8 dequant kernel targets `fp8e4nv` (Hopper/Ada e4m3); **sm_86 Ampere (A6000) lacks hardware e4m3** → the JIT compile dies. With triton on it grabs the **global** `--fp8_e4m3fn-text-enc` dequant, so every render (fp8 AND int8) dies upstream at the text-encode step — the int8 UNet path never ran, so the convrot-coverage caveat wasn't even the limiter. Reverted cleanly (~15s to healthy, image unchanged `sha256:94afb8ca`, sage intact, prod restored). **The parked cu130 rebuild won't fix it** (e4m3 = hardware format, not CUDA version). **DEFERRED to the Ada refresh** (operator: "ada is coming, we'll optimize then" — Ada sm_89 has native e4m3, so triton's fp8 path should compile there). **Mechanics:** `--enable-triton-backend` is a compose `environment:` var, so toggling it needs `docker compose up -d` (**recreate**), NOT `docker restart` (reuses the baked env, no-ops silently). Full: auto-memory `parked_triton_backend_ampere_fp8`.
_144 older entries archived to archival-memory.md._
_152 older entries archived to archival-memory.md._