memory: snapshot — NH3 outage recovered, VM 102 retired + RTX 2000 Ada installed (LXC decision), gx10 AC-restore unvalidated, 40 entries archived

This commit is contained in:
vh
2026-09-24 21:52:48 -07:00
parent 7cbd4c3781
commit eb46973051
28 changed files with 987 additions and 830 deletions
+65 -143
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-09-22 ~21:20 PT (⭐ BOTH carried decisions APPROVED — build the NRestarts flap sampler; the restic content-assertion ruling is ratified. safe-rm installed on nh3-dev and delegated fleet-wide to infra-hermes with the wiring ruled. D-0010/D-0011 were misrouted here by pane_find matching a rolling pane title — belayed, nothing lost. ⚠ restic/ana/esh-docker-vm has drifted 36h→44h against a 48h threshold.)_
_Last updated: 2026-09-24 ~2150 PT (NH3 power outage recovered: pbs-nh3 onboot set, NFS → automount. esh-pve: VM 102 retired, a 14-day hung VFIO process cleared, T400 → RTX 2000 Ada; ⭐ NEXT = provision it as LXC + host driver for embed/rerank. pfi-gx10 AC-restore patch applied but UNVALIDATED — box OFF until Prime's AC pull 09-25. elway root:root fix + fleet ownership audit. Miranda standing order + Prime callsign in CLAUDE.md. task-board mothballed.)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
@@ -29,7 +29,7 @@ Sister repos (separate gitea repos, deployed by playbooks here):
| Repo | Role | CI status |
|---|---|---|
| `vh/task-board` | MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) |
| `vh/task-board` | MCP + web dashboard for assistant task state (port 7878) — **[2026-09-24] MOTHBALLED** by Prime (superseded by the High Seat + ledger); container removed on ana-docker, data/image/compose kept (`e6da607`) | push-to-main → CI deploys (2026-04-29) |
| `vh/vor` | Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
| `vh/nevermore` | Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
| `vh/asset-engine` | Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
@@ -109,72 +109,84 @@ no longer deployed sidecars here. See Recent decisions.)
passwordless sudo); **`ssh infra-ops@ana-docker` HAS NOPASSWD root**. **→
For any sudo op on ana-docker, use `ssh infra-ops@ana-docker`.** `ssh
infra-ops@10.100.10.50` (nh3-dev) ALSO NOPASSWD sudo; on **nh3-extdev** infra-ops
is sudo-LESS by design (`ssh lkraven@10.100.50.42` is the NOPASSWD path). **irv-ml1:
is sudo-LESS by design (`ssh lkraven@10.100.50.42` is the NOPASSWD path) — **[2026-09-23] measured: infra-ops HAS NOPASSWD sudo on nh3-extdev** (fleet ownership audit; also CLAUDE.md 2026-09-05). **irv-ml1:
`ssh irv-ml1` = lkraven, docker-group (plain docker) but sudo needs a PASSWORD
(no NOPASSWD)** — stage model pulls to `/home`, not root-owned `/worktank`.
## Current state / in-flight
_As of 2026-09-22 ~21:20 PT._
_As of 2026-09-24 ~2150 PT._
### ✅ BOTH CARRIED DECISIONS APPROVED — build them
### ⭐ NEXT: provision the RTX 2000 Ada on esh-pve as an LXC + host NVIDIA driver
Operator, 2026-09-22 ~21:20: *"decide to do both pending."* They are no longer
decisions; they are work.
Prime's decision (2026-09-24, over a VFIO VM): the card serves **embedding +
reranking offload**. Installed tonight in esh-pve's single slot (`01:00.0`,
`10de:28b0`), no driver bound. Plan shape, not yet started:
1. **BUILD the `NRestarts` flap sampler.** Design and every trap already written
up in `services/althing-notify-failure/README.md`. Timer reads `NRestarts`
per unit, alarms on delta over a window, reuses the existing cooldown and
alert body. ⚠ Store a last-seen TIMESTAMP beside the count so a counter going
BACKWARDS registers as a reset rather than as quiet — `reset-failed` zeroes
it and sits on the remediation path of the other alarm. Tracking: `163bb97`.
2. **The restic content-assertion ruling is CONFIRMED.** It reached me relayed
by svos-dev rather than Miranda and I built it anyway as reversible; the
operator has now ratified it. Nothing to undo. Tracking: `ba60fda`.
1. NVIDIA driver on the PVE host (headers for the running `6.8.12-*-pve`
kernel; DKMS). Blacklist nouveau. Remove the stale T400 vfio ids from
`/etc/modprobe.d` (`10de:1ff2,10de:10fa`).
2. An LXC (Debian 12 template — PVE here rejects Debian 13) with the GPU
device nodes bound in and the SAME userspace driver version as the host;
docker + nvidia-container-toolkit inside.
3. Serve the SAME models the fleet already uses so vectors stay compatible:
`qwen3-embedding` (currently fv-ml1:8001 via LiteLLM) and the rerankers
(`reranker`, `reranker-a3-bge-v2-m3`). Confirm sizes from the live seats
before choosing an engine (TEI vs vLLM).
4. Wire as a LiteLLM failover/local deployment under the SAME model names; ESH
consumers (Open WebUI, Paperless) keep working if FV or the mesh is down.
### ✅ safe-rm is fleet-wide — 6/6, closed
⚠ esh-pve is ESH's only DNS and its only mesh route (esh-scale CT 108 lives
there). Driver work means reboots: do them when an ESH outage is acceptable,
and confirm power-off by the light, not by ping (my path in dies with esh-scale).
Delegated to infra-hermes and complete. Acceptance met on every host as
infra-ops over non-interactive ssh: `command -v rm` resolves to the wrapper, the
guarded probe printed `safe-rm: Skipping /home.`, normal deletes unaffected.
Backups at `/etc/bash.bashrc.bak-saferm` per host.
### pfi-gx10 is OFF — AC-pull test 2026-09-25 (Prime)
Wiring per my ruling — one self-guarded line above the `case $-` guard in
`/etc/bash.bashrc`, **not** `/etc/environment`: an rc can self-test with `-d`,
its failure blast radius is smaller, and it covers infra-ops-bash-over-ssh which
is the threat path. Accepted loss: cron and non-bash `sh`.
UEFI "Restore AC Power Loss" patch applied, **unvalidated**. Do NOT test it with
a shutdown (stays off by design). Outcomes and revert in
`servers/pfi-gx10/README.md`. → `persistent-memory.d/2026-09-24-gx10-ac-restore-patch.md`
⭐ **The acceptance test caught a real bug in flight.** His first insertion pass
failed silently on all six hosts (`$$` collapsed inside heredoc quoting), and
the five `NOT-GUARDED` lines were the guarded probe correctly refusing to test a
dead guard. A probe that cannot destroy what it tests also cannot lie about it.
### Still open from 09-22
⚠ **safe-rm does NOT cover the habit that prompted it.** Measured twice,
independently: it refuses `rm -rf /home` and deletes an unset-variable path
without complaint. Blacklist, not heuristic. `set -u` is the actual cover, and
anyone writing this up must say so or it will be trusted for a class it does not
protect.
- **Build the `NRestarts` flap sampler** — approved 2026-09-22, not started
(design + traps in `services/althing-notify-failure/README.md`; `163bb97`).
### Live threads
- ⚠ **`restic/ana/esh-docker-vm` is now 44h old** against a 48h threshold and
12h for every other repo. It was 36h this afternoon, so it is drifting, not
static. If it crosses 48h the check turns STALE and pages. Plausibly a missed
window from that host's forced reboot on 09-21; **not yet confirmed, and worth
confirming before it alarms.**
- **`talk.service` still `failed` (exit 143) while `:8092` serves 200.** The
unit is dead, its containers keep running, nothing manages talk. Unchanged
since this afternoon.
- `headscale-ddns` recovered on its own (`success`/`inactive`) after the
transient failure at 15:28; the diagnostics added then are untested against a
real recurrence.
### 20 commits unpushed
Push is the operator's call. Tree clean, nothing half-done.
- **Claude sessions on nh3-dev** died in the NH3 outage (only infra-ops,
infra-hermes, jekyll survived). Relaunch is Prime's call; not confirmed done.
- **Uncommitted change NOT mine:** `stacks/homepage/conf/services.yaml` renames
the Homepage card `infra-hermes seat` → `hermes-gateway seat` "per operator
handle-split ruling 2026-09-24" — a ruling I have no record of. Find the
author before committing or reverting; if infra-hermes was renamed, the
CLAUDE.md infra-hermes section is stale.
- **infra-hermes owns a daily 0110 job** proving the first High Seat report
(`~/.high-seat/reports/*.jsonl`) lands in a nh3-dev restic snapshot; it
replies to svos-dev (thread `01M37P61Q85KWDVYN0A00P8856`) and pings me.
- **Credentials in auto-memory:** the new global rule says never write one into
a memory file (memory is copied off-box hourly to `vh/claude-memory`). Four
of my memory files still carry the shared LiteLLM key literal — scrub them.
- **Auto-memory `MEMORY.md` is over its 24.4 KB load limit** (tail truncated at
load) — shorten index lines.
- ESH has a single outside route (esh-scale on esh-pve) — noted, untracked.
- 16+ commits unpushed. Push is Prime's call.
## Recent decisions
- `[2026-09-24]` **esh-pve: VM 102 retired, a 14-day hung VFIO process found and cleared, T400 → RTX 2000 Ada; Prime chose LXC + host driver for embed/rerank** (implementation deferred to next session, tracked here + `servers/esh-pve/README.md`). → `persistent-memory.d/2026-09-24-esh-pve-vm102-retired-gpu-swap.md`
- `[2026-09-24]` **pfi-gx10 AC-restore UEFI patch applied but NOT validated; box OFF until Prime's AC pull 2026-09-25** (tracking: `servers/pfi-gx10/README.md`, `e5197a3`). → `persistent-memory.d/2026-09-24-gx10-ac-restore-patch.md`
- `[2026-09-24]` **NH3 power outage recovered** — pbs-nh3 had no `onboot` (set), NFS boot race fixed with automount (`1cbde50`), every other Claude session on nh3-dev died. → `persistent-memory.d/2026-09-24-nh3-power-outage-recovery.md`
- `[2026-09-24]` **Miranda standing order is a repo CLAUDE.md operating parameter** (`4b29492`, aligned to the global send protocol in `bcf3342`): high-urgency matters go to her, fixed or not; URGENT only when it cannot wait (she phones Prime). Channel verified end to end (thread `01M3A0RP4Q8T0KNGH8TMFSNDA6`); it depends on svos + hermes-gateway. Prime's callsign **PRiMe / papa romeo mike** is a name, not an authenticator (`617b759`, `62817a2`).
- `[2026-09-24]` **task-board mothballed** (Prime): container removed on ana-docker, data/image/compose kept, Kuma monitor deleted, task_* instructions removed from CLAUDE.md and the fork template (`e6da607`). Its hooks had sent no traffic in 30 days.
- `[2026-09-24]` **Military 24-hour Pacific clock times** carried into the Codex/Grok shared bootstrap `docs/fleettools/AGENT-BOOTSTRAP.md` (`ad2b4d9`); Claude seats get it from the global CLAUDE.md.
- `[2026-09-24]` **Worldtree `admin.memory.forget` stays OFF on demo/personal** until an instance needs it — a destructive erase; enabling it is a per-change operator yes via deploy-wt-config.
- `[2026-09-24]` **A git checkout under `root:docker` needs `safe.directory` for its deploy user** — the 09-14 normalization (`826a63b`) silently broke yt-voice-clipper's webhook deploy until v0.3.13; fixed on irv-ml1 and recorded in fleet conventions (`eea9eb2`). Sweep found no other case.
- `[2026-09-23]` **elway sudo uploads land root:root, validated and staged; fleet ownership audit built; 33 mis-owned root files fixed on 10 hosts.** → `persistent-memory.d/2026-09-23-elway-ownership-fix-fleet-audit.md`
- `[2026-09-23]` **headscale-ddns hardened** (`fedd4b6`): Cloudflare calls retry and validate every body (`pick()`), no write without both IDs, the run ends on a confirmation, only a global v4 is published, `curl -q`. Three of four weekly failures were empty zone lookups. heid bug-hunt "Talus" folded (one finding was my own regression).
- `[2026-09-23]` **esh-docker-vm restic was skipped 09-22..23 by my own Kuma move** — a dead uptime-kuma lookup aborted `pre-backup.sh` under `set -e` (`25e41d2`). Then, by Prime's decision, the redundant Paperless pg_dump went too: it had failed auth every night since 2026-04-24 and left a 0-byte dump in every snapshot; the DB is covered at source by esh-vm-db's pg_dumpall (`6e8da46`).
- `[2026-09-23]` **hermes-gateway restart exit-1 is a Hermes race, not a crash** — the planned-stop watcher consumes the marker before SIGTERM re-runs the handler; drop-in `SuccessExitStatus=1` stops the false OnFailure page at zero coverage cost (`c2b0a05`). svos had its own stop-timeout (an open board SSE tab), fixed by svos-dev with `timeout_graceful_shutdown=5`.
- `[2026-09-23]` **Booth link board cleared to 14 durable links** (208 removed: drops, release posts, research links, token-bearing URLs; Scriberr reposted at `scriberr.fv.internal:8080`). Prime's rule: durable debug links only.
- `[2026-09-22]` **Both carried calls approved** — build the `NRestarts` flap sampler (`163bb97`); the restic content-assertion ruling is ratified and stays (`ba60fda`).
- `[2026-09-22]` ⭐ **safe-rm installed on nh3-dev and delegated fleet-wide to infra-hermes.** ⚠ The package installs INERT and looks fine — Debian's `/etc/zsh/zprofile` has 0 non-comment lines so the shipped `profile.d` hook never fires under zsh; verify with `command -v rm`, never `dpkg -l`. ⚠ And it does NOT cover the habit that prompted it: measured, it refuses `rm -rf /home` and deletes an unset-variable path without complaint. Blacklist, not heuristic; `set -u` is the actual cover. Wiring ruled: `/etc/bash.bashrc` above the `case $-` guard, not `/etc/environment` (an rc self-guards, smaller blast radius, covers bash-over-ssh).
- `[2026-09-22]` **The acceptance probe for a guard must not be able to destroy what it tests** (infra-hermes). `rm -rf /home` to prove safe-rm refuses is a test whose premise IS the thing under test. Guarded form is now standard: `[ "$(command -v rm)" = /usr/share/safe-rm/bin/rm ] && rm -rf /home || echo NOT-GUARDED`.
@@ -393,111 +405,21 @@ Push is the operator's call. Tree clean, nothing half-done.
- `[2026-09-11]` ⚠⚠ **4B is the FIRST rung to OVERFIT inside one epoch, which inverts my earlier "one epoch is right for this corpus" call.** Series 2.832 · 2.816 · **2.814** · 2.820 · 2.824 · 2.825 · 2.825 — minimum at ~step 75, then it TURNS and settles worse. 0.6B and 1.7B both plateaued with no turn, so **the optimal epoch count shrinks as the carrier grows** — 4B wants roughly half an epoch. ⚠ **Consequence: the shipped `adapter/` at `h02-4b-1ep/` is NOT the best checkpoint** (it is the end-of-run 2.825); the step-75 checkpoint at 2.814 is, and it exists only because `save_steps=25` was set. The voice test used the end-of-run adapter, so the booth understates 4B by ~0.011 nats. Re-cut the arms off the step-75 checkpoint before any adjudication.
- `[2026-09-11]` **The tone-override appears to close at 4B too.** On the operator's Abernathy frame prompt ("a *wonderful* story"), 1.7B held the frame on every seed but **2 of 4 killed the animals anyway**; 4B kept them alive on **2 of 2** and one seed did something new — the narrator *doubts Abernathy's story* ("I felt sure the thing was a lie"), then supplies a parallel childhood memory of his own puppy and his sister's kitten to explain the doubt. That is a narrator with an interior position on the tale being told. ⚠ n=2 per arm; directionally right, not established.
- `[2026-09-10]` **R49 rung 3 LAUNCHED: Qwen3-4B-Base, 1 epoch, seed 4919, same unwrapped corpus** — `gx10:~/r49-runs/h02-4b-1ep/`, 159 steps at ~37.8 s/it (**~100 min**), 252 adapted modules (vs 196 at 0.6B/1.7B). Last rung of the planned sweep; it tests whether **scene-level continuity** closes with carrier size. A two-arm voice test (4B base + 4B tuned, the nine prompts plus the operator's Abernathy frame) is **chained behind it**, gated on the adapter existing.
- `[2026-09-10]` **AN AUTHOR-VOICE ADAPTER TRANSFERS SUBJECT MATTER, NOT JUST STYLE — and that was invisible to my own test ** → `persistent-memory.d/2026-09-10-an-author-voice-adapter-transfers-subject.md`
- `[2026-09-10]` **R49 rung 2 COMPLETE, and the single-variable carrier effect is clean: 0.6B held-out 3.329 vs 1.7B 3.018, ** → `persistent-memory.d/2026-09-10-r49-rung-2-complete-and-the.md`
- `[2026-09-10]` **R49 rung 2 LAUNCHED: Qwen3-1.7B-Base, 1 epoch, seed 4919, on an UNWRAPPED corpus.** → `persistent-memory.d/2026-09-10-r49-rung-2-launched-qwen3-1.md`
- `[2026-09-10]` **BabyBronte H02 adapter: the VOICE transferred, the SENSE did not — operator's read, "it's all nonsense, b** → `persistent-memory.d/2026-09-10-babybronte-h02-adapter-the-voice-transferred.md`
- `[2026-09-10]` **mog-sec (sec/sec-reasoning, ana-ml2 GPU0 :8019) SETTLED at MOG_MAX_MODEL_LEN=163840 + MOG_KV_CACHE_ME** → `persistent-memory.d/2026-09-10-mog-sec-sec-sec-reasoning-ana.md`
- `[2026-09-10]` ⚠ **Near-miss on measurement discipline, worth keeping as a specimen.** The crash window logged `Avg Draft acceptance rate: 17.6%` and per-position rates of 0.049/0.024/0.015 for draft positions 5–7, which reads as an obvious "cut `num_speculative_tokens` 7 → 3, it is buying nothing." Across **180 samples** of the same counter over the container's life the real distribution is **median acceptance length 3.12 of 7 (range 1.83–6.75)** and **median draft acceptance 30.4% (range 11.9–82.1%)** — the crash window was near the *minimum*, not the norm, and cutting to 3 would cap the workloads that were accepting nearly the full 7-wide draft. **The n=1 window pointed the opposite way from the n=180 distribution.** Same session that wrote "a positive control is only worth what it can distinguish"; the lesson generalises to log lines.
- `[2026-09-10]` **R49 carrier SETTLED on dense `Qwen3-{0.6,1.7,4}B-Base`, overriding H02's own pin — the newest carrier was the SLOW one.** Dense 4.089 B trains 33% faster than hybrid 0.765 B; no fused SSM kernel installed. D1–D3 built, 1-epoch pilot beats the 3-epoch by 0.21 nats held-out. → `persistent-memory.d/2026-09-10-r49-babybronte-d1-d3-and-the-1-epoch-pilot.md`
- `[2026-09-10]` **R49 adjudication routed to infra-ops entirely** (operator, relayed by brokkr: *"leave babybronte to infra — concentrate on r50 and the memory mechanism"*). brokkr handed over the Delta instrument and stepped off. ⚠ I now grade my own run; brokkr's decision rule is **ratified verbatim and frozen before any adapted text existed** and must not be amended after seeing numbers. Their controls: real Charlotte 1.65–2.17, **Anne at 2.374** — so the absolute band decides, never `nearest`.
- `[2026-09-10]` **MeroMero A4B swapped onto the `erp-seat` seat as `char-rp-fast`; `Pfish-6` alias removed.** The A4B's FIRST quant used the dense recipe and 4-bit-quantized all 30 MoE routers — it passed its healthcheck and answered every request with the full token count decoding to the empty string, NaN logits the only tell. Re-quantized with the MoE recipe; live and verified (prose, vision, tool call, finite logprobs). Durable lesson: **a positive control must match the ARCHITECTURE CLASS** — the broken A4B was diffed against a good *dense* quant, which has no routers, so the clean result was meaningless. → playbook §3.15, §4.4
- `[2026-09-10]` **MeroMero: BOTH quants landed in-house at W4A16 — A4B first try, v2 dense on attempt 5.** Published quants are all W4A4 (our measured long-context collapse) or nonexistent for v2. Operator: *"pull both ablits bf16, run our own quant."* The durable lesson is **§3.17**: `pip install llmcompressor` silently pins transformers down a version, so attempt 4's error was a moved toolchain, not the malformed upload it looked like — a known-good positive control is what told them apart. Serve test still owed. → `persistent-memory.d/2026-09-10-meromero-quants-and-the-pinned-transformers-trap.md`
- `[2026-09-10]` **althing 3.6.2 deployed — post office + both heralds — and the fleet has TWO herald nodes, not seven.** Ask the post office's `nodes` table, not the box inventory. Cost a self-inflicted ~12 min bus outage. → `persistent-memory.d/2026-09-10-althing-362-rollout.md`
- `[2026-09-10]` **A grep over a log that records your greps counts itself.** I reported forseti's drop defect as reproducing here with 3 drops in 21 s; the session had **zero**. Searching transcripts writes the search term into them. Filter by `"type":"system"` provenance, never content. Generalises to any instrument that can see itself. Auto-memory `feedback_grep_over_a_log_that_records_your_greps`.
- `[2026-09-10]` **Operator-directed purges: 466 GB (qwopus + huihui 122B bf16) and 107.8 GB Docker on ana-ml2.** Serving/rollback artifacts and qwopus's MTP head verified intact after. ⚠ `/tank` is OUTSIDE restic, so both were final.
- `[2026-09-10]` **ana-docker disk pressure repaired: root 84% → 51%, 115 GiB free.** Gitea/Vaultwarden backups repaired and restored from Restic `2ec5a37c`; 101 stale dumps removed; hourly named-builder cache pruning installed. → `persistent-memory.d/2026-09-10-ana-docker-disk-repair.md`
- `[2026-09-09]` **Run 7 PURGED; pfi-gx10 declared an experimental/TRAINING box with no serving seat** — operator: *"gx10 is an experimental box, primarily for training … run 7 can be purged … no new run, we'll roll with run 6 for now."* ~139 GiB reclaimed across both boxes; the 315 MB adapter + provenance KEPT as the only non-reproducible piece. `Pfish-6` on ana-ml2 :8021 is the sole standing seat.
- `[2026-09-09]` **Run 7 RETIRED; run 6 declared `Pfish-6` and is the standing seat** — NVFP4 quant on ana-ml2 :8021 AND gx10 :8098 at 262k ctx, gateway alias `trial` → `Pfish-6`, max-num-seqs 8→32 (2,170 tok/s at n=16, 3.2x the old ceiling). ⚠ ana-ml2 measured **4.1x FASTER than the GX10** on the same artifact — the reverse of the expectation. → `persistent-memory.d/2026-09-09-run7-retired-pfish6.md`
- `[2026-09-09]` **The run-7 CSAM gate failure was a DETECTOR BUG** — HARD `child_term` matched the ADJECTIVE "minor"; operator-diagnosed, fixed `cc42d76` (nominal-use-only, selftest 24/24), retention wired so a hit can finally be adjudicated. ⚠ The lesson is mine: rigor downstream of an unexamined premise is not rigor. → `persistent-memory.d/2026-09-09-csam-detector-bug.md`
- `[2026-09-09]` **⚠ ERP RUN 7 FAILED THE SAFETY GATE — both seats stopped, nothing deleted.** → `persistent-memory.d/2026-09-09-erp-run-7-failed-the-safety.md`
- `[2026-09-09]` **run 7 quantized NVFP4A16 and serving as `trial`** — 49 GiB bf16 relayed gx10→ana-ml2 (16 min, 53 MB/s), quant 49→16 GiB via `services/erp-seat-quant/run_quant_erp_v7.sh` (dry-run gate passed: 11,725 targets / 11,520 experts, routers+vision BF16), seat on `:8021` under its TRUE name `erp-tune-v7-nvfp4a16`, LiteLLM `trial` repointed (config-file alias — `/model/update` REFUSES a config model, must edit `stacks/litellm/conf/config.yaml` + restart). Rollback: v6 artifact on disk + `/tmp/erp-seat-env.v6.bak`. ⚠ **`no direct path` was WRONG** — gx10↔ana-ml2 ROUTING is fine both ways; neither box holds a private key (only `authorized_keys`), so neither can *initiate*. `ssh -A` agent forwarding from nh3-dev gives a genuine direct path, verified. The relay costs nothing here anyway: both gx10 and nh3-dev are at NH3, so the WAN hop happens once either way.
- `[2026-09-09]` **Booth: partial ask answers are legal** (v0.1.15) — operator: the form failed when a question was left blank. `required` dropped from the radios; answered questions recorded, blanks land in `unanswered`, `complete` says whether the set is finished; refused only when there is no pick anywhere AND no notes. Reading sessions must check `complete`.
- `[2026-09-09]` **ERP run 7 COMPLETE and the base arm is serving.** 542/542 steps in 14h17m on pfi-gx10, adapter 13:23 PT, `train_loss` 3.205 / low 2.799, merge verified a sampled target actually changed (the silent-no-op check). `erp-seat-base-ara` up on `10.100.50.60:8098` for brokkr's floors, `erp-tune-v7` merged and staged pending his cue; Miranda notified for the operator. Runbook `docs/runbooks/gx10-run-07.md`.
- `[2026-09-09]` **Booth asks render INLINE in a custom report, placed by the author** (v0.1.14) — operator ruling: *"the asks should be inline with the artifacts, not on a separate page."* Placeholders `data-booth-ask="<stem>"` / `"<stem>:<key>"` / `data-booth-ask-submit`, plus `<!-- booth:ask … -->`; per-question fragments bind to ONE form via the HTML5 `form=` attribute so a four-voice audition submits every pick in a single POST. ⚠ The placeholder must sit OUTSIDE any grid/flex parent or it becomes a cell (measured on `redo-anchors`: a 224 px sixth grid cell). Unplaced questions + a missing submit block are appended, so a partially marked-up page can never yield an unsubmittable 400 — a test caught that as a real drop. `redo-anchors/index.html` was hand-marked-up on the LIVE copy; tts-dev told to move it into the generator or a regeneration loses it.
- `[2026-09-09]` **The Booth gained an ASKS primitive** (v0.1.12): a session drops `<stem>.ask.json` in a booth, the operator answers a radio form + notes in the browser, the pick lands as `<stem>.answer.json` the session reads (`booth ask|asks|answer --wait`). Multi-question form via a `questions` list. ⚠ Two defects found and fixed the same day: a booth serving its OWN `index.html` never rendered the panel (verbatim path returns early) → amber chip + standalone `/b/<name>/asks` page; and single-ask `title` was silently dropped. The `booth` CLI was ALSO not on PATH anywhere despite the global link-board convention telling every session to run it → symlinked to `~/.local/bin`. Global `CLAUDE.md` now teaches the primitive.
- `[2026-09-09]` **ana-ml2 pool actions LANDED (scrub 0 errors in 1h33 → `zpool clear`; nvme-cli + full-drive SMART table; zroot 91→73% via dangling-image + builder prune, tagged rollback seats kept) — and the missed-scrub mystery SOLVED: nvme7 (slot 0-5, `S47VNY0K600221`) was absent from every boot 04-23→09-05, tank was raidz2-DEGRADED for 4½ months, Debian's scrub/trim cron only touches `ONLINE` pools, and ZED's alert went to a root mailbox with no MTA.** nvme7's 2084 media errors did not move across the scrub → historical, keep + watch. Playbook `playbooks/ana-ml2-pool-health.yaml`; inventory in `servers/ana-ml2/README.md`. → `persistent-memory.d/2026-09-09-ana-ml2-pool-actions-done.md`
- `[2026-09-09]` **ana-ml2 `tank`: 2 CKSUM errors on nvme7n1 after a boot-time resilver, NO scrub since 04-12, zroot 91% — three actions DEFERRED to a clean-context session** (scrub → nvme-cli SMART → docker prune), operator ruling "we'll do all 3 on clean context"; tracked at commit `3e18a04` + the post-clear handoff. ESH 10G links measured clean (fiber run live on UDM SFP+2 ↔ USW-Pro-XG Media). → `persistent-memory.d/2026-09-09-ana-ml2-pool-actions-deferred.md`
- `[2026-09-08]` **ana-ml2 mesh return routes PERSISTED** as `/etc/network/if-up.d/mesh-routes` (Debian 13 ifupdown, no netplan) via `playbooks/ana-ml2-mesh-routes.yaml` (elway, verified) — operator: "persist the routes". Hook not yet exercised by a real reboot. `f923d6a`.
- `[2026-09-08]` **ERP run 7 LAUNCHED on pfi-gx10 23:06 PT** under `operator-2026-09-08-rnd-run7` — opening-split slot + mask union; free check passed with two explained deltas; first launch died on a missing recipe (zsh quoting). → `persistent-memory.d/2026-09-08-erp-run7-launched.md`
- `[2026-09-08]` **erp-tune-v6-nvfp4a16 quantized (data-free W4A16, ~90 s) and serving on ana-ml2 :8021; `trial` aliased to it ("no gate"); tool calling fixed where it can be** — `tool_choice:none` flag; forced tool_choice is prompt-driven on Gemma-4 by vLLM design, nightly `311b3513` raises it 1/9→6/9; json_schema is the deterministic path. → `persistent-memory.d/2026-09-08-erp-seat-nvfp4-trial-and-toolcalling.md`
- `[2026-09-08]` **Run-6 gate: CSAM level=review soft trip HALTED it; operator adjudicated GO ("baby is a pet name"); TRANSFERRED finalized without the tuned refusal leg; k=25 legs cut** — the flagged text exists nowhere by design. → `persistent-memory.d/2026-09-08-run6-gate-csam-adjudication.md`
- `[2026-09-08]` **ESH static-WAN follow-ups landed (FortiGate trusthost3, esh-ana IPsec rebind, UDP 41641 → mesh direct); YTVC chased back up (nh3-scale SOCKS, stale yt-dlp layer, punkt_tab) and v0.3.6 CrisperWhisper deployed; gitea webhook repointed off the dead wg0 IP with the HMAC secret re-applied.** → `persistent-memory.d/2026-09-08-esh-static-wan-followups-and-ytvc.md`
- `[2026-09-08]` **ERP run 5 = RESCUED (landmark R49.5)** — first capability-gate pass in the ERP-seat line; the 3.46%-loss dependency-forcing slot (GovReport+QMSum) broke the coupling runs 3c/4 couldn't. Seat `erp-tune-v5` served on gx10:8098, `trial` alias repointed 3c→v5. → `persistent-memory.d/2026-09-08-run5-rescued.md`
- `[2026-09-08]` **R47 base settled from bytes = STOCK `google/gemma-4-26B-A4B-it`** — three-way sha match (local == HF etag == stock LFS oid; commit `4d7ae498` == stock HEAD); the `-heretic` label is a naming error, all runs trained from stock. Accept-vs-swap now evidenced. → `persistent-memory.d/2026-09-08-base-provenance-stock.md`
- `[2026-09-08]` **yt-voice-clipper back UP** → `persistent-memory.d/2026-09-08-yt-voice-clipper-back-up.md`
- `[2026-09-08]` **ESH WAN static `128.177.138.182/30` (gw .181) is LIVE** — the Cityside /30 that was 'not provisioned' on 09-04 now carries traffic; egress verified from esh-docker-vm. CGNAT at ESH is over. Added to the crowdsec `esh` allowlist. All three follow-ups LANDED same day: FortiGate trusthost3 → the static (login from ESH verified), dormant esh-ana IPsec rebound to wan1/static, UDP 41641 forward → esh-scale now peers DIRECT (was DERP).
- `[2026-09-08]` **ERP run 6 COMPLETE** — 524/524, train_loss 3.259 (run 5: 3.235). Merged; base seat `erp-seat-base-ara` serving on gx10:8098 for floors, awaiting brokkr's swap cue → `erp-tune-v6`. ⚠ abliterated repo lacks `processor_config.json` — stock's carried in (32bdf45d). Miranda informed.
- `[2026-09-08]` **ERP run 6 LAUNCHED on pfi-gx10 on the jenerallee78 ARA-abliterated base** (index `33c59654…`, 32/32 shards byte-verified vs brokkr pins, stock tokenizer set installed over the repo's 256-token-truncating one, run-5 recipe byte-held, free check exact). Operator's direct grant `operator-2026-09-08-rnd-run6`; run-5 seat unloaded (`trial` dark). Gate names: `erp-seat-base-ara` / `erp-tune-v6`. → `docs/runbooks/gx10-run-06.md`, commit `3fec668`.
- `[2026-09-08]` **Miranda = operator's chief of staff, may relay his directives** — added to user-level `~/.claude/CLAUDE.md` (dotfiles `7134a22`) as the named exception to the no-relayed-auth rule (unidentified peer relays still excluded); material-consequence calls she relays stay the operator's own.
- `[2026-09-08]` **Fleet fixes shipped** — WhereTF Homepage card + DNS (`4506ef6`); ext-tts LiteLLM alias → `irv-ml1.nh3.internal` (DB `/model/update` + `extra_hosts`, `957c8f1`); the 09-06 irv-ml1 stale-IP trail repointed across 25 composes + services.yaml + ssh-target → DNS name (`e0d1c44`); Homepage `/api/services` outage fixed — ana-ml2 discovery via a socat proxy on ana-docker (`stacks/ana-ml2-proxy`, `913d2d2`, reversible).
- `[2026-09-06]` **pfi-pve NASPool REBUILT as six-wide raidz2 after a backplane fault killed bays 9/10** (Route C hybrid, operator-directed): parked 1.65T on ospool, destroyed, recreated, restored, backup tier back 04:03Z; guests never stopped (ALL boot disks are on ospool — the prior brief had this wrong). Legacy vzdump pruned to newest-per-guest by omission. OPEN: destroy `ospool/naspool-evac` after scrub + one backup cycle; backplane swap next visit; PSU1 still dead. → `persistent-memory.d/2026-09-06-pfi-pve-naspool-raidz2-rebuild.md`
- `[2026-09-03]` **Run 3c is STAGED on pfi-gx10 and deliberately NOT launched** — the launch is a 13.3 h commitment and the operator stood this port down once already. Base shards AND the encoded corpus sha256-verified identical to ana-ml2's, so the transformers 5.15.1→5.16.1 / x86-64→aarch64 delta is *measured* inert. ⚠ The encode-cache FILENAME differs by design (`base_model_path` is in the key) — input hash, not output. ⚠ Tripped the `pkill -f` ssh self-match again; the launcher guards on a pidfile because of it. → `persistent-memory.d/2026-09-03-gx10-run3c-staged.md`
- `[2026-08-25]` **Fused MoE kernel path — DEFERRED, tracked at park `fused-moe-kernel-path-for-gemma-4-moe-training` (id 47).** Operator: "note the fused MoE kernel for round two… if we nail it soon, the math has us wanting to restart the run anyway." Training MFU is **8.6%** (27.1 of a benchmarked 313.8 TFLOPS) because `transformers` runs the Gemma-4 experts in a Python loop — 128 experts × 30 layers, ~11,500 iterations per step under gradient checkpointing. ⚠ **The same fused 3-D expert layout that made bitsandbytes skip 88.5% of the model is exactly what a grouped GEMM wants** — the format is good for storage and for fused kernels, and hostile only to naive iteration. Two fixes: `group_by_length` (−29.9% compute, free, but breaks the seeded order manifest and re-opens a batch-composition call brokkr already made) and a grouped-GEMM/compiled MoE forward (the remaining ~10×). **Not applied to the live run** — restarting mid-flight to change batch ordering was judged a bad trade at step ~50 of 1,312.
- `[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`.
- `[2026-08-19]` **AI-tab Dormant regrouping BELAYED by the operator** — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong, Chatterbox Fast) show amber EXITED inside live groups rather than `AI - Dormant`. Fix is a label change + recreate per stack; needs the operator's read on which are retired vs temporarily down. `untracked by operator choice` (his words: "belay the ai dormant regrouping for now").
_30 older entries archived to archival-memory.md._
_70 older entries archived to archival-memory.md._
## Tried and abandoned
- `[2026-09-24]` **Testing "Restore AC Power Loss" with an OS shutdown** — a shutdown stays off BY DESIGN; only pulling and restoring AC tests it. A community README claimed the two are indistinguishable; trusting it cost Prime two trips to the pfi-gx10 power button.
- `[2026-09-23]` **`booth link --help`** — there is no help flag; it posts `--help` to the operator's link board as a link. Read `booth` with no args for usage.
- `[2026-09-21]` **Using directory mtime as a liveness test when pruning session scratchpads** — `find /tmp/claude-1000 -mindepth 2 -maxdepth 2 -type d -mtime +7` deleted an ACTIVE session's working directory. A directory's mtime does not change when files are written into its SUBDIRECTORIES, so a session writing continuously to `<id>/tasks/` looks 7+ days idle at `<id>/`. Four safety assertions passed; none of them asked whether the liveness test was sound. Use the deepest recent file, or cross-reference running `claude` PIDs.
- `[2026-09-18]` **Routing SearXNG's egress through a SOCKS5 proxy on esh-scale** — one day live, reverted. It fixed nothing, and the reason I gave for reverting it was itself wrong: the rollback was the counterfactual and it falsified my own published claim. Kept reverted on its own merits (no measurable gain, added a hard ESH dependency for all fleet search). → `persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md`