memory: ERP run 7 complete (542/542, adapter 13:23 PT) — erp-seat-base-ara serving for brokkr's floors, erp-tune-v7 merged and staged; Booth asks primitive + its two same-day defect fixes

This commit is contained in:
2026-09-09 13:59:31 -07:00
parent a40f979b7a
commit 8961ca078b
+9 -6
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-09-09 02:10 PT (fleet-ops: ana-ml2 pool actions LANDED — tank scrubbed clean + cleared, nvme7 SMART read [2084 lifetime media errors, zero growth], zroot 91→73%; ROOT CAUSE of the missed scrubs = nvme7 physically absent 04-23→09-05, pool DEGRADED and ZED mail to nowhere; ERP run 7 TRAINING on pfi-gx10 [~64/542 at 00:42 PT, adapter ~noon]; erp-tune-v6-nvfp4a16 serving as `trial`)_
_Last updated: 2026-09-09 14:05 PT (fleet-ops: ERP run 7 COMPLETE on pfi-gx10 — 542/542 steps, adapter 13:23 PT, merged to serve/merged-run07; `erp-seat-base-ara` SERVING on 10.100.50.60:8098 for brokkr's floors, `erp-tune-v7` staged awaiting his swap cue; Miranda notified. Booth gained the ASKS primitive [radio+notes -> answer sidecar, multi-question, verbatim-booth fix]. ana-ml2 pool actions landed. sox on nh3-dev for yt-voice-clipper-dev)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under an hour old, read it (it carries the in-flight
@@ -115,11 +115,12 @@ clean context"). Older in-flight blocks (09-08 morning, 09-05/06) are preserved
- **🔥 ERP RUN 7 TRAINING on pfi-gx10** — launched 2026-09-08 23:06 PT, pid in `~/erp-tune/run-07.pid`, 542 steps
at ~80 s/it, adapter ~noon 09-09. Watch: `ssh infra-ops@10.100.50.60 "tr '\r' '\n' < ~/erp-tune/run-07.log | tail"`.
Hourly cron in the old session reported stats; a fresh session checks by hand. When the adapter lands: post the
adapter-landed line to brokkr's thread (`01M20AHY9DY92RJK84YD24VSY9`), merge (`merge_lora.py --base <ARA dir>
--adapter run-07/adapter --out serve/merged-run07 --chat-template <stock>`), copy stock `processor_config.json`
into `merged-run07`, serve `erp-seat-base-ara` for floors → on brokkr's swap cue serve `erp-tune-v7` (run-5 flags,
:8098). Runbook `docs/runbooks/gx10-run-07.md`; run-6 choreography in `gx10-run-06.md`.
**DONE 2026-09-09 13:23 PT** — 542/542 steps, 14h17m, train_loss 3.205 (low 2.799 @ step 420), 410-tensor adapter,
flex_attention requested AND resolved, harness `0a6bd2e0` clean. Merged to `serve/merged-run07` (48.1 GiB, template
`ae53464b`, processor_config byte-identical to stock `32bdf45d`). **`erp-seat-base-ara` SERVING** on
`10.100.50.60:8098` (run-5/6 flags; health 200, round trip verified); **`erp-tune-v7` staged, NOT serving** —
awaiting brokkr's swap cue on thread `01M20AHY9DY92RJK84YD24VSY9`. Miranda notified for the operator. ⚠ Sampler
padding 17.1% (run 6: 0.0%) — the short opening-split rows pair badly; throughput only, not correctness.
- **✅ ana-ml2 pool actions DONE 2026-09-09 02:02 PT** (`playbooks/ana-ml2-pool-health.yaml`): tank scrub clean +
cleared, nvme-cli in, zroot 73%. ⚠ Open follow-ups, operator's call: (a) ZFS pool-health ALERTING — tank sat
DEGRADED 04-23→09-05 with nvme7 physically absent and nobody knew (ZED mails `root`, no MTA); (b) nvme7 / slot 0-5
@@ -146,6 +147,8 @@ clean context"). Older in-flight blocks (09-08 morning, 09-05/06) are preserved
## Recent decisions
- `[2026-09-09]` **ERP run 7 COMPLETE and the base arm is serving.** 542/542 steps in 14h17m on pfi-gx10, adapter 13:23 PT, `train_loss` 3.205 / low 2.799, merge verified a sampled target actually changed (the silent-no-op check). `erp-seat-base-ara` up on `10.100.50.60:8098` for brokkr's floors, `erp-tune-v7` merged and staged pending his cue; Miranda notified for the operator. Runbook `docs/runbooks/gx10-run-07.md`.
- `[2026-09-09]` **The Booth gained an ASKS primitive** (v0.1.12): a session drops `<stem>.ask.json` in a booth, the operator answers a radio form + notes in the browser, the pick lands as `<stem>.answer.json` the session reads (`booth ask|asks|answer --wait`). Multi-question form via a `questions` list. ⚠ Two defects found and fixed the same day: a booth serving its OWN `index.html` never rendered the panel (verbatim path returns early) → amber chip + standalone `/b/<name>/asks` page; and single-ask `title` was silently dropped. The `booth` CLI was ALSO not on PATH anywhere despite the global link-board convention telling every session to run it → symlinked to `~/.local/bin`. Global `CLAUDE.md` now teaches the primitive.
- `[2026-09-09]` **ana-ml2 pool actions LANDED (scrub 0 errors in 1h33 → `zpool clear`; nvme-cli + full-drive SMART table; zroot 91→73% via dangling-image + builder prune, tagged rollback seats kept) — and the missed-scrub mystery SOLVED: nvme7 (slot 0-5, `S47VNY0K600221`) was absent from every boot 04-23→09-05, tank was raidz2-DEGRADED for 4½ months, Debian's scrub/trim cron only touches `ONLINE` pools, and ZED's alert went to a root mailbox with no MTA.** nvme7's 2084 media errors did not move across the scrub → historical, keep + watch. Playbook `playbooks/ana-ml2-pool-health.yaml`; inventory in `servers/ana-ml2/README.md`. → `persistent-memory.d/2026-09-09-ana-ml2-pool-actions-done.md`
- `[2026-09-09]` **ana-ml2 `tank`: 2 CKSUM errors on nvme7n1 after a boot-time resilver, NO scrub since 04-12, zroot 91% — three actions DEFERRED to a clean-context session** (scrub → nvme-cli SMART → docker prune), operator ruling "we'll do all 3 on clean context"; tracked at commit `3e18a04` + the post-clear handoff. ESH 10G links measured clean (fiber run live on UDM SFP+2 ↔ USW-Pro-XG Media). → `persistent-memory.d/2026-09-09-ana-ml2-pool-actions-deferred.md`
- `[2026-09-08]` **ana-ml2 mesh return routes PERSISTED** as `/etc/network/if-up.d/mesh-routes` (Debian 13 ifupdown, no netplan) via `playbooks/ana-ml2-mesh-routes.yaml` (elway, verified) — operator: "persist the routes". Hook not yet exercised by a real reboot. `f923d6a`.