docs: refresh what today's work made stale — booth asks (inline placement promoted to its own section), ana-ml2 nvme7 settled by the scrub result, nh3-dev booth entry + the CLI-on-PATH fix, run-07 runbook outcome + serving state

This commit is contained in:
vh
2026-09-09 14:18:34 -07:00
parent 78c3a7c170
commit 6e0b85ba27
5 changed files with 137 additions and 29 deletions
+69
View File
@@ -55,3 +55,72 @@ tokenizer + template all unchanged from run 6.
Gate choreography after the adapter lands: brokkr-smithy `SFT-RECIPE-run7-SCOPE.md §5` (cells Gate choreography after the adapter lands: brokkr-smithy `SFT-RECIPE-run7-SCOPE.md §5` (cells
LENGTH-FIXED / LENGTH-MOVED / BOUGHT-T4 / FLAT); serve names `erp-seat-base-ara` then `erp-tune-v7`. LENGTH-FIXED / LENGTH-MOVED / BOUGHT-T4 / FLAT); serve names `erp-seat-base-ara` then `erp-tune-v7`.
## Outcome — COMPLETE 2026-09-09 13:23 PT
Ran to the end with no interventions, no resumes, no errors: **542/542 optimizer
steps, one epoch, 14h17m wall clock** (23:06 PT 09-08 → 13:23 PT 09-09).
| | |
|---|---|
| `train_loss` (run mean) | **3.205** |
| last logged points | 2.9366 @ step 530 (lr 3.5e-07), 2.9389 @ step 540 (lr 1.86e-08) |
| lowest logged | **2.7992 @ step 420** (from 7.1692 @ step 10) |
| step seconds | n=541, p50 **82.1**, mean 94.1, min 17.4, max 334.3, p99 268.3 |
| adapter | 410 tensors, 297 MB, `adapter_artifact` verified |
Provenance (`run-07/adapter/provenance.json`), the fields worth re-reading before
trusting any number: recipe `recipe-erp-seat-sft-r7.json` sha `2972559e4c2940dc`;
harness `0a6bd2e0`, `harness_dirty_at_launch False`; **`flex_attention` requested
AND resolved** (§4 of the training-throughput playbook — the backend delta that
voided two earlier conclusions); `quantized_base False`; torch 2.14.0+cu130,
transformers 5.16.1; chat template `ae53464b`; 8,661 encoded records from 8,505
source samples; 6,106 masked turns; `unfittable_disposition drop` (cwm-2602,
cwm-1826 — their loss moved into their opening-split rows, as designed).
Effective mix as the optimiser saw it, identical to the free check, so the single
variable landed: kvasir .3348 / fireball .3023 / c2-logs .1143 / bluemoon .0964 /
cwm .0942 / govreport .0289 / **opening-split .0235** / qmsum .0055.
**Grad-norm spikes, recorded rather than smoothed.** Four isolated points —
step 80 = 52.1, step 350 = 9.54, step 160 = 6.27, step 60 = 7.66 — against a body
at 0.34–1.9 from step 100 on. None moved the loss where it occurred, none
clustered, and a mid-run "every ~80 steps" hypothesis died at step 240 (0.398).
Read as occasional hard batches, not instability. Series shipped to brokkr and
written to `run-07/loss-series-r7.json` (sha `8322920f05ea6712`, 54 points, the
harness logs every 10 steps).
⚠ **Sampler padding 17.1%** (run 6: 0.0%) — the short opening-split rows pair
badly under the length-paired sampler. Throughput only; it does not touch what
the optimiser saw, and it is why the clock read 14h rather than 12h.
## Serving state after the run
merge ~/ml/.venv/bin/python merge_lora.py \
--base /home/infra-ops/models/gemma4-26b-a4b-it-ara-abliterated-jenerallee78-0631379a \
--adapter run-07/adapter --out serve/merged-run07 \
--chat-template run-07/adapter/chat_template.jinja
⚠ **Merge under `~/ml/.venv`, not `~/vllm-env`** — the serving venv has no `peft`
and the merge dies in the import. ⚠ **`--chat-template` must be passed
explicitly**: the script defaults to `/tank/aimodels/…`, which does not exist on
the GX10 (no NFS). The adapter's own `chat_template.jinja` is the right file and
its sha is `ae53464b`, the stock 390-line upstream template — verify before
trusting it.
Merged artifact `serve/merged-run07`: 48.1 GiB, 1013 tensors,
`Gemma4ForConditionalGeneration` bf16, template `ae53464b`, tokenizer `cc8d3a0c`
(stock), `processor_config.json` byte-identical to base and stock (`32bdf45d`).
The merge verified a sampled target actually changed — the check that catches a
silent no-op adapter.
serve ./gpu-clear.sh && vllm serve <ARA base> --served-model-name erp-seat-base-ara \
--host 0.0.0.0 --port 8098 --dtype bfloat16 --max-model-len 8192 \
--max-num-seqs 8 --gpu-memory-utilization 0.60 \
--enable-auto-tool-choice --tool-call-parser gemma4
log vllm-base-run07gate.log pid file vllm-base-run07gate.pid
`erp-seat-base-ara` came up 13:58 PT, health 200, `/v1/models` correct, round trip
verified at temperature 0. **`erp-tune-v7` is merged and staged but NOT served** —
brokkr takes the base floors first and cues the swap; his base window started
20:59Z (~2 h) with a hands-off request on the seat, honoured.
+2 -1
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management # Persistent memory — eshpfi-management
_Last updated: 2026-09-09 14:05 PT (fleet-ops: ERP run 7 COMPLETE on pfi-gx10 — 542/542 steps, adapter 13:23 PT, merged to serve/merged-run07; `erp-seat-base-ara` SERVING on 10.100.50.60:8098 for brokkr's floors, `erp-tune-v7` staged awaiting his swap cue; Miranda notified. Booth gained the ASKS primitive [radio+notes -> answer sidecar, multi-question, verbatim-booth fix]. ana-ml2 pool actions landed. sox on nh3-dev for yt-voice-clipper-dev)_ _Last updated: 2026-09-09 15:10 PT (fleet-ops: ERP run 7 COMPLETE — `erp-seat-base-ara` serving for brokkr's base window, `erp-tune-v7` staged awaiting his swap cue; Booth ASKS primitive shipped and then made INLINE per operator ruling; ana-ml2 pool actions landed; sox on nh3-dev; docs refreshed across booth/ana-ml2/nh3-dev/run-07 runbook)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its > **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under an hour old, read it (it carries the in-flight > `Written:` stamp is under an hour old, read it (it carries the in-flight
@@ -148,6 +148,7 @@ clean context"). Older in-flight blocks (09-08 morning, 09-05/06) are preserved
## Recent decisions ## Recent decisions
- `[2026-09-09]` **ERP run 7 COMPLETE and the base arm is serving.** 542/542 steps in 14h17m on pfi-gx10, adapter 13:23 PT, `train_loss` 3.205 / low 2.799, merge verified a sampled target actually changed (the silent-no-op check). `erp-seat-base-ara` up on `10.100.50.60:8098` for brokkr's floors, `erp-tune-v7` merged and staged pending his cue; Miranda notified for the operator. Runbook `docs/runbooks/gx10-run-07.md`. - `[2026-09-09]` **ERP run 7 COMPLETE and the base arm is serving.** 542/542 steps in 14h17m on pfi-gx10, adapter 13:23 PT, `train_loss` 3.205 / low 2.799, merge verified a sampled target actually changed (the silent-no-op check). `erp-seat-base-ara` up on `10.100.50.60:8098` for brokkr's floors, `erp-tune-v7` merged and staged pending his cue; Miranda notified for the operator. Runbook `docs/runbooks/gx10-run-07.md`.
- `[2026-09-09]` **Booth asks render INLINE in a custom report, placed by the author** (v0.1.14) — operator ruling: *"the asks should be inline with the artifacts, not on a separate page."* Placeholders `data-booth-ask="<stem>"` / `"<stem>:<key>"` / `data-booth-ask-submit`, plus `<!-- booth:ask … -->`; per-question fragments bind to ONE form via the HTML5 `form=` attribute so a four-voice audition submits every pick in a single POST. ⚠ The placeholder must sit OUTSIDE any grid/flex parent or it becomes a cell (measured on `redo-anchors`: a 224 px sixth grid cell). Unplaced questions + a missing submit block are appended, so a partially marked-up page can never yield an unsubmittable 400 — a test caught that as a real drop. `redo-anchors/index.html` was hand-marked-up on the LIVE copy; tts-dev told to move it into the generator or a regeneration loses it.
- `[2026-09-09]` **The Booth gained an ASKS primitive** (v0.1.12): a session drops `<stem>.ask.json` in a booth, the operator answers a radio form + notes in the browser, the pick lands as `<stem>.answer.json` the session reads (`booth ask|asks|answer --wait`). Multi-question form via a `questions` list. ⚠ Two defects found and fixed the same day: a booth serving its OWN `index.html` never rendered the panel (verbatim path returns early) → amber chip + standalone `/b/<name>/asks` page; and single-ask `title` was silently dropped. The `booth` CLI was ALSO not on PATH anywhere despite the global link-board convention telling every session to run it → symlinked to `~/.local/bin`. Global `CLAUDE.md` now teaches the primitive. - `[2026-09-09]` **The Booth gained an ASKS primitive** (v0.1.12): a session drops `<stem>.ask.json` in a booth, the operator answers a radio form + notes in the browser, the pick lands as `<stem>.answer.json` the session reads (`booth ask|asks|answer --wait`). Multi-question form via a `questions` list. ⚠ Two defects found and fixed the same day: a booth serving its OWN `index.html` never rendered the panel (verbatim path returns early) → amber chip + standalone `/b/<name>/asks` page; and single-ask `title` was silently dropped. The `booth` CLI was ALSO not on PATH anywhere despite the global link-board convention telling every session to run it → symlinked to `~/.local/bin`. Global `CLAUDE.md` now teaches the primitive.
- `[2026-09-09]` **ana-ml2 pool actions LANDED (scrub 0 errors in 1h33 → `zpool clear`; nvme-cli + full-drive SMART table; zroot 91→73% via dangling-image + builder prune, tagged rollback seats kept) — and the missed-scrub mystery SOLVED: nvme7 (slot 0-5, `S47VNY0K600221`) was absent from every boot 04-23→09-05, tank was raidz2-DEGRADED for 4½ months, Debian's scrub/trim cron only touches `ONLINE` pools, and ZED's alert went to a root mailbox with no MTA.** nvme7's 2084 media errors did not move across the scrub → historical, keep + watch. Playbook `playbooks/ana-ml2-pool-health.yaml`; inventory in `servers/ana-ml2/README.md`. → `persistent-memory.d/2026-09-09-ana-ml2-pool-actions-done.md` - `[2026-09-09]` **ana-ml2 pool actions LANDED (scrub 0 errors in 1h33 → `zpool clear`; nvme-cli + full-drive SMART table; zroot 91→73% via dangling-image + builder prune, tagged rollback seats kept) — and the missed-scrub mystery SOLVED: nvme7 (slot 0-5, `S47VNY0K600221`) was absent from every boot 04-23→09-05, tank was raidz2-DEGRADED for 4½ months, Debian's scrub/trim cron only touches `ONLINE` pools, and ZED's alert went to a root mailbox with no MTA.** nvme7's 2084 media errors did not move across the scrub → historical, keep + watch. Playbook `playbooks/ana-ml2-pool-health.yaml`; inventory in `servers/ana-ml2/README.md`. → `persistent-memory.d/2026-09-09-ana-ml2-pool-actions-done.md`
- `[2026-09-09]` **ana-ml2 `tank`: 2 CKSUM errors on nvme7n1 after a boot-time resilver, NO scrub since 04-12, zroot 91% — three actions DEFERRED to a clean-context session** (scrub → nvme-cli SMART → docker prune), operator ruling "we'll do all 3 on clean context"; tracked at commit `3e18a04` + the post-clear handoff. ESH 10G links measured clean (fiber run live on UDM SFP+2 ↔ USW-Pro-XG Media). → `persistent-memory.d/2026-09-09-ana-ml2-pool-actions-deferred.md` - `[2026-09-09]` **ana-ml2 `tank`: 2 CKSUM errors on nvme7n1 after a boot-time resilver, NO scrub since 04-12, zroot 91% — three actions DEFERRED to a clean-context session** (scrub → nvme-cli SMART → docker prune), operator ruling "we'll do all 3 on clean context"; tracked at commit `3e18a04` + the post-clear handoff. ESH 10G links measured clean (fiber run live on UDM SFP+2 ↔ USW-Pro-XG Media). → `persistent-memory.d/2026-09-09-ana-ml2-pool-actions-deferred.md`
+13 -2
View File
@@ -44,8 +44,19 @@ at import. While it was missing `tank` was DEGRADED, and Debian's `zfsutils-linu
(`/usr/lib/zfs-linux/{scrub,trim}`) only touches pools whose health is `ONLINE`, so tank (`/usr/lib/zfs-linux/{scrub,trim}`) only touches pools whose health is `ONLINE`, so tank
got **no scrub and no trim from 04-12 to 09-06**. ZED's `ZED_EMAIL_ADDR=root` has no got **no scrub and no trim from 04-12 to 09-06**. ZED's `ZED_EMAIL_ADDR=root` has no
MTA behind it, so the 4½-month degradation alerted nobody. `media_errors=2084` on MTA behind it, so the 4½-month degradation alerted nobody. `media_errors=2084` on
nvme7 is a lifetime counter; the 2026-09-09 scrub is the first fresh measurement nvme7 is a lifetime counter.
(baseline 2084 at 00:32 PT — compare after any future event, growth = replace).
**Settled by the 2026-09-09 scrub** (00:29–02:02 PT, `scrub repaired 0B in 01:32:44
with 0 errors`, then `zpool clear tank` → CKSUM 2 → 0): `media_errors` read **2084
before and 2084 after** a full 6.84 TiB verify, so the counter is prior-life
history, not an active fault, and the 2 CKSUM were the stale-block artefact of the
09-05 late resilver. **nvme7 stays in service; watch the counter at every visit and
replace on growth** (`zpool replace tank nvme7n1 <new>`; any PM1725b 1.6 TB or
larger). Slot 0-5 itself deserves a reseat / cable check at the next hands-on
visit — a bay that dropped a drive for 4½ months is the likelier fault than the
drive. Playbook: `playbooks/ana-ml2-pool-health.yaml` (idempotent; rerunning is a
no-op). ⚠ **Nothing alerts on this** — see the open follow-up in
`persistent-memory.d/2026-09-09-ana-ml2-pool-actions-done.md`.
## Key paths ## Key paths
+9
View File
@@ -25,6 +25,15 @@ local Bash already executes here — no SSH-to-self needed for non-privileged wo
- **The Booth** — ephemeral media drop board (`:8090`, `booth.service`), from - **The Booth** — ephemeral media drop board (`:8090`, `booth.service`), from
eshpfi `services/booth/`. Lets CC sessions surface A/B renders + smoke results eshpfi `services/booth/`. Lets CC sessions surface A/B renders + smoke results
(and browser uploads for pickup) to the operator; 24h TTL, Homepage-linked. (and browser uploads for pickup) to the operator; 24h TTL, Homepage-linked.
Since 2026-09-09 it also carries **asks** — a session poses a multiple-choice
question in a booth, the operator answers a radio form + notes in the browser,
and the pick lands as an answer sidecar the session reads (`booth ask` /
`booth answer --wait`). ⚠ The **`booth` CLI is on PATH via
`~/.local/bin/booth` → `services/booth/scripts/booth`**, symlinked 2026-09-09;
before that it was on no PATH at all, so every session following the global
link-board convention was hitting `command not found` unless it used the full
path. `~/.zshenv` puts `~/.local/bin` in PATH for non-interactive `ssh nh3-dev
'<cmd>'` too.
- **jackdaw-compose** — JackDAW AI Composer `/compose` backend (`:8787`, - **jackdaw-compose** — JackDAW AI Composer `/compose` backend (`:8787`,
`jackdaw-compose.service`), a thin stateless `bun server/index.ts` from `jackdaw-compose.service`), a thin stateless `bun server/index.ts` from
`~/development/jackdaw` → LiteLLM `gen`. Origin-gated (INV-BK04/BK05), reached `~/development/jackdaw` → LiteLLM `gen`. Origin-gated (INV-BK04/BK05), reached
+39 -21
View File
@@ -161,36 +161,54 @@ Rules of the primitive:
- **A broken ask is shown as broken**, not hidden: if the JSON does not - **A broken ask is shown as broken**, not hidden: if the JSON does not
validate, the page says why, so a session never thinks it posted a question validate, the page says why, so a session never thinks it posted a question
the operator cannot see. the operator cannot see.
- **A booth with its own `index.html` gets the ask INLINE, where you put it.** - Ask/answer files are not gallery items and do not count toward the booth's
That page is served verbatim, so the Booth substitutes placeholders in your item count; they render as the panel above the gallery. Answering bumps the
markup rather than rendering the panel above a gallery it does not have: booth's mtime, so it lives another TTL — the session has 24h to read it.
- Works with JavaScript off (plain form POST). No auth, same as everything here.
### Where the form renders
Two booth shapes, two placements. Either way the ask is never invisible — that
is the guarantee; markup only moves it somewhere better.
**Auto-gallery booth** (no `index.html` of its own): the asks panel renders
above the gallery, styled like the rest of the Booth. Nothing to do.
**A booth serving its own `index.html`**: that page is returned verbatim, so the
Booth substitutes **placeholders in your markup** rather than rendering a panel
above a gallery that does not exist. The question then sits with the artifact it
is about, which is the whole point (operator ruling 2026-09-09: *"the asks should
be inline with the artifacts, not on a separate page"*).
```html ```html
<div data-booth-ask="anchors"></div> <!-- the whole ask: every question + submit --> <div data-booth-ask="anchors"></div> <!-- the whole ask: every question + submit -->
<div data-booth-ask="anchors:lawson"></div> <!-- just that one question's radios --> <div data-booth-ask="anchors:lawson"></div> <!-- just that one question's radios -->
<div data-booth-ask-submit="anchors"></div> <!-- the notes field + submit button --> <div data-booth-ask-submit="anchors"></div> <!-- the notes field + submit button -->
<!-- booth:ask anchors:lawson --> <!-- comment form, same thing --> <!-- booth:ask anchors:lawson --> <!-- comment form, identical behaviour -->
``` ```
Per-question fragments bind to **one** form with the HTML5 `form=` attribute, Per-question fragments bind to **one** form via the HTML5 `form=` attribute, so a
so a four-voice audition can put each radio group under that voice's audio and four-voice audition puts each radio group under that voice's audio and still
still submit all four picks in a single POST — which is what a multi-question submits every pick in a single POST — which is what a multi-question ask
ask requires. The fragments ship their own scoped styles and inherit nothing requires. Fragments ship their own scoped styles, inherit nothing from your page,
from your page. No JavaScript. and use no JavaScript.
⚠ Put the placeholder **outside** any CSS grid or flex container, or it ⚠ **Put the placeholder outside any CSS grid or flex container**, or it becomes a
becomes a cell in it. A sibling of the block it belongs to is right. cell in it — measured on `redo-anchors`, where the first attempt rendered as a
224 px sixth grid cell wedged between two audio players. A sibling of the block
it belongs to is right.
Placement is optional: a page with no placeholders gets the whole ask appended The fallbacks, so a page can never strand a question:
at the end, so an ask is never invisible — markup only moves it somewhere
better. Mark up some questions and not others and the rest are appended too, | you marked up | what happens |
because a multi-question form that is missing a question is a 400 the operator |---|---|
would only meet after filling it in. The amber chip stays as a jump link to the | nothing | the whole ask is appended at the end of the page |
first open ask, and `/b/<name>/asks` still renders every ask on its own page. | some questions, no submit | the rest of the questions **and** a submit block are appended |
- Ask/answer files are not gallery items and do not count toward the booth's | a stem this booth does not have | your markup is left alone, untouched; the real ask is still appended |
item count; they render as the panel above the gallery. Answering bumps the
booth's mtime, so it lives another TTL — the session has 24h to read it. An amber `? N open asks` chip floats top-right as a jump link to the first open
- Works with JavaScript off (plain form POST). No auth, same as everything here. ask, and `GET /b/<name>/asks` still renders every ask on a plain page of its own
— useful when you want to hand someone only the question.
## Upload for pickup ## Upload for pickup