diff --git a/archival-memory.md b/archival-memory.md index d5252d4..c1e3f88 100644 --- a/archival-memory.md +++ b/archival-memory.md @@ -4,6 +4,13 @@ _Entries moved out of persistent-memory.md to keep the active file scannable. Re ## Recent decisions (archived) +- `[2026-08-25]` **Run 2's base is an OPEN OPERATOR DECISION, deliberately not staged** — four options with materially different safety postures, detailed in Current state. Tracked at althing thread `01M0WQ8W5574KMEVCHCEKEXNS5`. ⚠ Do not let it get filed as a config knob; it is a reversal of the trainee-selection decision. + _Archived 2026-09-13._ + +- `[2026-08-24]` **Serving the tuned ERP model: LoRA-on-NVFP4 PREFERRED, merged weights the expected fallback — and the recorded objection may be STALE.** Operator: "if you CAN load it as a lora, all the better, the issue is that we will want to run nvfp4 weights, which we had some serious trouble with loading loras on top of nvfp4." ⚠ **The archived root-cause says it was NOT NVFP4-specific**: `[2026-07-07]` vLLM 0.24.0 qwen3_5 LoRA application was a silent no-op (#47639, regression from #37912) — adapter loads HTTP 200, zero deltas at inference, proven **quant-agnostic (NVFP4 AND FP8 both inert)** and adapter-format-agnostic by a 3-peer dwarf panel. Fix PR #47640 was OPEN then. **ana-ml2 is FAR past 0.24.0 and the box runs a SPREAD, not one version** (measured 2026-08-24): `gen` on `nightly-311b3513` = **0.27.2rc1.dev150**, `mog-sec` on `nightly-e9d1398d` = 0.26.1rc1.dev1102, the small seats still on 0.24.0, and char-rp/trainee-bench pinned to v0.26.0. ⚠ **`vllm/vllm-openai:v0.27.1` is already ON DISK, unused** — a TAGGED release, which is the right retest target: no nightly variance, no pull, ~4 months past the diagnosis. So: RETEST hot-swap LoRA on **v0.27.1** before designing around merge — it is cheap, and if it works the post-tune gate can be two aliases on one engine. If it still no-ops, merged weights it is, which means the harness must EMIT merged weights and Eitri needs that in the contract while he is early. Tracked at this snapshot commit; settle it in the QLoRA sizing conversation. + _Archived 2026-09-13._ + + - `[2026-08-28]` **althing v3 flag day (U9b) executed, then six releases to 3.1.1 in one afternoon — and the post office MOVED to nh3-docker.** Every v2 command deleted; 73 handles seeded and verified by set difference; 5,043 orphaned wake FIFOs deleted (v2 named them per-session+PID, v3 per-handle). Image now registry-pulled, digest-pinned, under the `claude-bot` namespace. → `persistent-memory.d/2026-08-28-althing-v3-cutover.md` _Archived 2026-09-12._ diff --git a/persistent-memory.md b/persistent-memory.md index 348193e..8841664 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -1,6 +1,6 @@ # Persistent memory — eshpfi-management -_Last updated: 2026-09-13 ~00:10 PT (⚠ FV SITE DARK — UPS suspected overloaded/dead under two-card load; 19 gateway aliases down; visit 2026-09-14. Flash-Next seat + gen-large built and verified before the outage.)_ +_Last updated: 2026-09-13 ~00:40 PT (⚠⚠ FV COLO DARK since 06:56Z — undersized UPS under two-card load; 19 of 30 gateway aliases down, no local fallback; operator on site 2026-09-14 with PDU + ammeter. Flash-Next seat + gen-large built and verified before the outage.)_ > **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its > `Written:` stamp is under an hour old, read it (it carries the in-flight @@ -110,67 +110,76 @@ no longer deployed sidecars here. See Recent decisions.) (no NOPASSWD)** — stage model pulls to `/home`, not root-owned `/worktank`. ## Current state / in-flight -_As of 2026-09-12 ~19:30 PT._ +_As of 2026-09-13 ~00:40 PT._ -### ⚠⚠ FV colo — DARK since 2026-09-13 06:56Z. Site visit 2026-09-14. +### ⚠⚠ FV COLO IS DARK — site visit 2026-09-14. Nothing else matters until it is back. -Every Fountain Valley address is unreachable, **including the BMC** — the OOB path -goes through the same OPNsense gateway, which is also dark. All other sites healthy. -⭐ **Cause, per operator ruling 2026-09-13: the 1500 VA Eaton UPS was UNDERSIZED** for -this chassis and gave way under two-card GPU load. A Tailscale SNAT change on the FV -gateway 34 min earlier was considered and **DEMOTED** — it went effective, was verified -bidirectional, and then ran correctly for 20 minutes before the site died the moment -load was applied; a working config change does not spontaneously fail under someone -else's GPU load. The load correlation is tight, the NAT correlation merely adjacent. -⚠ **Do NOT use a UPS's surge-only outlets to exceed its rating** — both banks share one -NEMA 5-15P inlet rated 12 A total; the surge bank bypasses the inverter, not the current -limit. Bypass entirely instead. Operator bringing a **PDU + ammeter** 2026-09-14. -⚠ **`power.log` is GPU-ONLY** (nvidia-smi per-card; excludes CPU, 566 GB RAM, drives, -fans, PSU losses) — the ammeter at the plug is the primary instrument, power.log is a -cross-check on the GPU share. Measure idle / 1-card / **2-card** / **4-card**, and -⚠⚠ **capture PEAK not average** (UPS overload protection responds to short-term -overload; an average-only meter's figure is a FLOOR, not the draw). ⭐ The 4-card number -goes into `servers/fv-ml1/README.md` permanently — it closes the cutover's open -"circuit specced against half the real draw" question with a measurement. -⚠⚠ **Removing an undersized UPS does not remove the constraint, it promotes the next -one** (UPS ~900-1200 W → circuit ~1800 W @15 A / ~2400 W @20 A). If 4 cards + host -exceeds the circuit, NO UPS arrangement helps. ⭐ **The lever is `nvidia-smi -pl`** — -cap per-card TGP so the box fits its feed at a throughput cost instead of a rewiring -cost; read `nvidia-smi -q -d POWER` for the enforced range first, and **persist the cap** -(a limit that evaporates on reboot holds until the next power event and then does not). -⭐ Split power (firewall+BMC on UPS, chassis on PDU) survives a *breaker* trip on -battery — 25-40 W on a 1500 VA unit is hours — **but only delivers remote access if the -COLO'S HANDOFF survives too**; ask the facility whether the handoff is on our circuit or -theirs. Three questions for the visit: breaker rating, is the circuit dedicated, whose -gear is the handoff on. -**It will not self-recover; do not poll FV addresses.** 19 of 30 gateway aliases are -down with no local fallback (every free local model was on fv-ml1; irv-ml1 runs no -chat seat). On recovery: recover `power.log` first, mask Docker before the network, -then bring seats up one at a time with `compose up -d` — `gen` first, `flash-next` -last, and do NOT restart the MTP campaign. -→ `docs/runbooks/fv-site-dark-20260913.md` +Dark since **2026-09-13 06:56Z**. Every Fountain Valley address unreachable **including +the BMC** (its only route out is the same OPNsense gateway, also dark); all three other +sites healthy. **It will not self-recover — do not poll FV addresses.** -### FV colo — the pre-outage state (cutover done, one gap open) -- **fv-ml1** (ex ana-ml2) is racked at Fountain Valley, renamed, on `10.251.50.54`, mesh - node `100.64.0.7`; **vb-gateway** OPNsense on `10.251.50.1` / `100.64.0.8`; **BMC** on - `10.251.250.50`. Public `fv.phasefinal.com` → `172.83.89.66`. `tank` intact, all vLLM - seats healthy, inference verified through the Anaheim gateway. All three sites reach FV - by real IP. → `persistent-memory.d/2026-09-12-fv-cutover-executed.md` -- ⚠ **OPEN GAP: fv-ml1 cannot initiate to fleet LAN IPs** (10.100/10.250/10.0 all fail; - mesh IPs and internet fine, inbound fine). Return-path issue at the far gateways. Not - biting yet — DNS is MagicDNS, inference is inbound — but blocks fv-ml1 pulling from any - fleet LAN host. **Next concrete task.** -- ⚠ Plaintext creds to delete: `/tmp/opn.pw` (nh3-dev), `/tmp/io.pw` + `/tmp/key.io` (fv-ml1). - All four are vaulted under `fv-gateway/` and read-back verified. -- ⚠ FV WAN rule `InfraOps` is scoped to alias `fleet_egress` (NH3 70.230.226.88 / ANA - 38.120.12.42 / ESH 128.177.138.182). Operator wants it up a few days, then close. -- ⚠ headscale preauth keys `headscale/preauth-fv-{router,client}-7d-20260912` expire - **2026-09-19** — revoke after the build settles. +**19 of 30 gateway aliases are down and there is NO local fallback** — every free local +model lived on fv-ml1, and irv-ml1 runs no chat seat at all. The only non-fv chat +backends are paid (z.ai, Moonshot). Any paid coverage must be a NEW opt-in alias, never a +silent repoint of `summarizer`/`gen`/`classifier`. -### Next up — the VRAM the fleet didn't know it had -- fv-ml1 has **4× RTX PRO 6000 = 391 GB**, not the 196 GB every doc claimed. Operator: - "we have some fun things to do with the vram we now have." Revisit seat placement and - whether seats split across irv-ml1/gx10 can consolidate. Nothing decided yet. +Cause per operator ruling: **the 1500 VA Eaton UPS was undersized** and gave way under +two-card GPU load. Arithmetic: 2x Blackwell Max-Q @300 W ≈ 600 W of card + ~200-350 W +host ≈ **800-950 W against the unit's real ~900-1200 W** — at or over the line, which is +what explains a full day on ONE card and failure minutes into the SECOND. + +⭐ **Everything needed for the visit — the bring-list, the measurement protocol, the +staged recovery, the three circuit questions, the split-power proposal and the open +design gaps — is in `docs/runbooks/fv-site-dark-20260913.md`. Read that, not this.** +The two highest-value items from it: recover `/tank/aimodels/flash-next-mtp-bench/power.log` +BEFORE anything else (the only load measurement that exists, GPU-only, written locally +every 10 s), and **mask Docker before the network comes up** — ten +`restart: unless-stopped` vLLM containers loading at once is the largest transient the +box can make, into whatever just failed. + +### Built today, dark with the box, resumes on recovery +- **`stacks/flash-next-seat/`** — Qwen3.8-Flash-Next (abliterated) on fv-ml1 GPU 2 + `:8022`, the first seat whose weights do not fit its card: 176B total with a 51B n-gram + table pinned in host RAM and read over CUDA UVA. Measured and verified serving before + the outage — 74.36 GiB resident, 14.00 GiB KV = 560,654 tokens at the full 262,144 + context, `gen-large` answering through the gateway in 0.4 s. + → `persistent-memory.d/2026-09-13-flash-next-seat-and-fv-outage.md` +- **`services/flash-next-mtp-bench/`** — the MTP campaign, INCOMPLETE. One clean `off_A` + rep banked (75.5 / 212.3 / 387.8 tok/s at conc 1/4/8) before the power failed. The + hypothesis is live and unpublished: the head is ONE layer run autoregressively, the + vLLM recipe only ever tested k=3, so **k=1 may win.** ⚠ Do NOT restart it during + recovery — it is the prime suspect for the outage. +- **Power caps DECIDED, unexecuted** (box dark): fv-ml1 **250 W**/card, ana-ml3 **200 W**. + ⚠ `--kv-cache-memory`-style gotcha applies: `nvidia-smi -pl` caps BOARD power, so + 4x250 + host lands ~1310-1445 W at the plug against a 15 A circuit's 1440 W continuous + limit — **verify at the plug, fall back to 200 W if it reads near the limit.** +- **Stale homepage labels on 10 containers** — the compose files are fixed, the containers + died before recreation, and a plain power-on will NOT apply labels. The staged + `compose up -d` recovery sequence fixes them as a side effect. + +### FV pre-outage state, for reference +- fv-ml1 (ex ana-ml2) at `10.251.50.54` / mesh `100.64.0.7`; **vb-gateway** OPNsense + `10.251.50.1` / `100.64.0.8`; **BMC** `10.251.250.50`. Public `fv.phasefinal.com` → + `172.83.89.66`. ⭐ **4x RTX PRO 6000 Blackwell Max-Q @300 W = 391 GB VRAM** (every doc + said two cards until 2026-09-12). → `persistent-memory.d/2026-09-12-fv-cutover-executed.md` +- ✅ The old "fv-ml1 cannot initiate to fleet LAN IPs" gap is **CLOSED** — another session + landed a scoped OPNsense hybrid SNAT for fv-ml1→ANA at 06:22Z and verified Beszel 18/18. + → `docs/runbooks/fv-to-ana-nat.md` (that session's file, uncommitted, left alone) +- ⚠ Still open, blocked on the box: delete plaintext creds `/tmp/io.pw` + `/tmp/key.io` + (fv-ml1) and `/tmp/opn.pw` (nh3-dev) — all vaulted under `fv-gateway/`, read-back + verified. Revoke `headscale/preauth-fv-{router,client}-7d-20260912` (expire 2026-09-19). + Close the FV WAN `InfraOps` rule (scoped to alias `fleet_egress`) once done with it. +- ⚠ Eight stacks deliberately NOT pushed from canonical during the renumber because their + host copies have genuine drift; three of those (`mistral-medium-3.5`, `ms32-24b-angel`, + `qwen35-vl`) are **untracked host-only stacks** that should be brought into `stacks/`. + +### ana-ml3 — not yet deployed +- 2x **Ada Generation RTX 6000** @300 W, 2x48 = 96 GB. Lands in the **Anaheim** rack whose + breaker tripped 2026-08-26 and 2026-09-11, so the 200 W cap is remediation, not caution. + ⚠⚠ **sm_89 has native FP8 but NO NVFP4** — most in-house quants are NVFP4 and will not + run accelerated there; its seats want FP8 W8A8, or those checkpoints stay at FV. ⭐ It + unparks the triton-backend item (hard no on Ampere, `fp8e4nv` unsupported on sm_86, + explicitly deferred TO Ada). ### BabyYarros — complete; pair-corpus rebuild is the next step - Both arms trained + evaluated; Instruct renders beats 9/10 by a **lexical** metric that @@ -183,12 +192,12 @@ last, and do NOT restart the MTP campaign. paragraphs are 90-140w, so the pair unit must be a ~4-paragraph scene window (6,445 non-overlapping, 88% corpus coverage). Critical path is response-only loss masking in `train_voice_lora.py` (currently `labels = ids.clone()`), ~1 day. -- ⛔ Frozen adjudication still deferred (needs the gen seat, now back at FV). +- ⛔ Frozen adjudication still deferred — needs the gen seat, which is at FV and dark. ### Quants — cyber-preview to re-run - **sentinel-r3** NVFP4 complete at `/tank/aimodels/sentinel-r3-nvfp4-mixed`; acceptance A-B still needs a serving slot. **cyber-preview** died mid-quant in the Anaheim outage — - re-runnable now that FV is up. + re-runnable once FV is back. ### Jetson AGX Orin — 64 GB, in hand, unassigned - Operator has the 64 GB dev kit plus ~6 undeployed cameras, a depth camera and lidar for @@ -403,11 +412,9 @@ last, and do NOT restart the MTP campaign. - `[2026-09-01]` **Idle VRAM on this fleet is a RESERVED scratch pool, not waste.** Operator declined raising `vllm-mog-sec` from `gpu-memory-utilization 0.52`: single-user dev fleet, KV headroom nobody will consume is worth less than room for ephemeral models and small training runs. vLLM's "fully utilize gpu memory" startup hint does NOT apply here. Tracked in auto-memory `feedback_idle_vram_is_reserved_not_waste`. -- `[2026-08-25]` **Run 2's base is an OPEN OPERATOR DECISION, deliberately not staged** — four options with materially different safety postures, detailed in Current state. Tracked at althing thread `01M0WQ8W5574KMEVCHCEKEXNS5`. ⚠ Do not let it get filed as a config knob; it is a reversal of the trainee-selection decision. - `[2026-08-25]` **Fused MoE kernel path — DEFERRED, tracked at park `fused-moe-kernel-path-for-gemma-4-moe-training` (id 47).** Operator: "note the fused MoE kernel for round two… if we nail it soon, the math has us wanting to restart the run anyway." Training MFU is **8.6%** (27.1 of a benchmarked 313.8 TFLOPS) because `transformers` runs the Gemma-4 experts in a Python loop — 128 experts × 30 layers, ~11,500 iterations per step under gradient checkpointing. ⚠ **The same fused 3-D expert layout that made bitsandbytes skip 88.5% of the model is exactly what a grouped GEMM wants** — the format is good for storage and for fused kernels, and hostile only to naive iteration. Two fixes: `group_by_length` (−29.9% compute, free, but breaks the seeded order manifest and re-opens a batch-composition call brokkr already made) and a grouped-GEMM/compiled MoE forward (the remaining ~10×). **Not applied to the live run** — restarting mid-flight to change batch ordering was judged a bad trade at step ~50 of 1,312. -- `[2026-08-24]` **Serving the tuned ERP model: LoRA-on-NVFP4 PREFERRED, merged weights the expected fallback — and the recorded objection may be STALE.** Operator: "if you CAN load it as a lora, all the better, the issue is that we will want to run nvfp4 weights, which we had some serious trouble with loading loras on top of nvfp4." ⚠ **The archived root-cause says it was NOT NVFP4-specific**: `[2026-07-07]` vLLM 0.24.0 qwen3_5 LoRA application was a silent no-op (#47639, regression from #37912) — adapter loads HTTP 200, zero deltas at inference, proven **quant-agnostic (NVFP4 AND FP8 both inert)** and adapter-format-agnostic by a 3-peer dwarf panel. Fix PR #47640 was OPEN then. **ana-ml2 is FAR past 0.24.0 and the box runs a SPREAD, not one version** (measured 2026-08-24): `gen` on `nightly-311b3513` = **0.27.2rc1.dev150**, `mog-sec` on `nightly-e9d1398d` = 0.26.1rc1.dev1102, the small seats still on 0.24.0, and char-rp/trainee-bench pinned to v0.26.0. ⚠ **`vllm/vllm-openai:v0.27.1` is already ON DISK, unused** — a TAGGED release, which is the right retest target: no nightly variance, no pull, ~4 months past the diagnosis. So: RETEST hot-swap LoRA on **v0.27.1** before designing around merge — it is cheap, and if it works the post-tune gate can be two aliases on one engine. If it still no-ops, merged weights it is, which means the harness must EMIT merged weights and Eitri needs that in the contract while he is early. Tracked at this snapshot commit; settle it in the QLoRA sizing conversation. - `[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`. @@ -422,7 +429,7 @@ last, and do NOT restart the MTP campaign. -_28 older entries archived to archival-memory.md._ +_30 older entries archived to archival-memory.md._ ## Tried and abandoned