memory: snapshot — run 3c gated, run 4 training, DAC revert, GPU rebalance, NAS exposure

This commit is contained in:
vh
2026-09-05 06:36:56 -07:00
parent 5237efa299
commit 82158e7e36
13 changed files with 958 additions and 651 deletions
+63 -39
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-09-03_
_Last updated: 2026-09-05_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under an hour old, read it (it carries the in-flight
@@ -108,48 +108,76 @@ no longer deployed sidecars here. See Recent decisions.)
(no NOPASSWD)** — stage model pulls to `/home`, not root-owned `/worktank`.
## Current state / in-flight
_As of 2026-09-03 — **nothing is mid-action; the session closed clean.** Every
item below is a live commitment or a known-open risk, not work in progress._
_As of 2026-09-05 06:35 PDT — **ERP run 4 is TRAINING on pfi-gx10.** Everything else
below is a live commitment or a known-open risk._
- **Nothing is running.** No training, no deploys pending, no background jobs.
Commits sit unpushed on `main` — all docs, runbooks, memory and scripts;
push is the operator's call.
- **⚠ RUN 4 IS MID-FLIGHT — do not touch the GX10 GPU.** `~/erp-tune/run-04.pid`,
log `~/erp-tune/run-04.log`. At 06:32 it was **486/938 steps**, 6 h 17 m elapsed,
a genuinely settled **46.0 s/it** (unlike 3c, which climbed 52→70 — airoboros rows
are short and single-window, so there is no long tail for the sampler to find).
**~12.0 h total, finishing ~12:15 PDT 2026-09-05.** Loss ~2.04 at step 450,
gnorm well under 1, checkpoints every 50. **Ping brokkr-smithy-dev at completion**
— he takes base floors on the GX10 first, then the tuned arm, serially.
- **Run 3c is staged on pfi-gx10 and awaiting the operator's go.** Everything is
verified and one command away (`ssh infra-ops@10.100.50.60
'~/erp-tune/launch-run-03c.sh'`); ~13.3 h once started. Its train loop is the
one piece never exercised on sm_121 — watch the first three minutes.
- **Operator ruling on the GX10: training first, serving transiently.** *"it's mostly
for training, but can serve its trials. unless the box is needed for training work."*
So `trial` (= run 3c on :8098) is down for the duration and comes back when run 4
ends. I over-read an earlier version of this as "training-only" and had to correct
it to brokkr — his serial floors-then-arm plan on the GX10 was never wrong.
- **`web_search` needs a session restart to appear.** The SearXNG MCP server is
registered at user scope and `claude mcp list` reports it Connected, but MCP
servers load at session start — this session does not have the tool.
- **`trial` gateway alias is a live 404 while the seat is down** — expected, not a
fault. Restore with `~/erp-tune/relaunch-trial-seat.sh` on the GX10 (hand-run by
operator ruling: experimental, NOT a compose stack, does not survive a reboot).
- ⚠ **SearXNG general search is effectively single-engine.** `google cse`
carries it; `brave`/`duckduckgo`/`startpage` CAPTCHA even from NH3's
residential egress. If google cse breaks it goes quiet the same way it just
did. `scripts/searxng-health.sh` is the detector.
- ⚠ **Three dead gateway aliases return HTTP 500, not 404/503**: `trial`,
`gemma4-26b-a4b-it-base`, `erp-tune-v2`. A dead seat reporting an *internal error*
reads as an outage — brokkr checked his own work against mine because he could not
tell. Deregistration costs a ~60 s fleet-wide LiteLLM restart; batch it with the
next gateway change rather than spending a restart on tidying.
- ⚠ **The nh3-dev backup throughput cause is UNEXPLAINED.** Fleecing means a
slow `pbs-ana` can no longer wedge the guest, but a job that once ran at
- **`trial` is on the SHARED-KEY gateway with a measured −40pp selfharm/methods
regression.** Flagged to the operator twice (before adding, and after the gate
measured it); he has left it up. His direct endpoint `10.100.50.60:8098` gives the
same access with a blast radius of one. Settled — do not re-litigate.
- **ESH DAC: reverted to autoneg/1G, fiber going in at the weekend.** The operator
ran copper through a drilled floor 2x4 himself; recommendation was a 10Gtek
SR 2-pack + OM4 3 m LC-LC (~$40–65) because cable-vs-pull-damage was never resolved.
- ⚠ **Cityside fiber `/30` is NOT provisioned.** `128.177.138.182/30`, gw `.181`.
Static passes no traffic and DHCP still hands CGNAT `100.104.3.250`; operator
power-cycled both ends and opened a ticket. Cutover payloads stay staged:
`wan1-REVERT.json`, and the `esh-ana` IPsec fix (`ipsec_local_ip 100.104.3.250 →
128.177.138.182`) **which will otherwise silently break ESH→Anaheim restic backups.**
- ⚠ **esh-nas is effectively open to the whole ESH LAN** — twelve NFS exports rw to
`10.0.0.0/8` with `sec=sys`, and every SMB share but `backup` guest-writable.
Hardening offered, ~1 h, **operator has not ruled**. → `persistent-memory.d/2026-09-04-esh-nas-smb-and-exposure.md`
- ⚠ **The nh3-dev backup throughput cause is UNEXPLAINED.** A job that once ran at
941 MiB/s ran at 1.4 MiB/s with the link up and pbs-ana answering in 11 ms.
Nightly at 21:00, `all 1`. Worth its own investigation.
Nightly 21:00, `all 1`. Worth its own investigation.
- ⚠ **pfi-gx10 is now single-path.** Wired only on a reserved 10.100.50.60; the
Wi-Fi escape hatch is deliberately gone. A switch-port or reservation failure
is a rack visit — on the box run 3c was moved to.
- **ledger-dev asked about the Ledger→SVOS rename** (gitea repo rename cost, and
moving `nh3-dev/development/ledger/env.sh` in the vault — the `secret` CLI has no
rename, so it is re-put + delete). **Unanswered, no deadline**, nothing moves until
the operator fixes the strings.
- **Open commitment to vastblue-dev:** a dedicated CI runner, gated on their
first client-premises release cut (U10, unscheduled). They wrote the gate down
specifically so it would not quietly become never; ping expected when U10 is
scheduled.
- **Open commitment to vastblue-dev:** a dedicated CI runner, gated on their first
client-premises release cut (U10, unscheduled). Ping expected when U10 is scheduled.
- **Neither Mac nor the Studio is in `servers/` or `dns/internal.yaml`** —
deliberate, they are the operator's personal machines and registering them
implies PFI-managed. They now carry real config, so the omission is a choice
to revisit, not an oversight.
- **Neither Mac nor the Studio is in `servers/` or `dns/internal.yaml`** — deliberate;
they are the operator's personal machines. A choice to revisit, not an oversight.
## Recent decisions
- `[2026-09-05]` **A peer's "2.7x serving-stack effect" was a coin flip — the operator rejected it on instinct and the arithmetic backed him.** Each floor was `|b0-b1|` from n=2; the ratio is half-Cauchy, P=0.452. ⚠ The disconfirming evidence sat in brokkr's own sentence, and he named *why*: the claim was his and flattering. → `persistent-memory.d/2026-09-05-floor-claim-n2-retraction.md`
- `[2026-09-05]` **vLLM RUNS on sm_121 — the blocker was `ninja` off PATH, not the silicon** — and run 4 launched after two peer artifacts were rejected by reading the harness rather than accepting a "confirm this". ⚠ I declined to fill in missing `root_sha256` values I knew, because supplying both sides of a check makes it inert. → `persistent-memory.d/2026-09-05-vllm-on-sm121-and-run4.md`
- `[2026-09-04]` **ERP run 3c trained and GATED — the 20x LR cut erased the diversity gain and did NOT remove the −40pp selfharm regression.** LR-robust, so it comes from corpus content. CSAM clean on all three arms. ⚠ A pooled preserve-list test cannot see a single-axis collapse. → `persistent-memory.d/2026-09-04-run3c-trained-and-gated.md`
- `[2026-09-04]` **`gen` moved to ana-ml2 GPU0 to stop vllm-embed OOM-crashing (7 restarts) — and I sized it against vLLM's declared budget, not its runtime footprint.** Cost: gen KV down to 1.02x concurrency at 262K. → `persistent-memory.d/2026-09-04-ana-ml2-gpu-rebalance.md`
- `[2026-09-04]` **SMB account `dsp` created + vaulted for the Windows AudioGridder box — and esh-nas turns out to be wide open.** Twelve NFS exports rw to `10.0.0.0/8`, guest-writable SMB. → `persistent-memory.d/2026-09-04-esh-nas-smb-and-exposure.md`
- `[2026-09-04]` **SLZB-MR1U Zigbee coordinator moved to esh-iot (VLAN 90) at `10.0.90.10`, DHCP-reserved, DNS'd, handed to ha-dev.** ⚠ Home Assistant cannot resolve `.internal` at all (Docker's 127.0.0.11 upstream excludes the fleet AdGuard) — pre-existing; ha-dev declined the fix. Runbook `docs/runbooks/slzb-mr1u-zigbee-coordinator.md`, commits `fed29be`/`0bbdaf9`.
- `[2026-09-03]` **Run 3c is STAGED on pfi-gx10 and deliberately NOT launched** — the launch is a 13.3 h commitment and the operator stood this port down once already. Base shards AND the encoded corpus sha256-verified identical to ana-ml2's, so the transformers 5.15.1→5.16.1 / x86-64→aarch64 delta is *measured* inert. ⚠ The encode-cache FILENAME differs by design (`base_model_path` is in the key) — input hash, not output. ⚠ Tripped the `pkill -f` ssh self-match again; the launcher guards on a pidfile because of it. → `persistent-memory.d/2026-09-03-gx10-run3c-staged.md`
- `[2026-09-03]` **SearXNG returned ZERO results for every query while reporting `healthy` for 7 days — 4.5 months stale.** Moved to nh3-docker (residential egress beats the colo's CAPTCHA-gated 38.120.12.42), updated, and exposed to every CC session as the user-scope `web_search` MCP tool. ⚠ `/healthz` cannot tell you whether search works. → `persistent-memory.d/2026-09-03-searxng-nh3-move.md`
- `[2026-09-03]` **pfi-gx10 racked: VLAN 50 via a DHCP RESERVATION on the UDM, not a host static — operator ruling, so the box stays portable.** ⚠ The racked port arrived on the NATIVE VLAN; ⚠ `port_overrides` is a whole-array PUT; ⚠ prove inter-VLAN routing with `ping -I <wired>` BEFORE downing the Wi-Fi escape hatch. Now single-path. → `persistent-memory.d/2026-09-03-gx10-rack-network.md`
@@ -232,18 +260,12 @@ item below is a live commitment or a known-open risk, not work in progress._
- `[2026-08-22]` **`sec` retuned to util 0.52 / 420K after a runtime OOM at 0.55/480K** — `gpu-memory-utilization` is not a hard reservation; activation grows past the dummy-data profile and six vLLM containers share GPU1. Also measured: the KV pool varies ~6.6% between boots, so max-model-len must be sized against the *lower* observation. (`6e82899`)
- `[2026-08-22]` **Max-Q 1.8× spread does NOT apply to LLM decode — measured, not argued.** ana-ml2 draws 256–266 W of 300 W under sustained 100% decode with `SW Power Cap: Not Active` and clocks pinned. Corrected to brokkr-smithy-dev after I had lent the claim credibility; 122B figure (~90–93 tok/s at 262K) stands as a straight number.
- `[2026-08-21]` **ESH internal IPv6 live on two LANs; the Cityside v4 static is a CARRIER problem, proven.** A full gateway reboot forced a fresh DHCP DISCOVER and returned the identical CGNAT address. YaRN was already configured — "1M needs YaRN, absent" was false. → `persistent-memory.d/2026-08-22-dflash2-spec-decode.md` sibling entry in `ad21302`
- `[2026-08-21]` **speaches ASR live on irv-ml1 for Eyra — and `no_speech_prob` alone is a weak hallucination gate.** Silence and room tone both hallucinated "Thank you." under 0.11; `avg_logprob` separates ~6× better. Consumers should gate on a composite. (`aa5863c`, `c7e2187`)
- `[2026-08-20]` **Cold-Fusion abliteration — Robinson recipe captured; the fight was the environment, not the recipe.** Stock Cold-Fusion measured ~33% creative refusal → worth abliterating ourselves (supersedes waiting for DavidAU's heretic build). Recipe maps 1:1 (131 tensors); capture succeeded only in **fp32** — transformers' Qwen3.5 DeltaNet linear-attn NaNs nondeterministically in bf16 without the unbuildable `causal-conv1d` kernel (precision cancellation, not overflow). Direction finite at layer 22 but agreement 0.59 (vs Robinson's 0.99) → **calibration-set expansion is next.** → `persistent-memory.d/2026-08-20-coldfusion-abliteration-capture.md`
- `[2026-08-19]` **Fleet `.internal` DNS built and live — git-sourced, agent-managed, three resolvers.** Zone-scoped authority (ESH's hand-made `esteban.net` rewrites survive); the colo had no resolver at all; v6 column empty on purpose because SLAAC addresses rotate. → `persistent-memory.d/2026-08-19-fleet-internal-dns.md`
- `[2026-08-19]` **waterland studio containerised on irv-ml1 — three landmines, all measured.** cupy needs CUDA *headers* the host had by accident; `uv run` re-syncs and prunes cupy at RUNTIME; the A6000 is container-index 0, not the host's 1. → `persistent-memory.d/2026-08-19-waterland-studio-containerised.md`
- `[2026-08-19]` **Homepage cleaned up, then themed with Australis Skyfall + an Arbo-generated background.** Includes the hour lost to a self-healing tab-bar red herring, and the CSS-iteration loop that prevents it recurring. → `persistent-memory.d/2026-08-19-homepage-skyfall-theme.md`
- `[2026-08-19]` **Four unmanaged stacks found on live hosts — two quietly broken.** A dashboard card is a cheap census of what is actually running; check whether the stack is even in `stacks/` before debugging the symptom. → `persistent-memory.d/2026-08-19-unmanaged-stacks-searxng-seafile.md`
- `[2026-08-19]` **AI-tab Dormant regrouping BELAYED by the operator** — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong, Chatterbox Fast) show amber EXITED inside live groups rather than `AI - Dormant`. Fix is a label change + recreate per stack; needs the operator's read on which are retired vs temporarily down. `untracked by operator choice` (his words: "belay the ai dormant regrouping for now").
@@ -282,10 +304,12 @@ item below is a live commitment or a known-open risk, not work in progress._
_Older entries archived to archival-memory.md._
_242 older entries archived to archival-memory.md._
_248 older entries archived to archival-memory.md._
## Tried and abandoned
- `[2026-09-04]` **Forcing 10G on the ESH-Media DAC — it linked, then degraded over hours, and I reported a plateau at two minutes.** ⚠ A clean zero-error link at 1G does NOT rule out a marginal cable; autoneg's fallback was protecting something real. → `persistent-memory.d/2026-09-04-dac-forced-10g-failed.md`
- `[2026-08-25]` **Four throughput levers measured and killed — do not re-chase.** (1) **Fused MoE / `grouped_mm`** — 0.9% *slower* than the Python loop and dense GEMM is only 7.9% of the step, capping the whole category near 10%. (2) **CUDA graphs / `torch.compile` over the expert loop** — the two-term scaling fit closed with residuals under 3ms and needed NO constant term, so there is no fixed per-batch cost to amortise; 3,840 expert-GEMM launches per forward are not what we pay for. (3) **`liger` fused linear CE** — the chunked CE measured **1.1% of the step** forward, ~3% with recompute. A tidy-up, not a lever. (4) **Selective gradient checkpointing** — ~2% of a post-fix step, real bug surface. Also: **token-budget batching is dead by the same fit** — with no constant term, total time over a fixed set of widths is invariant to how you group them; only the widths matter, which is exactly why bucketing works and repacking does not.
- `[2026-08-25]` **`sample_packing` is NOT strictly better than bucketing on this model, and I told the operator it was before brokkr corrected me.** Packing needs FA2 varlen or a block-diagonal mask; FA2 is unavailable here (head_dim 512 > 256 cap), so packing means an explicit 4D mask on EVERY batch. Bucketing produces **78.3% exactly-zero-pad micro-batches** which recover the `is_causal` fast path on the 5 global layers — measured at 9.4% of step time. Packing forfeits that. ⚠ **The conclusion flips under `flex_attention`**, where a block-diagonal mask is just another BlockMask: do not carry "packing is bad" past the backend decision.
- `[2026-08-25]` **Merging a tune back toward STOCK to fix overfitting would UNDO the abliteration.** brokkr recommended a 50/50 merge-back, then retracted it himself: the published recipes merge into `google/gemma-4-*-it`, and following that literally re-installs exactly the refusal directions the abliteration removed — silently, because the merged model looks *healthier* on general benchmarks. Any merge-back must target the SAME abliterated base. Wider lesson: **recipe cards are per-checkpoint artifacts, not per-family** — the advice came from a card for a DENSE STOCK 31B applied to a MoE ABLITERATED 26B-A4B, three axes apart on a shared name.