memory: snapshot — speech seat live + gen-small OOM incident (mitigated, fix tasked); leftovers deleted; no Scriberr upstream; repos pushed; 32 entries archived

This commit is contained in:
vh
2026-10-01 04:24:44 -07:00
parent 6b66207b0c
commit f83e35b94b
32 changed files with 1189 additions and 1165 deletions
+32 -83
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-09-30 ~1800 PT (U11a legacy off on both Worldtree instances, U11b gate armed; SemIf → intern-decision with Jev /v1/systemone at 32k; Scriberr → GPU 3 with overlap slicer + gap retry; Parakeet seat switch to unified-en APPROVED, next session; 26 old entries archived.)_
_Last updated: 2026-10-01 ~0420 PT (Parakeet seat → unified-en under NeMo LIVE + audited; gen-small util 0.36 → .env 0.33; leftover bench weights + spike dirs deleted; Scriberr no-upstream; eshpfi + worldtree-instance-configs pushed. Prior: U11a off, SemIf → intern-decision, Scriberr GPU 3 + patches.)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
@@ -115,53 +115,44 @@ no longer deployed sidecars here. See Recent decisions.)
## Current state / in-flight
_As of 2026-09-30 ~1800 PT._
_As of 2026-10-01 ~0420 PT._
### Parakeet speech seat switched to unified-en under NeMo (Prime 2026-09-30 ~1755: "reasonable terms, ship the switch")
### Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01)
- **DONE: LIVE since ~0126 PT 2026-10-01** as `parakeet-nemo` (`stacks/parakeet-nemo`, image `local/parakeet-nemo:nemo-0.1.0`, infra-hermes de6ea32 + 41d2014) on :8300; LiteLLM untouched. **infra-ops AUDIT PASSED 0137.**
- Latency on GPU 0, p50 for 1–3 / 3–8 / 8–20 / 20–60 s: 33 / 36 / 42 / 71 ms, against 187 / 308 / 626 ms before. Through LiteLLM a 2.5 s clip takes 80–98 ms.
- **⚠ INCIDENT 04:21 PT 2026-10-01, MITIGATED, root fix in flight:**
- `vllm-gen-small`'s EngineCore CUDA-OOM'd when it needed a 394 MiB runtime workspace and GPU 0 had 388 MiB Free. The parakeet seat was parked at its 3,582 MiB window cache.
- vLLM grows ~0.8 GB at runtime beyond its preallocation; my audit checked gen-small's boot margin, not its runtime growth. gen-small auto-restarted, healthy at 04:23.
- I restarted parakeet-nemo to drop its cache (rest 2,084 MiB). gen-small then answered 3/3 via LiteLLM, and GPU 0 Free is ~1,075 MiB.
- **infra-hermes is tasked with nemo-0.1.1:** `empty_cache` after windowed requests, plus a hard memory ceiling so the seat 503s instead of starving gen-small. It also measures gen-small's runtime growth.
- **Awaiting Prime:** trim gen-small's KV pin (8 → 7 GiB frees ~1 GiB; 670k → ~586k tokens), or move the seat to GPU 3.
- Until fixed, a long transcription can re-grow the seat's cache and starve gen-small.
- **LIVE since ~0126 PT 2026-10-01** as `parakeet-nemo` (`stacks/parakeet-nemo`, `local/parakeet-nemo:nemo-0.1.0`, built by infra-hermes) on fv-ml1 GPU 0 :8300, with LiteLLM `ext-stt`/`whisper-1` unchanged. **infra-ops audit PASSED 0137.**
- p50 on GPU 0 for 1–3 / 3–8 / 8–20 / 20–60 s: 33 / 36 / 42 / 71 ms, against 187 / 308 / 626 ms for the old seat.
- WER: LibriSpeech clean 1.965, other 3.026.
- Fixed: a 714–726 s file returns 200 (long files go in 360 s windows, because NeMo builds the full T×T attention mask even under local attention); no drop after a pause.
- Rollback: `docker stop parakeet-nemo && docker start parakeet` (the old container is stopped, not removed).
- **gen-small: util 0.48 → 0.36** (0.46 and 0.40 failed its boot check). Its KV is byte-pinned (`--kv-cache-memory 8 GiB`), still 670,142 tokens / 2.56×, so there was no KV cost. It was down ~34 min (0012–0046 PDT) while the util was iterated; zero LiteLLM errors.
- ⚠ **Steady state is 3,582 MiB for the seat (its cached window peak); GPU 0 Free is 385 MiB.** That leaves gen-small's restart boot-check margin at only ~0.45 GiB. infra-hermes was asked to set gen-small util to 0.33 in the .env WITHOUT restarting, so it applies at the next restart. **Nothing else fits on GPU 0.**
- Seat invariants are in its README: cast to bf16 AFTER change_attention_model; uvicorn pinned with `--http h11`, because httptools 0.8.0 writes `HTTP/1.1 200\x00OK` and LiteLLM/httpx rejects it; the 360 s window.
- The original brief, for reference:
- **The task:** replace the live speech seat with `nvidia/parakeet-unified-en-0.6b` under NeMo 3.0.0 with bf16 weights. The **NVIDIA Open Model License is ACCEPTED** for internal use.
- The live seat today: container `parakeet` on fv-ml1 GPU 0, port :8300, sherpa-onnx int8 `parakeet-tdt-0.6b-v3`, reached through LiteLLM as `ext-stt` and `whisper-1`; its caller is `talk`.
- Evidence, `docs/pfi/parakeet-seat-ab-2026-09-30.md` (a6c1d3c): end-to-end p50 for 1–3 / 3–8 / 8–20 s clips goes 144 / 260 / 565 → 23 / 27 / 33 ms, and WER is lower on every set.
- **Kit:** the wrapper `services/parakeet-ab-2026-09-30/code/serve_nemo.py` keeps the seat's endpoints and text, and matched NeMo's own transcribe on 400/400. The weights are pinned on fv-ml1 in `/tank/aimodels/huggingface` (rev `fe53cd88`). A working NeMo 3.0.0 env for reference is under `/tank/spikes/parakeet-ab`. **No image is built yet.**
- **What the image needs:**
- a warm-up at the longest served length;
- a bf16 cast BEFORE `.to(cuda)`, which avoids a +1.5 GB load spike;
- local attention for long files (a 30-min file took 2.6 s in one request).
- **Room:** it needs about +1.1 GB while serving (+1.5 GB at load) over the seat's 1,690 MiB, and GPU 0 has ~100 MiB free.
- The plan is to trim `vllm-gen-small` `--gpu-memory-utilization` from 0.48 to about 0.46 at a quiet moment (a 2–3 min restart).
- ⚠ The util value does NOT predict resident VRAM: on 09-15 gen-small at 0.48 held 36,942 MiB and cyberprev at 0.40 held 47,124. MEASURE nvidia-smi Free after the change; do not compute it. My 09-30 "0.01 ≈ 0.95 GB" estimate is unverified.
- GPU 1's ~6.6 GB free is intern-decision's 32k headroom, so it is not available.
- **Cut-over:** keep the old seat as the rollback, and leave LiteLLM alone unless the port changes. Re-measure live on GPU 0: latency per length bin against the old seat, a WER spot-check, and memory.
- **Live-seat defects until then:**
- HTTP 500 above ~400 s of audio;
- long-form dropouts;
- after a 1.5 s digital-silence pause it can drop the rest of the utterance (6 of 40).
- Files longer than 6 min run in 360 s windows. That avoids NeMo's T×T attention mask; a seam can lose a space or a word.
- **Rollback:** `docker stop parakeet-nemo && docker start parakeet`. The old container and image are kept.
- **GPU 0 is FULL:**
- The seat's steady state is **3,582 MiB** (its cached window peak); Free is **385 MiB**.
- `vllm-gen-small` runs at util 0.36, and its `.env` holds **0.33** for the next restart (~3 GiB of boot-check margin). Its KV is byte-pinned: 670,142 tokens / 2.56×.
- Before restarting any vLLM seat on this card, check that util × 95.6 GiB ≤ measured Free + the seat's own resident memory. Do not trial-boot. A trial-boot sequence took gen-small down for 34 min on 2026-10-01.
- Seat invariants (in its README): cast to bf16 AFTER change_attention_model; uvicorn pinned with `--http h11` (httptools 0.8.0 emits `HTTP/1.1 200\x00OK`, which LiteLLM/httpx rejects).
- NVIDIA Open Model License accepted for internal use. → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md`
### Worldtree U11 memory cutover (demo + personal)
- **Legacy plane OFF since 0115/0120 PT 2026-09-30** (config repo 63cf268; personal /metrics in b6fdd81). → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
- **Daily gate batches:** infra-hermes runs them from 2026-10-01 with `scripts/wt-memory-gate-batch` and copies me on every verdict. The count is **1 of 3** consecutive PASS at off (20260930T090608Z); a FAIL restarts it.
- **Legacy plane OFF since 0115/0120 PT 2026-09-30** (config repo 63cf268; personal /metrics b6fdd81; the repo is pushed). → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
- **Daily gate batches:** infra-hermes runs `scripts/wt-memory-gate-batch` from 2026-10-01 and copies me on every verdict. The count is **1 of 3** consecutive PASS at off (20260930T090608Z); a FAIL restarts it.
- **⚠ U11b STEP 5 IS MINE, triggered by the 3rd consecutive PASS:**
1. Run `docker exec -i <api> python - < scripts/wt-h2-count.py` VERBATIM, right before each instance's deletion. Exit 2 means STOP and send worldtree-dev the output.
2. Delete LIVE, with the api running, using literal paths only: `agents/{forseti,lofn,mimir}/memory/<agent>.chroma` and `memory/context_promotion`, on BOTH instances.
3. Send worldtree-dev the stamp; b193 ships after it.
- A b192 restart re-creates an empty schema-only `ledger.db`. That is residue, not memory data: say so in the stamp, and remove it after b193.
- **Legacy archive:** DESTROY it whole by 2026-10-30, or at retirement-done, or on any subject-erasure request, whichever comes first. The runbook is in the detail file.
- **TODO:** re-sweep both api logs after real traffic, grepping `legacy executor|Traceback|ERROR|will not be remembered`. After the b193 push, remove the retired config keys at my pace.
- A b192 restart re-creates an empty schema-only `ledger.db`. That is residue: say so in the stamp and remove it after b193.
- **Legacy archive:** DESTROY it whole by **2026-10-30**, or at retirement-done, or on a subject-erasure request, whichever comes first. The runbook is in the detail file.
- **TODO:** re-sweep both api logs after real traffic. After the b193 push, remove the retired config keys.
### fv-ml1 GPU layout (as of 2026-09-30)
### fv-ml1 GPU layout (as of 2026-10-01)
- **GPU 0:** cyberprev (47.1 GB), gen-small (37.5 GB), voices (10.8 GB), the parakeet seat (1.7 GB); ~100 MiB free.
- **GPU 0:** cyberprev (47.1 GB), gen-small (35.3 GB), voices (10.8 GB), parakeet-nemo (3.6 GB steady). Free 385 MiB, FULL.
- **GPU 1:** vllm-coder, erp-seat, meromero-rp, plus intern-decision (cap 14.4 GiB, 32k tokens, peak 15,220 of a 15,437 MiB budget). FULL.
- **GPU 3:** the full-size-seat reserve (Flash-Next is parked). On-demand tenants: Blender, and Scriberr (0 idle, ~5.5 GB per job). When a full-size seat claims GPU 3, Scriberr steps aside to **irv-ml1's A6000**, not back to GPU 1.
@@ -172,17 +163,7 @@ _As of 2026-09-30 ~1800 PT._
### Scriberr (fv-ml1 GPU 3)
- **LIVE `scriberr:local-blackwell-a353078-dropout2`:** upstream a353078 plus patch 0001 (overlap slicer) and patch 0002 (gap retry, `PARAKEET_MODEL_PATH`). v3 stays (Prime: no NeMo 3.0.0 surgery). → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md`
- **Awaiting Prime:**
- Open the slicer upstream PR, and choose which GitHub account (`stacks/scriberr/patches/upstream-pr/`).
- Delete the leftovers:
- the candidate weights in `/tank/aimodels/huggingface`, EXCEPT unified-en, which the seat switch needs;
- `/tank/spikes/scriberr-slicer`, including `private/`, which holds Prime's recordings (mode 700);
- `/tank/spikes/parakeet-ab` (~25 GB), but only after the switch.
### irv-ml1 /storetank: CLOSED (2026-10-01)
- Prime ruled, via comfy-dev: "delete unused weights + old staged files". comfy-dev executed it himself, 84 files, logged at irv-ml1 `~/3d-dl/deleted-2026-10-01.log`. Free went from 260 to 299 GB (86% → 84%). Tiers B/C/D got no ruling. infra-hermes closed it (thread `01M3TCSYRSFNAPA9BPQTMFQ6KJ`).
- **LIVE `scriberr:local-blackwell-a353078-dropout2`:** upstream a353078 plus patch 0001 (overlap slicer) and patch 0002 (gap retry, `PARAKEET_MODEL_PATH`), carried LOCALLY ONLY (Prime 2026-10-01: no upstream). v3 stays. `scripts/scriberr-rebuild` re-applies both. → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md`
### nh3-pve + nh3-ml1: post-visit, all live (2026-09-25/26)
@@ -305,7 +286,7 @@ _As of 2026-09-30 ~1800 PT._
### Live threads
- git: origin/main is at `128d1d8`, pushed 2026-09-30 1047 by someone other than infra-ops (presumably Prime). Local is ahead with unpushed commits, this snapshot included. `worldtree-instance-configs` has 3 unpushed commits (0a1387e, 63cf268, b6fdd81). Pushing is Prime's call. ⚠ The working tree AND index are shared with infra-hermes and subagents: commit with `git commit -- <paths>` (auto-memory `feedback_shared_git_index_commit_pathspecs`). `graphify-out/GRAPH_REPORT.md` stays modified and uncommitted on purpose: it is auto-regenerated.
- git: **eshpfi-management pushed to `6b66207` and worldtree-instance-configs to `b6fdd81` (2026-10-01 ~0418, Prime's go).** Anything after that is unpushed. ⚠ The working tree AND index are shared with infra-hermes and subagents: commit with `git commit -- <paths>` (auto-memory `feedback_shared_git_index_commit_pathspecs`). `graphify-out/GRAPH_REPORT.md` stays modified and uncommitted on purpose: it is auto-regenerated.
- nh3-dev root disk was cleaned 2026-09-30 1704 (uv prune, dangling images, old build cache): 86% → 82%. The Beszel 85% alert flaps near the line.
- Booth submit-all fix (Prime's report) is LIVE since 2026-09-27 1705, via booth-dev (booth `50bfc7b`).
Pushing it is booth's call, per Prime; it is not ours.
@@ -313,6 +294,8 @@ _As of 2026-09-30 ~1800 PT._
## Recent decisions
- `[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push.** DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; **unified-en KEPT, the live seat mounts it**), plus `/tank/spikes/scriberr-slicer` (including the private copies of Prime's recordings) and `/tank/spikes/parakeet-ab`. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed.
- `[2026-10-01]` **irv-ml1 /storetank reclaim done:** Prime ruled through comfy-dev, which deleted 84 files of its own (260 → 299 GB free). Tiers B/C/D got no ruling (infra-hermes thread `01M3TCSYRSFNAPA9BPQTMFQ6KJ`).
- `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. DONE 2026-10-01 0126 by infra-hermes; infra-ops audit passed 0137.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md`
- `[2026-09-30]` **Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30.** → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
- `[2026-09-30]` **SemIf replaced by intern-decision (Intern-Decision-4B, the Jev bench pick): semif-compatible plus Jev `/v1/systemone` at 32k tokens on GPU 1, with a Triton warm-up cache volume.** → `persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md`
@@ -432,49 +415,15 @@ _As of 2026-09-30 ~1800 PT._
- `[2026-09-18]` **Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** — (operator-approved). → `persistent-memory.d/2026-09-18-miranda-s-hermes-plugin-install-is-now-a-symlink-to-the.md`
- `[2026-09-18]` **Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal… → `persistent-memory.d/2026-09-18-worldtree-s-env-sh-secrets-are-vaulted.md`
- `[2026-09-17]` ⭐⭐ **The next voice seat was MEASURED, not chosen by taste — and the corpus size ranking INVERTS the voice ranking at the top.** Our two largest authors are Stephen King (76 works, 12.1M words) and Agatha Christie (72, 5.5M); neither should get a seat, Christie being the Krakauer failure mode exactly (genius in plot architecture, prose deliberately transparent, invisible to a char-bigram Delta). Picks, in order: **Faulkner** (~15 pure novels, ~1.6M words — highest voice signal in the catalogue, AND he is McCarthy's stylistic ancestor, so training him next supplies the **hard-negative sister the gate has lacked since the Brontë record named it missing**); **Morrison** (11 novels after pruning criticism/anthology, ~818k — 11 val units, beating Hemingway's 10); **Chandler** (7 novels + a 409k short-story omnibus — fills the first-person hardboiled gap, Hemingway-class corpus size). ⚠ Faulkner's catalogue rows carry a 446k-word Snopes omnibus that duplicates novels also present individually — the Hemingway 90-96% collection-duplication trap, needs the containment pass first. → `persistent-memory.d/2026-09-17-next-voice-seats.md`
- `[2026-09-17]` ⭐⭐ **Romantasy IS a real register, we already trained its most distinctive member, and the obvious next pick is its worst.** Measured on the gate's own instrument (char-bigram Burrows's Delta, ~120k words/author from mid-work), with within-author floors and cross-genre positive controls. Cluster median pair **0.537 = 1.2x the worst floor** against controls at 1.4-1.9x — tighter than cross-genre but NOT collapsed. Two findings survive either floor reading: **Yarros is the cluster OUTLIER** (4 of the 5 largest pair distances involve her), so a second romantasy seat buys measurably less than the first did; and **Maas is the centroid** (the two smallest distances in the matrix are hers), so the obvious commercial pick is the least distinctive. If the lane gets a seat it is **Kenyon** — furthest from Yarros at 0.674 and **27 works = 27 val units, the best-powered gate the line could build** (Hemingway 10, McCarthy 6, Brontë 4). ⚠ Sensitivity floor stated: one sample per pair, no repeat draws; the rank ordering is indicative, fine gaps are not resolvable. → `persistent-memory.d/2026-09-17-romantasy-register-measured.md`
- `[2026-09-17]` **`dragonfireacoustics.com` expires 2026-10-30 — six weeks — at eNom with NO transfer lock, and its sibling domain was already lost exactly this way.** → `persistent-memory.d/2026-09-17-dragonfireacoustics-com-expires-2026-10-30-six-weeks-at.md`
⚠ **If it is ever transferred, DNS does NOT come with the registration** — the nameservers are eNom's `name-services.com` and the zone must be recreated first or mail dies. The whole zone is two facts plus a landmine: `*` (WILDCARD) → 199.250.192.76 which is **dead** (no HTTP at all, and it is what the apex/mail/webmail/admin/ftp all answer with), `www` → 38.120.12.45 (us), and **7 Google Workspace MX records that must not be lost**. No DNSSEC (`delegationSigned: false`), so no transfer complication. ⚠ Also found: **no SPF and no DMARC** at the apex on a Google Workspace domain — a live deliverability problem independent of everything else. Transfer gate is the **TAC/EPP code from the eNom account**, not the lock; the missing lock is not authorization. 60-day rule is satisfied (last changed 2025-10-24).
- `[2026-09-17]` **`dragonfireacoustics.com` IS configured on `pfi-ana-webhost`, and the whole thing is dead — a forgotten public-facing VM.** → `persistent-memory.d/2026-09-17-dragonfireacoustics-com-is-configured-on-pfi-ana-webhost.md`
- `[2026-09-17]` **headscale now split-DNSes `nh3.phasefinal.com` to the three AdGuards, so mesh clients can resolve the internal-only wildcard** → `persistent-memory.d/2026-09-17-headscale-now-split-dnses-nh3-phasefinal-com-to-the-three.md`
- `[2026-09-17]` **ESH is back on the Cityside static `128.177.138.182/30` and the site is healthy — confirmed on four axes, not one.** UDM WAN1 `wan_type` is `static` again (switched back from the DHCP set during the 09-17 outage), `stat/health` names Cityside Fiber with 0 disconnected and Verizon-5G idle at failover priority 2, esh-docker-vm's egress EQUALS the WAN ip so nothing is behind CGNAT, and colo→ESH reads **5.0 ms / 0% loss** at 2005/2142 Mbps (Cityside CGNAT was 9 ms, Verizon failover 33–37 ms). ⭐ The FortiGate `infra-ops` trusthost3 pin un-broke itself and that was VERIFIED: from ESH, ana-gw tcp/22 is open and offers a password prompt, which a trusthost mismatch would never do. ⚠ The two 7-day crowdsec entries are being left to expire 2026-09-23 on purpose — Cityside failed twice in six hours, so they are cheap insurance. → `persistent-memory.d/2026-09-17-esh-fiber-outages.md`
- `[2026-09-17]` **Operator ruled "leave it" on lv-hemingway's 3 separator-hidden names.** So `leak_gate.py` exits 1 on a SHIPPED tree by design; a future session seeing that red result should read this line, not start fixing. lv-bronte re-ran clean.
- `[2026-09-17]` ⭐⭐⭐ **The leak gate PASSED lv-mccarthy while five protagonist names sat in all six copies, and the blind spot generalises to every corpus in the line.** `\b(Surface)\b` cannot match a name with a character inserted in it, so a mangled occurrence is unrenameable AND unreportable: `B ell`, `C higurh`, `M oss`, `T oadvine` (a small-caps drop cap kept as its own token) and `Toad-vine`, `Glan-ton` (a print line-break hyphen). Every VISIBLE occurrence had been renamed, which is what made it invisible. Same family as lv-bronte's `_Antigua_`, now generalised: **any separator inside a name blinds a word-boundary scan.** Fixed in the corpus builder (rules 4+5, counted), and `leak_gate.py` now runs a separator-tolerant pass with its own controls that FAILS the gate — validated against the pre-fix tree. ⚠ Its fragment filter is load-bearing: a naive scan returns 18 false positives on Hemingway (`God damn`, `I run`) against 3 real. Whole D1→D3 chain reproduced byte-identically before and after. Commit `c559664`. → `persistent-memory.d/2026-09-17-mccarthy-split-name-leak.md`
- `[2026-09-17]` **The SHIPPED lv-bronte adapter emits mid-sentence line breaks at 12.46 per 1k chars, and nothing downstream looks for it.** → `persistent-memory.d/2026-09-17-the-shipped-lv-bronte-adapter-emits-mid-sentence-line.md`
- `[2026-09-17]` **The `mccarthy` register names the punctuation ON PURPOSE, and that is a gate-design call made before any McCarthy number existed.** → `persistent-memory.d/2026-09-17-the-mccarthy-register-names-the-punctuation-on-purpose-and.md`
- `[2026-09-17]` **lv-mccarthy's D1→D3 chain was RECOVERED, not remembered — there was no runbook and the commands went over non-interactive ssh, so no history survived.** → `persistent-memory.d/2026-09-17-lv-mccarthy-s-d1d3-chain-was-recovered-not-remembered-there.md`
- `[2026-09-17]` **Measured and DELIBERATELY not changed, three of them.** — The oversize-passage drop is 13.9% of McCarthy's train words, between Hemingway's 10.0% and the shipped… → `persistent-memory.d/2026-09-17-measured-and-deliberately-not-changed-three-of-them.md`
- `[2026-09-17]` ⭐⭐ **A unit splitter must choose by SIZE, not count — and the val split scales with WORK COUNT, not corpus size.** `scripts/r49-corpus/split_units.py` + a multi-index `--holdout-chapter`. The inherited most-units rule gave Cities of the Plain 4 units of 22,312w (the book's PARTS); the single-index holdout would have given McCarthy a Brontë-class 18k-word val reference on a 588k corpus. Both fixed, both caught by controls. → `persistent-memory.d/2026-09-17-mccarthy-d1-d3.md`
- `[2026-09-17]` ⭐ **lv-mccarthy D1–D3 complete on gx10, leak gate PASSED (0 of 75 renameable, 0 of 37 sub-threshold, both controls green).** Three McCarthy-specific calls, each forced by a measurement: corpus-scoped rename (the Border Trilogy shares 9 surfaces across books), a new `mccarthy` name preset (Hemingway's carries it_IT/fr_FR and McCarthy writes neither), and `--min-cap 5` to match the entity map's floor — the first gate run failed with 45 survivors purely because rename's floor was 8 and the map's was 5. → `persistent-memory.d/2026-09-17-mccarthy-d1-d3.md`
- `[2026-09-17]` **PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the…** → `persistent-memory.d/2026-09-17-parked-lv-krakauer-and-the-reason-is-a-selection-criterion.md`
- `[2026-09-17]` **The althing route-declaring SessionStart hook is documented but NOT installed on nh3-dev** — `dev_launch.py` has zero occurrences of "route", no hook declares one, and every live route was hand-declared… → `persistent-memory.d/2026-09-17-the-althing-route-declaring-sessionstart-hook-is-documented.md`
- `[2026-09-17]` **Hemingway ships as-is: operator ruled "ship stands" on both measured corpus defects** — the 0.96% beat contamination and the 130 non-name entity-map surfaces. → `persistent-memory.d/2026-09-17-hemingway-ships-as-is-operator-ruled-ship-stands-on-both.md`
- `[2026-09-17]` **PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the…** → `persistent-memory.d/2026-09-17-parked-lv-krakauer-and-the-reason-is-a-selection-criterion-2.md`
- `[2026-09-17]` **A unit splitter must choose by SIZE, not by count — the inherited rule silently produced 22,000-word "chapters".** → `persistent-memory.d/2026-09-17-a-unit-splitter-must-choose-by-size-not-by-count-the.md`
- `[2026-09-17]` **lv-mccarthy D1 built — 167 units, 584,756 words — and the whole job was protecting a style that reads as damage.** → `persistent-memory.d/2026-09-17-lv-mccarthy-d1-built-167-units-584-756-words-and-the-whole.md`
- `[2026-09-17]` **lv-krakauer D1 built — 126 units, 422,880 words — and its name guard caught three defects nothing else would have reported.** → `persistent-memory.d/2026-09-17-lv-krakauer-d1-built-126-units-422-880-words-and-its-name.md`
- `[2026-09-17]` **`triage_disposition = 'accepted'` in the Kvasir catalogue does NOT mean the extraction succeeded.** — Blood Meridian's epub row holds 1,167 words of a 117,000-word book, The Crossing's 222 of 150,000 — real… → `persistent-memory.d/2026-09-17-triagedisposition-accepted-in-the-kvasir-catalogue-does-not.md`
- `[2026-09-17]` ⭐⭐⭐ **lv-hemingway SHIPPED (ckpt850) with the line's strongest voice result — and the memorisation control it passed turned out to be the WRONG control.** Voice +0.413 delta_cb at 6.4x the floor, closing 73.8% of the achievable span (lv-bronte closed 48%). ⚠ `memorization_check.py` uses the base-unadapted arm as its negative control, but base writes 18,035 words of summary against the adapted arms' 27,413 of pastiche — **text that does not imitate the register cannot collide with its n-grams**, so a 0.00 there means "different register", not "did not memorise". The right reference is the author himself: **held-out Hemingway against the train split collides at 0.01 while the adapter does at 0.07**, so the comfortable "his plain register makes collisions inevitable" story is FALSE and was refuted rather than assumed. All 19 matched runs were READ: stock dialogue, max **9 words**, no proper noun — shorter than the 10-word run unseen Hemingway shares with the train split by coincidence. ⭐ **A negative control that differs from the candidate in a way correlated with the metric is not a control.** → `persistent-memory.d/2026-09-17-lv-hemingway-gate.md`
- `[2026-09-17]` **The v2 voice floor is now PAIRWISE, and it retroactively passes lv-bronte.** — lv-bronte's ckpt475 shipped as a voice-axis FAILURE at +0.193 against a 0.251 floor contributed entirely by… → `persistent-memory.d/2026-09-17-the-v2-voice-floor-is-now-pairwise-and-it-retroactively.md`
- `[2026-09-17]` **The beat-contamination leak IS present in Hemingway — 70 of 7,094 train beats (0.96%), 0 of 200 val.** → `persistent-memory.d/2026-09-17-the-beat-contamination-leak-is-present-in-hemingway-70-of-7.md`
- `[2026-09-17]` **`audit_entity_map.py` — the rename can DAMAGE the prose and no gate will ever say so.** — Mirror of `audit_stoplist.py`: surfaces wrongly held IN the map rather than out of it. → `persistent-memory.d/2026-09-17-auditentitymap-py-the-rename-can-damage-the-prose-and-no.md`
- `[2026-09-17]` **The two-epoch recipe is now 0 for 2 and should stop being carried forward.** — Hemingway's eval minimum is step 1750, but step 850 is +0.0040 against a 0.0044 median neighbour jitter … → `persistent-memory.d/2026-09-17-the-two-epoch-recipe-is-now-0-for-2-and-should-stop-being.md`
- `[2026-09-17]` **gitea was reaching the PUBLIC route from every repo on nh3-dev** — brokkr-smithy, sleipnir, Galdrabok, kvasir — and brokkr-smithy is pushed several times a week, so the… → `persistent-memory.d/2026-09-17-gitea-was-reaching-the-public-route-from-every-repo-on-nh3.md`
- `[2026-09-17]` **`servers/fv-ml1/ssh-target` was bare `10.251.50.54`, so `deploy-stack.sh` connected as `lkraven` and could not write the infra-ops-owned…** → `persistent-memory.d/2026-09-17-servers-fv-ml1-ssh-target-was-bare-10-251-50-54-so-deploy.md`
- `[2026-09-17]` ⭐⭐ **A leak class the corpus gate structurally CANNOT see: the beat-writing model recognises the book and restores the author's real character names.** 1.8% of Brontë beats named Rochester/Jane/Brocklehurst while 0 responses did. Worst for public-domain classics; Yarros and Hemingway's clean runs are NOT evidence they are immune. Patched as a `sourcename` reject + `--source-entities`. → `persistent-memory.d/2026-09-17-beat-contamination-leak.md`
- `[2026-09-17]` **A stoplist entry is an assertion the leak gate can no longer check** — stoplisting removes a surface from the entity map, so a wrongly stoplisted CHARACTER is an undetectable leak. → `persistent-memory.d/2026-09-17-a-stoplist-entry-is-an-assertion-the-leak-gate-can-no.md`
- `[2026-09-17]` **ESH: Cityside Fiber failed TWICE (19:09 and ~01:06); operator switched WAN1 to DHCP to restore service and has a ticket for the static.** crowdsec `esh` allowlist carries both failover egresses with 7-day expiries — the rotation-fragility is live. → `persistent-memory.d/2026-09-17-esh-fiber-outages.md`
- `[2026-09-17]` ⭐⭐ **lv-bronte SHIPPED on voices-seat (ckpt475) DESPITE failing the v2 VOICE axis — additive, reversible, safety-axis clean.** Both candidates closed 48–52% of the achievable distance to Brontë but +0.193/+0.210 sit under a 0.251 noise floor set by ONE outlier seed in the arm not being shipped; cause is structural (81 val pairs vs Hemingway's 200) and not cheaply fixable. ckpt475 is the pick if it ships. The two-epoch recipe did NOT transfer. → `persistent-memory.d/2026-09-17-lv-bronte-gate.md`
- `[2026-09-16]` ⭐⭐ **Instruction-pair SFT BEATS raw-text for author voice, and the raw-text incumbent never cleared its own control (+0.141 against a 0.153 floor).** `lv-yarros` shipped; the v1 decision rule was amended by the operator after measurement showed it gated on axes the unadapted carrier already maxes. → `persistent-memory.d/2026-09-16-lv-voices-line.md`
- `[2026-09-16]` ⭐ **voices-seat live: one carrier, N `lv-<author>` LoRA adapters, hot-swap measured at 0.24 s.** LoRA costs 24.3% of decode against a 0.1% A-vs-A floor and is worth paying; `--gpu-memory-utilization` is a request against TOTAL VRAM and only a pinned KV makes it predictive. → `persistent-memory.d/2026-09-16-voices-seat-lora.md`
- `[2026-09-16]` ⭐ **lv-hemingway corpus gated at 994,760 words — and half the catalogue had to be EXCLUDED.** 169,759 words of measured 90–96% collection duplication, a Sherwood Anderson parody, and the author's own name 95 times in publisher back matter; the gender resolver needed a corpus base-rate correction to stop reading women as men. → `persistent-memory.d/2026-09-16-lv-hemingway-corpus.md`
- `[2026-09-16]` **Grok token broker built then SHELVED — operator ruled "keep the jail stop the a/b", so the renewal feature has no consumer.** ⛔ Do NOT arm `probe-rotation`: the risk did not shrink (it reaches BOTH Gróa transports through one shared session) and the payoff went to zero. → `persistent-memory.d/2026-09-16-grok-broker-shelved.md`
- `[2026-09-15]` ⚠⚠ **`--gpu-memory-utilization` DOES NOT PREDICT RESIDENT VRAM — measure it, never compute it.** Wrong in **both** directions on fv-ml1: `vllm-cyberprev` util 0.40 (expect ~39,155 MiB) holds **47,124** (+8 GB over); `vllm-gen-small` util 0.48 (expect ~46,986) holds **36,942** (−10 GB under). Planning a placement off the fractions would have been 8 GB wrong. Read `nvidia-smi --query-compute-apps`. Full per-seat residency table + the breeze shuffle arithmetic → `persistent-memory.d/2026-09-15-breeze-placement-sizing.md`
- `[2026-09-15]` **breeze-tts stays on irv-ml1; the TTS-stack move to fv-ml1 is PARKED (park id 75, `move-the-tts-stack-breeze-tts-bragi-tts-gateway`), triggered on evacuating embed/rerank/reward.** ⚠ Trigger as stated says "gpu0" but those three are on **GPU 1** (~0.16 util, ~15.7 GB; GPU 1 is the tight card at 0.975 / 4,336 MiB free) — confirm which he meant before executing. All three services move together because only `breeze-tts` is GPU-resident (~10.3 GiB, **growing**) while `bragi` and `tts-gateway` are CPU proxies, and co-location is what avoids a cross-site hop per TTS call. **breeze-tts sizing — original recommendation NOT to move it.** ~**10.3 GiB** measured under load at 53 min uptime, **up from 9.2 GiB** shortly after warm-up (it grows; n=2, plateau unmeasured) — so GPU 0's 11,982 MiB free is a **1.7 GB margin and shrinking**, on the live chat serving path. ⚠ Two measurement traps: it reports **nothing at idle on the wrong card** (`BREEZE_GPU_DEVICES=0` = the **3090**, not the A6000), and an early reading understates it. ⭐ The real objection is **topology**: `tts-gateway` is on irv-ml1 and reaches it same-box, so moving breeze alone adds a cross-site hop to every TTS call against a 478 ms first-sample budget. GPU 3 would fit it but spends the reserve. → `persistent-memory.d/2026-09-15-breeze-placement-sizing.md`
@@ -487,7 +436,7 @@ _As of 2026-09-30 ~1800 PT._
- `[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`.
_135 older entries archived to archival-memory.md._
_167 older entries archived to archival-memory.md._
## Tried and abandoned