memory: snapshot — lv-bronte shipped with a failed voice axis, next goal is landing lv-hemingway

In-flight rewritten for the next goal. lv-hemingway is TRAINED and nothing else
has been done to it: ship candidate is checkpoint-1750 (ep 1.97, eval 2.2783),
the end-of-run adapter is 0.0763 worse, and the v2 gate has not been run. Every
instrument it needs was parameterised during the lv-bronte run tonight and the
in-flight section names all four with their traps.

New detail files:
  2026-09-17-lv-bronte-gate.md            shipped, voice axis failed, why anyway
  2026-09-17-beat-contamination-leak.md   the leak the corpus gate cannot see
  2026-09-17-esh-fiber-outages.md         two Cityside failures, rotation fragility

Also commits the memorization_check.py parameterisation, which was left
uncommitted: its hardcoded Yarros defaults would have compared a Hemingway arm
against the Yarros corpus and reported a meaningless clean zero.

Auto-archival: index was 415 lines pre-run, over the 300 cap. Only five entries
cleared the 14-day age guard, and three of those carry open deferred pointers
(fused MoE park 47, nconnect=8, AI-tab belayed) and are referenced by in-flight.
A fourth — every CI job on pfi-fleet runs as root on ana-docker — is a live
security property rather than settled history, so it is held back deliberately.
One entry archived. The file stays over cap, which is the guard working: an
over-cap file that keeps live decisions beats a scannable one that lost them.
This commit is contained in:
vh
2026-09-17 01:29:09 -07:00
parent 61840f3131
commit c445ce9e93
5 changed files with 160 additions and 36 deletions
+3
View File
@@ -4,6 +4,9 @@ _Entries moved out of persistent-memory.md to keep the active file scannable. Re
## Recent decisions (archived)
- `[2026-09-02]` **`vastblue` gitea org created (id 8, private, owner `vh`) with empty repo `vastblue/platform`** — third entity namespace alongside `corviduo` and `pfi`; most repos still live under `vh/`. Home of VastBlueDocumentAI + the anchor healthcare-billing SPA (signed 3-yr client contract). **Org scope was the decision**: org-level runner registration and secrets are inherited free by the DocumentAI repo when it splits out, and that is the only binding expensive to retrofit. Deliberately NOT set: org runner (instance-wide `ana-docker-runner` already serves it; org scope is for the DEDICATED runner, deferred to U10) and org secrets (none exist yet; a guessed secret looks bound). ⚠ **Dedicated runner is gated on the first client-premises release cut**, not on the first green pipeline — the risk is another repo's CI sharing a root-level daemon with a build that ships to a healthcare client, see the runner entry above. Push needs no credential: vastblue-dev is on nh3-dev and git-SSH there auths as `vh`. → `stacks/gitea-runner/README.md`
_Archived 2026-09-17._
- `[2026-09-01]` **Ops boundary ruled by the operator: worldtree-dev writes the bridge code; infra-ops OPERATES the Worldtree/Matrix instances and may change them.** Corrects a mis-route where infra-ops asked worldtree-dev to provision an account on a box it does not run. Tracked at `931bac8` + althing `01M1F4PK796EDGDCBKZ9W3JC0S`.
_Archived 2026-09-15._
@@ -0,0 +1,39 @@
# `[2026-09-17]` The beat-writing model restores the author's real names — a leak the corpus gate structurally cannot see
⭐⭐ **THE RENAME AND ITS GATE ARE BOTH WORKING, AND THE TRAINING DATA IS STILL CONTAMINATED.**
The corpus gate reads the corpus and the renamed copies. **It never reads the generated
instruction beats.** Those are written by an LLM that just read the passage — and if it
recognises the book, it supplies the canonical names out of its own training.
**Measured on the first 714 lv-bronte pairs, before the filter existed:**
- **13 beats (1.8%)** named source characters — `Rochester` ×6, `Jane` ×3, `Brocklehurst` ×2,
`Beck`, `Fairfax`, `Helen`, `Burns`, `Eyre`, `Reed`, `Rivers`.
- **0 of 714 RESPONSES did.** The rename was perfect; the instruction side was not.
- One beat read *"Saoirse confirms Rochester's flaws, then agrees in English to marry him"* —
a renamed name and a canonical one in the same sentence, which is the mechanism in miniature.
**Why it matters more than 1.8% sounds:** the beat is the INSTRUCTION half of the pair, so
training on it re-teaches exactly the inventions the rename pipeline exists to remove.
⚠⚠ **EXPOSURE SCALES WITH HOW WELL THE GENERATOR KNOWS THE BOOK.** It is worst for
public-domain classics and mildest for recent work. That is precisely why the Yarros and
Hemingway runs came up clean and Brontë did not — **their clean runs are NOT evidence they are
immune.** Both should be re-verified, and regenerated with `--source-entities`, before their
pairs are trusted again.
**The fix.** `vet()` in `scripts/yarros-corpus/build_sft_pairs.py` gained a `sourcename`
reject plus `--source-entities <entities.json>`, taking the UNRENAMED entity map. Fired at
~3% of attempts on the Brontë rebuild. Commit `533cc0c`.
**The end-to-end guard that proves it.** The chain now verifies every built pair — beat,
response and context — against every source surface before spending GPU hours:
`[verify] 3858 pairs vs 368 source surfaces -> 0 leaks`.
⚠ A guard stricter than the gate cries wolf. The first verify pass excluded nothing and
flagged `Monsieur’` ×14 and `Qu’est-ce` ×4 — French grammar, not leaks — because
`--fold-clitics` leaves apostrophe keys the gate deliberately skips. Mirror `leak_gate.py`'s
own predicate; a guard that fails on false positives gets disabled, which is worse than the
leak it guarded.
Related: [[2026-09-17-lv-bronte-gate]], [[2026-09-16-lv-hemingway-corpus]], [[2026-09-16-lv-voices-line]].
@@ -0,0 +1,41 @@
# `[2026-09-17]` ESH: Cityside Fiber failed twice in six hours; site ran on Verizon failover
**Timeline (PDT).**
```
19:09:07 Cityside dies. UDM fails over to Verizon 5G (WAN2). Site stays up at ~33 ms.
19:51 Verified healthy on failover.
20:01:37 esh-scale drops off the headscale mesh; 10.0.0.0/16 withdrawn; whole site dark
from the colo. Beszel fires on all five ESH hosts.
20:11-15 Service restored. Operator had switched WAN1 to DHCP to get Cityside working at
all; it came back on CGNAT, not the static. Latency back to 9 ms.
01:06:23 Cityside fails AGAIN. Failover to Verizon. Site up, ~37 ms.
```
⭐ **The 20:01 blackout was most likely the operator's own WAN reconfiguration**, not ISP
instability — switching the WAN type bounces the interface, esh-scale loses its path,
headscale withdraws the route, and the site vanishes from the colo's view until it settles.
An earlier session theory ("Cityside came back half-provisioned and the UDM failed back into
an unstable session") is retired.
⚠ **The diagnostic that mattered: physical link stayed UP at 2.5 GE with zero errors
throughout, while the ISP's next-hop `128.177.138.181` was unresponsive.** So "the ONT is
fine, it is upstream of the ONT" — the line to give Cityside. Traceroute from NH3 reached
`209.249.146.170` (one hop short) before dying, so the prefix was still routed.
⚠ **CROWDSEC ROTATION FRAGILITY IS LIVE.** The `esh` allowlist on ana-docker carries the now-
dark static `128.177.138.182` (never-expiry), plus `97.190.18.88` (Verizon failover) and
`23.164.40.174` (Cityside CGNAT), both **7-day expiry**. ESH is on a rotating carrier-NAT
egress until the static is restored — the exact regime the 09-08 static purchase was meant to
end, and the class that once blackholed the whole site via a false ban. **If ESH loses colo
access, check `curl -s4 ifconfig.me` from esh-docker-vm FIRST** and allowlist the new address.
**Still pinned to the dark static and broken until it returns:** FortiGate `infra-ops`
trusthost3 = `128.177.138.182`, so logins to ana-gw from ESH are refused. The dormant
`esh-ana` IPsec is bound to wan1/static (disabled, so no impact).
⭐ **The mesh was NOT degraded on CGNAT** — tailscale hole-punched straight through
(`direct 23.164.40.174:41641`), which is why latency read 9 ms rather than a DERP figure. An
expectation of relay-on-CGNAT was wrong.
Related: [[2026-09-06-headscale-cutover]], [[2026-09-08-esh-static-wan-followups-and-ytvc]].
+58 -30
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-09-16 ~16:15 PT (lv-yarros SHIPPED — instruction-pair SFT beats raw-text on voice; voices-seat live on fv-ml1 GPU 0 :8027 with measured 24.3% LoRA cost; lv-hemingway corpus gated and training, ~1h out; Grok token broker built then SHELVED by the keep-the-jail ruling.)_
_Last updated: 2026-09-17 ~01:30 PT (lv-bronte SHIPPED with a FAILED voice axis on the record; a beat-contamination leak class the corpus gate cannot see, found and patched; audit_stoplist added; next goal is landing lv-hemingway, which is trained but un-gated.)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
@@ -115,43 +115,74 @@ no longer deployed sidecars here. See Recent decisions.)
## Current state / in-flight
_As of 2026-09-16 ~16:15 PT._
_As of 2026-09-17 ~01:30 PT._
### One job in flight: `lv-hemingway` training on pfi-gx10, ~1 hour out.
### Next goal: land `lv-hemingway`. It is TRAINED; nothing else has been done to it.
**IN FLIGHT — `lv-hemingway` pair-SFT.** `gx10:~/r49-runs/hemingway-4b-pairs-3ep`, 7,094 pairs,
2,661 steps at ~3.68 s/it, launched ~15:05 PT. **Take the checkpoint at the LOSS MINIMUM, not the
end-of-run adapter** — the recipe is two epochs on a three-epoch schedule. When it lands: run the
v2 gate (voice `delta_cb` vs base control beyond the noise floor · 8-gram overlap near the
never-saw-it control · overshoot) using `scripts/yarros-corpus/{score_beats,memorization_check}.py`
and the 30-beat in-genre fixture, then ship to
`/tank/aimodels/voice-adapters/lv-hemingway-4b-v1/` and add it to `stacks/voices-seat/compose.yaml`.
**Ship candidate is `gx10:~/r49-runs/hemingway-4b-pairs-3ep/checkpoints/checkpoint-1750`**
(epoch 1.97, eval_loss 2.2783 — the loss minimum). The run finished 2026-09-16 18:27, 2,661
steps in 3:17:34, clean. The end-of-run `adapter/` is **0.0763 worse** (2.3546) — epoch 3
overfits and plateaus. Do NOT ship `adapter/`; the run's own log says so.
**SHIPPED — `lv-yarros`.** `gx10:~/adapters/lv-yarros-4b-v1/` (sha `63fda6cc61f380b2`) and live on
`vllm-voices`, fv-ml1 GPU 0 :8027, alongside `voices-base`.
**The v2 gate has NOT been run on it.** Everything it needs now exists and is parameterised,
built during the lv-bronte run tonight — reuse it rather than rebuilding:
**OPEN, operator's call, nothing blocked:**
- **Brontë has never been through the automated leak gate** — its "0 of 203" was a HAND COUNT and
the gate did not exist yet. On Yarros the same instrument read **212 surviving where a hand count
said 86**. Its gated corpus survives at `gx10:~/r49-corpus-renamed-unwrapped/` and is
pair-buildable; `lv-bronte` exists only as RAW-TEXT arms, i.e. the arm that never cleared its own
control. Re-gate before building pairs, or accept the hand count.
- **xAI billing** — did the stopped HTTP A/B bill on top of the coding plan? Needs the operator's
console access; this fleet holds no xAI credential.
- **A `BabyYarros → lv-yarros` pointer** in memory, so historical entries stay findable under the
new name. Historical entries were deliberately left as dated records.
1. `scripts/r49-corpus/build_beat_fixture.py --pairs <hemingway val pairs> --out beats-hemingway-30.json --sidecar ... -n 30` — refuses any split but `val`.
2. `scripts/r49-corpus/gen_beats_chat_yarros.py --system-from <pairs>.provenance.json` — **the `--system-from` flag is mandatory**; the harness's hardcoded SYS is Yarros's and driving a Hemingway arm with it confounds the carrier change with a prompt change.
3. `scripts/yarros-corpus/memorization_check.py --eval-dir ... --corpus ... --glob 'beats5.*.jsonl'` — **now takes paths**; its old hardcoded Yarros defaults would have compared a Hemingway arm against the Yarros corpus and reported a meaningless clean zero.
4. `scripts/r49-corpus/voice_distance.py <renamed-corpus> <eval-dir>` — needs `voice.<arm>.jsonl` files with a `continuation` field, and identifies the control by the substring **`unadapted`** in the arm name. Adapt with a script like `gx10:~/lv-bronte/voice-prep.py`.
⚠ **Hemingway's pairs were built BEFORE the `--source-entities` beat filter existed.** Measured
on Brontë: 1.8% of generated beats named the author's real characters, because the beat-writing
model recognises the book and restores canonical names — a leak the corpus gate structurally
cannot see. Exposure scales with how famous the book is. **Re-verify the Hemingway pairs against
its entity map before trusting them**, and regenerate with `--source-entities` if they leak.
⚠ **Do not assume the two-epoch recipe.** It held for Yarros and Hemingway and did NOT hold for
Brontë. Read the eval curve; Hemingway's minimum genuinely is at 1750, which is already known.
### SHIPPED tonight — `lv-bronte`, with a FAILED voice axis on the record
Live on `vllm-voices` (fv-ml1 GPU 0 :8027) as `lv-bronte` beside `voices-base` and `lv-yarros`.
Checkpoint-475. **It did not pass its voice gate** (+0.193 delta_cb against a 0.251 measured
floor); shipped because it is additive, reversible, and clean on the safety axis (8-gram overlap
0.00, identical to the never-saw-it control) on a public-domain corpus. The caveat is written
into the commit, the compose file, and a README beside the adapter on NFS.
→ `persistent-memory.d/2026-09-17-lv-bronte-gate.md`
### ESH is on Verizon failover — Cityside Fiber failed TWICE tonight
19:09–20:15 and again from ~01:06. The operator switched WAN1 to DHCP to get service back at all
and has a ticket in to restore the static `128.177.138.182/30`. ESH egress is currently
`97.190.18.88`; crowdsec's `esh` allowlist carries it plus `23.164.40.174`, both with 7-day
expiries set 2026-09-16 ~19:31 and ~01:xx. ⚠ **When those expire, or when the CGNAT egress
rotates, ESH loses colo access** — that is the false-ban class that blackholed the site before.
Check `curl -s4 ifconfig.me` from esh-docker-vm first if ESH goes dark.
### OPEN, operator's call, nothing blocked
- **Henge id 79** (`catalog-all-robotics-bits-and-bobs`) was filed under source
`eshpfi-management` by the skill's repo inference; it is a personal idea. Re-park with
`--personal` if wanted.
- **xAI billing** — did the stopped HTTP A/B bill on top of the coding plan? Needs the
operator's console; this fleet holds no xAI credential.
- **A `BabyYarros → lv-yarros` pointer** in memory so historical entries stay findable.
- Older deferred set, unchanged: AI-tab Dormant regrouping (**belayed**), `nconnect=8` on
`/mnt/smithy` (**declined in scope**), fused MoE kernel path (**parked, id 47**), TTS-stack move
to fv-ml1 (**parked, id 75**).
`/mnt/smithy` (**declined in scope**), fused MoE kernel path (**parked, id 47**), TTS-stack
move to fv-ml1 (**parked, id 75**).
⚠ **fv-ml1 GPU 0 is at 96.0 of 97.9 GB.** GPU 1 ~5.7 free, GPU 2 ~2.4, GPU 3 a HELD RESERVE. The
next seat needing room on fv-ml1 requires a placement decision.
⚠ **fv-ml1 GPU 0 is at 96.0 of 97.9 GB.** Adding `lv-bronte` cost nothing measurable (96092 →
96090 MiB) because a LoRA rides inside the existing seat — but a new SEAT still needs a
placement decision. GPU 3 is a HELD RESERVE.
**Uncommitted:** `graphify-out/GRAPH_REPORT.md` and `scripts/seat-inventory.py` were modified
before this session began. ⚠ **Do NOT commit them** — untouched and deliberately left alone.
## Recent decisions
- `[2026-09-17]` ⭐⭐ **A leak class the corpus gate structurally CANNOT see: the beat-writing model recognises the book and restores the author's real character names.** 1.8% of Brontë beats named Rochester/Jane/Brocklehurst while 0 responses did. Worst for public-domain classics; Yarros and Hemingway's clean runs are NOT evidence they are immune. Patched as a `sourcename` reject + `--source-entities`. → `persistent-memory.d/2026-09-17-beat-contamination-leak.md`
- `[2026-09-17]` ⭐ **A stoplist entry is an assertion the leak gate can no longer check** — stoplisting removes a surface from the entity map, so a wrongly stoplisted CHARACTER is an undetectable leak. Three were wrong on Brontë (Leaven, Pierrot, Samuel); `scripts/r49-corpus/audit_stoplist.py` finds them by honorific and now gates the pipeline. Commit `8bb7686`.
- `[2026-09-17]` **ESH: Cityside Fiber failed TWICE (19:09 and ~01:06); operator switched WAN1 to DHCP to restore service and has a ticket for the static.** crowdsec `esh` allowlist carries both failover egresses with 7-day expiries — the rotation-fragility is live. → `persistent-memory.d/2026-09-17-esh-fiber-outages.md`
- `[2026-09-17]` ⭐⭐ **lv-bronte SHIPPED on voices-seat (ckpt475) DESPITE failing the v2 VOICE axis — additive, reversible, safety-axis clean.** Both candidates closed 48–52% of the achievable distance to Brontë but +0.193/+0.210 sit under a 0.251 noise floor set by ONE outlier seed in the arm not being shipped; cause is structural (81 val pairs vs Hemingway's 200) and not cheaply fixable. ckpt475 is the pick if it ships. The two-epoch recipe did NOT transfer. → `persistent-memory.d/2026-09-17-lv-bronte-gate.md`
- `[2026-09-16]` ⭐⭐ **Instruction-pair SFT BEATS raw-text for author voice, and the raw-text incumbent never cleared its own control (+0.141 against a 0.153 floor).** `lv-yarros` shipped; the v1 decision rule was amended by the operator after measurement showed it gated on axes the unadapted carrier already maxes. → `persistent-memory.d/2026-09-16-lv-voices-line.md`
- `[2026-09-16]` ⭐ **voices-seat live: one carrier, N `lv-<author>` LoRA adapters, hot-swap measured at 0.24 s.** LoRA costs 24.3% of decode against a 0.1% A-vs-A floor and is worth paying; `--gpu-memory-utilization` is a request against TOTAL VRAM and only a pinned KV makes it predictive. → `persistent-memory.d/2026-09-16-voices-seat-lora.md`
@@ -382,9 +413,6 @@ before this session began. ⚠ **Do NOT commit them** — untouched and delibera
- `[2026-09-03]` **nh3-dev wedged for ~40 min and it was the BACKUP, not the disk — a stalled cross-site vzdump holding ever** → `persistent-memory.d/2026-09-03-nh3-dev-wedged-for-40-min.md`
- `[2026-09-02]` **`vastblue` gitea org created (id 8, private, owner `vh`) with empty repo `vastblue/platform`** — third entity namespace alongside `corviduo` and `pfi`; most repos still live under `vh/`. Home of VastBlueDocumentAI + the anchor healthcare-billing SPA (signed 3-yr client contract). **Org scope was the decision**: org-level runner registration and secrets are inherited free by the DocumentAI repo when it splits out, and that is the only binding expensive to retrofit. Deliberately NOT set: org runner (instance-wide `ana-docker-runner` already serves it; org scope is for the DEDICATED runner, deferred to U10) and org secrets (none exist yet; a guessed secret looks bound). ⚠ **Dedicated runner is gated on the first client-premises release cut**, not on the first green pipeline — the risk is another repo's CI sharing a root-level daemon with a build that ships to a healthcare client, see the runner entry above. Push needs no credential: vastblue-dev is on nh3-dev and git-SSH there auths as `vh`. → `stacks/gitea-runner/README.md`
- `[2026-09-02]` **Every CI job on the shared `pfi-fleet` runner is root on ana-docker — and `container.valid_volumes: []` does NOT prevent it.** Measured: a job container is uid 0, `/var/run/docker.sock` is mounted by act_runner independently of that list, `docker ps` returns all 49 host containers (gitea itself, synapse, phasefinal-web, adguardhome), `docker compose v2.33.0` on PATH. ⚠ **LOAD-BEARING** — `vh/Worldtree`, `vh/soong-lab`, `vh/skaldsong`, `vh/wt-matrix-bridge` all drive buildx through that socket, so it cannot simply be closed; **isolate sensitive builds onto a dedicated runner instead.** Also measured the same night: `services:` containers work (Postgres 16), and **full-URL `uses: https://gitea.phasefinal.com/actions/checkout@v4` resolves from the local mirrors** — the un-parked half of the github-independence work, needing neither `DEFAULT_ACTIONS_URL=self` nor the act_runner auth path that blocked it on 2026-08-05. Prompted by vastblue-dev's CI-posture question for a client-funded healthcare repo. → `stacks/gitea-runner/README.md`
@@ -398,7 +426,7 @@ before this session began. ⚠ **Do NOT commit them** — untouched and delibera
- `[2026-08-19]` **AI-tab Dormant regrouping BELAYED by the operator** — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong, Chatterbox Fast) show amber EXITED inside live groups rather than `AI - Dormant`. Fix is a label change + recreate per stack; needs the operator's read on which are retired vs temporarily down. `untracked by operator choice` (his words: "belay the ai dormant regrouping for now").
_9 older entries archived to archival-memory.md._
_10 older entries archived to archival-memory.md._
_5 older entries archived to archival-memory.md._
+19 -6
View File
@@ -10,12 +10,25 @@ Instrument: longest and mean maximal verbatim n-gram shared with the TRAINING co
generation. Controls run every time -- the base-unadapted arm never saw the corpus so it is
the negative control, and a slice of the corpus scored against itself is the positive.
"""
import json, pathlib, re, sys
import argparse, json, pathlib, re, sys
from collections import Counter
EVAL = pathlib.Path("/home/infra-ops/r49-runs/yarros-eval")
CORP = pathlib.Path("/home/infra-ops/yarros-corpus-renamed/copies")
N = 8
# ⚠ These were hardcoded to Yarros. Pointed at a Brontë arm they would have compared
# it against the YARROS corpus and reported a clean zero — a negative that means
# "different book", not "did not memorise". Defaults are unchanged so every Yarros
# number already recorded stays reproducible byte for byte.
_ap = argparse.ArgumentParser()
_ap.add_argument("--eval-dir", default="/home/infra-ops/r49-runs/yarros-eval")
_ap.add_argument("--corpus", default="/home/infra-ops/yarros-corpus-renamed/copies",
help="the RENAMED copies the adapter actually trained on")
_ap.add_argument("--glob", default="beats5.*.jsonl", help="arm files inside --eval-dir")
_ap.add_argument("--strip", default="beats5.", help="prefix trimmed to name the arm")
_ap.add_argument("-n", type=int, default=8, help="n-gram length")
_a = _ap.parse_args()
EVAL = pathlib.Path(_a.eval_dir)
CORP = pathlib.Path(_a.corpus)
N = _a.n
def norm(t): return re.findall(r"[a-z']+", t.lower())
@@ -44,8 +57,8 @@ def longest_match(words):
print(f"{'arm':<22} {'gens':>5} {'hit-rate':>9} {'mean-longest':>13} {'max':>5}")
print("-" * 60)
for f in sorted(EVAL.glob("beats5.*.jsonl")):
arm = f.stem.replace("beats5.", "")
for f in sorted(EVAL.glob(_a.glob)):
arm = f.stem.replace(_a.strip, "")
rows = [json.loads(l) for l in f.read_text(encoding="utf-8").splitlines()]
longs = [longest_match(norm(r["raw"])) for r in rows]
hits = sum(1 for x in longs if x >= N)