diff --git a/archival-memory.md b/archival-memory.md index 8a8268e..238e906 100644 --- a/archival-memory.md +++ b/archival-memory.md @@ -4,6 +4,9 @@ _Entries moved out of persistent-memory.md to keep the active file scannable. Re ## Recent decisions (archived) +- `[2026-09-02]` **`vastblue` gitea org created (id 8, private, owner `vh`) with empty repo `vastblue/platform`** — third entity namespace alongside `corviduo` and `pfi`; most repos still live under `vh/`. Home of VastBlueDocumentAI + the anchor healthcare-billing SPA (signed 3-yr client contract). **Org scope was the decision**: org-level runner registration and secrets are inherited free by the DocumentAI repo when it splits out, and that is the only binding expensive to retrofit. Deliberately NOT set: org runner (instance-wide `ana-docker-runner` already serves it; org scope is for the DEDICATED runner, deferred to U10) and org secrets (none exist yet; a guessed secret looks bound). ⚠ **Dedicated runner is gated on the first client-premises release cut**, not on the first green pipeline — the risk is another repo's CI sharing a root-level daemon with a build that ships to a healthcare client, see the runner entry above. Push needs no credential: vastblue-dev is on nh3-dev and git-SSH there auths as `vh`. → `stacks/gitea-runner/README.md` + _Archived 2026-09-17._ + - `[2026-09-01]` **Ops boundary ruled by the operator: worldtree-dev writes the bridge code; infra-ops OPERATES the Worldtree/Matrix instances and may change them.** Corrects a mis-route where infra-ops asked worldtree-dev to provision an account on a box it does not run. Tracked at `931bac8` + althing `01M1F4PK796EDGDCBKZ9W3JC0S`. _Archived 2026-09-15._ diff --git a/persistent-memory.d/2026-09-17-beat-contamination-leak.md b/persistent-memory.d/2026-09-17-beat-contamination-leak.md new file mode 100644 index 0000000..1e95342 --- /dev/null +++ b/persistent-memory.d/2026-09-17-beat-contamination-leak.md @@ -0,0 +1,39 @@ +# `[2026-09-17]` The beat-writing model restores the author's real names — a leak the corpus gate structurally cannot see + +⭐⭐ **THE RENAME AND ITS GATE ARE BOTH WORKING, AND THE TRAINING DATA IS STILL CONTAMINATED.** +The corpus gate reads the corpus and the renamed copies. **It never reads the generated +instruction beats.** Those are written by an LLM that just read the passage — and if it +recognises the book, it supplies the canonical names out of its own training. + +**Measured on the first 714 lv-bronte pairs, before the filter existed:** + +- **13 beats (1.8%)** named source characters — `Rochester` ×6, `Jane` ×3, `Brocklehurst` ×2, + `Beck`, `Fairfax`, `Helen`, `Burns`, `Eyre`, `Reed`, `Rivers`. +- **0 of 714 RESPONSES did.** The rename was perfect; the instruction side was not. +- One beat read *"Saoirse confirms Rochester's flaws, then agrees in English to marry him"* — + a renamed name and a canonical one in the same sentence, which is the mechanism in miniature. + +**Why it matters more than 1.8% sounds:** the beat is the INSTRUCTION half of the pair, so +training on it re-teaches exactly the inventions the rename pipeline exists to remove. + +⚠⚠ **EXPOSURE SCALES WITH HOW WELL THE GENERATOR KNOWS THE BOOK.** It is worst for +public-domain classics and mildest for recent work. That is precisely why the Yarros and +Hemingway runs came up clean and Brontë did not — **their clean runs are NOT evidence they are +immune.** Both should be re-verified, and regenerated with `--source-entities`, before their +pairs are trusted again. + +**The fix.** `vet()` in `scripts/yarros-corpus/build_sft_pairs.py` gained a `sourcename` +reject plus `--source-entities `, taking the UNRENAMED entity map. Fired at +~3% of attempts on the Brontë rebuild. Commit `533cc0c`. + +**The end-to-end guard that proves it.** The chain now verifies every built pair — beat, +response and context — against every source surface before spending GPU hours: +`[verify] 3858 pairs vs 368 source surfaces -> 0 leaks`. + +⚠ A guard stricter than the gate cries wolf. The first verify pass excluded nothing and +flagged `Monsieur’` ×14 and `Qu’est-ce` ×4 — French grammar, not leaks — because +`--fold-clitics` leaves apostrophe keys the gate deliberately skips. Mirror `leak_gate.py`'s +own predicate; a guard that fails on false positives gets disabled, which is worse than the +leak it guarded. + +Related: [[2026-09-17-lv-bronte-gate]], [[2026-09-16-lv-hemingway-corpus]], [[2026-09-16-lv-voices-line]]. diff --git a/persistent-memory.d/2026-09-17-esh-fiber-outages.md b/persistent-memory.d/2026-09-17-esh-fiber-outages.md new file mode 100644 index 0000000..c76add8 --- /dev/null +++ b/persistent-memory.d/2026-09-17-esh-fiber-outages.md @@ -0,0 +1,41 @@ +# `[2026-09-17]` ESH: Cityside Fiber failed twice in six hours; site ran on Verizon failover + +**Timeline (PDT).** + +``` +19:09:07 Cityside dies. UDM fails over to Verizon 5G (WAN2). Site stays up at ~33 ms. +19:51 Verified healthy on failover. +20:01:37 esh-scale drops off the headscale mesh; 10.0.0.0/16 withdrawn; whole site dark + from the colo. Beszel fires on all five ESH hosts. +20:11-15 Service restored. Operator had switched WAN1 to DHCP to get Cityside working at + all; it came back on CGNAT, not the static. Latency back to 9 ms. +01:06:23 Cityside fails AGAIN. Failover to Verizon. Site up, ~37 ms. +``` + +⭐ **The 20:01 blackout was most likely the operator's own WAN reconfiguration**, not ISP +instability — switching the WAN type bounces the interface, esh-scale loses its path, +headscale withdraws the route, and the site vanishes from the colo's view until it settles. +An earlier session theory ("Cityside came back half-provisioned and the UDM failed back into +an unstable session") is retired. + +⚠ **The diagnostic that mattered: physical link stayed UP at 2.5 GE with zero errors +throughout, while the ISP's next-hop `128.177.138.181` was unresponsive.** So "the ONT is +fine, it is upstream of the ONT" — the line to give Cityside. Traceroute from NH3 reached +`209.249.146.170` (one hop short) before dying, so the prefix was still routed. + +⚠ **CROWDSEC ROTATION FRAGILITY IS LIVE.** The `esh` allowlist on ana-docker carries the now- +dark static `128.177.138.182` (never-expiry), plus `97.190.18.88` (Verizon failover) and +`23.164.40.174` (Cityside CGNAT), both **7-day expiry**. ESH is on a rotating carrier-NAT +egress until the static is restored — the exact regime the 09-08 static purchase was meant to +end, and the class that once blackholed the whole site via a false ban. **If ESH loses colo +access, check `curl -s4 ifconfig.me` from esh-docker-vm FIRST** and allowlist the new address. + +**Still pinned to the dark static and broken until it returns:** FortiGate `infra-ops` +trusthost3 = `128.177.138.182`, so logins to ana-gw from ESH are refused. The dormant +`esh-ana` IPsec is bound to wan1/static (disabled, so no impact). + +⭐ **The mesh was NOT degraded on CGNAT** — tailscale hole-punched straight through +(`direct 23.164.40.174:41641`), which is why latency read 9 ms rather than a DERP figure. An +expectation of relay-on-CGNAT was wrong. + +Related: [[2026-09-06-headscale-cutover]], [[2026-09-08-esh-static-wan-followups-and-ytvc]]. diff --git a/persistent-memory.md b/persistent-memory.md index 30355e7..3c09cb4 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -1,6 +1,6 @@ # Persistent memory — eshpfi-management -_Last updated: 2026-09-16 ~16:15 PT (lv-yarros SHIPPED — instruction-pair SFT beats raw-text on voice; voices-seat live on fv-ml1 GPU 0 :8027 with measured 24.3% LoRA cost; lv-hemingway corpus gated and training, ~1h out; Grok token broker built then SHELVED by the keep-the-jail ruling.)_ +_Last updated: 2026-09-17 ~01:30 PT (lv-bronte SHIPPED with a FAILED voice axis on the record; a beat-contamination leak class the corpus gate cannot see, found and patched; audit_stoplist added; next goal is landing lv-hemingway, which is trained but un-gated.)_ > **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its > `Written:` stamp is under **8 hours** old, read it (it carries the in-flight @@ -115,43 +115,74 @@ no longer deployed sidecars here. See Recent decisions.) ## Current state / in-flight -_As of 2026-09-16 ~16:15 PT._ +_As of 2026-09-17 ~01:30 PT._ -### One job in flight: `lv-hemingway` training on pfi-gx10, ~1 hour out. +### Next goal: land `lv-hemingway`. It is TRAINED; nothing else has been done to it. -**IN FLIGHT — `lv-hemingway` pair-SFT.** `gx10:~/r49-runs/hemingway-4b-pairs-3ep`, 7,094 pairs, -2,661 steps at ~3.68 s/it, launched ~15:05 PT. **Take the checkpoint at the LOSS MINIMUM, not the -end-of-run adapter** — the recipe is two epochs on a three-epoch schedule. When it lands: run the -v2 gate (voice `delta_cb` vs base control beyond the noise floor · 8-gram overlap near the -never-saw-it control · overshoot) using `scripts/yarros-corpus/{score_beats,memorization_check}.py` -and the 30-beat in-genre fixture, then ship to -`/tank/aimodels/voice-adapters/lv-hemingway-4b-v1/` and add it to `stacks/voices-seat/compose.yaml`. +**Ship candidate is `gx10:~/r49-runs/hemingway-4b-pairs-3ep/checkpoints/checkpoint-1750`** +(epoch 1.97, eval_loss 2.2783 — the loss minimum). The run finished 2026-09-16 18:27, 2,661 +steps in 3:17:34, clean. The end-of-run `adapter/` is **0.0763 worse** (2.3546) — epoch 3 +overfits and plateaus. Do NOT ship `adapter/`; the run's own log says so. -**SHIPPED — `lv-yarros`.** `gx10:~/adapters/lv-yarros-4b-v1/` (sha `63fda6cc61f380b2`) and live on -`vllm-voices`, fv-ml1 GPU 0 :8027, alongside `voices-base`. +**The v2 gate has NOT been run on it.** Everything it needs now exists and is parameterised, +built during the lv-bronte run tonight — reuse it rather than rebuilding: -**OPEN, operator's call, nothing blocked:** -- **Brontë has never been through the automated leak gate** — its "0 of 203" was a HAND COUNT and - the gate did not exist yet. On Yarros the same instrument read **212 surviving where a hand count - said 86**. Its gated corpus survives at `gx10:~/r49-corpus-renamed-unwrapped/` and is - pair-buildable; `lv-bronte` exists only as RAW-TEXT arms, i.e. the arm that never cleared its own - control. Re-gate before building pairs, or accept the hand count. -- **xAI billing** — did the stopped HTTP A/B bill on top of the coding plan? Needs the operator's - console access; this fleet holds no xAI credential. -- **A `BabyYarros → lv-yarros` pointer** in memory, so historical entries stay findable under the - new name. Historical entries were deliberately left as dated records. +1. `scripts/r49-corpus/build_beat_fixture.py --pairs --out beats-hemingway-30.json --sidecar ... -n 30` — refuses any split but `val`. +2. `scripts/r49-corpus/gen_beats_chat_yarros.py --system-from .provenance.json` — **the `--system-from` flag is mandatory**; the harness's hardcoded SYS is Yarros's and driving a Hemingway arm with it confounds the carrier change with a prompt change. +3. `scripts/yarros-corpus/memorization_check.py --eval-dir ... --corpus ... --glob 'beats5.*.jsonl'` — **now takes paths**; its old hardcoded Yarros defaults would have compared a Hemingway arm against the Yarros corpus and reported a meaningless clean zero. +4. `scripts/r49-corpus/voice_distance.py ` — needs `voice..jsonl` files with a `continuation` field, and identifies the control by the substring **`unadapted`** in the arm name. Adapt with a script like `gx10:~/lv-bronte/voice-prep.py`. + +⚠ **Hemingway's pairs were built BEFORE the `--source-entities` beat filter existed.** Measured +on Brontë: 1.8% of generated beats named the author's real characters, because the beat-writing +model recognises the book and restores canonical names — a leak the corpus gate structurally +cannot see. Exposure scales with how famous the book is. **Re-verify the Hemingway pairs against +its entity map before trusting them**, and regenerate with `--source-entities` if they leak. + +⚠ **Do not assume the two-epoch recipe.** It held for Yarros and Hemingway and did NOT hold for +Brontë. Read the eval curve; Hemingway's minimum genuinely is at 1750, which is already known. + +### SHIPPED tonight — `lv-bronte`, with a FAILED voice axis on the record + +Live on `vllm-voices` (fv-ml1 GPU 0 :8027) as `lv-bronte` beside `voices-base` and `lv-yarros`. +Checkpoint-475. **It did not pass its voice gate** (+0.193 delta_cb against a 0.251 measured +floor); shipped because it is additive, reversible, and clean on the safety axis (8-gram overlap +0.00, identical to the never-saw-it control) on a public-domain corpus. The caveat is written +into the commit, the compose file, and a README beside the adapter on NFS. +→ `persistent-memory.d/2026-09-17-lv-bronte-gate.md` + +### ESH is on Verizon failover — Cityside Fiber failed TWICE tonight + +19:09–20:15 and again from ~01:06. The operator switched WAN1 to DHCP to get service back at all +and has a ticket in to restore the static `128.177.138.182/30`. ESH egress is currently +`97.190.18.88`; crowdsec's `esh` allowlist carries it plus `23.164.40.174`, both with 7-day +expiries set 2026-09-16 ~19:31 and ~01:xx. ⚠ **When those expire, or when the CGNAT egress +rotates, ESH loses colo access** — that is the false-ban class that blackholed the site before. +Check `curl -s4 ifconfig.me` from esh-docker-vm first if ESH goes dark. + +### OPEN, operator's call, nothing blocked + +- **Henge id 79** (`catalog-all-robotics-bits-and-bobs`) was filed under source + `eshpfi-management` by the skill's repo inference; it is a personal idea. Re-park with + `--personal` if wanted. +- **xAI billing** — did the stopped HTTP A/B bill on top of the coding plan? Needs the + operator's console; this fleet holds no xAI credential. +- **A `BabyYarros → lv-yarros` pointer** in memory so historical entries stay findable. - Older deferred set, unchanged: AI-tab Dormant regrouping (**belayed**), `nconnect=8` on - `/mnt/smithy` (**declined in scope**), fused MoE kernel path (**parked, id 47**), TTS-stack move - to fv-ml1 (**parked, id 75**). + `/mnt/smithy` (**declined in scope**), fused MoE kernel path (**parked, id 47**), TTS-stack + move to fv-ml1 (**parked, id 75**). -⚠ **fv-ml1 GPU 0 is at 96.0 of 97.9 GB.** GPU 1 ~5.7 free, GPU 2 ~2.4, GPU 3 a HELD RESERVE. The -next seat needing room on fv-ml1 requires a placement decision. +⚠ **fv-ml1 GPU 0 is at 96.0 of 97.9 GB.** Adding `lv-bronte` cost nothing measurable (96092 → +96090 MiB) because a LoRA rides inside the existing seat — but a new SEAT still needs a +placement decision. GPU 3 is a HELD RESERVE. **Uncommitted:** `graphify-out/GRAPH_REPORT.md` and `scripts/seat-inventory.py` were modified before this session began. ⚠ **Do NOT commit them** — untouched and deliberately left alone. ## Recent decisions +- `[2026-09-17]` ⭐⭐ **A leak class the corpus gate structurally CANNOT see: the beat-writing model recognises the book and restores the author's real character names.** 1.8% of Brontë beats named Rochester/Jane/Brocklehurst while 0 responses did. Worst for public-domain classics; Yarros and Hemingway's clean runs are NOT evidence they are immune. Patched as a `sourcename` reject + `--source-entities`. → `persistent-memory.d/2026-09-17-beat-contamination-leak.md` +- `[2026-09-17]` ⭐ **A stoplist entry is an assertion the leak gate can no longer check** — stoplisting removes a surface from the entity map, so a wrongly stoplisted CHARACTER is an undetectable leak. Three were wrong on Brontë (Leaven, Pierrot, Samuel); `scripts/r49-corpus/audit_stoplist.py` finds them by honorific and now gates the pipeline. Commit `8bb7686`. +- `[2026-09-17]` **ESH: Cityside Fiber failed TWICE (19:09 and ~01:06); operator switched WAN1 to DHCP to restore service and has a ticket for the static.** crowdsec `esh` allowlist carries both failover egresses with 7-day expiries — the rotation-fragility is live. → `persistent-memory.d/2026-09-17-esh-fiber-outages.md` - `[2026-09-17]` ⭐⭐ **lv-bronte SHIPPED on voices-seat (ckpt475) DESPITE failing the v2 VOICE axis — additive, reversible, safety-axis clean.** Both candidates closed 48–52% of the achievable distance to Brontë but +0.193/+0.210 sit under a 0.251 noise floor set by ONE outlier seed in the arm not being shipped; cause is structural (81 val pairs vs Hemingway's 200) and not cheaply fixable. ckpt475 is the pick if it ships. The two-epoch recipe did NOT transfer. → `persistent-memory.d/2026-09-17-lv-bronte-gate.md` - `[2026-09-16]` ⭐⭐ **Instruction-pair SFT BEATS raw-text for author voice, and the raw-text incumbent never cleared its own control (+0.141 against a 0.153 floor).** `lv-yarros` shipped; the v1 decision rule was amended by the operator after measurement showed it gated on axes the unadapted carrier already maxes. → `persistent-memory.d/2026-09-16-lv-voices-line.md` - `[2026-09-16]` ⭐ **voices-seat live: one carrier, N `lv-` LoRA adapters, hot-swap measured at 0.24 s.** LoRA costs 24.3% of decode against a 0.1% A-vs-A floor and is worth paying; `--gpu-memory-utilization` is a request against TOTAL VRAM and only a pinned KV makes it predictive. → `persistent-memory.d/2026-09-16-voices-seat-lora.md` @@ -382,9 +413,6 @@ before this session began. ⚠ **Do NOT commit them** — untouched and delibera - `[2026-09-03]` **nh3-dev wedged for ~40 min and it was the BACKUP, not the disk — a stalled cross-site vzdump holding ever** → `persistent-memory.d/2026-09-03-nh3-dev-wedged-for-40-min.md` -- `[2026-09-02]` **`vastblue` gitea org created (id 8, private, owner `vh`) with empty repo `vastblue/platform`** — third entity namespace alongside `corviduo` and `pfi`; most repos still live under `vh/`. Home of VastBlueDocumentAI + the anchor healthcare-billing SPA (signed 3-yr client contract). **Org scope was the decision**: org-level runner registration and secrets are inherited free by the DocumentAI repo when it splits out, and that is the only binding expensive to retrofit. Deliberately NOT set: org runner (instance-wide `ana-docker-runner` already serves it; org scope is for the DEDICATED runner, deferred to U10) and org secrets (none exist yet; a guessed secret looks bound). ⚠ **Dedicated runner is gated on the first client-premises release cut**, not on the first green pipeline — the risk is another repo's CI sharing a root-level daemon with a build that ships to a healthcare client, see the runner entry above. Push needs no credential: vastblue-dev is on nh3-dev and git-SSH there auths as `vh`. → `stacks/gitea-runner/README.md` - - - `[2026-09-02]` **Every CI job on the shared `pfi-fleet` runner is root on ana-docker — and `container.valid_volumes: []` does NOT prevent it.** Measured: a job container is uid 0, `/var/run/docker.sock` is mounted by act_runner independently of that list, `docker ps` returns all 49 host containers (gitea itself, synapse, phasefinal-web, adguardhome), `docker compose v2.33.0` on PATH. ⚠ **LOAD-BEARING** — `vh/Worldtree`, `vh/soong-lab`, `vh/skaldsong`, `vh/wt-matrix-bridge` all drive buildx through that socket, so it cannot simply be closed; **isolate sensitive builds onto a dedicated runner instead.** Also measured the same night: `services:` containers work (Postgres 16), and **full-URL `uses: https://gitea.phasefinal.com/actions/checkout@v4` resolves from the local mirrors** — the un-parked half of the github-independence work, needing neither `DEFAULT_ACTIONS_URL=self` nor the act_runner auth path that blocked it on 2026-08-05. Prompted by vastblue-dev's CI-posture question for a client-funded healthcare repo. → `stacks/gitea-runner/README.md` @@ -398,7 +426,7 @@ before this session began. ⚠ **Do NOT commit them** — untouched and delibera - `[2026-08-19]` **AI-tab Dormant regrouping BELAYED by the operator** — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong, Chatterbox Fast) show amber EXITED inside live groups rather than `AI - Dormant`. Fix is a label change + recreate per stack; needs the operator's read on which are retired vs temporarily down. `untracked by operator choice` (his words: "belay the ai dormant regrouping for now"). -_9 older entries archived to archival-memory.md._ +_10 older entries archived to archival-memory.md._ _5 older entries archived to archival-memory.md._ diff --git a/scripts/yarros-corpus/memorization_check.py b/scripts/yarros-corpus/memorization_check.py index bd3ce72..7d5fdf3 100644 --- a/scripts/yarros-corpus/memorization_check.py +++ b/scripts/yarros-corpus/memorization_check.py @@ -10,12 +10,25 @@ Instrument: longest and mean maximal verbatim n-gram shared with the TRAINING co generation. Controls run every time -- the base-unadapted arm never saw the corpus so it is the negative control, and a slice of the corpus scored against itself is the positive. """ -import json, pathlib, re, sys +import argparse, json, pathlib, re, sys from collections import Counter -EVAL = pathlib.Path("/home/infra-ops/r49-runs/yarros-eval") -CORP = pathlib.Path("/home/infra-ops/yarros-corpus-renamed/copies") -N = 8 +# ⚠ These were hardcoded to Yarros. Pointed at a Brontë arm they would have compared +# it against the YARROS corpus and reported a clean zero — a negative that means +# "different book", not "did not memorise". Defaults are unchanged so every Yarros +# number already recorded stays reproducible byte for byte. +_ap = argparse.ArgumentParser() +_ap.add_argument("--eval-dir", default="/home/infra-ops/r49-runs/yarros-eval") +_ap.add_argument("--corpus", default="/home/infra-ops/yarros-corpus-renamed/copies", + help="the RENAMED copies the adapter actually trained on") +_ap.add_argument("--glob", default="beats5.*.jsonl", help="arm files inside --eval-dir") +_ap.add_argument("--strip", default="beats5.", help="prefix trimmed to name the arm") +_ap.add_argument("-n", type=int, default=8, help="n-gram length") +_a = _ap.parse_args() + +EVAL = pathlib.Path(_a.eval_dir) +CORP = pathlib.Path(_a.corpus) +N = _a.n def norm(t): return re.findall(r"[a-z']+", t.lower()) @@ -44,8 +57,8 @@ def longest_match(words): print(f"{'arm':<22} {'gens':>5} {'hit-rate':>9} {'mean-longest':>13} {'max':>5}") print("-" * 60) -for f in sorted(EVAL.glob("beats5.*.jsonl")): - arm = f.stem.replace("beats5.", "") +for f in sorted(EVAL.glob(_a.glob)): + arm = f.stem.replace(_a.strip, "") rows = [json.loads(l) for l in f.read_text(encoding="utf-8").splitlines()] longs = [longest_match(norm(r["raw"])) for r in rows] hits = sum(1 for x in longs if x >= N)