memory: snapshot — lv-yarros shipped, voices-seat live, lv-hemingway training, Grok broker shelved

This commit is contained in:
Vuong Hoang
2026-09-16 16:19:58 -07:00
parent cf9d167453
commit e8086941e2
9 changed files with 813 additions and 560 deletions
+36 -18
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-09-15 ~09:45 PT (Parakeet STT live on fv-ml1 GPU 0; `svos_miranda` LIVE in Hermes; talk v10 deployed; irv-ml1 dead-address sweep COMPLETE; secrets-broker concurrency fixed; breeze stays on irv-ml1, TTS-stack move parked at id 75. Nothing blocked, nothing mid-flight.)_
_Last updated: 2026-09-16 ~16:15 PT (lv-yarros SHIPPED — instruction-pair SFT beats raw-text on voice; voices-seat live on fv-ml1 GPU 0 :8027 with measured 24.3% LoRA cost; lv-hemingway corpus gated and training, ~1h out; Grok token broker built then SHELVED by the keep-the-jail ruling.)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
@@ -115,26 +115,47 @@ no longer deployed sidecars here. See Recent decisions.)
## Current state / in-flight
_As of 2026-09-15 ~09:45 PT._
_As of 2026-09-16 ~16:15 PT._
### Nothing is in flight and nothing is blocked. Both of the previous session's named jobs closed, plus six unplanned pieces of work.
### One job in flight: `lv-hemingway` training on pfi-gx10, ~1 hour out.
**Closed this session:**
1. **Parakeet STT live** — fv-ml1 **GPU 0**, port 8300, v3 int8 25-language model, behind LiteLLM `ext-stt` / `whisper-1`. ⚠ Placed on GPU 3 first; operator corrected it — a ~800 MiB seat belongs on the card with the most uncommitted headroom, not on the one pristine 96 GB card, because vLLM sizes KV against TOTAL VRAM. **GPU 3 is a deliberate reserve at 2 MiB.**
2. **`svos_miranda` LIVE in Hermes** — gateway restarted, 29 toolsets, Miranda scoped to exactly 8, operator's own surface intact at 46. ⚠ `agent.disabled_toolsets` is DELETED and stays out (operator ruling); svos-dev fixed their roster check at `c9d2a96`. SVOS restarted itself; both roster lines verified.
3. **talk v10 deployed** on nh3-dev :8092 — push-to-talk STT through `ext-stt`, barge-in. First consumer of the Parakeet seat.
4. **irv-ml1 dead-address sweep DONE** — 0 of 112 Homepage cards on `10.100.79.3`, was 9. Found and fixed four live breakages on OTHER hosts (Open WebUI TTS, asset-engine, skaldsong x2).
5. **`secret` concurrency bug fixed** — parallel `secret get` returned empty with exit 0. Command-level lock + empty-value guard + `find()` no longer coercing empty stdout to `[]`. `~/.local/bin/secret` is now a symlink, was a stale copy.
6. **Retired:** irv-ml1 parakeet (lost tts-dev's bench) and voice-studio (dots obsoleted by Breeze).
**IN FLIGHT — `lv-hemingway` pair-SFT.** `gx10:~/r49-runs/hemingway-4b-pairs-3ep`, 7,094 pairs,
2,661 steps at ~3.68 s/it, launched ~15:05 PT. **Take the checkpoint at the LOSS MINIMUM, not the
end-of-run adapter** — the recipe is two epochs on a three-epoch schedule. When it lands: run the
v2 gate (voice `delta_cb` vs base control beyond the noise floor · 8-gram overlap near the
never-saw-it control · overshoot) using `scripts/yarros-corpus/{score_beats,memorization_check}.py`
and the 30-beat in-genre fixture, then ship to
`/tank/aimodels/voice-adapters/lv-hemingway-4b-v1/` and add it to `stacks/voices-seat/compose.yaml`.
**Settled at the end of the session:** breeze-tts **stays on irv-ml1**; moving it plus `bragi` and `tts-gateway` to fv-ml1 is **parked at id 75**, triggered on evacuating embed/rerank/reward. Full fv-ml1 per-seat residency measured — see the breeze detail file.
**SHIPPED — `lv-yarros`.** `gx10:~/adapters/lv-yarros-4b-v1/` (sha `63fda6cc61f380b2`) and live on
`vllm-voices`, fv-ml1 GPU 0 :8027, alongside `voices-base`.
**Open, all operator-deferred, none blocking:** the AI-tab Dormant regrouping (belayed), `nconnect=8` on /mnt/smithy (deferred), fused MoE kernel path (park id 47), TTS-stack move (park id 75). `speaches`'s label claims `:8204`, which is breeze-tts's live port — a latent conflict if anyone starts it.
**OPEN, operator's call, nothing blocked:**
- **Brontë has never been through the automated leak gate** — its "0 of 203" was a HAND COUNT and
the gate did not exist yet. On Yarros the same instrument read **212 surviving where a hand count
said 86**. Its gated corpus survives at `gx10:~/r49-corpus-renamed-unwrapped/` and is
pair-buildable; `lv-bronte` exists only as RAW-TEXT arms, i.e. the arm that never cleared its own
control. Re-gate before building pairs, or accept the hand count.
- **xAI billing** — did the stopped HTTP A/B bill on top of the coding plan? Needs the operator's
console access; this fleet holds no xAI credential.
- **A `BabyYarros → lv-yarros` pointer** in memory, so historical entries stay findable under the
new name. Historical entries were deliberately left as dated records.
- Older deferred set, unchanged: AI-tab Dormant regrouping (**belayed**), `nconnect=8` on
`/mnt/smithy` (**declined in scope**), fused MoE kernel path (**parked, id 47**), TTS-stack move
to fv-ml1 (**parked, id 75**).
**Uncommitted:** `graphify-out/GRAPH_REPORT.md` and `scripts/seat-inventory.py` were modified before this session began; untouched and deliberately not committed.
⚠ **fv-ml1 GPU 0 is at 96.0 of 97.9 GB.** GPU 1 ~5.7 free, GPU 2 ~2.4, GPU 3 a HELD RESERVE. The
next seat needing room on fv-ml1 requires a placement decision.
**Uncommitted:** `graphify-out/GRAPH_REPORT.md` and `scripts/seat-inventory.py` were modified
before this session began. ⚠ **Do NOT commit them** — untouched and deliberately left alone.
## Recent decisions
- `[2026-09-16]` ⭐⭐ **Instruction-pair SFT BEATS raw-text for author voice, and the raw-text incumbent never cleared its own control (+0.141 against a 0.153 floor).** `lv-yarros` shipped; the v1 decision rule was amended by the operator after measurement showed it gated on axes the unadapted carrier already maxes. → `persistent-memory.d/2026-09-16-lv-voices-line.md`
- `[2026-09-16]` ⭐ **voices-seat live: one carrier, N `lv-<author>` LoRA adapters, hot-swap measured at 0.24 s.** LoRA costs 24.3% of decode against a 0.1% A-vs-A floor and is worth paying; `--gpu-memory-utilization` is a request against TOTAL VRAM and only a pinned KV makes it predictive. → `persistent-memory.d/2026-09-16-voices-seat-lora.md`
- `[2026-09-16]` ⭐ **lv-hemingway corpus gated at 994,760 words — and half the catalogue had to be EXCLUDED.** 169,759 words of measured 90–96% collection duplication, a Sherwood Anderson parody, and the author's own name 95 times in publisher back matter; the gender resolver needed a corpus base-rate correction to stop reading women as men. → `persistent-memory.d/2026-09-16-lv-hemingway-corpus.md`
- `[2026-09-16]` **Grok token broker built then SHELVED — operator ruled "keep the jail stop the a/b", so the renewal feature has no consumer.** ⛔ Do NOT arm `probe-rotation`: the risk did not shrink (it reaches BOTH Gróa transports through one shared session) and the payoff went to zero. → `persistent-memory.d/2026-09-16-grok-broker-shelved.md`
- `[2026-09-15]` ⚠⚠ **DO NOT carry "a client-side timeout is not a cancellation" as a rule — it is FALSE as stated.** A clean abandon cancels itself ~6 s later (measured); yet six requests genuinely orphaned on `vllm-erp-seat`. Some propagate, some do not, **boundary unknown** — which argues for a detector, not a rule. ⭐⭐ The durable artifact: **a serving engine's KV cache CYCLES, an orphaned one only CLIMBS** — request count and throughput are ambiguous between loaded and wedged, and I called the seat healthy twice off them (correctly, on the evidence). ⚠ A `max_tokens` ceiling would NOT have prevented it: the worst offender had 16384 set, hit it, and returned 24,594 chars of whitespace. → `persistent-memory.d/2026-09-15-client-abandon-cancellation-boundary.md`
- `[2026-09-15]` ⚠⚠ **`--gpu-memory-utilization` DOES NOT PREDICT RESIDENT VRAM — measure it, never compute it.** Wrong in **both** directions on fv-ml1: `vllm-cyberprev` util 0.40 (expect ~39,155 MiB) holds **47,124** (+8 GB over); `vllm-gen-small` util 0.48 (expect ~46,986) holds **36,942** (−10 GB under). Planning a placement off the fractions would have been 8 GB wrong. Read `nvidia-smi --query-compute-apps`. Full per-seat residency table + the breeze shuffle arithmetic → `persistent-memory.d/2026-09-15-breeze-placement-sizing.md`
@@ -359,19 +380,14 @@ _As of 2026-09-15 ~09:45 PT._
- `[2026-09-03]` **nh3-dev wedged for ~40 min and it was the BACKUP, not the disk — a stalled cross-site vzdump holding ever** → `persistent-memory.d/2026-09-03-nh3-dev-wedged-for-40-min.md`
- `[2026-09-02]` **althing deploy is SIX surfaces, and #6 is outside the althing repo: ~/.claude/settings.json crossSessionI** → `persistent-memory.d/2026-09-02-althing-deploy-is-six-surfaces-and.md`
- `[2026-09-02]` **`vastblue` gitea org created (id 8, private, owner `vh`) with empty repo `vastblue/platform`** — third entity namespace alongside `corviduo` and `pfi`; most repos still live under `vh/`. Home of VastBlueDocumentAI + the anchor healthcare-billing SPA (signed 3-yr client contract). **Org scope was the decision**: org-level runner registration and secrets are inherited free by the DocumentAI repo when it splits out, and that is the only binding expensive to retrofit. Deliberately NOT set: org runner (instance-wide `ana-docker-runner` already serves it; org scope is for the DEDICATED runner, deferred to U10) and org secrets (none exist yet; a guessed secret looks bound). ⚠ **Dedicated runner is gated on the first client-premises release cut**, not on the first green pipeline — the risk is another repo's CI sharing a root-level daemon with a build that ships to a healthcare client, see the runner entry above. Push needs no credential: vastblue-dev is on nh3-dev and git-SSH there auths as `vh`. → `stacks/gitea-runner/README.md`
- `[2026-09-02]` **althing 3.3.0 deployed — the cc channel, and a plugin-cache false green.** → `persistent-memory.d/2026-09-02-althing-3-3-0-deployed-the.md`
- `[2026-09-02]` **Every CI job on the shared `pfi-fleet` runner is root on ana-docker — and `container.valid_volumes: []` does NOT prevent it.** Measured: a job container is uid 0, `/var/run/docker.sock` is mounted by act_runner independently of that list, `docker ps` returns all 49 host containers (gitea itself, synapse, phasefinal-web, adguardhome), `docker compose v2.33.0` on PATH. ⚠ **LOAD-BEARING** — `vh/Worldtree`, `vh/soong-lab`, `vh/skaldsong`, `vh/wt-matrix-bridge` all drive buildx through that socket, so it cannot simply be closed; **isolate sensitive builds onto a dedicated runner instead.** Also measured the same night: `services:` containers work (Postgres 16), and **full-URL `uses: https://gitea.phasefinal.com/actions/checkout@v4` resolves from the local mirrors** — the un-parked half of the github-independence work, needing neither `DEFAULT_ACTIONS_URL=self` nor the act_runner auth path that blocked it on 2026-08-05. Prompted by vastblue-dev's CI-posture question for a client-funded healthcare repo. → `stacks/gitea-runner/README.md`
- `[2026-09-02]` **pfi-gx10 BASELINED: 79.36 s/it median on the run-3c shape, and the training stack works on aarch64/sm_121.** Median across 10 timed steps, 0.19% spread, **peak 75.1 / 121.6 GiB — 46 GiB spare**, `attn_resolved: flex_attention`. **6× slower than ana-ml2 where compute predicts 2.7×** → likely memory-bandwidth-bound; **capacity box, not throughput box.** Ruled **bare metal, not Proxmox** (no aarch64 PVE; the GPU is on-package and cache-coherent, so passthrough would partition the unified memory that is the whole point). ⚠ `sm_121` is NOT in torch's arch list — everything JITs from sm_120 PTX, so **warm up before timing anything** (an unwarmed bench read 27 TFLOP/s against a true 93). (context archived → `archival-memory.md`)
- `[2026-09-02]` **I priced a failure in the units I happened to be measuring — operator overruled me, correctly.** Recommended run 3c to ana-ml2 by costing a breaker trip as "≤50 steps ≈ 11 min of recompute". It is a **40-minute drive each way** with **13 Anaheim hosts dark, three of them SureFire CLIENT machines**. `save_steps` caps the recompute, never the outage. ⚠ **General form: a metric in hand will volunteer itself as the unit of risk.** (context archived → `archival-memory.md`)
- `[2026-09-02]` **althing 3.2.0→3.2.4 deployed, and ALTHING DEPLOY IS FOUR SURFACES not three.** The fourth (plugin) had no runbook step and was frozen at Aug 28 — **missing the SessionStart/SessionEnd hooks and `pane-route.sh` entirely**, so "CC seats re-declare automatically" was never true here. Now one command (`scripts/deploy-althing.sh`). ⚠ `uv tool install .` **without `--force` is a silent no-op**. ⚠ **A missing deploy surface presents as "the migration needs manual work", not as an error.** → `persistent-memory.d/2026-09-01-althing-320-deploy.md`
- `[2026-08-25]` **Fused MoE kernel path — DEFERRED, tracked at park `fused-moe-kernel-path-for-gemma-4-moe-training` (id 47).** Operator: "note the fused MoE kernel for round two… if we nail it soon, the math has us wanting to restart the run anyway." Training MFU is **8.6%** (27.1 of a benchmarked 313.8 TFLOPS) because `transformers` runs the Gemma-4 experts in a Python loop — 128 experts × 30 layers, ~11,500 iterations per step under gradient checkpointing. ⚠ **The same fused 3-D expert layout that made bitsandbytes skip 88.5% of the model is exactly what a grouped GEMM wants** — the format is good for storage and for fused kernels, and hostile only to naive iteration. Two fixes: `group_by_length` (−29.9% compute, free, but breaks the seeded order manifest and re-opens a batch-composition call brokkr already made) and a grouped-GEMM/compiled MoE forward (the remaining ~10×). **Not applied to the live run** — restarting mid-flight to change batch ordering was judged a bad trade at step ~50 of 1,312.
@@ -383,6 +399,8 @@ _As of 2026-09-15 ~09:45 PT._
_9 older entries archived to archival-memory.md._
_5 older entries archived to archival-memory.md._
## Tried and abandoned
- `[2026-09-15]` ⚠⚠ **Probing OPNsense API endpoints by POSTing at them — one was `/api/core/system/reboot` and it took the FV site dark for 3.5 min.** Endpoints are ACTIONS; a 200 means it ran. The call I wanted was documented in this repo's own `docs/pfi/opnsense-api-reference.md`. → `persistent-memory.d/2026-09-15-opnsense-api-reboot.md`