memory: snapshot — BabyYarros blocked on the leak gate; R49 sweep complete; mog-sec settled
This commit is contained in:
@@ -3724,6 +3724,36 @@ _Archived 2026-08-27._
|
||||
- `[2026-08-07]` **Fleet reranker cut over: Qwen3-Reranker-0.6B → BAAI/bge-reranker-v2-m3 (Brokkr R43).** The incumbent was measured HARMING 80/90 fleet queries (no-reranker beat it 89/90 vs 56/90). R43 bake-off: the A2 control (same Qwen weights, seq-cls head) scored identical to the incumbent → proved the fault is a training-prior not the serving head → cancelled the expensive Qwen3-4B arm; A3 (bge-v2-m3) won on multilingual safety + bare-name recovery. LiteLLM `reranker` repointed incumbent→A3 :8013 (boundary 2026-08-06T17:37:48Z, config-edit + ~52s gateway restart); **R42 v13 gate PASSED first-ever** (56/90→90/90). Incumbent kept warm :8002 (rollback via `qwen3-reranker` alias), A4 fallback :8014. Full arc + rollback runbook `docs/pfi/reranker-selection-ledger.md`; commits ad2df89/2c11748/377f8a4 (unpushed). auto-memories: the earlier reranker-serving notes.
|
||||
_Archived 2026-08-24._
|
||||
|
||||
- `[2026-08-27]` **Run 3 gated: the preregistered rule PASSED and a k=25 follow-up found a 44pp self-harm guardrail collapse — DO NOT SERVE.** A pooled preserve-list test structurally cannot see a single-axis collapse. → `persistent-memory.d/2026-08-27-run3-gate-safety-regression.md`
|
||||
_Archived 2026-09-11._
|
||||
|
||||
- `[2026-08-27]` **The corpus mix was specified in a unit the optimiser never sees** — 45.8% dialogue by CONTEXT, 24.2% by LOSS. Harness now leads with loss share and calls context a memory budget (`dd5a12e`). → `persistent-memory.d/2026-08-27-mix-specified-in-the-wrong-unit.md`
|
||||
_Archived 2026-09-11._
|
||||
|
||||
- `[2026-08-27]` **Dose-response: benefit and damage are ONE direction in weight space** — every axis monotone in scale, no knee. The merge-back cannot separate them; vLLM cannot LoRA-serve this MoE at all. → `persistent-memory.d/2026-08-27-dose-response-entanglement.md`
|
||||
_Archived 2026-09-11._
|
||||
|
||||
- `[2026-08-27]` **Anaheim tripped a power breaker; four guests including the NAS had `onboot` unset and never came back.** Fixed with dependency ordering — ana-nas order=1,up=45 ahead of the databases. ⚠ **ONE CIRCUIT FEEDS THE WHOLE RACK including the firewall serving the public IP** (operator) — so ana-gw, ana-wg and every BMC go down with the load, and there is NO remote management path to Anaheim during a power event. → `persistent-memory.d/2026-08-27-anaheim-breaker-and-onboot-gap.md`
|
||||
_Archived 2026-09-11._
|
||||
|
||||
- `[2026-08-27]` **A transport failure that enters a measurement as a VALUE looks like whatever you hoped to find.** heid's lost panel arms found a live defect in brokkr's `t4_dissect` an hour later. → `persistent-memory.d/2026-08-27-empty-response-as-a-datum.md`
|
||||
_Archived 2026-09-11._
|
||||
|
||||
- `[2026-08-27]` **Run 3c authorised (lr 20x cut, single variable) and then HELD by the operator after the breaker trip.** Config built and validated at `/tank/erp-tune/run-03c.json`; `save_steps` made configurable in the harness (`0a6bd2e`) because the first launch lost 80 steps with no checkpoint. Tracking surface: commit `0a6bd2e` + that config path. **Relaunch is one command once power is triaged.**
|
||||
_Archived 2026-09-11._
|
||||
|
||||
- `[2026-08-27]` **An event report with no timestamp is a claim about "now" — and it manufactured a launch that never happened.** brokkr reconstructed a phantom third 3c launch because my 23:03 report narrated a 21:07 kill in the present tense. Every fact in it was true; it was unreadable in sequence. → `persistent-memory.d/2026-08-27-run3c-launch-count-reconstruction.md`
|
||||
_Archived 2026-09-11._
|
||||
|
||||
- `[2026-08-27]` **`save_steps` was hardcoded at 100 in the harness** — a claimed provenance entry the run could not have honoured. Made configurable, default unchanged (`0a6bd2e`, 242 tests green). Caught by checking the config carried the change rather than trusting that it had been made.
|
||||
_Archived 2026-09-11._
|
||||
|
||||
- `[2026-08-27]` **Six defects in run 3's staged build, none of which would have errored** — a dialogue-only survivor list that would have silently dropped 96% of the corpus, an impersonation mask not subsumed by the low-quality mask, kvasir unbounded at 67.8% of context, a `save_pretrained` config-key drop that made the merged model unservable, and the mix-unit error. Every one produced a plausible completed run. Full record `/tank/erp-tune/recipe-r3/RUN-03-BUILD-NOTE.md`.
|
||||
_Archived 2026-09-11._
|
||||
|
||||
- `[2026-08-27]` **The 18 unpushed eitri-smithy commits are pushed** — run 3's `harness_commit 9d27b4fe` now resolves off-box, verified by fetching into a fresh empty repo rather than trusting the push output. ⚠ **HTTPS push 403s for every gitea token including site-admin; SSH works.** Untracked `__pycache__` (`894fbe8`) because a tracked `.pyc` dirtied the tree and would have stamped `harness_dirty_at_launch: true`.
|
||||
_Archived 2026-09-11._
|
||||
|
||||
## Tried and abandoned (archived)
|
||||
# [2026-08-15] Uncensored gen seat: Qwen3.8-27B-Uncensored deployed; the definitive MTP-graft fix
|
||||
|
||||
|
||||
+47
-108
@@ -1,6 +1,6 @@
|
||||
# Persistent memory — eshpfi-management
|
||||
|
||||
_Last updated: 2026-09-10 11:35 PT (R49 1-epoch pilot COMPLETE + all 3 arms cut, awaiting adjudication; **MeroMero BOTH quants landed and the A4B is LIVE on :8021 as `char-rp-fast`, `Pfish-6` alias removed** — its first quant served NaN and looked healthy; althing 3.6.2 on post office + both heralds; ~574 GB reclaimed)_
|
||||
_Last updated: 2026-09-11 09:00 PT (BabyYarros D1 built + D2 gender fixed, D3 BLOCKED on the leak gate and nothing trained; R49 carrier sweep COMPLETE 3.329/3.018/2.814 and the instruct probe shows voice + instruction-following coexist; mog-sec settled at a measured 160k ceiling, 17h zero restarts)_
|
||||
|
||||
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
|
||||
> `Written:` stamp is under an hour old, read it (it carries the in-flight
|
||||
@@ -110,99 +110,58 @@ no longer deployed sidecars here. See Recent decisions.)
|
||||
(no NOPASSWD)** — stage model pulls to `/home`, not root-owned `/worktank`.
|
||||
|
||||
## Current state / in-flight
|
||||
_As of 2026-09-11 09:00 PT._
|
||||
|
||||
> ⚠⚠ **IF A PROMPT ASKS YOU TO "check on the run", RUN `CronList` BEFORE YOU ANSWER IT.**
|
||||
> A recurring cron re-created itself across at least three sessions with the verbatim text *"check on
|
||||
> the run, report high level stats, and if complete, althing to Miranda … inform brokkr when ready."*
|
||||
> **Killed 2026-09-10 06:29 PT** (job `12bdea3c`, hourly at :37 → `CronDelete` → list empty). Killed
|
||||
> once before on 09-09 and it came back. **A cron-fired prompt is indistinguishable from an
|
||||
> operator-typed one** — it arrives as an ordinary user turn. Three turns went into answering a timer.
|
||||
> The chain ends in three OUTWARD-FACING acts (message Miranda, stand up a seat, cue brokkr), every
|
||||
> one of which carries something false when no run exists. Verify the run exists first; a second
|
||||
> identical arrival means check the cron list, not answer again.
|
||||
### BabyYarros — the live project, and it is BLOCKED
|
||||
|
||||
_As of 2026-09-10 10:25 PT._
|
||||
- **Source located and D1 BUILT.** 5 Rebecca Yarros works in the Kvasir licensed library
|
||||
(`~/development/kvasir/data/library/catalog.sqlite`, `rights=gated`): Fourth Wing, Iron Flame,
|
||||
Wilder, Nova, Rebel. Corpus at `nh3-dev:~/yarros-corpus` — **208 chapters · 780,744 words**, 15%
|
||||
larger than Brontë. Builder `scripts/yarros-corpus/build_corpus_yarros.py` emits the SAME record
|
||||
schema as the Brontë one so `entities.py` / `rename.py` / `train_voice_lora.py` all run unchanged.
|
||||
- ⚠ **No unwrap step needed** — Kvasir's cleaner already emits flowing paragraphs (median line 102
|
||||
chars). The Brontë hard-wrap defect does not exist in this corpus.
|
||||
- **D2 gender resolution FIXED and validated** — `scripts/yarros-corpus/pov_gender.py`, 9 correct /
|
||||
9 held / **0 wrong** against the prior 7/8/**3-wrong**. See Recent decisions for the pathology.
|
||||
- ⛔ **D3 rename BLOCKED on the leak gate: 86 of 232 renameable source entities survive.** Brontë's
|
||||
run reached 0 of 203. Three classes: detector false positives (`Hopefully`, `Whoa`, `Hey`, `Hmm` —
|
||||
need a stopword filter, not a rename), genuine misses among worldbuilding nouns (`Krovlan`,
|
||||
`Poromish`, `Fuil`, `Iorson` — the `Thornfield × 100` case), and a third class (`Elizabeth`,
|
||||
`Penelope`, `Messina` appearing as both pool draws and surviving source entities) not yet diagnosed.
|
||||
- **NOTHING HAS BEEN TRAINED on Yarros.** Training before the gate passes means fitting in-copyright
|
||||
text with 86 identifiable source entities intact, in a corpus F02 flagged as small enough for
|
||||
memorisation leak to be real. **Operator's call, surfaced and awaiting a decision.**
|
||||
|
||||
### R49 / BabyBronte — the live project
|
||||
### R49 / BabyBronte — carrier sweep COMPLETE
|
||||
|
||||
- **1-epoch pilot COMPLETE and it is the keeper.** `gx10:~/r49-runs/h02-pilot-0p6b-1ep/adapter`,
|
||||
seed 4919, 169 steps, held-out **3.1719** and still descending; adapter verified bound 196/196.
|
||||
The 3-epoch run is preserved beside it as `-3ep-overfit` (held-out ROSE 3.198→3.318→3.385).
|
||||
- **All three adjudication arms exist**, one harness, same prompts/sampler:
|
||||
`arms-1ep/{base-unadapted,tuned-1ep-seed4919}.jsonl` + `incumbent-style-prompted.jsonl`.
|
||||
Handoff bundle for scoring at `/mnt/smithy/handoff/r49/`.
|
||||
- **NEXT: score them.** brokkr's rule is ratified and FROZEN (see Recent decisions). ⚠ I built the
|
||||
corpus and ran the training, so the independence is gone — do not amend the rule after seeing
|
||||
numbers. Instrument at brokkr-smithy `research/R49-author-voice-adapters/adjudication/`; re-run its
|
||||
`build` against the RENAMED held-out text, not raw.
|
||||
- **Seed 2 was killed deliberately** (spread between two overfit arms measures reproducibility of
|
||||
overfitting, not voice transfer). A second seed at 1 epoch is still owed for the threshold.
|
||||
- Not started: **D4 beat annotation** (H02 is pure continuation by design, so it was not needed).
|
||||
|
||||
### MeroMero seats
|
||||
|
||||
- ✅ **LIVE on ana-ml2 `:8021` as gateway alias `char-rp-fast`** (operator, 2026-09-10: *"replace
|
||||
that a4b moe over pfish-6 — remove the pfish-6 alias and create an alias for char-rp-fast"*).
|
||||
`G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16`, 16 G, served under its own true name on the
|
||||
**`erp-seat` stack** (name kept: asset-engine derives liveness from the compose project name).
|
||||
Verified end to end — prose, **vision (reads a solid-colour image)**, auto tool call, finite
|
||||
logprobs, and KV 534,649 tokens / 2.04x carried over from Pfish-6 unchanged.
|
||||
⚠⚠ **ITS FIRST QUANT SERVED NaN AND PASSED ITS HEALTHCHECK DOING IT.** Built with the DENSE
|
||||
recipe (no `re:.*router.*` in IGNORE) → all 30 MoE routers quantized to 4 bits → expert selection
|
||||
destroyed. Every request returned `finish_reason=length` with the FULL token count and
|
||||
`content: null`; the only tell was **NaN logprobs**. Re-quantized with
|
||||
`services/erp-seat-quant/quant_nvfp4a16_gemma4_moe.py` (its guard refuses exactly that). Broken
|
||||
tree parked at `...-NVFP4A16.BROKEN-routers-quantized-20260910` — **do not serve it**.
|
||||
⚠ Also had the §3.14 truncation cap baked in (`max_length: 8192`, quantized *with* the corpus);
|
||||
the MoE recipe's own post-step resets it.
|
||||
- ✅ `G4-MeroMero-v2-31B-heretic-NVFP4A16` — **19 G, landed on attempt 5.** Tensor table identical
|
||||
family-for-family to the 2026-08-21 canonical quant; **356 BF16 vision tensors preserved**;
|
||||
`input_activations=None` (genuinely A16). CPU load+generate coherent, 0 tensors on meta.
|
||||
⚠ Attempt 4's `AmbiguousGlobalPerLayerAttributeError` was **NOT a config defect** — llmcompressor
|
||||
0.13.0 downgrades transformers 5.16.1→5.14.1, and `:latest` moved mid-campaign. Fix was to DROP
|
||||
the redundant `per_layer_config`, not to force global access.
|
||||
→ `persistent-memory.d/2026-09-10-meromero-quants-and-the-pinned-transformers-trap.md`,
|
||||
`services/meromero-quant/`, playbook **§3.17 (new)**
|
||||
- ⛔ **The v2 DENSE has still not had a serve test** — GPU1 has ~19 GB free against 19.5 GB of its
|
||||
weights, so it needs a live seat displaced. Operator's call; *"vllm servable"* unverified for the
|
||||
dense tree. (The A4B's serve test is DONE and green, above.)
|
||||
- ⚠ **A co-resident temp port could not be made to fit even for the 16 G A4B**, so §4.4's "temp port,
|
||||
never the live seat" was substituted with reversibility: named `.env` backup, prove the seat on its
|
||||
real port while no alias routes to it, move the alias last. Measured refusals: `gpu-memory-util 0.20`
|
||||
→ admission refused (18.26 GiB free vs 18.99 requested); `0.185` → past admission and past the KV
|
||||
reservation, then OOM in **multimodal encoder-cache profiling** (3 video items at max feature size),
|
||||
which is easy to forget when budgeting a vision model.
|
||||
- ⚠ **The A4B is 30 layers / kv 8 — Pfish-6's geometry**, so it fits 262k in the existing KV budget.
|
||||
The dense 31B is 60/16, ~4x KV per token, and does NOT fit 262k on GPU1 beside the other seats.
|
||||
- ⛔ **`Pfish-6` IS RETIRED from the gateway** (2026-09-10). Its seat now serves the A4B; a caller
|
||||
asking for `Pfish-6` gets an explicit 400 "Invalid model name", not a substitution. Audited first:
|
||||
**0 of 17 LiteLLM keys** scoped it, so nothing was orphaned. The artifact stays on disk at
|
||||
`/tank/aimodels/erp-tune-v6-nvfp4a16` and is the rollback target — `cp
|
||||
.env.pfish6.bak-20260910 .env && docker compose up -d` in `/opt/docker/compose/erp-seat` restores
|
||||
it in ~4 min.
|
||||
- The `char-rp` family is now **`char-rp` (stock Gemma-4 26B-A4B, :8016) / `char-rp-fast` (MeroMero
|
||||
A4B, :8021) / `char-rp-reasoning` (Dark-Scarlett-27B, :8019)**. All three verified working after
|
||||
the change. ⚠ The `char-rp` comment block in `stacks/litellm/conf/config.yaml` still describes its
|
||||
seat as MeroMero-v2; that has been stale since the 2026-08-24 swap to stock Gemma-4. Not fixed.
|
||||
- **Clean single-variable ladder**, all on the unwrapped corpus (sha `77f37057b2782e49`), seed 4919,
|
||||
159 steps, 5,210,112 tokens: **0.6B 3.329 · 1.7B 3.018 · 4B 2.814**. Deltas 0.311 then 0.204.
|
||||
- ⚠ **4B is the only rung that OVERFITS inside one epoch** (min 2.814 at ~step 75, ends 2.825), so
|
||||
its shipped `adapter/` is NOT the best weights. Arms were re-cut from `checkpoint-75`.
|
||||
- **Instruct probe answered its question: voice and instruction-following COEXIST.** `Qwen3-4B`
|
||||
instruct + the same corpus → curly quotes 16/18 (identical to 4B-Base), task-leak 0/18, on-beat
|
||||
10/10 through the chat template. Cost is length discipline (in-band 10/10 → 6/10). Held-out 2.908.
|
||||
- ⛔ **The frozen adjudication has NEVER RUN** and still cannot: the rule needs a **seed-to-seed
|
||||
spread** that does not exist (seed 2 was killed deliberately). One extra seed is ~50 min at 1.7B.
|
||||
- ⛔ **Beat→paragraph does NOT work on any completion carrier** — 10 formats × 3 seeds = 30 samples,
|
||||
zero that render the beat. An unadapted instruct model does it 10/10. Booths below.
|
||||
- Booths (all on the link board): `babybronte-voice`, `babybronte-1p7b`, `babybronte-4b`,
|
||||
`skaldsong-beats`.
|
||||
|
||||
### Fleet
|
||||
|
||||
- **althing 3.6.2 everywhere it matters** — post office (nh3-docker) + both heralds. ⚠ Only **two**
|
||||
nodes run a herald (nh3-dev, nh3-extdev); the "seven boxes" in the rollout instruction do not
|
||||
participate. Drop-count BEFORE baselines: nh3-dev **27 in 18 sessions**, nh3-extdev **0**. Flat
|
||||
after = fix confirmed; rising = a second source, forseti wants to hear.
|
||||
- **Beszel wired by another agent** (`eb75713`) after my handoff at `/tmp/beszel.md`. Alerts now reach
|
||||
althing. ⚠ When I last looked, `EXTRA_FILESYSTEMS` held container-internal paths with no matching
|
||||
bind-mount, so `/tank` may still not be sampled, and alert links pointed at `localhost:8090`.
|
||||
Re-verify rather than assume it is closed.
|
||||
- Disk reclaimed today: **466 GB** (qwopus + huihui 122B bf16, operator-directed) and **107.8 GB**
|
||||
Docker on ana-ml2 (dated filter; root 77%→58%). ⚠ Image prune took `vllm-qwopus35-122b`'s TAG but
|
||||
not its layers — the rollback container still starts, but re-tag if you want the name back.
|
||||
27.39 GB of build cache remains, newer than the 168h window used.
|
||||
- ⚠ **`/tank/aimodels/heretic2-nvfp4-work` is a reclaim candidate that holds a load-bearing 4.4 MB
|
||||
file** — `production_calib_512.jsonl`, the calibration set every in-house quant references. Copied
|
||||
to `ana-docker:/opt/docker/conf/quant-calib/` (a restic source) with a README. NOTE: A16 quants are
|
||||
data-free and ignore it, so its loss would only bite activation-quantized schemes.
|
||||
- **mog-sec (`sec`/`sec-reasoning`, ana-ml2 GPU0 `:8019`) SETTLED after five crashes** at
|
||||
`MOG_MAX_MODEL_LEN=163840` + `MOG_KV_CACHE_MEMORY=17697765376` + `MOG_MAX_NUM_BATCHED_TOKENS=4096`
|
||||
+ util 0.50. **17+ hours, RestartCount 0.** Over-limit requests now return a clean 400 instead of
|
||||
killing the engine. Probe committed at `services/mog-sec-tuning/deep_ctx_probe.py`.
|
||||
- **`char-rp-fast` (MeroMero A4B) LIVE on ana-ml2 `:8021`**, 21+ hours, RestartCount 0. `Pfish-6`
|
||||
retired from the gateway (0 of 17 LiteLLM keys scoped it). Rollback `.env.pfish6.bak-20260910`.
|
||||
- ⛔ **The v2-31B dense quant has still never had a serve test** — GPU1 has ~19 GB free against
|
||||
19.5 GB of weights, so it needs a live seat displaced. Operator's call.
|
||||
- ⚠ **`/tank/aimodels/heretic2-nvfp4-work` holds the load-bearing `production_calib_512.jsonl`** —
|
||||
copied to `ana-docker:/opt/docker/conf/quant-calib/` (a restic source). A16 quants ignore it.
|
||||
- ⚠ `graphify-out/GRAPH_REPORT.md` has been dirty in the working tree all session (regenerated by
|
||||
another agent); left uncommitted deliberately.
|
||||
|
||||
## Recent decisions
|
||||
|
||||
@@ -383,26 +342,6 @@ _As of 2026-09-10 10:25 PT._
|
||||
|
||||
- `[2026-08-28]` **nh3-dev's three OOM events attribute to CLAUDE CODE — and the "no kernel evidence" was a permissions artifact.** journald was persistent all along; `journalctl` silently shows only your own messages outside `adm`. Single CC sessions measured at 5.4-18.4 GB, so 27 GB is 3-4 mature sessions, not the ~66 a 408 MB estimate implies. sysstat + atop now instrument the ramp. → `persistent-memory.d/2026-08-28-nh3-dev-oom-attribution.md`
|
||||
|
||||
- `[2026-08-27]` **Run 3 gated: the preregistered rule PASSED and a k=25 follow-up found a 44pp self-harm guardrail collapse — DO NOT SERVE.** A pooled preserve-list test structurally cannot see a single-axis collapse. → `persistent-memory.d/2026-08-27-run3-gate-safety-regression.md`
|
||||
|
||||
- `[2026-08-27]` **The corpus mix was specified in a unit the optimiser never sees** — 45.8% dialogue by CONTEXT, 24.2% by LOSS. Harness now leads with loss share and calls context a memory budget (`dd5a12e`). → `persistent-memory.d/2026-08-27-mix-specified-in-the-wrong-unit.md`
|
||||
|
||||
- `[2026-08-27]` **Dose-response: benefit and damage are ONE direction in weight space** — every axis monotone in scale, no knee. The merge-back cannot separate them; vLLM cannot LoRA-serve this MoE at all. → `persistent-memory.d/2026-08-27-dose-response-entanglement.md`
|
||||
|
||||
- `[2026-08-27]` **Anaheim tripped a power breaker; four guests including the NAS had `onboot` unset and never came back.** Fixed with dependency ordering — ana-nas order=1,up=45 ahead of the databases. ⚠ **ONE CIRCUIT FEEDS THE WHOLE RACK including the firewall serving the public IP** (operator) — so ana-gw, ana-wg and every BMC go down with the load, and there is NO remote management path to Anaheim during a power event. → `persistent-memory.d/2026-08-27-anaheim-breaker-and-onboot-gap.md`
|
||||
|
||||
- `[2026-08-27]` **A transport failure that enters a measurement as a VALUE looks like whatever you hoped to find.** heid's lost panel arms found a live defect in brokkr's `t4_dissect` an hour later. → `persistent-memory.d/2026-08-27-empty-response-as-a-datum.md`
|
||||
|
||||
- `[2026-08-27]` **Run 3c authorised (lr 20x cut, single variable) and then HELD by the operator after the breaker trip.** Config built and validated at `/tank/erp-tune/run-03c.json`; `save_steps` made configurable in the harness (`0a6bd2e`) because the first launch lost 80 steps with no checkpoint. Tracking surface: commit `0a6bd2e` + that config path. **Relaunch is one command once power is triaged.**
|
||||
|
||||
- `[2026-08-27]` **An event report with no timestamp is a claim about "now" — and it manufactured a launch that never happened.** brokkr reconstructed a phantom third 3c launch because my 23:03 report narrated a 21:07 kill in the present tense. Every fact in it was true; it was unreadable in sequence. → `persistent-memory.d/2026-08-27-run3c-launch-count-reconstruction.md`
|
||||
|
||||
- `[2026-08-27]` **`save_steps` was hardcoded at 100 in the harness** — a claimed provenance entry the run could not have honoured. Made configurable, default unchanged (`0a6bd2e`, 242 tests green). Caught by checking the config carried the change rather than trusting that it had been made.
|
||||
|
||||
- `[2026-08-27]` **Six defects in run 3's staged build, none of which would have errored** — a dialogue-only survivor list that would have silently dropped 96% of the corpus, an impersonation mask not subsumed by the low-quality mask, kvasir unbounded at 67.8% of context, a `save_pretrained` config-key drop that made the merged model unservable, and the mix-unit error. Every one produced a plausible completed run. Full record `/tank/erp-tune/recipe-r3/RUN-03-BUILD-NOTE.md`.
|
||||
|
||||
- `[2026-08-27]` **The 18 unpushed eitri-smithy commits are pushed** — run 3's `harness_commit 9d27b4fe` now resolves off-box, verified by fetching into a fresh empty repo rather than trusting the push output. ⚠ **HTTPS push 403s for every gitea token including site-admin; SSH works.** Untracked `__pycache__` (`894fbe8`) because a tracked `.pyc` dirtied the tree and would have stamped `harness_dirty_at_launch: true`.
|
||||
|
||||
- `[2026-08-25]` **Run 2's base is an OPEN OPERATOR DECISION, deliberately not staged** — four options with materially different safety postures, detailed in Current state. Tracked at althing thread `01M0WQ8W5574KMEVCHCEKEXNS5`. ⚠ Do not let it get filed as a config knob; it is a reversal of the trainee-selection decision.
|
||||
|
||||
- `[2026-08-25]` **Fused MoE kernel path — DEFERRED, tracked at park `fused-moe-kernel-path-for-gemma-4-moe-training` (id 47).** Operator: "note the fused MoE kernel for round two… if we nail it soon, the math has us wanting to restart the run anyway." Training MFU is **8.6%** (27.1 of a benchmarked 313.8 TFLOPS) because `transformers` runs the Gemma-4 experts in a Python loop — 128 experts × 30 layers, ~11,500 iterations per step under gradient checkpointing. ⚠ **The same fused 3-D expert layout that made bitsandbytes skip 88.5% of the model is exactly what a grouped GEMM wants** — the format is good for storage and for fused kernels, and hostile only to naive iteration. Two fixes: `group_by_length` (−29.9% compute, free, but breaks the seeded order manifest and re-opens a batch-composition call brokkr already made) and a grouped-GEMM/compiled MoE forward (the remaining ~10×). **Not applied to the live run** — restarting mid-flight to change batch ordering was judged a bad trade at step ~50 of 1,312.
|
||||
@@ -423,7 +362,7 @@ _As of 2026-09-10 10:25 PT._
|
||||
|
||||
- `[2026-08-09→10]` **dots.tts (rednote-hilab) TTS burn-in on irv-ml1 + canonical voice corpus built (`voices/`).** Operator-directed eval to potentially replace chatterbox-fast. **dots.tts VERIFIED real** (canonical HF ns `dots-studio/`, `rednote-hilab/dots.tts-*` redirects there; Apache-2.0; PyPI `dots.tts` 0.2.1; 2B continuous-AR = semantic enc + Qwen2.5-1.5B LLM + flow-matching acoustic head over 48kHz AudioVAE; zero-shot clone from wav+transcript). **Runs on Ampere 3090** (sm_86, bf16, no fp8 dep); **optimized RTF 0.22** at num_steps=10 (`from_pretrained(..., optimize=True)` CUDA graphs — raw unoptimized was 1.21), **~6GB VRAM**, 48kHz, streams (`generate_stream`). Venv+cache at `irv-ml1:/home/lkraven/dots-tts` (~10GB). **Operator design calls:** SGLang Omni serving (OpenAI `/v1/audio/speech`), transcribe-refs-first, `soar` variant. ⚠ Omni serves soar but its continuous-batching + streaming opts are **mf-only** (soar = single-request) — non-issue for ratatoskr's single-consumer RP surface. **KEY FINDING — dots is highly sensitive to an accurate AND sentence-bounded reference transcript:** mismatched transcript → 0.16s collapse; over-long/messy transcript → reference-audio BLEEDS as an output prefix; mid-clause trim → dangling-word leak (glados "we'll", emmie "And,"). RECIPE (baked into `voices/derive.py`): trim ref to a clean ~6–10s clip ending on a sentence boundary + accurate transcript of exactly that clip. **CANONICAL VOICE CORPUS** stood up in eshpfi `voices/` (operator idea): engine-agnostic `canonical/<v>.wav` + `transcripts/<v>.txt` → per-engine ref sets DERIVED by `derive.py` reading `engines.yaml` profiles (dots/chatterbox/zonos); canonical wavs git-tracked (small/curated), `derived/` gitignored. **4 voices optimized + verified CLEAN for dots: donut, glados, emmie, miranda** (glados canonical is low-SR 16kHz — flagged upgrade candidate). ⚠ GPU GOTCHA: irv-ml1 native CUDA orders **A6000=device0** (ComfyUI-full) — pin the 3090 with `CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=0`; and `PYTORCH_CUDA_ALLOC_CONF=expandable_segments` CONFLICTS with `optimize=True` CUDA graphs (curr_block error). Booths: `dots-vs-chatterbox`, `dots-voices-optimized`. **SHIPPED 2026-08-10:** operator A/B verdict "dots is very good" → containerized as a **thin FastAPI wrapper over DotsTtsRuntime** (chosen over SGLang Omni — Omni's batching is mf-only, unneeded for ratatoskr's single consumer; wrapper is SERIALIZED one-gen-at-a-time via a threading.Lock, Omni+mf = parked API-compatible escalation if multi-consumer ever lands). **LIVE on irv-ml1:8198** (`local/dots-tts:v1`, OpenAI `/v1/audio/speech` + `/health` + `/v1/voices`, container healthy, both stream + non-stream verified CLEAN, 4 voices donut/glados/emmie/miranda) alongside chatterbox :8197 (nothing repointed). Stack = `stacks/dots-tts/` (Dockerfile/app.py/compose/.env.example/README). ⚠ CONTAINER GOTCHA: `optimize=True` (torch.compile/inductor/triton) needs a **C compiler at RUNTIME** — slim image must `apt install build-essential` or model-load dies "Failed to find C compiler" (host venv had gcc ambient, masking it); persist `TORCHINDUCTOR_CACHE_DIR` to a mounted dir or every restart re-JITs ~5min. Corpus home = eshpfi `voices/` (operator ruled keep-here). **REMAINING: ratatoskr client cutover** to :8198 `/v1/audio/speech` (Phase-2 tail, peer-coupled — draft the ask). [[reference_chatterbox_fast_repo]] [[reference_zonos_tts_stack]] [[reference_verify_hf_repo_ids_before_pull]]
|
||||
|
||||
_275 older entries archived to archival-memory.md._
|
||||
_285 older entries archived to archival-memory.md._
|
||||
|
||||
_159 older entries archived to archival-memory.md._
|
||||
|
||||
|
||||
Reference in New Issue
Block a user