diff --git a/persistent-memory.md b/persistent-memory.md index d7f0171..eeb0580 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -114,7 +114,7 @@ no longer deployed sidecars here. See Recent decisions.) - **šŸ”“ā†’šŸŸ¢ ESH OUTAGE 2026-08-19 — esh-pve hard-froze 03:34, ~4.5h, whole house lost DNS.** Presented as "wifi/routing issues"; internet was healthy throughout. Cause: `esh-userland` (VLAN 10, the `PVC` SSID) handed out **one** resolver, `10.0.50.45` (AdGuard on `esh-docker-vm`) — cross-VLAN, no secondary — and its hypervisor died. **Recovered by hand power-cycle; all VMs, cluster quorum and DNS restored.** Two fixes landed: gateway `10.0.10.1` added as secondary DNS on esh-userland (operator-approved, first confirmed WRITE on the ESH UDM key), and **`softdog` → `iTCO_wdt` hardware watchdog owned by systemd** (`playbooks/esh-pve-hardware-watchdog.yaml`, idempotent, verified armed) so a repeat self-recovers in 60s. **VM 102 pinned off** (`onboot: 0`) as the passthrough suspect. **ā³ OPEN:** (a) the watchdog is armed but **has not been proven to fire** — needs a deliberate wedge to confirm; (b) AMT/vPro still unusable until an onboard **RJ45** is cabled (the MS-01 is SFP+-only on the network and AMT cannot ride it); (c) kernel `6.8.12-16` rollback held in reserve if it freezes again. → `persistent-memory.d/2026-08-19-esh-pve-freeze-dns-spof.md` -_As of 2026-08-20 23:30 — **the Heretic-300 session** (see the šŸ”“ entry above; full epic + roadmap in `persistent-memory.d/2026-08-20-heretic-300-epic.md`). Landed tonight: 300-trial Heretic search → **8/100 refusals @ KL 0.0136**, hand-verified coherent, beating the `absolute-heresy` bar 3.6Ɨ; MTP head grafted back after PR #317 silently dropped it; **NVFP4 quant COMPLETE and verified** (`qwen38-27b-coldfusion-h300-nvfp4-mixed`, 21 GB, 1968 tensors, `re:^mtp.*` re-injected); **āœ… GEN-SEAT CUTOVER DONE 23:05** — 7/7 aliases green, vision intact, MTP 59.7%, 118.4 tok/s, KV 401,550 tok / 1.53Ɨ (baseline 403k/1.54Ɨ). ā­ **Only PPL remains** — blocked on a spec-decode-free probe seat (both GPUs ~96% full). ⚠ **The self-harm guardrail is GONE on this build** — operator is handling restoration directly and does not want dwarf analysis on it; four-dwarf panel formally STOOD DOWN. Earlier in the day: KL harness (28.4Ɨ selectivity on L35), esh-nas-pve corruption resolved (171 files recoverable, error scrub clean, pool ONLINE), `/tank` on ana-ml2 still DEGRADED 69 days (parked, operator going to colo). ⚠ MANY commits unpushed._ +_As of 2026-08-21 00:35 — **the Heretic-300 session, and its reversal** (see the šŸ”“ entry above; full epic + roadmap in `persistent-memory.d/2026-08-20-heretic-300-epic.md`). Landed tonight: 300-trial Heretic search → **8/100 refusals @ KL 0.0136**, hand-verified coherent, beating the `absolute-heresy` bar 3.6Ɨ; MTP head grafted back after PR #317 silently dropped it; **NVFP4 quant COMPLETE and verified** (`qwen38-27b-coldfusion-h300-nvfp4-mixed`, 21 GB, 1968 tensors, `re:^mtp.*` re-injected); gen-seat cutover done 23:05 and every gate passed (7/7 aliases, vision intact, MTP 59.7%). **ā›” THEN ROLLED BACK 00:28 ON OPERATOR DIRECTIVE.** Lobe surfaced an unterminated-`` leak; measurement showed the **Cold-Fusion base** carries 18.5% first-token mass on it and our abliteration only added +3.7 — so the operator called it: **abandon h300 AND the Cold-Fusion base**, since no rollback inside that family escapes the leak. Gen seat is back on **`qwen38-27b-heresy-nvfp4-mixed`** (KV 403,065 / 1.54Ɨ = its exact documented baseline; 7/7 aliases; vision intact; `` not even in the top-20 at <0.002 vs Cold-Fusion's 0.185). **The methodology is the deliverable and it stands** — see the Heretic-300 entry. ⚠ **The self-harm guardrail is GONE on this build** — operator is handling restoration directly and does not want dwarf analysis on it; four-dwarf panel formally STOOD DOWN. Earlier in the day: KL harness (28.4Ɨ selectivity on L35), esh-nas-pve corruption resolved (171 files recoverable, error scrub clean, pool ONLINE), `/tank` on ana-ml2 still DEGRADED 69 days (parked, operator going to colo). ⚠ MANY commits unpushed._ - **🟢 FLEET `.internal` DNS — LIVE 2026-08-19.** `..internal`, sites `ana`/`esh`/`nh3`. `dns/internal.yaml` is the source of truth; `scripts/dns-sync.py` reconciles the three AdGuard resolvers (diff → prompt → apply, idempotent). 42 names resolving from all three sites. **Colo got its first resolver ever** (`stacks/adguard-ana/`, API on **8053** not 8080, no blocklists by design) — before this, ana-docker resolved straight against `1.1.1.1`. Auth = a dedicated `infra-ops` AdGuard user, password vaulted `nh3-dev/adguard-infra-ops-password`. **ā³ TWO OPEN, both operator's to schedule:** (a) colo hosts still point at `1.1.1.1` so they do not yet *use* the new resolver — repointing a site's DNS is a separate change; (b) the static-v6 convention (each server at its site's `/64` with low bits echoing the v4 octet, `esh-docker-vm` → `…::45`) is **proposed, not ruled on**. The `v6:` column is empty and correct — no fleet host has a global v6 address yet. → `persistent-memory.d/2026-08-19-fleet-internal-dns.md` @@ -122,7 +122,16 @@ _As of 2026-08-20 23:30 — **the Heretic-300 session** (see the šŸ”“ entry abov - **🟢 HOMEPAGE — cleaned + themed (Australis Skyfall).** Three real defects fixed (UltraSeedbox on all tabs, Uptime Kuma double-rendered, fiction column counts), AI tab reordered by clickability, then themed from the operator's Skyfall handoff bundle with a background generated by **Arbo** (`t2i-ui-background`, job `13f0891f4e42`). ⚠ **After any recreate the tab bar/wallpaper/i18n vanish for up to ~an hour and then return on their own — do not chase it.** ⚠ CSS is served per-request: a theme change needs a **reload**, not a recreate, and candidate CSS can be injected live via Playwright for seconds-long iteration. → `persistent-memory.d/2026-08-19-homepage-skyfall-theme.md` -- **šŸ”“ GEN SEAT DEFECT 2026-08-21 — the h300 build emits an UNTERMINATED `` into `content`, ~27% of the time, on any temp>0 alias. Operator-reported via Lobe ("sends CoT, never completes the turn").** +- **ā›” COLD-FUSION ABANDONED — GEN SEAT ROLLED BACK TO `heresy` 2026-08-21 00:28 (operator directive).** The operator's call, made in advance of the result: *"If it's the base, we abandon h300 AND the base and chalk it up to a very powerful and useful learning experience. Our heretic methodology will definitely translate in the future."* The measurement came back **base**, so the condition fired. + - **LIVE GEN SEAT = `/tank/aimodels/qwen38-27b-heresy-nvfp4-mixed`** (`MuXodious/Qwen3.8-27B-absolute-heresy` through our mixed NVFP4+FP8 recipe). Restored from `.env.bak-coldfusion-L35-20260820`; the h300 env is preserved at `.env.bak-h300-abandoned-20260821`. Verified: healthy, **KV 403,065 tok / 1.54Ɨ — its exact documented baseline**, 7/7 aliases 200, vision intact (red circle / blue square / green rectangle). + - **ā˜… The clincher: `` is not even in heresy's top-20 first tokens (<0.002), against Cold-Fusion's 0.185.** That is a >100Ɨ gap — the two families are categorically different on this axis, and it is why no rollback *inside* Cold-Fusion (L35 or stock) would have helped. + - **WHY NOT just apply the `chat_template_kwargs` fix?** It worked (8/30 → 0/30) but it is a **workaround for a base the operator no longer wants**: it forces `gen` to become a thinking deployment to paper over a finetune whose whole purpose is reasoning compression. Rolling back removes the defect at the root and restores a build already operator-confirmed "working very well" in real multi-turn use (2026-08-17, coherent through 60k tokens). + - **COST, stated plainly:** we give up **8/100 refusals** (h300) and go back to **29/100** (heresy's own bar) — a 3.6Ɨ regression on the refusal axis, which was the entire point of the Heretic-300 run. Also lost: the in-band-vs-pristine MTP comparison stays academic. **Accepted deliberately** — a seat that breaks the operator's daily client is worth less than one that occasionally refuses. + - **ā˜… WHAT CARRIES FORWARD (the operator's point, and it is right).** None of the Heretic-300 learning was in the Cold-Fusion weights. Still valid and model-agnostic: `direction_scope=0` beats per-layer on a merged base (8/100 vs 52/100); aggression is not the lever (r=āˆ’0.561); **PR #317 silently drops the MTP head on save** — always diff tensor keys; the MPOA/sink-screen reasoning; `graft_mtp.py`, `kl_divergence.py`, `catatonia_gate.py`, `heretic_export.py`; **a pristine MTP graft accepts as well as an in-band edit (59.7% vs 59.1%)**; and the new `think_prior.py` probe. **The methodology is the deliverable; the base was the wrong substrate.** + - **ā˜…ā˜… NEW ACCEPTANCE GATE, earned here — add a FORMAT-COMPLIANCE check to every abliteration/base evaluation, and run it on the STOCK BASE BEFORE spending a GPU-week.** `think_prior.py` on the stock candidate is a ~10s CPU measurement that would have disqualified Cold-Fusion **before** the 300-trial study ever ran. Screen candidate bases for it. (Related: Heretic's objective has no format term at all — same blindness that removed the self-harm guardrail.) + - **⚠ NOTHING DELETED.** `h300-nvfp4-mixed`, `h300-mtp-bf16`, `coldfusion-L35-*`, `coldfusion-bf16`, the 300-trial Optuna journal and `catatonia-T260.json` are all intact on disk. "Abandon" = stop serving, not `rm`. + +- **⚪ GEN SEAT DEFECT 2026-08-21 — RESOLVED BY THE ROLLBACK ABOVE. The h300 build emitted an UNTERMINATED `` into `content`, ~27% of the time, on any temp>0 alias. Operator-reported via Lobe ("sends CoT, never completes the turn").** - **Mechanism.** With `enable_thinking:false` the chat template appends a **pre-closed** `\n\n\n\n` to the prompt (jinja L165-166). The h300 model **opens a fresh `` anyway and never closes it** — verified raw: `has : False`, `finish_reason: stop`, reasoning *and* answer in one `content` blob starting `Ok, let's figure this out:`. vLLM's `qwen3` reasoning parser can't catch it: the prompt already closed the block, so the parser isn't in reasoning state and the tag is just text (`reasoning_content` empty, `reasoning_tokens: 0`). **The client is blameless** — Lobe correctly treats an unterminated `` as still-thinking, so it renders an endless thought bubble and never shows the answer. - **ā˜… It is a SAMPLING event, and the trigger is TEMPERATURE — not presence_penalty.** n=12 per arm on the reproducer: `pp 1.5` → 4 leaks, `pp 0.0` → 4, `pp 0.5` → 3 (all the same), **`temperature 0` → 0**. āš ļø **This FALSIFIES the standing "presence_penalty 1.5 is the first dial to move" hypothesis** recorded in the litellm config comment and by the operator 2026-08-16 — it is not this bug's cause. Leave that dial alone for this symptom. - **Blast radius = exactly the two temp-0.7 aliases.** `gen` and `summarizer-large` leak (~17-27%); `summarizer`, `classifier`, `image-judge`, `qwen-image-bench` are all **temp=0 and clean at 0/12** — so **nevermore's `summarizer` path is NOT affected**. `gen-reasoning` doesn't leak (its think block is legitimately open) but shows the *other* symptom, empty `content`, at ~1/12. @@ -142,7 +151,7 @@ _As of 2026-08-20 23:30 — **the Heretic-300 session** (see the šŸ”“ entry abov - **FIX, validated n=30 over 4 prompt types + a 3-turn conversation:** `chat_template_kwargs: {enable_thinking: true, reasoning_effort: low}` on `gen` → **0/30 leaks** (current config: **8/30**, worst on prose 4/6). Give the model a legitimately open `` and it closes it properly, the parser does its job, `content` comes out clean. Cost ~+27% completion tokens (257 vs 202 avg) and a residual **1/30 empty-content**. Semantic change: `gen` stops being a non-thinking deployment — **operator's call, not applied.** - **ā˜… PROCESS LESSON: the 7/7 alias smoke test structurally CANNOT catch this.** Trivial prompts ("Reply with exactly: OK-gen") never invite reasoning, so they never sample the leaking token. Same shape as the compose file's own warning that single-turn probes missed the xhigh budget bug. **Probe with a reasoning-inviting prompt at n≄12, and grep the raw `content` for `` — never just check HTTP 200.** The tell was sitting in my own `eval_coldfusion_h300.json` output the night of the cutover and I read past it. -- **🟢 GEN SEAT — LIVE = COLD-FUSION HERETIC-300 (cut over 2026-08-20 23:05).** `GEN_MODEL=/tank/aimodels/qwen38-27b-coldfusion-h300-nvfp4-mixed` on `gen-seat`/`vllm-gen`, ana-ml2 GPU0 `:8015`, served-name unchanged (`qwen3.8-27b-uncensored` / `-thinking`) so all 7 LiteLLM aliases route without a gateway edit. **Verified end to end:** healthy in 5.5 min; KV **401,550 tok / 1.53Ɨ** (baseline 403k/1.54Ɨ — within noise); **7/7 aliases green** through LiteLLM; **VISION INTACT** (correctly enumerated colour/form/position of 3 shapes — the surface that had never been exercised after abliteration → MTP-dropping export → graft → quant); **MTP acceptance 59.7% median @ 118.37 tok/s** (`bench/mtp_coldfusion_h300.json`) — statistically identical to L35's **59.1% @ 118.71** on the same instrument, so **the roadmap's "~47% for a pristine graft" prediction was WRONG — a pristine graft accepts as well as an in-band one.** ⚠ A single long-prose sample read 47.5%; the 8-run spread is 49.0–65.4%, so **one sample cannot characterize acceptance** — always use `quickbench.py`. Deterministic quality gens all correct (heat-pump, primes=77, 14:20→17:05, prose); abliteration survival **4/4 compliance**. **ā³ PPL NOT MEASURED** — `eval_quality.py` aborts with "prompt_logprobs look uniform" because the seat runs `--speculative-config`; the documented workaround is a spec-decode-free probe seat (`bench/serve_probe.sh`, :8017), and **there is no VRAM for one** (GPU0 5.2 GB free, GPU1 2.4 GB free). Comparison target = heresy's **6.910 mean / 5.625 median**. **ROLLBACK (one line):** `sudo cp /opt/docker/compose/gen-seat/.env.bak-pre-h300-20260820 /opt/docker/compose/gen-seat/.env && cd /opt/docker/compose/gen-seat && sudo docker compose up -d vllm-gen` → back to `qwen38-27b-coldfusion-L35-nvfp4-mixed`. **Do NOT delete** `-L35-nvfp4-mixed` or `qwen38-27b-heresy-nvfp4-mixed`. ⚠ **The self-harm guardrail is GONE on this build** (operator's own next work item). Also normalized the quant dir from root:0600 to `llmuser:llmuser` 0664 to match every other model dir. **Confirmed the right weights are mounted on TWO discriminating views** — mtime and a 64 MB head-hash both match h300 and differ from L35; `config.json` sha256 is **identical** across both builds and therefore useless as a discriminator (it carries no weight-specific content — don't reach for it again). +- **⚪ GEN SEAT — h300 WAS LIVE 2026-08-20 23:05 → 2026-08-21 00:28, then ABANDONED (see the ā›” entry above). Historical record of that window:** `GEN_MODEL=/tank/aimodels/qwen38-27b-coldfusion-h300-nvfp4-mixed` on `gen-seat`/`vllm-gen`, ana-ml2 GPU0 `:8015`, served-name unchanged (`qwen3.8-27b-uncensored` / `-thinking`) so all 7 LiteLLM aliases route without a gateway edit. **Verified end to end:** healthy in 5.5 min; KV **401,550 tok / 1.53Ɨ** (baseline 403k/1.54Ɨ — within noise); **7/7 aliases green** through LiteLLM; **VISION INTACT** (correctly enumerated colour/form/position of 3 shapes — the surface that had never been exercised after abliteration → MTP-dropping export → graft → quant); **MTP acceptance 59.7% median @ 118.37 tok/s** (`bench/mtp_coldfusion_h300.json`) — statistically identical to L35's **59.1% @ 118.71** on the same instrument, so **the roadmap's "~47% for a pristine graft" prediction was WRONG — a pristine graft accepts as well as an in-band one.** ⚠ A single long-prose sample read 47.5%; the 8-run spread is 49.0–65.4%, so **one sample cannot characterize acceptance** — always use `quickbench.py`. Deterministic quality gens all correct (heat-pump, primes=77, 14:20→17:05, prose); abliteration survival **4/4 compliance**. **ā³ PPL NOT MEASURED** — `eval_quality.py` aborts with "prompt_logprobs look uniform" because the seat runs `--speculative-config`; the documented workaround is a spec-decode-free probe seat (`bench/serve_probe.sh`, :8017), and **there is no VRAM for one** (GPU0 5.2 GB free, GPU1 2.4 GB free). Comparison target = heresy's **6.910 mean / 5.625 median**. **ROLLBACK (one line):** `sudo cp /opt/docker/compose/gen-seat/.env.bak-pre-h300-20260820 /opt/docker/compose/gen-seat/.env && cd /opt/docker/compose/gen-seat && sudo docker compose up -d vllm-gen` → back to `qwen38-27b-coldfusion-L35-nvfp4-mixed`. **Do NOT delete** `-L35-nvfp4-mixed` or `qwen38-27b-heresy-nvfp4-mixed`. ⚠ **The self-harm guardrail is GONE on this build** (operator's own next work item). Also normalized the quant dir from root:0600 to `llmuser:llmuser` 0664 to match every other model dir. **Confirmed the right weights are mounted on TWO discriminating views** — mtime and a 64 MB head-hash both match h300 and differ from L35; `config.json` sha256 is **identical** across both builds and therefore useless as a discriminator (it carries no weight-specific content — don't reach for it again). - **⚪ PRIOR GEN SEAT — `absolute-heresy` 2026-08-17 (validated, promoted; superseded by L35 then h300 on 2026-08-20).** Live gen = `/tank/aimodels/qwen38-27b-heresy-nvfp4-mixed` — **MuXodious/Qwen3.8-27B-absolute-heresy** (Heretic v1.4.0 + **SOMPOA**, trial T377, pin `c2374593`) put through our own mixed NVFP4+FP8 recipe. Chosen because it beats the incumbent on **both** axes at once: author refusals 2/101 vs 12/100, first-token KL 0.0759 vs 0.1191. **Gate (probe :8017, pinned nightly, seat-matched flags): MTP 47.2% (inc. 48.2%), decode 103.5 tok/s (96.4), prefill 6618/5403 @6.7k/27k (6334/5085), PPL 6.910 (7.059 — 2.1% BETTER), surface 6/6, abliteration 4/4, and 0/55 refusals on our battery-instruct arm with ZERO EMPTY (no catatonia).** ⚠ speed deltas are **image-confounded** (probe on the pinned nightly, incumbent numbers from an earlier image) — read as "not worse", not a clean win. All 7 LiteLLM aliases verified end-to-end; GPU0 at 91.3/97.9 GB with meromero healthy (more headroom than the old build's 96.8). ⚠ **RC1, 2 days old, ~348 downloads.** **Operator-confirmed "working very well" in real use 2026-08-17**, same evening as the cutover — the signal the synthetic gates structurally cannot give (multi-turn degeneration is stochastic; four synthetic tests once validated three non-fixes). Not yet the 60k-token bar the prior seat cleared, so **keep watching and do NOT delete the rollback weights yet**. **ROLLBACK:** `sudo cp /opt/docker/compose/gen-seat/.env.bak-heresy-20260817 /opt/docker/compose/gen-seat/.env && cd /opt/docker/compose/gen-seat && sudo docker compose up -d vllm-gen`; incumbent weights UNTOUCHED at `qwen38-27b-uncensored-nvfp4-mixed` — **do NOT delete** until this holds. Runbook `services/gen-seat-mixed-quant/RUNBOOK-heresy-swap.md`. diff --git a/services/gen-seat-mixed-quant/bench/think-leak/README.md b/services/gen-seat-mixed-quant/bench/think-leak/README.md index 1746126..076e2cb 100644 --- a/services/gen-seat-mixed-quant/bench/think-leak/README.md +++ b/services/gen-seat-mixed-quant/bench/think-leak/README.md @@ -105,3 +105,35 @@ Run it: ⚠ Abliteration and post-quant outputs are written root-owned `0600` and are unreadable to `llmuser`; normalize to `llmuser:llmuser 0664` first. The failure surfaces as a misleading `FileNotFoundError`, not a permission error. + +## Resolution — Cold-Fusion abandoned, seat rolled back to `heresy` (2026-08-21) + +The operator's call, made before the result was in: *"If it's the base, we abandon +h300 AND the base."* The measurement said base, so it fired. The gen seat is back +on `qwen38-27b-heresy-nvfp4-mixed`. + +The clincher is the same instrument, pointed at heresy: + +| build | P(``) at first token | +|---|---| +| Cold-Fusion stock | 0.1850 | +| Cold-Fusion L35 | 0.2048 | +| Cold-Fusion h300 | 0.2216 | +| **`heresy` (restored)** | **not in the top 20 — <0.002** | + +A >100x gap. The two families are categorically different here, which is exactly +why no rollback *inside* Cold-Fusion would have helped. + +Verified after the rollback, same probes as before: + +- `final_validate.py` — **0/30 leaks, 0 empty**, with the *existing* + `enable_thinking: false` config. h300 scored 8/30 on this same instrument. +- KV pool 403,065 tok / 1.54x — heresy's exact documented baseline. +- 7/7 gateway aliases 200; vision intact. +- **No LiteLLM config change was needed.** The `chat_template_kwargs` fix + developed above is left unapplied: it worked, but it was a workaround for a base + we no longer serve. + +**The gate this earns: run `think_prior.py` on a candidate's STOCK weights before +committing GPU time to it.** It is a ~10s CPU measurement, and it would have +disqualified Cold-Fusion before the 300-trial Heretic study ever started.