feat(gen-seat): cut over to Heretic-300 — 7/7 aliases, vision intact, MTP 59.7%
Live GEN_MODEL is now qwen38-27b-coldfusion-h300-nvfp4-mixed (ana-ml2 GPU0 :8015). Served-name left unchanged so all 7 LiteLLM aliases route without a gateway edit. Verification: KV pool 401,550 tok / 1.53x (baseline 403k / 1.54x) LiteLLM aliases 7/7 green vision 3/3 shapes, colour+form+position correct MTP acceptance 59.7% median @ 118.37 tok/s quality gens 4/4 correct abliteration 4/4 compliance PPL NOT measured (see below) The roadmap predicted ~47% acceptance for a pristine MTP graft versus L35's 59.1% in-band edit. Measured 59.7% on the same harness: there is no acceptance penalty, which removes the throughput argument for reimplementing MPOA. A single long-prose generation read 47.5% off the same counters -- below the 8-run minimum of 49.0% -- and would have "confirmed" the prediction by coincidence. Acceptance must be read from quickbench.py, never one sample. PPL is blocked on VRAM, not on the model: eval_quality.py aborts with "prompt_logprobs look uniform" under --speculative-config, and the probe-seat workaround needs ~22 GB while both cards sit at ~96% committed. Also normalizes the quant dir from root:0600 to llmuser:llmuser 0664 to match every other model dir, and records that config.json sha256 is byte-identical across the h300 and L35 quants and is therefore useless for confirming which weights are mounted (mtime and a head-hash are the discriminating views). Rollback is one line to .env.bak-pre-h300-20260820.
This commit is contained in:
@@ -95,6 +95,71 @@ the 1.5 cap, kernels centred ~41–42 vs population median ~49) — a basin, not
|
||||
Log-trial 262 sits 5.6% away in normalised parameter space: the same basin, **not**
|
||||
independent confirmation.
|
||||
|
||||
## ✅ CUTOVER + VERIFICATION `[2026-08-20 23:05]`
|
||||
|
||||
The gen seat is live on `qwen38-27b-coldfusion-h300-nvfp4-mixed`. Served-name unchanged
|
||||
(`qwen3.8-27b-uncensored`), so no gateway edit was needed. Healthy in 5.5 min.
|
||||
|
||||
| gate | h300 | comparator | verdict |
|
||||
|---|---|---|---|
|
||||
| KV pool | 401,550 tok / 1.53× | 403k / 1.54× baseline | within noise ✓ |
|
||||
| LiteLLM aliases | 7/7 green | — | ✓ |
|
||||
| **vision** | 3/3 shapes, colour+form+position correct | never before exercised | ✓ |
|
||||
| MTP acceptance | **59.7%** median | L35 in-band **59.1%** | ✓ — *prediction wrong* |
|
||||
| decode | 118.37 tok/s median | L35 118.71 | equal ✓ |
|
||||
| quality gens | 4/4 correct | — | ✓ |
|
||||
| abliteration survival | 4/4 compliance | — | ✓ |
|
||||
| PPL | **not measured** | heresy 6.910 / 5.625 | ⏳ blocked |
|
||||
|
||||
### ★ The ~47% prediction was wrong — a pristine graft accepts as well as in-band
|
||||
|
||||
Finding 4 / the roadmap predicted **~47%** for the pristine MTP graft, versus 59.1% for
|
||||
L35's in-band edit, and treated ~12 points of acceptance as the price of not having
|
||||
MPOA. Measured on the same instrument (`bench/quickbench.py`, 8×400 tok): **59.7%.**
|
||||
There is no acceptance penalty. This weakens — but does not kill — the case for
|
||||
reimplementing MPOA (roadmap item 6); its remaining justification is prior art and
|
||||
in-band elegance, **not ~12 points of throughput.**
|
||||
|
||||
⚠️ **A single sample cannot characterize acceptance.** One long-prose generation read
|
||||
**47.5%** by hand off the same `spec_decode_num_{draft,accepted}_tokens_total` counters
|
||||
quickbench uses — which is *below the 8-run min of 49.0%* and would have "confirmed" the
|
||||
47% prediction by coincidence. The 8-run spread is 49.0–65.4%. Always use the harness.
|
||||
|
||||
### ⏳ PPL is blocked on VRAM, not on the model
|
||||
|
||||
`eval_quality.py` aborts every passage with *"prompt_logprobs look uniform (median rank
|
||||
…); re-run against a seat started WITHOUT --speculative-config"* — the documented
|
||||
spec-decode logprobs trap (playbook; also banked in the `[2026-08-15]` mixed-requant
|
||||
entry). Passage 1's `ppl 2142183.691` is **garbage from that same cause, not a result** —
|
||||
do not quote it. The fix is the probe-seat path (`bench/serve_probe.sh`, :8017), which
|
||||
needs ~22 GB, and both cards are ~96% committed. Cheapest window is stopping
|
||||
`vllm-fablefusion-probe` (43.4 GB on GPU1, nearly idle).
|
||||
|
||||
### Traps that fired, and one that did not
|
||||
|
||||
- **`config.json` sha256 is BYTE-IDENTICAL between the h300 and L35 quants** — same
|
||||
architecture, same recipe, same ignore list, no weight-specific content. It is a
|
||||
**non-discriminating** probe; it neither confirms nor contradicts which weights are
|
||||
mounted. Discriminating views that *did* work: **mtime** (h300 22:52:44.351659025 vs
|
||||
L35 10:05:35.761199352) and a **64 MB head hash** (container == h300). Reached for the
|
||||
hash first out of "two views must agree" discipline; the right lesson is that a view
|
||||
must be *discriminating* before agreement means anything.
|
||||
- **The quant dir was written root-owned `0600`** while every other model dir is
|
||||
`llmuser:llmuser 0664`. vLLM runs as root so it would have loaded fine, but it also
|
||||
made the files unreadable to `infra-ops` (the L35 head-hash comparison failed on
|
||||
EACCES). Normalized to match convention.
|
||||
- **PR #317 did not re-fire**: 15 `mtp.*` tensors present in the index, all BF16, all in
|
||||
`model-mtp.safetensors`, `re:^mtp.*` in `quantization_config.ignore`, 333 visual
|
||||
tensors intact. `post_quant.py` did its job.
|
||||
|
||||
### Rollback
|
||||
|
||||
```
|
||||
sudo cp /opt/docker/compose/gen-seat/.env.bak-pre-h300-20260820 /opt/docker/compose/gen-seat/.env
|
||||
cd /opt/docker/compose/gen-seat && sudo docker compose up -d vllm-gen # -> L35
|
||||
```
|
||||
`-L35-nvfp4-mixed` and `qwen38-27b-heresy-nvfp4-mixed` both intact. **Do not delete.**
|
||||
|
||||
## 🗺️ ROADMAP — where to pick up
|
||||
|
||||
**Immediate (in flight at session end)**
|
||||
|
||||
Reference in New Issue
Block a user