feat(gen-seat): CPU-only MTP head check so the gate costs no second seat
Operator ruled the probe port for validation; runbook updated to match. The bf16 MTP acceptance gate is ~56 GB resident, which on a full 97.9 GB card means downing meromero-charrp as well as gen -- freeing gen's 0.43 (~42 GB) alone is not enough. Two seats down to answer one question. compare_mtp_head.py answers the common case for free. The Qwen3_5ForConditionalGeneration wrapper never loads the MTP head, so PEFT merges, Heretic runs, and llm-compressor passes all leave mtp.* as it came from the base. It hashes a candidate's 15 mtp.* tensors against the incumbent's grafted-verbatim head -- the one measured at 47.7% acceptance in production through this exact pipeline. Identical means the acceptance question is already answered; different means the head was edited and the real gate is warranted; missing means it was dropped. CPU only, reads just the shard holding mtp.*. The runbook states the residual risk plainly: an identical head proves the head is intact, not that the abliterated body still drafts well with it -- which the Stage-3 acceptance measurement on the 22 GB quantized build catches anyway.
This commit is contained in:
@@ -27,17 +27,59 @@ Detached container `hf-pull-heresy` on ana-ml2 →
|
||||
|
||||
---
|
||||
|
||||
## Stage 1 — bf16 MTP acceptance gate (DO THIS BEFORE SPENDING A QUANT)
|
||||
## Stage 1 — the MTP gate, WITHOUT downing a second seat
|
||||
|
||||
**Presence is not acceptance.** The incumbent's own provenance still reads
|
||||
`MTP acceptance UNVERIFIED` — the author grafted 15 `mtp.*` verbatim and we
|
||||
shipped on the assumption. It happened to measure 47.7% later; that was luck,
|
||||
not method. Playbook: *test MTP on bf16 FIRST, isolate before deleting.*
|
||||
shipped on the assumption. It measured 47.7% later; that was luck, not method.
|
||||
|
||||
Gate: **acceptance ≳40%**. Below that, the mixed quant is not worth the GPU
|
||||
time — the whole +18% decode story rests on MTP working.
|
||||
The literal reading of the playbook rule (*test MTP on bf16 FIRST*) is
|
||||
**expensive here**: bf16 is ~56 GB resident, and on a full 97.9 GB card that
|
||||
means downing `meromero-charrp` as well as `gen` — freeing gen's 0.43 (~42 GB)
|
||||
alone is not enough. Two seats down to answer one question.
|
||||
|
||||
bf16 needs ~56 GB, so this needs a GPU window on its own (see § GPU reality).
|
||||
### 1a. Free, CPU-only: is the head one we have already measured?
|
||||
|
||||
The wrapper class `Qwen3_5ForConditionalGeneration` **does not load the MTP
|
||||
head**, so PEFT merges, Heretic runs, and llm-compressor passes all leave
|
||||
`mtp.*` exactly as it came from the base. If the candidate's 15 `mtp.*` tensors
|
||||
are numerically identical to a head we have measured in production, the
|
||||
acceptance question is already answered.
|
||||
|
||||
```bash
|
||||
# runs in the vllm image (has torch + safetensors); reads only the mtp shard
|
||||
sudo docker run --rm -v /tank/aimodels:/models --entrypoint python3 \
|
||||
vllm/vllm-openai:latest /models/compare_mtp_head.py \
|
||||
/models/qwen38-27b-heresy-bf16 /models/qwen38-27b-uncensored-bf16
|
||||
```
|
||||
|
||||
Reference is the incumbent's `model-mtp.safetensors` — grafted verbatim from
|
||||
base per its PROVENANCE, and **measured at 47.7% acceptance** on the live seat
|
||||
through this exact pipeline. Known-good, not assumed.
|
||||
|
||||
- **IDENTICAL** → skip the bf16 gate entirely. Go to Stage 2 and verify
|
||||
acceptance on the **quantized** build (~22 GB) inside the freed `gen` budget.
|
||||
No second seat down.
|
||||
- **DIFFERENT** → the head was edited. Not automatically bad — a deliberate
|
||||
`mtp.*` abliteration is legitimate (hotdogs/Qwen3.8-27B-abliterated edits 2
|
||||
mtp tensors on purpose) — but it is no longer a head we have measured, so
|
||||
1b becomes necessary.
|
||||
- **MISSING** → dropped; needs a graft.
|
||||
|
||||
### 1b. Only if 1a says DIFFERENT: the real bf16 gate
|
||||
|
||||
Gate: **acceptance ≳40%**. Below that the mixed quant is not worth the GPU time
|
||||
— the whole +18% decode case rests on MTP working. Costs `gen` **and**
|
||||
`meromero-charrp` for the window.
|
||||
|
||||
**Residual risk of taking the 1a shortcut, stated plainly:** an identical head
|
||||
proves the *head* is intact, not that the surrounding model still drafts well
|
||||
with it. The MTP head reads hidden states from a body that abliteration *did*
|
||||
change, so acceptance could in principle move even with a byte-identical head.
|
||||
That is exactly what the Stage-3 acceptance measurement on the quantized build
|
||||
catches — at 22 GB instead of 56, and after only ~20 min of quant spend rather
|
||||
than a second seat's downtime. The shortcut trades a small, bounded risk for a
|
||||
real saving; it does not skip the measurement.
|
||||
|
||||
---
|
||||
|
||||
@@ -148,8 +190,8 @@ option.
|
||||
It isn't, quite. Downing `vllm-gen` frees its budget either way, so the
|
||||
*disruption* is identical whether the candidate comes up on `:8017` or on the
|
||||
live `:8015`. The difference is exposure: in situ points all 7 aliases at an
|
||||
unvalidated RC the moment it starts. **Recommended sequence — same downtime,
|
||||
much lower risk:**
|
||||
unvalidated RC the moment it starts. **Operator ruled 2026-08-17: use the probe
|
||||
port.** Sequence — same downtime, much lower risk:
|
||||
|
||||
1. `compose down vllm-gen` (aliases are down either way for the window)
|
||||
2. bring the candidate up via `serve_probe.sh` on `:8017`, its own served-name
|
||||
|
||||
Reference in New Issue
Block a user