feat(coldfusion-abliteration): first-token KL measured — 28.4x selectivity, harmless median 0.0211

Adds `kl_divergence.py`: first-token KL(stock || abliterated) over the full
248,320-token vocabulary, bf16 vs bf16, scored separately for held-out harmless
and reserved-harmful prompts.

Result (L35, 256 harmless / 104 harmful, answer mode):

  harmless  median 0.0211  mean 0.0364  top-1 agreement 89.8%
  harmful   median 0.5996  mean 0.6992  top-1 agreement 55.8%
  selectivity 28.4x (72.8x in think mode)

Self-KL noise floor is exactly 0.0, and all 720 per-prompt values are
bit-identical between a single-process and a two-process run, so the figures are
signal rather than bf16 jitter. Reverse KL on harmful/answer is 1.43 vs forward
0.70 — the mass-where-stock-had-none asymmetry expected of a refusal-direction
removal. Against the Heretic reference figures (0.1191 prior seat, 0.0759 the
live absolute-heresy seat) this is materially gentler, but those are the other
tool's optimizer output on a different base with its own harmless set and
template — order-of-magnitude, not head-to-head. KL remains a fidelity number;
the viability gate is still MTP acceptance (59.1%).

Method notes:
- Prompt classes are reported separately by design. A single averaged KL over a
  mixed corpus is close to meaningless, since the metric is meant to be large on
  harmful prompts and small on benign ones; the ratio carries the information.
- The harmless evaluation set is drawn from the alpaca pool minus calibration's
  own draw, reconstructed by replaying that draw rather than remembered, and
  asserted disjoint on text. The harmful set is the reserved test split.
- `render` is imported from abliterate.py rather than copied, so the measurement
  cannot drift from the rendering the direction was captured against.
- Batch size 1 with logits_to_keep=1: no padding semantics, ~0.6 MB of logits.

Three corrections to the runbook, each of which cost time:
- "bf16 is 50 GB, only gen must go" was 50.10 GiB mislabelled. Text-only weights
  are 51,300 MiB; freeing either GPU0 seat alone leaves ~50,933 MiB. Both must
  stop. VRAM is now sized from the safetensors headers at run time.
- A 27B model cannot be released in-process: `del` + gc + empty_cache left free
  VRAM at 45,287 MiB, and so did confining the model to an inner frame that
  exits. Only process exit returned the card (96,689 MiB). The first run
  completed only because the allocator hit OOM, collected, and retried. Each
  model now gets its own process, handing log-probs to disk between stages.
- The residency gate read hf_device_map, which transformers leaves empty when the
  model fits on one device — it reported "(unsharded)" whether or not anything
  was wrong, so it could never fail. It now reads parameter devices directly.

Model-agnostic lessons promoted to the quant playbook (new 3.12).
This commit is contained in:
vh
2026-08-20 13:00:39 -07:00
parent 8c354a0e79
commit 1b3fb270e7
7 changed files with 823 additions and 10 deletions
+134 -9
View File
@@ -124,8 +124,10 @@ RUN="sudo -u llmuser env HF_HUB_OFFLINE=1 CUDA_VISIBLE_DEVICES=0 \
# confirms the recipe maps onto THIS checkpoint's names.
$RUN --dry-run
# --- capture needs GPU0 to itself: bf16 is 50 GB, so only gen must go ---
sudo docker stop -t 60 vllm-gen
# --- capture needs GPU0 to itself. The text weights are 51,300 MiB, and
# freeing either seat alone leaves ~50,900 MiB -- BOTH must go. See
# gotcha 3; the "only gen" line that used to be here was a unit error. ---
sudo docker stop -t 60 vllm-gen vllm-meromero-rp
cp $M/refusal-direction.pt $M/refusal-direction.pt.bak # capture overwrites it
# 2. CONTROL RUN — the legacy 8/8 set. Reproduces layer 22, |cos| 0.5944, sink
@@ -252,6 +254,101 @@ a separate operator decision needing the full Stage-3 gate (PPL, prefill, surfac
delete-too-early / multi-day-degeneration lesson). The thesis is proven; the
cutover is a distinct call.
## ✅ KL RESULT — the surgery is highly selective (2026-08-20)
`kl_divergence.py` measures **first-token KL(stock ‖ abliterated)** over the full
248,320-entry vocabulary, bf16 vs bf16, on prompts the direction was never fitted
on. Both classes are scored separately because a single mixed average would hide
the only thing worth knowing: the divergence is supposed to be *large* on harmful
prompts (that is the effect) and *small* on benign ones (that is the damage).
| mode | class | n | median | mean | p95 | max | top-1 agreement |
|---|---|---|---|---|---|---|---|
| **answer** | harmless (held out) | 256 | **0.0211** | **0.0364** | 0.1219 | 0.2654 | 89.8% |
| **answer** | harmful (reserved test) | 104 | **0.5996** | 0.6992 | 1.6937 | 1.9920 | 55.8% |
| think | harmless (held out) | 256 | 0.0042 | 0.0066 | 0.0205 | 0.0392 | 94.5% |
| think | harmful (reserved test) | 104 | 0.3068 | 0.3186 | 0.4689 | 0.5298 | 57.7% |
**Selectivity — harmful/harmless median KL — is 28.4× in answer mode and 72.8× in
think mode.** The direction moves the model hard exactly where it is meant to and
leaves benign behaviour close to untouched: on held-out harmless prompts the
abliterated model still picks the *same first token* 89.8% of the time.
**Noise floor: exactly 0.0** in both modes (32 prompts re-run through the same
model, self-KL). This stack is bit-deterministic here, so every digit above is
signal — none of it is bf16 jitter. It also validates the scoring path end to end:
a bug in the KL code would almost certainly have shown up as a non-zero floor.
**The reverse-KL asymmetry is the abliteration's signature.** On harmful prompts
in answer mode, KL(stock‖abl) is 0.70 but KL(abl‖stock) is **1.43** — the
abliterated model puts substantial mass where the stock model put almost none.
That is precisely what removing a refusal direction does, and it is a sanity check
that the surgery did the intended thing rather than merely adding noise.
### Against the Heretic reference figures — favourable, with a caveat
| model | first-token KL, harmless | abliteration method |
|---|---|---|
| `JonathanColetti/Qwen3.8-27B-Uncensored` (prior gen seat) | 0.1191 | Heretic, out-of-band MTP |
| `absolute-heresy` (**current** gen seat) | 0.0759 | Heretic v1.4.0 + SOMPOA |
| **Cold-Fusion L35 (ours)** | **0.0211 median / 0.0364 mean** | Robinson, in-band MTP |
⚠️ **Not a head-to-head.** The two reference numbers are Heretic's own optimizer
output on a *different base model*, with *its own* harmless prompt set and
template. Same metric, different measurement conditions — read this as
order-of-magnitude ("ours is not worse, and looks materially gentler"), not as a
ranking. A true head-to-head would mean re-measuring the incumbent through this
same script, which is one more GPU window if the cutover decision ever needs it.
Also note what this does **not** cover: the MTP head (`AutoModelForCausalLM` is
text-only, so this is the main head only — MTP is gated on acceptance, measured at
**59.1%**), quantization damage (both sides are bf16), and anything past the first
token. Consistent with `reference_abliteration_mtp_lessons`, KL is reported here
as a *fidelity* number, not as the viability gate.
**Reproducibility: exact.** The measurement was run twice — once single-process,
once through the two-process design below — and **all 720 per-prompt KL values are
bit-identical** between them. Combined with the 0.0 self-KL floor, the numbers
above are stable across processes, not just within one.
Artifacts: `kl-L35.json` (+ `kl-L35-rerun.json`, the reproducibility check) and the
two `.ref.pt` / `.cand.pt` log-prob caches, beside the harness on ana-ml2. Run
cost: **2m40s** single-process, **3m26s** two-process, both seats down.
```bash
# free, no GPU, safe with the seats up — run this first
$V $P/kl_divergence.py --ref $M --cand $A --out $P/kl-L35.json --dry-run
# the real thing: needs BOTH GPU0 seats stopped (see gotcha 3)
$V $P/kl_divergence.py --ref $M --cand $A --out $P/kl-L35.json
```
**Why it runs one process per model.** The default `--stage all` re-execs itself
once per checkpoint (`--stage ref`, then `--stage cand`), each writing its
first-token log-probs to a ~682 MiB `.pt` cache, then scores from the caches.
This is not tidiness — **it is the only teardown that works.** Measured, free VRAM
after the reference model:
| teardown | free VRAM |
|---|---|
| `del model` + `gc.collect()` + `empty_cache()` | 45,287 MiB |
| the same, model confined to an inner frame that exits | 45,287 MiB |
| **the process exits** | **96,689 MiB** |
The weights survive both in-process teardowns. The very first run only completed
because PyTorch's allocator hit OOM on the second load, collected, and retried —
the second model landed on the card *by rescue, not by design*, and on this
architecture a silent CPU offload does not raise, it returns confident garbage
(gotcha 1). The headroom gate (`exit 10`) is what turned that from an invisible
near-miss into a loud failure. Side benefit: the `ref` cache is reusable, so
measuring a different candidate against the same stock model skips a stage
entirely (`--stage cand` then `--stage score`).
⚠️ **The old residency gate could not fail.** It read `hf_device_map`, which
transformers leaves **empty** when the whole model fits on one device — so it
printed "(unsharded)" both when everything was fine and when there was nothing to
inspect. It now reads `{p.device for p in model.parameters()}` and prints the real
placement (`all parameters on cuda:0`).
## Why the write is shard surgery, not `model.save_pretrained`
The `--out` path edits the 18 safetensors shards directly and never instantiates
@@ -318,13 +415,41 @@ for headroom; it buys corruption. Gated (exit 9).
layers and is **deterministic**. This did neither. *If a NaN doesn't
propagate, debug memory, not math.*
**3. bf16 fits on one GPU — so the window is small now.** 50.1 GB of a 96 GB
card, which means a capture needs only **`vllm-gen` stopped**, not all three
seats. (`--capture-dtype float32` remains as an escape hatch; it needs 111 GB, so
it also needs `--max-layer 46` to fit on one card. The two agree to 0.0005, so
there is no reason to reach for it.) **Restore after:** start
`vllm-meromero-rp` **first**, then `vllm-gen` — gen grabs a fraction of *free*
VRAM at startup and will starve meromero if it goes first.
**3. bf16 fits on one GPU — but it needs BOTH GPU0 seats stopped, not one.**
> ⚠️ **CORRECTED 2026-08-20.** This section used to read "50.1 GB … a capture
> needs only `vllm-gen` stopped." **The unit was wrong and the conclusion that
> rode on it was wrong.** The real figure is **50.10 GiB = 51,300 MiB = 53.8 GB**
> of text-only weights, measured from the safetensors headers rather than read off
> a `/1e9` print:
>
> | | GB | GiB | MiB |
> |---|---|---|---|
> | checkpoint total | 55.56 | 51.75 | 52,989 |
> | vision (not loaded by `AutoModelForCausalLM`) | 0.92 | 0.86 | 879 |
> | MTP (not loaded either) | 0.85 | 0.79 | 810 |
> | **text-only — what actually lands on the card** | **53.79** | **50.10** | **51,300** |
>
> GPU0's two tenants are meromero (50,072 MiB) and gen (46,304 MiB), and
> **freeing either one alone leaves at most 50,933 MiB — about 400 MiB short.**
> A run that assumes one seat is enough will stop a service, sit at the edge, and
> then OOM. Stop **both**. Recompute this table if the checkpoint changes; do not
> trust a remembered gigabyte figure.
Both seats down leaves ~97,200 MiB, so the fit is comfortable rather than
marginal. Gate the run on **observing** the free VRAM (`nvidia-smi
--query-gpu=memory.free`) rather than sleeping after `docker stop`, and put the
restore in a `trap ... EXIT` so an abort hands the seats back — the 2026-08-20
aborted window did exactly that and cost nothing but two minutes.
(`--capture-dtype float32` remains as an escape hatch; it needs 111 GB, so it
also needs `--max-layer 46` to fit on one card. The two agree to 0.0005, so there
is no reason to reach for it.) **Restore after:** start `vllm-meromero-rp`
**first**, then `vllm-gen`. (The stated reason — "gen grabs a fraction of *free*
VRAM" — is not what the configs do: both seats pass `--gpu-memory-utilization` as
a fraction of **total** (`MEROMERO_GPU_MEM_UTIL=0.52`, `GEN_GPU_MEM_UTIL=0.43`),
so restore order is not actually load-bearing. Kept as the runbook order anyway;
it costs nothing.)
**4. fla is irrelevant here — but harmless.** `fla` + `einops` are `--target`
-installed to `/tank/aimodels/coldfusion-abliteration/pylibs` and reached via