feat(coldfusion-abliteration): first-token KL measured — 28.4x selectivity, harmless median 0.0211
Adds `kl_divergence.py`: first-token KL(stock || abliterated) over the full 248,320-token vocabulary, bf16 vs bf16, scored separately for held-out harmless and reserved-harmful prompts. Result (L35, 256 harmless / 104 harmful, answer mode): harmless median 0.0211 mean 0.0364 top-1 agreement 89.8% harmful median 0.5996 mean 0.6992 top-1 agreement 55.8% selectivity 28.4x (72.8x in think mode) Self-KL noise floor is exactly 0.0, and all 720 per-prompt values are bit-identical between a single-process and a two-process run, so the figures are signal rather than bf16 jitter. Reverse KL on harmful/answer is 1.43 vs forward 0.70 — the mass-where-stock-had-none asymmetry expected of a refusal-direction removal. Against the Heretic reference figures (0.1191 prior seat, 0.0759 the live absolute-heresy seat) this is materially gentler, but those are the other tool's optimizer output on a different base with its own harmless set and template — order-of-magnitude, not head-to-head. KL remains a fidelity number; the viability gate is still MTP acceptance (59.1%). Method notes: - Prompt classes are reported separately by design. A single averaged KL over a mixed corpus is close to meaningless, since the metric is meant to be large on harmful prompts and small on benign ones; the ratio carries the information. - The harmless evaluation set is drawn from the alpaca pool minus calibration's own draw, reconstructed by replaying that draw rather than remembered, and asserted disjoint on text. The harmful set is the reserved test split. - `render` is imported from abliterate.py rather than copied, so the measurement cannot drift from the rendering the direction was captured against. - Batch size 1 with logits_to_keep=1: no padding semantics, ~0.6 MB of logits. Three corrections to the runbook, each of which cost time: - "bf16 is 50 GB, only gen must go" was 50.10 GiB mislabelled. Text-only weights are 51,300 MiB; freeing either GPU0 seat alone leaves ~50,933 MiB. Both must stop. VRAM is now sized from the safetensors headers at run time. - A 27B model cannot be released in-process: `del` + gc + empty_cache left free VRAM at 45,287 MiB, and so did confining the model to an inner frame that exits. Only process exit returned the card (96,689 MiB). The first run completed only because the allocator hit OOM, collected, and retried. Each model now gets its own process, handing log-probs to disk between stages. - The residency gate read hf_device_map, which transformers leaves empty when the model fits on one device — it reported "(unsharded)" whether or not anything was wrong, so it could never fail. It now reads parameter devices directly. Model-agnostic lessons promoted to the quant playbook (new 3.12).
This commit is contained in:
@@ -124,8 +124,10 @@ RUN="sudo -u llmuser env HF_HUB_OFFLINE=1 CUDA_VISIBLE_DEVICES=0 \
|
||||
# confirms the recipe maps onto THIS checkpoint's names.
|
||||
$RUN --dry-run
|
||||
|
||||
# --- capture needs GPU0 to itself: bf16 is 50 GB, so only gen must go ---
|
||||
sudo docker stop -t 60 vllm-gen
|
||||
# --- capture needs GPU0 to itself. The text weights are 51,300 MiB, and
|
||||
# freeing either seat alone leaves ~50,900 MiB -- BOTH must go. See
|
||||
# gotcha 3; the "only gen" line that used to be here was a unit error. ---
|
||||
sudo docker stop -t 60 vllm-gen vllm-meromero-rp
|
||||
cp $M/refusal-direction.pt $M/refusal-direction.pt.bak # capture overwrites it
|
||||
|
||||
# 2. CONTROL RUN — the legacy 8/8 set. Reproduces layer 22, |cos| 0.5944, sink
|
||||
@@ -252,6 +254,101 @@ a separate operator decision needing the full Stage-3 gate (PPL, prefill, surfac
|
||||
delete-too-early / multi-day-degeneration lesson). The thesis is proven; the
|
||||
cutover is a distinct call.
|
||||
|
||||
## ✅ KL RESULT — the surgery is highly selective (2026-08-20)
|
||||
|
||||
`kl_divergence.py` measures **first-token KL(stock ‖ abliterated)** over the full
|
||||
248,320-entry vocabulary, bf16 vs bf16, on prompts the direction was never fitted
|
||||
on. Both classes are scored separately because a single mixed average would hide
|
||||
the only thing worth knowing: the divergence is supposed to be *large* on harmful
|
||||
prompts (that is the effect) and *small* on benign ones (that is the damage).
|
||||
|
||||
| mode | class | n | median | mean | p95 | max | top-1 agreement |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **answer** | harmless (held out) | 256 | **0.0211** | **0.0364** | 0.1219 | 0.2654 | 89.8% |
|
||||
| **answer** | harmful (reserved test) | 104 | **0.5996** | 0.6992 | 1.6937 | 1.9920 | 55.8% |
|
||||
| think | harmless (held out) | 256 | 0.0042 | 0.0066 | 0.0205 | 0.0392 | 94.5% |
|
||||
| think | harmful (reserved test) | 104 | 0.3068 | 0.3186 | 0.4689 | 0.5298 | 57.7% |
|
||||
|
||||
**Selectivity — harmful/harmless median KL — is 28.4× in answer mode and 72.8× in
|
||||
think mode.** The direction moves the model hard exactly where it is meant to and
|
||||
leaves benign behaviour close to untouched: on held-out harmless prompts the
|
||||
abliterated model still picks the *same first token* 89.8% of the time.
|
||||
|
||||
**Noise floor: exactly 0.0** in both modes (32 prompts re-run through the same
|
||||
model, self-KL). This stack is bit-deterministic here, so every digit above is
|
||||
signal — none of it is bf16 jitter. It also validates the scoring path end to end:
|
||||
a bug in the KL code would almost certainly have shown up as a non-zero floor.
|
||||
|
||||
**The reverse-KL asymmetry is the abliteration's signature.** On harmful prompts
|
||||
in answer mode, KL(stock‖abl) is 0.70 but KL(abl‖stock) is **1.43** — the
|
||||
abliterated model puts substantial mass where the stock model put almost none.
|
||||
That is precisely what removing a refusal direction does, and it is a sanity check
|
||||
that the surgery did the intended thing rather than merely adding noise.
|
||||
|
||||
### Against the Heretic reference figures — favourable, with a caveat
|
||||
|
||||
| model | first-token KL, harmless | abliteration method |
|
||||
|---|---|---|
|
||||
| `JonathanColetti/Qwen3.8-27B-Uncensored` (prior gen seat) | 0.1191 | Heretic, out-of-band MTP |
|
||||
| `absolute-heresy` (**current** gen seat) | 0.0759 | Heretic v1.4.0 + SOMPOA |
|
||||
| **Cold-Fusion L35 (ours)** | **0.0211 median / 0.0364 mean** | Robinson, in-band MTP |
|
||||
|
||||
⚠️ **Not a head-to-head.** The two reference numbers are Heretic's own optimizer
|
||||
output on a *different base model*, with *its own* harmless prompt set and
|
||||
template. Same metric, different measurement conditions — read this as
|
||||
order-of-magnitude ("ours is not worse, and looks materially gentler"), not as a
|
||||
ranking. A true head-to-head would mean re-measuring the incumbent through this
|
||||
same script, which is one more GPU window if the cutover decision ever needs it.
|
||||
|
||||
Also note what this does **not** cover: the MTP head (`AutoModelForCausalLM` is
|
||||
text-only, so this is the main head only — MTP is gated on acceptance, measured at
|
||||
**59.1%**), quantization damage (both sides are bf16), and anything past the first
|
||||
token. Consistent with `reference_abliteration_mtp_lessons`, KL is reported here
|
||||
as a *fidelity* number, not as the viability gate.
|
||||
|
||||
**Reproducibility: exact.** The measurement was run twice — once single-process,
|
||||
once through the two-process design below — and **all 720 per-prompt KL values are
|
||||
bit-identical** between them. Combined with the 0.0 self-KL floor, the numbers
|
||||
above are stable across processes, not just within one.
|
||||
|
||||
Artifacts: `kl-L35.json` (+ `kl-L35-rerun.json`, the reproducibility check) and the
|
||||
two `.ref.pt` / `.cand.pt` log-prob caches, beside the harness on ana-ml2. Run
|
||||
cost: **2m40s** single-process, **3m26s** two-process, both seats down.
|
||||
|
||||
```bash
|
||||
# free, no GPU, safe with the seats up — run this first
|
||||
$V $P/kl_divergence.py --ref $M --cand $A --out $P/kl-L35.json --dry-run
|
||||
# the real thing: needs BOTH GPU0 seats stopped (see gotcha 3)
|
||||
$V $P/kl_divergence.py --ref $M --cand $A --out $P/kl-L35.json
|
||||
```
|
||||
|
||||
**Why it runs one process per model.** The default `--stage all` re-execs itself
|
||||
once per checkpoint (`--stage ref`, then `--stage cand`), each writing its
|
||||
first-token log-probs to a ~682 MiB `.pt` cache, then scores from the caches.
|
||||
This is not tidiness — **it is the only teardown that works.** Measured, free VRAM
|
||||
after the reference model:
|
||||
|
||||
| teardown | free VRAM |
|
||||
|---|---|
|
||||
| `del model` + `gc.collect()` + `empty_cache()` | 45,287 MiB |
|
||||
| the same, model confined to an inner frame that exits | 45,287 MiB |
|
||||
| **the process exits** | **96,689 MiB** |
|
||||
|
||||
The weights survive both in-process teardowns. The very first run only completed
|
||||
because PyTorch's allocator hit OOM on the second load, collected, and retried —
|
||||
the second model landed on the card *by rescue, not by design*, and on this
|
||||
architecture a silent CPU offload does not raise, it returns confident garbage
|
||||
(gotcha 1). The headroom gate (`exit 10`) is what turned that from an invisible
|
||||
near-miss into a loud failure. Side benefit: the `ref` cache is reusable, so
|
||||
measuring a different candidate against the same stock model skips a stage
|
||||
entirely (`--stage cand` then `--stage score`).
|
||||
|
||||
⚠️ **The old residency gate could not fail.** It read `hf_device_map`, which
|
||||
transformers leaves **empty** when the whole model fits on one device — so it
|
||||
printed "(unsharded)" both when everything was fine and when there was nothing to
|
||||
inspect. It now reads `{p.device for p in model.parameters()}` and prints the real
|
||||
placement (`all parameters on cuda:0`).
|
||||
|
||||
## Why the write is shard surgery, not `model.save_pretrained`
|
||||
|
||||
The `--out` path edits the 18 safetensors shards directly and never instantiates
|
||||
@@ -318,13 +415,41 @@ for headroom; it buys corruption. Gated (exit 9).
|
||||
layers and is **deterministic**. This did neither. *If a NaN doesn't
|
||||
propagate, debug memory, not math.*
|
||||
|
||||
**3. bf16 fits on one GPU — so the window is small now.** 50.1 GB of a 96 GB
|
||||
card, which means a capture needs only **`vllm-gen` stopped**, not all three
|
||||
seats. (`--capture-dtype float32` remains as an escape hatch; it needs 111 GB, so
|
||||
it also needs `--max-layer 46` to fit on one card. The two agree to 0.0005, so
|
||||
there is no reason to reach for it.) **Restore after:** start
|
||||
`vllm-meromero-rp` **first**, then `vllm-gen` — gen grabs a fraction of *free*
|
||||
VRAM at startup and will starve meromero if it goes first.
|
||||
**3. bf16 fits on one GPU — but it needs BOTH GPU0 seats stopped, not one.**
|
||||
|
||||
> ⚠️ **CORRECTED 2026-08-20.** This section used to read "50.1 GB … a capture
|
||||
> needs only `vllm-gen` stopped." **The unit was wrong and the conclusion that
|
||||
> rode on it was wrong.** The real figure is **50.10 GiB = 51,300 MiB = 53.8 GB**
|
||||
> of text-only weights, measured from the safetensors headers rather than read off
|
||||
> a `/1e9` print:
|
||||
>
|
||||
> | | GB | GiB | MiB |
|
||||
> |---|---|---|---|
|
||||
> | checkpoint total | 55.56 | 51.75 | 52,989 |
|
||||
> | vision (not loaded by `AutoModelForCausalLM`) | 0.92 | 0.86 | 879 |
|
||||
> | MTP (not loaded either) | 0.85 | 0.79 | 810 |
|
||||
> | **text-only — what actually lands on the card** | **53.79** | **50.10** | **51,300** |
|
||||
>
|
||||
> GPU0's two tenants are meromero (50,072 MiB) and gen (46,304 MiB), and
|
||||
> **freeing either one alone leaves at most 50,933 MiB — about 400 MiB short.**
|
||||
> A run that assumes one seat is enough will stop a service, sit at the edge, and
|
||||
> then OOM. Stop **both**. Recompute this table if the checkpoint changes; do not
|
||||
> trust a remembered gigabyte figure.
|
||||
|
||||
Both seats down leaves ~97,200 MiB, so the fit is comfortable rather than
|
||||
marginal. Gate the run on **observing** the free VRAM (`nvidia-smi
|
||||
--query-gpu=memory.free`) rather than sleeping after `docker stop`, and put the
|
||||
restore in a `trap ... EXIT` so an abort hands the seats back — the 2026-08-20
|
||||
aborted window did exactly that and cost nothing but two minutes.
|
||||
|
||||
(`--capture-dtype float32` remains as an escape hatch; it needs 111 GB, so it
|
||||
also needs `--max-layer 46` to fit on one card. The two agree to 0.0005, so there
|
||||
is no reason to reach for it.) **Restore after:** start `vllm-meromero-rp`
|
||||
**first**, then `vllm-gen`. (The stated reason — "gen grabs a fraction of *free*
|
||||
VRAM" — is not what the configs do: both seats pass `--gpu-memory-utilization` as
|
||||
a fraction of **total** (`MEROMERO_GPU_MEM_UTIL=0.52`, `GEN_GPU_MEM_UTIL=0.43`),
|
||||
so restore order is not actually load-bearing. Kept as the runbook order anyway;
|
||||
it costs nothing.)
|
||||
|
||||
**4. fla is irrelevant here — but harmless.** `fla` + `einops` are `--target`
|
||||
-installed to `/tank/aimodels/coldfusion-abliteration/pylibs` and reached via
|
||||
|
||||
Reference in New Issue
Block a user