feat(coldfusion-abliteration): first-token KL measured — 28.4x selectivity, harmless median 0.0211

Adds `kl_divergence.py`: first-token KL(stock || abliterated) over the full
248,320-token vocabulary, bf16 vs bf16, scored separately for held-out harmless
and reserved-harmful prompts.

Result (L35, 256 harmless / 104 harmful, answer mode):

  harmless  median 0.0211  mean 0.0364  top-1 agreement 89.8%
  harmful   median 0.5996  mean 0.6992  top-1 agreement 55.8%
  selectivity 28.4x (72.8x in think mode)

Self-KL noise floor is exactly 0.0, and all 720 per-prompt values are
bit-identical between a single-process and a two-process run, so the figures are
signal rather than bf16 jitter. Reverse KL on harmful/answer is 1.43 vs forward
0.70 — the mass-where-stock-had-none asymmetry expected of a refusal-direction
removal. Against the Heretic reference figures (0.1191 prior seat, 0.0759 the
live absolute-heresy seat) this is materially gentler, but those are the other
tool's optimizer output on a different base with its own harmless set and
template — order-of-magnitude, not head-to-head. KL remains a fidelity number;
the viability gate is still MTP acceptance (59.1%).

Method notes:
- Prompt classes are reported separately by design. A single averaged KL over a
  mixed corpus is close to meaningless, since the metric is meant to be large on
  harmful prompts and small on benign ones; the ratio carries the information.
- The harmless evaluation set is drawn from the alpaca pool minus calibration's
  own draw, reconstructed by replaying that draw rather than remembered, and
  asserted disjoint on text. The harmful set is the reserved test split.
- `render` is imported from abliterate.py rather than copied, so the measurement
  cannot drift from the rendering the direction was captured against.
- Batch size 1 with logits_to_keep=1: no padding semantics, ~0.6 MB of logits.

Three corrections to the runbook, each of which cost time:
- "bf16 is 50 GB, only gen must go" was 50.10 GiB mislabelled. Text-only weights
  are 51,300 MiB; freeing either GPU0 seat alone leaves ~50,933 MiB. Both must
  stop. VRAM is now sized from the safetensors headers at run time.
- A 27B model cannot be released in-process: `del` + gc + empty_cache left free
  VRAM at 45,287 MiB, and so did confining the model to an inner frame that
  exits. Only process exit returned the card (96,689 MiB). The first run
  completed only because the allocator hit OOM, collected, and retried. Each
  model now gets its own process, handing log-probs to disk between stages.
- The residency gate read hf_device_map, which transformers leaves empty when the
  model fits on one device — it reported "(unsharded)" whether or not anything
  was wrong, so it could never fail. It now reads parameter devices directly.

Model-agnostic lessons promoted to the quant playbook (new 3.12).
This commit is contained in:
vh
2026-08-20 13:00:39 -07:00
parent 8c354a0e79
commit 1b3fb270e7
7 changed files with 823 additions and 10 deletions
@@ -148,3 +148,70 @@ def load_calibration(name: str, n_harmful: int, n_harmless: int, seed: int):
"harmful_pool": len(harmful_pool), "harmless_pool": len(harmless_pool),
"heldout_reserved": len(heldout),
}
def load_evaluation(n_harmless: int, n_harmful: int, seed: int,
calib_harmless_n: int, calib_harmless_seed: int):
"""Held-out evaluation prompts. Returns (harmless, harmful, provenance).
This is the *measurement* corpus — deliberately disjoint from anything the
refusal direction was fitted on, because a divergence measured on the fitting
set answers a different (and much easier) question than a divergence measured
on prompts the surgery never saw.
- **harmful** is the reserved `harmful_behaviors[test]` split (104 prompts,
overlap 0 with train by construction). `load_calibration` refuses to hand
these out as calibration, so they are still virgin here.
- **harmless** is drawn from `harmless_alpaca[train]` *minus the indices
calibration already consumed*. The exclusion has to be reconstructed
rather than remembered: calibration samples with
`random.Random(calib_harmless_seed).sample(range(pool), calib_harmless_n)`,
so replaying that exact draw recovers the used index set. Both the seed and
the n must match the capture that produced the direction under test, which
is why they are explicit parameters and not constants — a future capture at
a different n would otherwise silently leak its calibration into this set.
Disjointness is asserted on the returned *text*, not just on indices, so a
duplicated row in the alpaca pool cannot sneak a calibration prompt back in.
"""
root = _datasets_root()
harmless_pool = _arrow_rows("harmless", "train", root)
harmful = _arrow_rows("harmful", "test", root)
if calib_harmless_n > len(harmless_pool):
raise ValueError(
f"calibration claimed {calib_harmless_n} harmless prompts but the pool "
f"holds {len(harmless_pool)} — the exclusion set cannot be reconstructed")
used_idx = set(random.Random(calib_harmless_seed).sample(
range(len(harmless_pool)), calib_harmless_n))
used_text = {harmless_pool[i] for i in used_idx}
free_idx = [i for i in range(len(harmless_pool)) if i not in used_idx]
if n_harmless > len(free_idx):
raise ValueError(
f"asked for {n_harmless} held-out harmless prompts but only "
f"{len(free_idx)} remain after excluding the {calib_harmless_n} "
f"calibration drew")
idx = sorted(random.Random(seed).sample(free_idx, n_harmless))
harmless = [harmless_pool[i] for i in idx]
leaked = sorted(set(harmless) & used_text)
if leaked:
raise AssertionError(
f"{len(leaked)} evaluation prompt(s) are byte-identical to a calibration "
f"prompt — the harmless pool has duplicate rows and the index-level "
f"exclusion was not enough. First: {leaked[0]!r}")
if n_harmful > len(harmful):
raise ValueError(
f"asked for {n_harmful} harmful eval prompts but the reserved test split "
f"holds {len(harmful)}")
harmful = harmful[:n_harmful] if n_harmful else harmful
return harmless, harmful, {
"eval_source": "mlabonne/harmless_alpaca[train] minus calibration draw + "
"mlabonne/harmful_behaviors[test]",
"n_harmless": len(harmless), "n_harmful": len(harmful), "seed": seed,
"harmless_pool": len(harmless_pool),
"excluded_calibration": {"n": calib_harmless_n, "seed": calib_harmless_seed},
}