Files
esh-pfi-infrastructure/docs/pfi/abliteration-recipe-qwen38.md
T
vh ccb56a0a51 docs(pfi): capture the RobinsonLabs Qwen3.8-27B abliteration recipe
Reference recipe (not a deployed artifact) for MTP-aware, vision-preserving
single-direction abliteration of Qwen3.8-27B -- the base family the gen seat
runs. Captures the two things this recipe gets right that naive abliterations
of this architecture miss:

- The MTP head is abliterated in-band (its two residual-write matrices, glue
  left alone), so speculative acceptance does not collapse on the prompts
  abliteration exists to fix -- directly relevant to the gen seat's MTP>=40%
  gate.
- The vision tower is preserved byte-identical (333 tensors, max delta 0).

Plus the two calibration traps specific to this base: the twice-captured
refusal direction (layer 26, |cos| 0.99) and the attention-sink dimension 3994
that bricks the model if orthogonalized out. Documents the coverage gate
(o_proj 16 + linear_out 48 == 64 layers) that catches a half-abliterated
model before it writes a byte, and the foot-gun that the GGUF imatrix does not
cover the MTP block. Links into model-quantization-playbook.md for the quant
half of the pipeline.
2026-08-19 22:02:14 -07:00

7.4 KiB
Raw Blame History

Abliteration recipe — Qwen3.8-27B (MTP-aware, vision-preserving)

Captured 2026-08-19 from RobinsonLabs/Qwen3.8-27B-abliterated (base pinned at commit 1d4bf0f2, Apache-2.0). It is the cleanest public abliteration of the Qwen3.8-27B architecture we have found — the base family our gen seat runs (see auto-memory reference_abliteration_mtp_lessons, reference_gen_qwopus_122b lineage). This is a reference recipe, not a deployed artifact: the value is the method, and specifically the two things it gets right that most abliterations of this architecture get wrong.

Companion: docs/pfi/model-quantization-playbook.md owns the quant half of the pipeline; this owns the abliteration half. When an abliteration lesson is model-agnostic it lands here; when it is specific to one checkpoint's tensor names it stays with that checkpoint.

Why this architecture is the hard case

Qwen3.8-27B (model_type: qwen3_5, Qwen3_5ForConditionalGeneration) is not a plain transformer. Abliterating it correctly means touching three surfaces a naïve layer-loop misses:

  1. A hybrid attention trunk. 64 language layers, most using DeltaNet linear attention (linear_attn.out_proj), with full attention at every 4th layer (self_attn.o_proj). A refusal-direction orthogonalization that only knows about self_attn.o_proj edits 16 of 64 layers and silently leaves the model 75% un-abliterated on the attention path.
  2. A multi-token-prediction (MTP) head (mtp.layers.0) used for speculative decode. The generic 64-layer loop never reaches it.
  3. A vision tower (model.visual.*, 333 tensors) that must survive untouched or the model stops being multimodal.

The two things this recipe gets right

1. The MTP head is abliterated in-band

This is the finding that matters most to us, because our gen seat gates on MTP acceptance ≳40% (reference_abliteration_mtp_lessons).

Most abliterations orthogonalize the trunk and leave mtp.layers.0 untouched. The consequence is subtle and nasty: the draft head keeps proposing refusal-prefix tokens that the abliterated trunk then rejects, so speculative acceptance collapses on exactly the prompts abliteration exists to fix. You get a model that is abliterated and slow, and the slowness is worst precisely where you wanted the behaviour change.

The fix is to orthogonalize the MTP block's two residual-write matrices (self_attn.o_proj, mlp.down_proj) with the same refusal direction as the trunk. The MTP glue — mtp.fc, mtp.norm, mtp.pre_fc_norm_* — is left alone, because those are norms and an input projection, not residual writers. Editing them would corrupt the draft path without removing any refusal.

2. The vision tower is preserved byte-identical

All 333 model.visual.* tensors pass through unmodified — verified by direct tensor diff (max delta 0.000000), not asserted. An mmproj is published so the vision half is actually usable, not just nominally intact.

The edit set (131 tensors)

Single-direction weight orthogonalization, Arditi et al. style, applied to every matrix that writes the residual stream:

scope tensor count
model.language_model.layers.* (64) mlp.down_proj 64
linear_attn.out_proj (DeltaNet) 48
self_attn.o_proj (full-attn, interval 4) 16
mtp.layers.0 o_proj + down_proj 2
model.language_model embed_tokens 1
edited total 131
model.visual.* preserved byte-identical 333

Hard coverage gate before writing a byte: o_proj(16) + linear_out(48) == 64 == num_hidden_layers. This is the check that catches a partial tensor-name match — the failure mode that otherwise ships a quietly half-abliterated model that passes a smoke test and fails in the field. Adopt this gate in any re-derivation.

Two calibration traps specific to this base

Refusal-direction selection

The direction was captured twice, from two structurally different chat-template renderings:

  • one with enable_thinking=false
  • one with thinking on at reasoning_effort=xhigh (which injects an extra system block and shifts every token position)

The two agree at |cos| 0.96–0.99 across layers 18–45, peaking 0.9925 at layer 26 — the layer used. Two different prompt distributions converging on the same vector is the evidence that the direction encodes refusal semantics rather than template formatting. A single-template capture cannot distinguish the two.

The attention-sink dimension — the one that bricks the model

Qwen3.8-27B's massive-activation dimension is 3994. It carries 19–21% of the direction's energy at layers 1–3, and orthogonalizing it out of every residual writer produces a model that loads, runs, and emits garbage. Layer 26 was chosen partly because it carries only 0.06% of its energy in dim 3994.

Any re-derivation MUST screen for this. It is the single most likely way to waste a GPU afternoon on this architecture and mistake the result for a failed abliteration when it is actually an attention-sink blowout.

Measured behaviour (their numbers, for reference)

Base vs abliterated, same session/harness/prompts, both at Q4_K_M:

prompt set base abliterated
in-distribution (24, from capture set) 96% (23/24) 8% (2/24)
held-out (40, disjoint, overlap=0) 100% (40/40) 8% (3/40)

Capability axes (reasoning / code / math / factual / instruction-following / creative-RP coherence): no regression on any axis. Held-out train/test split was 416/104 with overlap 0, so the 8% held-out figure is generalization, not a reshuffle of calibration prompts.

Note the design point: 8% is deliberate. Harm guardrails are retained — self-harm prompts still redirect (988) rather than comply. This is a creative-content abliteration shipped "at the ceiling where capability and guardrails both survive," explicitly not a jailbreak. That makes it a milder abliteration than our incumbent gen seat (absolute-heresy, ~2% author refusals, aggressive Heretic). Adopt the method here; the ceiling is a separate call.

How this maps onto our pipeline

The recipe is a drop-in for the front half of the House quant pipeline:

  1. Pull bf16 master to NFS (verify repo id first — reference_verify_hf_repo_ids_before_pull).
  2. Baseline MTP acceptance on bf16 before any surgery — the standing rule.
  3. Orthogonalize per the edit set above; enforce the coverage gate; screen dim 3994; gate the result on MTP acceptance ≳40%, not KL (KL misled us once — reference_abliteration_mtp_lessons).
  4. Verify vision byte-identical, refusals down, PPL not blown, no catatonia.
  5. NVFP4-quantize in-house (mixed W4A4 + FP8-attn/lm_head — model-quantization-playbook.md). Foot-gun the GGUF card itself flags: the imatrix does not cover the MTP block — so a GGUF requant path leaves MTP uncalibrated. Our NVFP4 path must calibrate it explicitly.

Provenance

  • Recipe: RobinsonLabs README, fetched verbatim 2026-08-19. Authored with their "ModelForge" manufacturing system-of-record (not public).
  • Method lineage: Arditi et al., single-direction refusal orthogonalization.
  • Our prior art: reference_abliteration_mtp_lessons (modest abliteration preserves MTP; test MTP on bf16 first; gate on acceptance not KL), and the gen-seat quant recipe in model-quantization-playbook.md.