feat(gen-seat): quant orcarouter — its MTP head is already Robinson-abliterated in-band

Pulled orcarouter/Qwen3.8-27B-Uncensored at rev 9878936b (55.5 GB, gated, our
token has access) and built /tank/aimodels/qwen38-27b-orcarouter-nvfp4-mixed
(23.4 GB, mixed NVFP4+FP8). Verified, not yet cut over.

The operator asked whether we could apply the Robinson path to the MTP head. We
cannot, because the author already did. compare_mtp_head.py against the verbatim
base graft: 13 of 15 tensors byte-identical, exactly 2 differ --
mtp.layers.0.self_attn.o_proj.weight and mtp.layers.0.mlp.down_proj.weight, which
are precisely the two residual writers our own abliterate.py targets
(EXPECT_MTP_WRITERS = 2).

Reverse-engineered the edit from the weights alone (mtp_delta.py, added here):

  sigma2/sigma1 = 0.0164 on BOTH tensors    rank-1, a single-direction projection
  |cos| between the two recovered dirs = 1.0000   ONE shared direction
  ||delta||/||W|| = 1.42% and 1.41%         a gentle, consistent projection
  sink energy dim 3994 = 0.0000%            sink-clean; Heretic's was 6.18%

That is the Robinson in-band MTP abliteration, already applied, with a direction
that passes our sink screen outright. Nothing to do but preserve it, and the quant
carries it byte-identically. This is the configuration the entire Cold-Fusion
experiment was designed to test and never cleanly delivered.

The new format screen paid for itself on its first real use: think_prior.py on the
bf16 BEFORE any GPU time gave P(<think>) = 1.23e-06 at rank 52, against
Cold-Fusion stock 0.1850 and h300 0.2216. Roughly 150,000x cleaner.

Two durable findings about the pipeline itself:

The quant needs ~17 GB, not a whole card. It ran entirely in GPU1's spare 16 GB
with ZERO production seats stopped -- the h300 run's "stop BOTH GPU0 seats" was
never necessary, it simply had a free card by coincidence. The first attempt OOM'd
by 2.37 GiB at layer 64 of 65 with 3.57 GiB reserved-but-unallocated, which is
fragmentation, and PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True closed it.

post_quant.py now builds a missing output index from the safetensors headers.
A sub-23 GB quant saves one bare shard with no index, and post_quant needs one;
this has broken three separate rounds and been hand-fixed every time. The header
is read by struct-unpacking the u64 length and parsing the JSON -- never
safe_open, which mmaps the whole 22 GB shard and ENOMEMs on ZFS.

Artifact verified: mixed-precision, 1968 tensors, 15 mtp, 333 visual, re:^mtp.*
present in the ignore list (llm-compressor pruned it as always), preproc restored.
Imatrix deferred per operator; the log confirms the usual uniform-MSE fallback, so
this build stays apples-to-apples with heresy's PPL 6.910.
This commit is contained in:
vh
2026-08-21 01:25:37 -07:00
parent bf65d0254d
commit c8f128bdff
3 changed files with 83 additions and 0 deletions
@@ -60,6 +60,36 @@ def main():
src_idx = json.load(open(os.path.join(src, "model.safetensors.index.json")))
mtp_keys = [k for k in src_idx["weight_map"] if k.startswith("mtp")]
out_idx_p = os.path.join(out, "model.safetensors.index.json")
# A quant that lands under ~23 GB fits in ONE shard, and llm-compressor then
# writes a bare `model.safetensors` with NO index at all. Every step below
# needs one, so build it here rather than failing.
#
# Read the safetensors HEADER directly -- the first 8 bytes are a
# little-endian u64 header length, followed by that many bytes of JSON
# keyed by tensor name. Do NOT use safe_open() for this: it mmaps the whole
# shard and ENOMEMs on ZFS against a 22 GB file.
#
# This has now bitten THREE separate rounds (2026-08-15, -08-20, -08-21),
# each time fixed by hand and never in the script. Fixed in the script.
if not os.path.exists(out_idx_p):
import struct
weight_map, total = {}, 0
for fn in sorted(f for f in os.listdir(out) if f.endswith(".safetensors")):
path = os.path.join(out, fn)
total += os.path.getsize(path)
with open(path, "rb") as fh:
n = struct.unpack("<Q", fh.read(8))[0]
header = json.loads(fh.read(n))
for key in header:
if key != "__metadata__":
weight_map[key] = fn
json.dump({"metadata": {"total_size": total}, "weight_map": weight_map},
open(out_idx_p, "w"), indent=2)
print(f"BUILT missing output index from safetensors headers: "
f"{len(weight_map)} tensors across "
f"{len(set(weight_map.values()))} shard(s), {total/1e9:.1f} GB")
out_idx = json.load(open(out_idx_p))
added = 0
for k in mtp_keys: