Files
esh-pfi-infrastructure/scripts/training-probes/probe_carrier.py
T
vh 7db6c44bcd feat(r49-prep): author-voice LoRA regime prep on gx10 — carriers staged, throughput measured, adapters secured
Prep for the BabyBronte / brokkr-smithy R49 author-voice adapter regime, plus
the operator's "keep the adapter" ruling made durable.

Measured on pfi-gx10 (GB10, sm_121), n=10 per arm after 3 warmup steps, seq
4096, LoRA r=32 on q/k/v/o + MLP, bf16, sdpa, grad-checkpointing on:

  Qwen3-0.6B-Base    dense    0.616 B   1.707 s/step   2,399 tok/s
  Qwen3-1.7B-Base    dense    1.755 B   2.895 s/step   1,415 tok/s
  Qwen3.5-0.8B-Base  hybrid   0.765 B   7.581 s/step     540 tok/s

The dense 1.755 B carrier trains 2.6x faster than the hybrid 0.765 B one on 2.3x
the parameters (~6x per parameter), with more LoRA modules adapted (196 vs 96).
Spreads of 0.6-2.6% put instrument noise an order of magnitude below the effect.
Cause: Qwen3.5 is 18 linear-attention (SSM) layers to 6 attention, and no fused
linear-attention kernel is installed on the box. Grad checkpointing is not the
culprit (19%, and saves 2.6x memory). Batching is not the lever for either
family -- both sit at this box's roofline at batch 1.

Projected per voice on a Brontë-scale corpus: dense 0.6B 2.7 h, dense 1.7B
4.6 h, hybrid 0.8B 12 h. The hybrid would take longer than the 7 h 26B-A4B tune
the regime exists to replace, so the carrier family is now an open decision with
a recommendation for the dense Qwen3 line -- the design doc's original pin.

Two further Qwen3.5 findings, both measured rather than read off the config: the
Base checkpoints ship a vision tower (153/297 model.visual.* Linear tensors that
target_modules="all-linear" would train on text) and an MTP head, both dropped
for free by loading through AutoModelForCausalLM -- which renames modules
relative to the vLLM serving path, so adapter binding needs the
sampled-target-changed check on the serving side; and cross-document packing is
unsafe because SSM state ignores the attention mask, breaking the per-copy
name-consistency invariant the design doc calls sacred. Neither exists on dense.

Adapter disposition, per the operator's ruling: all five gx10-resident ERP
adapters (run-03c/04/05/06/07) mirrored to ana-ml2:/tank/erp-tune/run-<N>/adapter
matching the layout runs 01-03 already used, byte-totals identical both sides and
sha256 matching on every adapter_model.safetensors. /tank/* is deliberately
excluded from ana-ml2's restic sources, so the profile gains one documented
carve-out for /tank/erp-tune/run-*/adapter, verified by resticprofile --dry-run
to expand to exactly those eight paths.

Nothing is training and nothing is queued.
2026-09-09 22:41:47 -07:00

70 lines
3.2 KiB
Python

"""Read-only structural probe of an R49 candidate carrier.
Answers, by measurement rather than by reading the config:
* does transformers on this box load the checkpoint at all,
* which module paths are nn.Linear (the only LoRA-attachable leaves),
* how the parameter budget splits across text body / vision tower / MTP /
embeddings, so "0.8B carrier" can be reported honestly,
* which attention implementations the class accepts.
Loads on CPU in bf16. No training, no GPU, nothing written but stdout.
"""
import json, sys, collections, re
import torch
from transformers import AutoConfig, AutoModelForCausalLM
path = sys.argv[1]
print(f"== {path}")
cfg = AutoConfig.from_pretrained(path, trust_remote_code=False)
print(" config class :", type(cfg).__name__)
print(" architectures :", getattr(cfg, "architectures", None))
tc = getattr(cfg, "text_config", None)
if tc is not None:
lt = getattr(tc, "layer_types", None) or []
print(" text layers :", getattr(tc, "num_hidden_layers", "?"),
"| full_attention:", lt.count("full_attention"),
"| linear_attention:", lt.count("linear_attention"))
print(" hidden/inter :", getattr(tc, "hidden_size", "?"), "/", getattr(tc, "intermediate_size", "?"))
print(" vocab :", getattr(tc, "vocab_size", "?"), "| tied:", getattr(tc, "tie_word_embeddings", "?"))
try:
model = AutoModelForCausalLM.from_pretrained(
path, dtype=torch.bfloat16, device_map="cpu",
attn_implementation="sdpa",
)
except Exception as e:
print(" LOAD FAILED:", type(e).__name__, str(e)[:400])
raise SystemExit(1)
print(" model class :", type(model).__name__)
print(" attn impl :", getattr(model.config, "_attn_implementation", "?"))
# --- parameter budget ------------------------------------------------------
buckets = collections.Counter()
def bucket(name):
if ".visual." in name or name.startswith("visual."): return "vision_tower"
if name.startswith("mtp.") or ".mtp." in name: return "mtp_head"
if "embed_tokens" in name or name.endswith("lm_head.weight"): return "embeddings"
if "linear_attn" in name: return "text_linear_attn"
if "self_attn" in name: return "text_full_attn"
if ".mlp." in name: return "text_mlp"
return "text_other"
for n, p in model.named_parameters():
buckets[bucket(n)] += p.numel()
total = sum(buckets.values())
print(f" TOTAL params : {total/1e9:.3f} B")
for k, v in sorted(buckets.items(), key=lambda kv: -kv[1]):
print(f" {k:<18} {v/1e6:9.1f} M ({100*v/total:5.1f}%)")
# --- LoRA-attachable leaves ------------------------------------------------
lin = collections.defaultdict(list)
for name, mod in model.named_modules():
if isinstance(mod, torch.nn.Linear):
lin[bucket(name + ".weight")].append(name)
print(" nn.Linear leaves by region:")
for region in sorted(lin):
names = lin[region]
tmpl = sorted({re.sub(r"\.\d+\.", ".N.", n) for n in names})
print(f" {region:<18} {len(names):4d} modules, {len(tmpl)} distinct shapes")
for t in tmpl:
print(f" {t}")