7db6c44bcd
Prep for the BabyBronte / brokkr-smithy R49 author-voice adapter regime, plus the operator's "keep the adapter" ruling made durable. Measured on pfi-gx10 (GB10, sm_121), n=10 per arm after 3 warmup steps, seq 4096, LoRA r=32 on q/k/v/o + MLP, bf16, sdpa, grad-checkpointing on: Qwen3-0.6B-Base dense 0.616 B 1.707 s/step 2,399 tok/s Qwen3-1.7B-Base dense 1.755 B 2.895 s/step 1,415 tok/s Qwen3.5-0.8B-Base hybrid 0.765 B 7.581 s/step 540 tok/s The dense 1.755 B carrier trains 2.6x faster than the hybrid 0.765 B one on 2.3x the parameters (~6x per parameter), with more LoRA modules adapted (196 vs 96). Spreads of 0.6-2.6% put instrument noise an order of magnitude below the effect. Cause: Qwen3.5 is 18 linear-attention (SSM) layers to 6 attention, and no fused linear-attention kernel is installed on the box. Grad checkpointing is not the culprit (19%, and saves 2.6x memory). Batching is not the lever for either family -- both sit at this box's roofline at batch 1. Projected per voice on a Brontë-scale corpus: dense 0.6B 2.7 h, dense 1.7B 4.6 h, hybrid 0.8B 12 h. The hybrid would take longer than the 7 h 26B-A4B tune the regime exists to replace, so the carrier family is now an open decision with a recommendation for the dense Qwen3 line -- the design doc's original pin. Two further Qwen3.5 findings, both measured rather than read off the config: the Base checkpoints ship a vision tower (153/297 model.visual.* Linear tensors that target_modules="all-linear" would train on text) and an MTP head, both dropped for free by loading through AutoModelForCausalLM -- which renames modules relative to the vLLM serving path, so adapter binding needs the sampled-target-changed check on the serving side; and cross-document packing is unsafe because SSM state ignores the attention mask, breaking the per-copy name-consistency invariant the design doc calls sacred. Neither exists on dense. Adapter disposition, per the operator's ruling: all five gx10-resident ERP adapters (run-03c/04/05/06/07) mirrored to ana-ml2:/tank/erp-tune/run-<N>/adapter matching the layout runs 01-03 already used, byte-totals identical both sides and sha256 matching on every adapter_model.safetensors. /tank/* is deliberately excluded from ana-ml2's restic sources, so the profile gains one documented carve-out for /tank/erp-tune/run-*/adapter, verified by resticprofile --dry-run to expand to exactly those eight paths. Nothing is training and nothing is queued.
70 lines
3.2 KiB
Python
70 lines
3.2 KiB
Python
"""Read-only structural probe of an R49 candidate carrier.
|
|
|
|
Answers, by measurement rather than by reading the config:
|
|
* does transformers on this box load the checkpoint at all,
|
|
* which module paths are nn.Linear (the only LoRA-attachable leaves),
|
|
* how the parameter budget splits across text body / vision tower / MTP /
|
|
embeddings, so "0.8B carrier" can be reported honestly,
|
|
* which attention implementations the class accepts.
|
|
|
|
Loads on CPU in bf16. No training, no GPU, nothing written but stdout.
|
|
"""
|
|
import json, sys, collections, re
|
|
import torch
|
|
from transformers import AutoConfig, AutoModelForCausalLM
|
|
|
|
path = sys.argv[1]
|
|
print(f"== {path}")
|
|
cfg = AutoConfig.from_pretrained(path, trust_remote_code=False)
|
|
print(" config class :", type(cfg).__name__)
|
|
print(" architectures :", getattr(cfg, "architectures", None))
|
|
tc = getattr(cfg, "text_config", None)
|
|
if tc is not None:
|
|
lt = getattr(tc, "layer_types", None) or []
|
|
print(" text layers :", getattr(tc, "num_hidden_layers", "?"),
|
|
"| full_attention:", lt.count("full_attention"),
|
|
"| linear_attention:", lt.count("linear_attention"))
|
|
print(" hidden/inter :", getattr(tc, "hidden_size", "?"), "/", getattr(tc, "intermediate_size", "?"))
|
|
print(" vocab :", getattr(tc, "vocab_size", "?"), "| tied:", getattr(tc, "tie_word_embeddings", "?"))
|
|
|
|
try:
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|
path, dtype=torch.bfloat16, device_map="cpu",
|
|
attn_implementation="sdpa",
|
|
)
|
|
except Exception as e:
|
|
print(" LOAD FAILED:", type(e).__name__, str(e)[:400])
|
|
raise SystemExit(1)
|
|
print(" model class :", type(model).__name__)
|
|
print(" attn impl :", getattr(model.config, "_attn_implementation", "?"))
|
|
|
|
# --- parameter budget ------------------------------------------------------
|
|
buckets = collections.Counter()
|
|
def bucket(name):
|
|
if ".visual." in name or name.startswith("visual."): return "vision_tower"
|
|
if name.startswith("mtp.") or ".mtp." in name: return "mtp_head"
|
|
if "embed_tokens" in name or name.endswith("lm_head.weight"): return "embeddings"
|
|
if "linear_attn" in name: return "text_linear_attn"
|
|
if "self_attn" in name: return "text_full_attn"
|
|
if ".mlp." in name: return "text_mlp"
|
|
return "text_other"
|
|
for n, p in model.named_parameters():
|
|
buckets[bucket(n)] += p.numel()
|
|
total = sum(buckets.values())
|
|
print(f" TOTAL params : {total/1e9:.3f} B")
|
|
for k, v in sorted(buckets.items(), key=lambda kv: -kv[1]):
|
|
print(f" {k:<18} {v/1e6:9.1f} M ({100*v/total:5.1f}%)")
|
|
|
|
# --- LoRA-attachable leaves ------------------------------------------------
|
|
lin = collections.defaultdict(list)
|
|
for name, mod in model.named_modules():
|
|
if isinstance(mod, torch.nn.Linear):
|
|
lin[bucket(name + ".weight")].append(name)
|
|
print(" nn.Linear leaves by region:")
|
|
for region in sorted(lin):
|
|
names = lin[region]
|
|
tmpl = sorted({re.sub(r"\.\d+\.", ".N.", n) for n in names})
|
|
print(f" {region:<18} {len(names):4d} modules, {len(tmpl)} distinct shapes")
|
|
for t in tmpl:
|
|
print(f" {t}")
|