The playbook covered why a run is SLOW. It did not cover the more expensive
failure: a run that COMPLETES, reports plausible numbers, and is wrong about
itself. Seven of those turned up on the Gemma-4 ERP/RP tune between 08-24 and
08-26 and not one raised an error.
New §4, seven landmines plus a pre-launch checklist:
4.1 a cache key must cover the MEANING of the cached thing. The encode
cache missed the impersonation mask; run 2 would have reused run 1's
unmasked encodings and written impersonation_mask_sha256 into its own
manifest while doing it. No error, no count change, normal loss curve.
4.2 validating a VALUE is not validating the PARAMETER. warmup_ratio was
in range and deleted from transformers 5. Build kwargs as data and
diff the NAMES against the installed signature -- you cannot check the
argument list of a call you have already made.
4.3 record what the run RESOLVED to, never what it requested. Run 1
recorded no attention backend, so an MFU panel profiled the serving
seat under sdpa and recommended adopting flex_attention for a run that
was already using it.
4.4 never train from a dirty tree; harness_commit will name a commit that
does not describe the run. Annotate afterwards, never edit the shipped
artifact -- and state what is NOT wrong, or the note casts doubt on
every field it omits.
4.5 a watchdog whose pgrep pattern appears in its own argv can only ever
return "alive". The inert-gate shape in a liveness check.
4.6 an instrument nobody runs is not an instrument. Mutation-check any
test guarding a property that fails silently.
4.7 fix a stale measurement at the source. "~4.3 HOURS to rebuild the
encode cache" (really 145.5 s) was copied into a new launcher by the
same person who had just measured the real number.
4.8 the pre-launch honesty checklist, ten minutes.
Also:
- Header and framing widened. The file is now a training playbook with a
throughput half and an integrity half; the filename stays for inbound links.
- Sections 4-7 renumbered to 5-8. External refs are all to §1.1 and §3.4 and
are unaffected.
- Four rows added to the superseded-claims table, including the kernel table /
68% quadratic / 8.6% MFU set, which describe the serving seat rather than
the training run.
- gemma4-erp-tune-sizing.md §6 carries a correction banner with the explicit
falls/survives split, because that is the doc someone actually reads before
a run.
44 KiB
Training playbook — spending a training window without wasting it
Sibling to model-quantization-playbook.md.
That one is for making a model small; this one is for spending a training
window well. Same contract: model-agnostic lessons live here, model-specific
ones stay in the per-model artifact and link up.
First written 2026-08-24 as a throughput playbook, out of the Gemma-4 26B-A4B
ERP/RP tune that ran at 8.6% MFU and cost a four-model frontier panel and most
of a night to explain. §1–§3 are still that: how to find where the step
time went. §4 is the other half, added 2026-08-26 — how to keep a run's
artifacts from lying about what they are. The filename still says
training-throughput-playbook.md because things link to it; the scope is
wider than the name.
The worked example in §8 is that first run. The lessons above it are not about Gemma-4.
Read §1–§3 before hypothesising about kernels. The single most expensive failure in that investigation was not a wrong hypothesis. It was four people, including four frontier models, reasoning confidently from arithmetic instead of spending ten minutes on a measurement that settled it. Two of the panel's conclusions were retracted by their own authors within the hour. Every retraction was a derivation; every survivor was a measurement.
Read §4 before you launch. Every failure in it produced a run that completed, reported plausible numbers, and was wrong about itself. None raised an error. Two of them cost a panel and a night after the fact, chasing a configuration the run was already in.
1. The 10-minute triage — do this FIRST, always
Before you profile, before you read a modelling file, before you ask anyone: measure the step's scaling curve. Three sequence lengths, fixed batch, fwd+bwd, best-of-2 after a warmup.
t(w) = A·w + B·w² w = per-sequence length
Fit two parameters to three points. The residuals tell you which regime you are in, and the regime tells you which lever exists:
observed t(4w)/t(w) |
regime | the lever |
|---|---|---|
| ~4× | linear — per-token work dominates | fewer tokens; fused elementwise |
| ~16× | quadratic — attention dominates | attention backend / kernel |
| ~1× | launch-bound — fixed per-batch cost | CUDA graphs, torch.compile, bigger batch |
If the two-term fit closes with residuals under ~1%, launch-bound is refuted. You did not need a constant term, so there is not a meaningful one. This is the cheapest possible refutation of the most seductive wrong answer, and it costs one extra data point.
⚠ 1.1 ⭐⭐ Three points minimum. A two-point fit with three plausible terms is UNDETERMINED
This is the lesson that cost the most. A two-point fit over {quadratic, linear, fixed} has infinitely many solutions, and which one you land on is decided by whichever per-step number you happened to quote. In the worked example a peer produced two confident, opposite conclusions from the same method inside an hour — "attention is ~5 s of 35" and then "attention is 21–33 s of 35" — because the inputs drifted between attempts.
Three points, two parameters, and check the residuals. If they do not close, you have a third term and you need a fourth point.
⚠ 1.2 ⭐⭐ Benchmark the shape you RUN, not the worst case you can construct
The quadratic share is strongly shape-dependent — in the worked example it
ran 20.9% at w=2,048, 51.3% at w=8,192, 67.8% at w=16,384. A synthetic
max_seq_len benchmark therefore measures the shape where attention looks
worst, and generalising from it overstates the attention prize by ~1.3×.
Get the real distribution off the encode cache and weight by it:
E[t] = A·E[w] + B·E[w²]
E[w²] is not E[w]². For a bimodal length distribution they can differ
by 2× or more, and a quadratic term is dominated by the rare long batches that
an E[w]² shortcut averages away. In the worked example E[n²]/E[n]² was
2.08.
Sanity check the weighted prediction against the observed s/it before you
trust any of it.
2. The reference probe set
Committed at scripts/training-probes/.
Run them in this order; each is minutes and none needs the real checkpoint
except the profiler.
| probe | what it settles | needs GPU? |
|---|---|---|
step0_mask.py |
mask band structure + which layers keep the fast path | no |
step2_padding.py |
padding waste, length distribution, CE chunk sizing | no |
step_bucket.py |
bucketing gain, bucket-size sweep, root diversity | no |
step1_profile.py |
scaling fit, padding penalty, CE wall clock, kernel table | yes |
step1_profile.py loads the real model but reuses the harness's own
discover_target_modules and compute_loss, so it measures the thing that
actually runs rather than a re-implementation. Keep that property when you
adapt it — a probe that reimplements the training step measures the probe.
3. The recurring landmines
3.1 ⭐⭐ Right-padding is a compute tax AND a backend tax
Everyone knows padding wastes tokens. The second effect is the one that gets
missed: an explicit padding mask can knock fast-path-eligible layers off
is_causal.
scaled_dot_product_attention takes is_causal=True or an attn_mask,
never both usefully. HF sets is_causal=True only when attention_mask is None. Right-pad a batch and you hand it a 2D mask, it materialises a 4D
tensor, and every layer that could have taken the clean causal route now takes
a masked dense one.
Measured, same width, same n, only the mask differing:
no padding 35.244 s 26,048 loss targets
50% pad on one row 38.567 s 19,640 loss targets
9.4% slower for 24% less work. Verify this on your own stack with
step0_mask.py — it prints whether create_causal_mask returns None or a
tensor for each mask case.
3.2 ⭐⭐ Length-bucket to PAIR, shuffle micro-batches to MIX — and the bucket should be TIGHT
Naive length-bucketing has a real hazard: length correlates with data source, so length-homogeneous batches are source-homogeneous batches, and an accumulation window can end up drawing its entire gradient from one root.
The fix costs nothing: form micro-batches within length buckets, then shuffle the resulting micro-batches globally. Padding efficiency is a property of the pairing alone, so all of the saving survives the shuffle.
The non-obvious part — bucket size is not a diversity knob. Measured across a 256× range of bucket sizes, roots per accumulation window stayed flat at 3.54–3.61 (against 3.68 for a pure shuffle). The global micro-batch shuffle does all of the mixing; the bucket contributes nothing to diversity and only costs padding. So use the tightest bucket you can — which in the limit is a full length sort.
| bucket | padding waste | zero-pad micro-batches | roots/window |
|---|---|---|---|
| current (shuffle) | 29.9% | 0.1% | 3.68 |
| 2 | 0.0% | 78.3% | 3.56 |
| 32 | 0.1% | 41.9% | 3.55 |
| 512 | 2.4% | 4.1% | 3.61 |
Note the zero-pad micro-batches column — that is §3.1 compounding. A tight
bucket does not merely cut tokens, it puts most batches back on the causal
fast path.
Peak memory does not rise. padded = batch × max(len), so one long record
forces a full-width batch regardless of its partner. Bucketing pairs long
records with each other, which roughly halves the number of worst-case
batches.
3.3 ⭐⭐ Check the kernel GENERATION, not just the backend name
The backend name (EFFICIENT_ATTENTION, FLASH_ATTENTION, …) is not the whole
story. Read the actual kernel symbols out of the profiler:
fmha_cutlassF_bf16_aligned_32x128_gmem_sm80
fmha_cutlassB_bf16_aligned_128x64_k65536_sm80
^^^^
sm80 is Ampere. Those were running on an sm_120 Blackwell card, on the
dominant cost centre of the step. A backend can be "selected correctly" and
still be a generation behind, and nothing in the config surface tells you.
Also read the variant suffix: gmem on the forward kernel is the
global-memory fallback tier of the memory-efficient path, chosen when the
working set will not fit in shared memory. Wrong backend and that backend's
slow path.
3.4 ⭐ key_averages() double-counts — use device-kernel rows only
torch.profiler's key_averages() table lists both the ATen op and the CUDA
kernel it launched, each carrying the same self_device_time_total. Summing
the whole table gives you roughly 2× the real step time.
The tell is exact equality between an aten:: row and a kernel row:
aten::_efficient_attention_backward 30 16144.6
fmha_cutlassB_bf16_aligned_128x64_k65536 30 16144.6
Filter to device kernels (void …, fmha_…, cutlass::…, Memcpy…) and
sanity-check the total against the measured wall clock. In the worked example
the filtered total came to 89.5% of the step, which is the right shape; the
unfiltered total came to 202%.
3.5 ⭐ Time the loss forward AND account for its backward recompute separately
If the loss head is gradient-checkpointed, a CUDA-event window around the
forward loop measures half the story at best — the recompute happens inside
.backward(), outside your window.
State the caveat explicitly when you report the number. In the worked example the CE forward measured 374 ms of a 35.3 s step (1.1%); even at 3× for recompute-plus-backward it is ~3%, which was enough to kill a proposed dependency swap — but "1.1%" alone would have been an unearned claim.
3.6 ⭐⭐ Assert mask band structure directly; never infer it from performance
transformers can silently skip mask creation and pass
attention_mask=None when a mask function is not registered. If that fires on
a sliding-window model, the windowed layers do full causal attention — not a
speed bug, a different model from the one you will serve.
There is a tempting alibi: "if constraints were dropped we would be on the fast path and fast; we are slow, therefore correct." It is decent evidence and it is not an assertion. Materialise the mask once and count allowed positions per row:
sliding_attention max 1,024 allowed/row, saturates at row 1,023 PASS
Thirty seconds, on CPU, no weights. Do it before every run that changes the masking path, and before believing any optimisation result.
3.7 ⭐ "Bit-identical output from a different backend" — ask could this have disagreed?
A backend flag that produces max_abs_diff == 0.0 against the reference is
either (a) legitimately the same GEMMs behind a different launcher, or (b) a
flag that never took. Argument cannot separate these — in the worked
example three frontier models split 2–1 on it and the majority was not
obviously right.
Do not resolve it by vote. Count kernel launches. A per-expert loop leaves
n_experts dispatches per layer visible; a grouped path leaves one. That is
unambiguous and falls out of a trace you are running anyway.
Related trap: a trace of the default path does not test the flag. If the run was relaunched without the flag set, the profile tells you what the default does and nothing about the flag. Say so rather than over-claiming.
3.8 ⭐ MFU is a denominator argument waiting to happen — report the decomposition instead
MFU invites an unwinnable fight about what counts as a FLOP (active vs dense params for MoE, whether checkpoint recompute counts, whether frozen-base skipped GEMMs count). That fight consumed an hour of the worked example and produced nothing.
Report these beside MFU, not instead of it:
- tokens/s, and real (unpadded) tokens/s separately
- achieved hardware FLOPs straight from the profiler
- the time decomposition (attention / GEMM / elementwise / other)
Then the denominator stops mattering.
The reading that actually diagnosed it was not an MFU number at all:
100% SM utilisation at 279–292 W, running 27 TFLOPS, on a card that does 304 TFLOPS on a dense GEMM at the same power.
SM-busy, tensor-core-idle. The chip is fully occupied doing work that is not matrix multiplication. No FLOP-counting convention changes that, and it points straight at the kernel table.
3.9 Frozen-base LoRA is ~4ND, not ~6ND — and the arithmetic intensity does NOT drop
A claim that circulated and was wrong: "frozen-base LoRA has structurally lower arithmetic intensity, so a dense-GEMM ceiling is unreachable in principle."
The correct accounting: forward is 2ND, input-gradient backward through the frozen weights is 2ND, and only the weight-gradient (~2ND) is skipped. So ~4ND against ~6ND — two-thirds of the work, at the same arithmetic intensity per remaining GEMM. You do fewer GEMMs; the ones you do are exactly as dense.
Gradient checkpointing is a separate, real ~⅓ recompute tax. Account for it separately rather than folding it into an intensity story.
3.10 ⭐⭐ Prove the SERVING path before you spend the training window
Playbook-for-quants §4.1 says prove the quantization targets before spending GPU time. The same rule applies one step later and is easier to skip: prove you can serve the artifact, in the shape you intend to serve it, before you train it.
Worked failure, 2026-08-25. A ~7-hour LoRA run was built on the assumption that the adapter could be hot-swapped onto a quantized base at serve time. The sizing doc had flagged this correctly — "serving the result is not settled… if it still no-ops, the harness must emit merged weights, and Eitri needs that requirement while he is early, not after the run" — and then the check was deferred rather than run. Tested after the fact:
AttributeError: To support LoRA for MoE model,
'get_expert_mapping' must be implemented
One grep would have found it. vllm/lora/utils.py::process_packed_modules_mapping
branches on is_moe_model(), and the model class in question implements zero
occurrences of get_expert_mapping. Static fact about the serving stack,
available months before the run.
The check is cheap and mechanical:
# does the serving engine's model class support what you plan to do?
grep -c "SupportsLoRA\|get_expert_mapping" <engine>/model_executor/models/<arch>.py
# and: start the engine with the feature flag ONLY (no adapter needed).
# --enable-lora alone forces the machinery to initialise, which is where it fails.
Two generalisations worth carrying:
- Feature support is per-architecture, not per-family. LoRA worked for the dense sibling of this exact model family and not for the MoE one. "Model X is supported" is not a statement about X's variants.
- A capability gap in the serving engine is not fixable by the training side. No harness change, no quantization choice, and no adapter scoping works around it — the adapter here never touched experts and was refused anyway, because the refusal keys on the model being MoE.
The recovery is usually fine (merge instead of hot-swap, at ~35 min per tune). The point is that it should be a decision made before the window, not a discovery made after — because the alternative it forecloses may be an architecture choice, and by then you have already trained.
3.11 Base-viability pre-flight — three greps, before you pick
Run this on any candidate base BEFORE committing a training window. Each check is minutes; skipping them cost a night in 2026-08.
1. Does it fit for TRAINING? BF16 weights on one card, with room for the real peak — not the weight figure.
ana-ml2 reference: Gemma-4 26B-A4B is 48.1 GiB of weights and peaks at
79.7 GiB at micro-batch 2 / seq 16,384. So ~48 GB of weights is close to
the practical ceiling for a 97.9 GiB card at that shape.
⚠ Model-line names lie about size. "Mistral Small 4" is 119 B — 238 GB in
BF16, more than both cards combined. Read params.json / config.json, never
the name.
⚠ QLoRA is NOT an escape hatch for MoE. bitsandbytes walks nn.Linear, and
fused 3-D expert parameters are not that — see quantization playbook §3.15.
2. If MoE — does the serving engine implement the expert mapping?
grep -c "def get_expert_mapping" <engine>/model_executor/models/<arch>.py
Zero means LoRA cannot be served at all and merged weights are mandatory.
Measured: gemma4*.py → 0 (refuses); deepseek_v2.py, mixtral.py,
glm4_moe.py, ernie45_moe.py → present.
3. Does the model class support LoRA? ⚠ Grep the class, not the file — capability is usually INHERITED and a file-level grep misses it entirely:
from vllm.model_executor.models.<mod> import <Class> as C
print([c.__name__ for c in C.__mro__])
print(hasattr(C, "get_expert_mapping"), getattr(C, "supports_lora", None))
mistral.py greps as SupportsLoRA=0 and is fully LoRA-capable — it inherits
from LlamaForCausalLM. mistral_large_3.py greps as 0 for both and inherits
get_expert_mapping from DeepseekV3ForCausalLM. Both file greps are wrong;
only MRO resolution is right. (Same failure as asserting a substring instead of
an effective value.)
Worked results, 2026-08-25:
| base | fits (1) | MoE mapping (2) | LoRA (3) | verdict |
|---|---|---|---|---|
| Gemma-4 26B-A4B | ✅ 48 GB | ❌ absent | n/a | trainable, merge-only |
| Mistral Small 4 119B | ❌ 238 GB | ✅ via DeepSeek-V3 | ✅ | servable w/ hot-swap, not trainable here |
| Ministral 3 14B | ✅ ~28 GB | n/a (dense) | ✅ inherited | passes all three |
Architecture shape is worth a fourth glance, because it predicts how much
of this playbook you will need. Uniform head_dim ≤ 128 with no sliding window
means flash AND cuDNN are both reachable and §3.1/§3.3 simply do not apply.
Mixed head dims plus a sliding window — Gemma-4's shape — is what forces dense
O(n²) attention on Ampere-generation kernels and costs 65% of the step.
3.12 ⭐⭐ Merging a tune "back toward the base" can UNDO an abliteration
A common community remedy for an overfit tune is a partial merge back toward the base — 50/50 or similar — to recover general capability. On an abliterated base that remedy is an undo, not a heal, and it is silent.
The published recipes that recommend it merge back into the stock instruct
checkpoint (e.g. google/gemma-4-*-it). If you follow them literally on a
project whose base was abliterated, you re-introduce exactly the refusal
directions the abliteration was run to remove. The tune still looks "healthier"
on general benchmarks while the property the seat exists for quietly returns.
Rule: any merge-back must target the SAME base the LoRA was trained against. Never the upstream stock weights, however similar the name.
⚠ The generalisation is wider than merge-back. Community recipe cards are per-checkpoint artifacts and their findings do not transfer across:
- dense vs MoE — different training dynamics, different memory profile, different everything
- stock vs abliterated/uncensored — different alignment surface
- size variants of the same family — different optima
Real 2026-08-25 example: a recommendation was carried across from a recipe card for a dense, stock 31B onto a MoE, abliterated 26B-A4B, on the strength of the shared model family name. The overfitting warning on that card came from the right architecture; the five-stage pipeline, reward stacks and merge-back came from the wrong one. Same family, three axes apart.
Before quoting any recipe card at a decision, state which checkpoint it was written for and which axes differ from yours. If the answer is "same family" that is not an answer.
3.13 ⭐⭐ Measure refusal retention on the axis the ABLITERATION targeted
Two distinct lessons from 2026-08-25, both about measuring the wrong thing confidently.
A tune can re-install what an abliteration removed, and no capability gate will see it. If you tune AFTER abliterating, the tune has every training token as an opportunity to walk the abliteration back. A reasoning/craft/memorisation gate measures none of that: a tune that gains 41 items of contradiction detection and quietly restores refusals is a failed seat that passes every check. Add a compliance axis explicitly — it will not fall out of the others.
But measure the axis the abliteration was FOR. This is the trap, and it is easy to walk into precisely because a general harm set is sitting right there, cached, with a recorded baseline.
abliteration run so the model engages EXPLICIT FICTION
probe used: mlabonne/harmful_behaviors (weapons, malware, fraud)
Those are different refusal surfaces and a model moves on them independently. The measured result — 29/100 general-harm refusals on a tune whose prose the operator was actively praising — is not obviously a defect and may be desirable: general-harm refusals returning while domain compliance holds is close to the ideal shape for an internal creative seat. The number was real; its relevance was assumed.
Read the interesting cell. In 29 hard / 0 deflect / 71 comply, the
load-bearing figure is 71. Stock refused 100/100; anything near that would
mean the abliteration was undone. 71 complying says "partially walked back on
one axis", which is a completely different finding — and only one of the two
threatens the seat.
⚠ A baseline from a different harness is not a baseline. The recorded "3/100" for that base came from the abliteration tool's own scorer, which works off first-token probability distributions. A probe that generates 256 tokens and regexes them is a different instrument; the two can disagree in both directions. Run your own probe against BOTH arms on the SAME seat, or report the number alone and say the comparison is missing.
⚠ A refusal regex undercounts — models decline by redirecting, with no refusal token present. Classify three ways (hard / deflect / comply). And note the free discriminator: if both arms return zero deflections the model is binary; if only one does, the regex is fine and the difference is real. An instrument artifact does not care which arm it runs against.
4. When the artifact lies about itself
§1–§3 are about a run that is slow. This section is about a run that completes, reports plausible numbers, and is wrong — and about the derived artifacts that go on repeating the wrong thing afterwards.
Every failure below was found on the Gemma-4 ERP/RP tune between 2026-08-24 and 2026-08-26. Not one of them raised an error. They are ordered by how much they cost.
The shape they share, stated by brokkr-smithy-dev on 2026-08-26 after the third instance in a day: when you change what an artifact means, every derived artifact keyed on the old meaning is now a liar. Caches, logs, comments, manifests, benchmark write-ups. The fix is always the same — put the meaning in the key — and the reason it keeps happening is that the old artifact still loads, still parses, and still looks right.
⚠ 4.1 ⭐⭐ A cache key must cover the MEANING of the cached thing, not just its inputs
The encode cache for the ERP corpus was keyed on corpus identity, max_seq_len,
base-model path and chat-template sha. Run 2 then added an impersonation
loss-mask — 813 turns whose labels change from trained to ignored.
The mask was not in the key. Run 2 would have hit run 1's 609 MB cache, reused
its unmasked encodings, trained the impersonation straight back in, and
written impersonation_mask_sha256 into its own provenance manifest while doing
it.
Nothing downstream could have caught it:
| signal | what it would have shown |
|---|---|
| error / exception | none — a cache hit is the happy path |
| sample count | unchanged — 0 samples were fully masked |
| record count | unchanged — 20,982 either way |
| loss curve | normal |
| provenance manifest | asserts the mask was applied |
A seven-hour run whose artifact claims a property it does not have, invisible from every number anyone would think to check.
The rule. A cache key covers every input that can change the semantics of the output, not just the ones that change its shape. Concretely:
- Hash the content of every auxiliary file (mask, filter list, label map), not its path. A path is stable across an edit; that is the whole problem.
- Carry an explicit
ENCODE_VERSIONinteger and bump it whenever the encoder's output can change for identical inputs. Cheap, blunt, and it catches the cases you did not think to hash. - Extract the key computation into a named, tested function. Ours became
core.encode_cache_keyspecifically so a test could assert that a masked run and an unmasked run get different keys — and that a legitimate resume still gets the same one. An inline dict comprehension inside the loader cannot be tested and will not be.
def test_encode_cache_key_separates_a_masked_run_from_an_unmasked_one():
unmasked = encode_cache_key(ordered, **common)
masked = encode_cache_key(ordered, **common, impersonation_mask_sha256="d"*64)
assert unmasked != masked
# and the cache must still HIT on a legitimate resume
assert masked == encode_cache_key(ordered, **common, impersonation_mask_sha256="d"*64)
Generalises past caches. Any memoised, derived, or checkpointed artifact has this problem: encode caches, tokenised datasets, precomputed embeddings, distillation logits, eval-result files. If it was derived under one meaning and is reused under another, it is now a liar.
⚠ 4.2 ⭐⭐ Validating a VALUE is not validating the PARAMETER
Run 2 died after the full encode and after all 1,013 weight shards had loaded:
TypeError: TrainingArguments.__init__() got an unexpected keyword
argument 'warmup_ratio'
warmup_ratio exists in transformers 4. It is gone in 5.15.1, which keeps
only warmup_steps. The harness had careful config-level validation — it
checked 0.0 <= warmup_ratio < 1.0 and rejected an unknown scheduler name — and
none of it could have caught this. The value was in range. The parameter had
been deleted.
The check that catches it has to run against the installed library, not against your own schema:
def assert_training_arguments_accepts(kwargs: dict, cls) -> None:
accepted = set(inspect.signature(cls.__init__).parameters)
unknown = sorted(set(kwargs) - accepted)
if unknown:
raise TypeError(f"{cls.__name__} does not accept {unknown}. "
f"It accepts: {', '.join(sorted(accepted))}")
Which requires building the kwargs as data first — you cannot check the argument list of a call you have already made. That single structural change is the lesson; the signature diff is trivial once the kwargs are a dict.
Three riders:
- Keep the portable unit in config, convert at the call site. We kept
warmup_ratioin the config file and convert towarmup_stepson the wire. A ratio survives a change in corpus size; a step count silently becomes a different fraction of the run. - Mirror the framework's own arithmetic exactly. HF ceilings twice — records into micro-batches, then micro-batches into optimizer steps. One combined division gave 1,311 steps for our corpus where the trainer reports 1,312, which would have put warmup a step short. Pin it with a test against a step count a real run actually printed.
- The failure mode is the expensive one: late. Anything checkable from the config, the library signature, or a file on disk belongs before the tokenizer, the encode, and the model load. Ours now dies in under a second.
⚠ 4.3 ⭐⭐ Record what the run RESOLVED to, never what it requested
Run 1's provenance recorded no attention backend at all. Six weeks of conclusions rested on the answer.
An MFU investigation had profiled the serving seat with
attn_implementation="sdpa" explicitly set, produced a kernel table
(fmha_cutlass*_sm80, EFFICIENT_ATTENTION, attention 65.2% of step), an 8.6%
MFU figure, and a headline recommendation: adopt flex_attention for round
two.
Training had been running flex_attention the entire time.
ATTN_IMPLEMENTATION = "flex_attention" # module constant
...from_pretrained(..., attn_implementation=ATTN_IMPLEMENTATION)
A panel, a kernel profile and three rounds of arithmetic went toward a configuration the run was already in. The single biggest round-two optimisation identified was a thing training already had.
The correction has to be stated as a split, because some of the work survives. After the flex finding, on this investigation:
FALLS — describes the sdpa seat, not the training run
the three-point scaling fit and its 68% quadratic share
the kernel table
the 8.6% MFU number
"adopt flex_attention" as the round-two headline lever
SURVIVES — measured on the live training run
the padding/bucketing win, 44.3 -> 20.1 s/it
the zero-pad fast-path second-order effect
the eval-battery noise-floor work (a different instrument entirely)
⚠ Do not assume the direction of the correction. Training's real MFU is now unmeasured, not obviously better. Flex with a BlockMask ought to beat dense-masked sdpa — but that is a prediction, and predictions have done badly here.
What to record, and why two fields:
"attn_implementation_requested": ATTN_IMPLEMENTATION, # a constant
"attn_implementation_resolved": model.config._attn_implementation, # what happened
"torch_version": torch.__version__,
"transformers_version": transformers.__version__,
"dynamo_counters": _dynamo_counters(), # best-effort, nullable, never raises
Requested and resolved are different claims. The requested value is a line in your source; the resolved value is what the library settled on after checking availability, and only that one describes the run. A framework that silently downgrades an unavailable backend will make them differ, and that difference is exactly what you want on the record.
The dynamo counters matter for the same reason: dynamo's fallback to uncompiled flex is silent, roughly 20× slower, and documented not to work correctly through the backward pass. The counters are the only in-band evidence it did not happen.
When the run is already going and the field is missing — as ours was — you can often still answer it, but only forensically. For us:
- Source, for what was requested (unconditional constant, no fallback flag).
- A config-only side probe on the same library versions, for what it
resolves to:
AutoConfig.from_pretrained(path, attn_implementation=...)then read_attn_implementation. No weights, no GPU, seconds. - The step-time distribution, for whether it stayed compiled. Run 1's 1,445 logged steps ran min 11.84 / p50 19.75 / p99 30.52 / max 45.79 s/it, the maximum being step 1's compile. A dynamo fallback would sit in the hundreds of seconds per step. Nothing in the trace approaches it.
That is three sources of evidence to replace one recorded field, and it only worked because the source tree happened to still be on disk. Record the field.
⚠ 4.4 ⭐⭐ Never train from a dirty tree — the provenance commit will be a lie
Run 1's manifest recorded harness_commit 35a4e8e. The working tree carried
224 uncommitted lines across two modules for the entire run, so the recorded
commit predates the code that trained the adapter. The artifact is not
reproducible from the commit it names, and nothing says so.
_git_commit() calling git rev-parse HEAD is not wrong; it is incomplete.
It answers "what is HEAD" when the question is "what code ran."
Add a cleanliness assertion to the pre-flight, beside the corpus and holdout checks:
git diff --quiet && git diff --cached --quiet || {
echo "REFUSING: working tree is dirty; harness_commit would not describe this run"
exit 1
}
Record git describe --always --dirty rather than a bare SHA if you want a
softer version, but an outright refusal is better: a run long enough to be worth
provenance is long enough to be worth one commit first.
Correcting it after the fact — annotate, never edit. We left
provenance.json untouched and wrote a PROVENANCE-NOTE.md beside it. Editing
a shipped artifact so it says something it did not say is the worse failure.
⚠ State what is NOT wrong. A bare correction note casts doubt over every field it does not mention, and the next reader has no way to tell which. Ours ends by listing what remains accurate — recipe sha, root shas, base model path and revision, template sha, the override triple, the dropped ids.
⚠ 4.5 ⭐ A watchdog whose pattern matches its own argv can only ever return "alive"
The training monitor polled liveness with:
while pgrep -f "erp_sft_harness --config" >/dev/null; do sleep 60; done
echo "PROCESS EXITED"
pgrep -f matches full command lines — including the monitor's own, because
the pattern is right there in its argv. The loop matched itself, so the exit
branch was unreachable for every possible input. The run crashed and the
watchdog reported nothing, because it was watching itself.
It also blocked the recovery: the launcher's already-running guard used the same pattern, matched the monitor, and refused to start the replacement run.
This is the inert-gate shape in a liveness check — a test that cannot return the failing verdict for any input. It is the same defect as an assertion comparing a value to itself, and it hides better, because a watchdog that never fires looks exactly like a system that never breaks.
Fixes, in order of preference:
RUNPID=$(pgrep -f "<pattern>" | head -1) # resolve ONCE, from a clean shell
while kill -0 "$RUNPID" 2>/dev/null; do sleep 60; done
- Poll a captured PID, not a pattern.
kill -0cannot self-match. - If you must pattern-match, break the literal so it is absent from your own
argv (
"erp_sft_harn""ess --config"), and put the guard in a file rather than an inlinessh host '...'— the invoking command line is argv too. - Never
pkill -fa shared pattern: it is handle-blind and kills every other agent's monitor on the box along with yours.
Test the negative. Whatever the gate is, construct the input that should make it fire and confirm that it does. Every gate in this project that has ever caught anything was one somebody deliberately broke first.
⚠ 4.6 ⭐ An instrument nobody runs is not an instrument
The harness test suite was 10 passed / 4 failed, and there was no pytest installed in the training venv at all — so nothing had ever run it. A run had already shipped an adapter through it.
The good version of that news: all four failures were stale tests, not broken
code. Each asserted a behaviour that had been deliberately changed —
chat_template_path became required, the provenance schema grew three keys,
persona trimming gave way to the unfittable path. The bad version: nobody knew
that, because the instrument was dark.
- Installing the test runner is part of standing up the training venv, not a later nicety. It is three pure-Python packages and touches nothing in the torch/transformers stack.
- Repair stale tests to the current contract; do not delete them. Each rewrite is an opportunity to write down why the behaviour changed — ours now carry the corpus measurement showing the removed persona-trimming costs zero samples on this corpus.
- Mutation-check any test guarding an invisible property. A test for a silent failure is itself silent when it is vacuous. Break the code deliberately and confirm the test fails:
| deliberate break | test that must fail |
|---|---|
| OR-merge → last-wins | whole-run masking |
| drop the sha assertion | corpus/mask disagreement |
| resolve indices after the merge | source-index resolution |
| re-add the removed kwarg | installed-signature check |
If breaking it changes nothing, you have documentation, not a test.
⚠ 4.7 ⭐ Fix a stale measurement at the SOURCE, or the copy carries it forward
A launcher script carried the comment "the 609 MB encode cache (2.5 min to reuse, ~4.3 HOURS to rebuild)." That figure predated the encoder's own parallelisation. Measured on run 2: 145.5 seconds on 32 workers. Off by a factor of 106.
It was believed twice — once to project an 11.7-hour window for a 7.6-hour job, and once when it was copied verbatim into a new launcher written by the same person who had just measured the real number. The stale figure propagated in the same motion that was supposed to retire it.
- When a measurement in a comment is superseded,
grepthe repo for the number before you fix the one in front of you. - Copying a header wholesale copies its claims wholesale. Re-read boilerplate you paste for facts that have expired.
- Prefer a dated measurement in prose — "145.5 s on 32 workers, measured 2026-08-26" — over a bare figure. A dated claim invites a re-check; a bare one reads as timeless.
4.8 The pre-launch honesty checklist
Ten minutes, before the window opens. Every item is something that produced a completed, plausible, wrong run above.
[ ] working tree committed git diff --quiet && git diff --cached --quiet
[ ] test suite green and the runner is actually installed
[ ] cache key covers the change bump the version integer; hash aux file CONTENT
[ ] kwargs checked by NAME against the installed library signature
[ ] config validated before tokenizer, encode and model load
[ ] provenance records RESOLVED backend, library versions, aux-file shas
[ ] present-and-null, not absent a run that claims nothing must say so explicitly
[ ] watchdog tested negative kill something and confirm it fires
[ ] stale numbers grepped for the figure, repo-wide, not just in view
⚠ The present-and-null line is load-bearing and the least obvious. Emit
every provenance key always, null when unused. A manifest that omits a key
when there was nothing to report is indistinguishable from one written by a
harness too old to know the key exists — so an adapter trained without your
safeguard becomes byte-indistinguishable from one trained with it. Present-and-
null is a statement; absent is an accident.
5. Panel / consult discipline for perf work
Perf investigations are unusually good at generating confident wrong answers, because the arithmetic is easy and the ground truth is expensive. Specific guards, learned the hard way:
- Every arm's claim gets a measurement or an expiry date. In the worked example the panel produced four self-retractions in ninety minutes. The measurements produced zero.
- Treat cross-arm agreement as weak evidence. Ask arms to attack a hypothesis rather than extend it; agreement among similarly-primed readers of the same artifact is not independent confirmation.
- A dispute about what a specific dispatcher does is a question of fact. Do not put it to a panel. Instrument it.
- When an arm says "you missed X," check what they read. If your settled artifact was not on their reading list, the "miss" is usually restatement-of-a-settled-prior, not a genuine gap.
6. Superseded claims — do not follow these
| claim | status | replaced by |
|---|---|---|
"Explicit mask → EFFICIENT_ATTENTION" is over-specific; Blackwell defaults to CUDNN_ATTENTION |
WRONG (2026-08-24) | Measured: sm_120 selects fmha_cutlass*_sm80, i.e. EFFICIENT_ATTENTION. The original claim was right. |
| Attention's quadratic share is ~5 s of a 35 s step | WRONG (2026-08-24) | Measured 22.8 s / 65.2% at w=16,384; 67.8% by independent scaling fit |
| Frozen-base LoRA has structurally lower arithmetic intensity | WRONG (2026-08-24) | ~4ND vs 6ND at unchanged intensity — see §3.9 |
| The chunked CE is a 2–5× under-estimated cost centre | WRONG (2026-08-24) | Measured 1.1% of step forward, ≲3% with recompute |
attn_implementation="flash_attention_2" is the per-layer lever |
NOT A FLAG (2026-08-24) | All-or-nothing at from_pretrained; per-layer needs a custom fn on ALL_ATTENTION_FUNCTIONS. FA2 also caps head_dim at 256. |
| Bucket size ~256 is needed to preserve source diversity | UNNECESSARY (2026-08-24) | Diversity is flat in bucket size; the global micro-batch shuffle does that work — see §3.2 |
| The 2026-08-24 kernel table / 68% quadratic share / 8.6% MFU describe the TRAINING run | WRONG (2026-08-26) | They describe the serving seat, benchmarked with attn_implementation="sdpa" set explicitly. Training ran flex_attention throughout. Training's real MFU is now unmeasured — see §4.3 for the full falls/survives split, and do not assume the correction's direction. |
"Adopt flex_attention" is the round-two headline lever |
ALREADY BANKED (2026-08-26) | It was live in round one. A panel, a kernel profile and three rounds of arithmetic went toward a configuration the run was already in — see §4.3 |
| The encode cache takes ~4.3 hours to rebuild | WRONG (2026-08-26) | 145.5 s on 32 workers, measured on run 2. The stale figure predated the encoder's own parallelisation, was used to project an 11.7 h window for a 7.6 h job, and was then copied verbatim into a new launcher — see §4.7 |
warmup_ratio is a valid TrainingArguments kwarg |
REMOVED IN transformers 5 (2026-08-26) | 5.15.1 keeps only warmup_steps. Keep the ratio in config, convert at the call site, and diff kwarg NAMES against the installed signature — see §4.2 |
7. Measured negatives — don't re-chase
- Fused MoE kernel (
grouped_mm) as the throughput fix. Measured 0.9% slower than the Python loop and bit-identical. Independently, dense GEMM is only 7.9% of the step, so the whole category is capped near 10%. - CUDA graphs /
torch.compileover the expert loop. The two-term scaling fit closed without a constant term, so there is no meaningful fixed per-batch cost to amortise. ~3,840 expert-GEMM launches per forward are not what you are paying for. liger-kernelfused linear CE. Real and correct, but a ~1–3% lever on this shape. Not a project.- FlashAttention-4 on sm_120. Public reports are sour — one measurement of 1.07× over FA2, and an sm_120 patch people could not get working that fell back to torch SDPA. Do not bet a round on it.
8. Worked example — Gemma-4 26B-A4B ERP/RP tune, 2026-08-24
Model-specific detail lives in
gemma4-erp-tune-sizing.md §6. The short version,
because the shape of the investigation is the transferable part:
Symptom. 8.6% MFU, ~35–46 s/it, 1,312 steps, ~13.9 h ETA.
What the panel produced. Four frontier arms plus an orchestrator, over ninety minutes: a sliding-window hypothesis, a retraction of it, a retraction of the retraction, a correctness scare that resolved itself, two mutually contradictory readings of one dispatcher, and four self-corrections.
What settled it, in about twenty minutes of GPU time:
scaling fit (3 points, 2 params, residuals <3 ms over an 8× range)
A = 6.87e-4 s/token B = 8.85e-8 s/token²
quadratic share: 20.9% @ w=2,048 → 67.8% @ w=16,384
kernel table (device rows only)
attention 22,835.8 ms 65.2% fmha_cutlass*_sm80
dense GEMM 2,774.0 ms 7.9%
other 5,739.0 ms 16.4%
Two independent methods, 2.6 points apart. Attention was the answer, on Ampere-generation kernels, with the forward on a global-memory fallback tier.
The largest actionable win was not the attention kernel. It was a sampler change — bucket-to-pair, shuffle-to-mix — worth 29.9% of tokens and ~35.5% of wall clock, with no new dependency, no kernel work, and unchanged peak memory. It also wins under every branch of the diagnosis, which is why it was recommended while the rest was still unresolved.
The transferable ordering:
- Assert correctness (mask band structure). Everything downstream assumes it.
- Scaling curve. Names the regime in ten minutes.
- Kernel table. Names the cost centre.
- Data-side levers first (padding, bucketing) — they need no dependency and they multiply into every other cost.
- Kernel/backend levers last, gated on 2 and 3.