docs(training-playbook): §4 — when the artifact lies about itself
The playbook covered why a run is SLOW. It did not cover the more expensive
failure: a run that COMPLETES, reports plausible numbers, and is wrong about
itself. Seven of those turned up on the Gemma-4 ERP/RP tune between 08-24 and
08-26 and not one raised an error.
New §4, seven landmines plus a pre-launch checklist:
4.1 a cache key must cover the MEANING of the cached thing. The encode
cache missed the impersonation mask; run 2 would have reused run 1's
unmasked encodings and written impersonation_mask_sha256 into its own
manifest while doing it. No error, no count change, normal loss curve.
4.2 validating a VALUE is not validating the PARAMETER. warmup_ratio was
in range and deleted from transformers 5. Build kwargs as data and
diff the NAMES against the installed signature -- you cannot check the
argument list of a call you have already made.
4.3 record what the run RESOLVED to, never what it requested. Run 1
recorded no attention backend, so an MFU panel profiled the serving
seat under sdpa and recommended adopting flex_attention for a run that
was already using it.
4.4 never train from a dirty tree; harness_commit will name a commit that
does not describe the run. Annotate afterwards, never edit the shipped
artifact -- and state what is NOT wrong, or the note casts doubt on
every field it omits.
4.5 a watchdog whose pgrep pattern appears in its own argv can only ever
return "alive". The inert-gate shape in a liveness check.
4.6 an instrument nobody runs is not an instrument. Mutation-check any
test guarding a property that fails silently.
4.7 fix a stale measurement at the source. "~4.3 HOURS to rebuild the
encode cache" (really 145.5 s) was copied into a new launcher by the
same person who had just measured the real number.
4.8 the pre-launch honesty checklist, ten minutes.
Also:
- Header and framing widened. The file is now a training playbook with a
throughput half and an integrity half; the filename stays for inbound links.
- Sections 4-7 renumbered to 5-8. External refs are all to §1.1 and §3.4 and
are unaffected.
- Four rows added to the superseded-claims table, including the kernel table /
68% quadratic / 8.6% MFU set, which describe the serving seat rather than
the training run.
- gemma4-erp-tune-sizing.md §6 carries a correction banner with the explicit
falls/survives split, because that is the doc someone actually reads before
a run.
This commit is contained in:
@@ -364,6 +364,38 @@ Model-agnostic lessons from this investigation are in
|
||||
probes are at [`scripts/training-probes/`](../../scripts/training-probes/).
|
||||
What follows is Gemma-4-specific.
|
||||
|
||||
> ## ⚠⚠ CORRECTION 2026-08-26 — MUCH OF THIS SECTION MEASURES THE WRONG PROCESS
|
||||
>
|
||||
> **The benchmarks below were run against the SERVING seat with
|
||||
> `attn_implementation="sdpa"` set explicitly. Training was running
|
||||
> `flex_attention` the whole time.** `ATTN_IMPLEMENTATION = "flex_attention"`
|
||||
> was a module constant passed unconditionally into `from_pretrained`, and
|
||||
> run 1's step-time distribution (n=1,445; min 11.84 / p50 19.75 / p99 30.52 /
|
||||
> max 45.79 s/it, the max being step 1's compile) confirms it stayed compiled —
|
||||
> a dynamo fallback sits in the hundreds of seconds per step.
|
||||
>
|
||||
> **FALLS** — describes sdpa, not the training run:
|
||||
> the three-point scaling fit and its 68% quadratic share; the kernel table
|
||||
> (`fmha_cutlassF/B` sm80, `EFFICIENT_ATTENTION`, attention 65.2%); the **8.6%
|
||||
> MFU** figure quoted above and throughout; the projection that elementwise
|
||||
> becomes the largest line item post-fix; and "adopt `flex_attention`" as the
|
||||
> round-two headline lever — **which round one already had.**
|
||||
>
|
||||
> **SURVIVES** — measured on the live training run:
|
||||
> the padding/bucketing win (44.3 → 20.1 s/it); the zero-pad fast-path
|
||||
> second-order effect; the eval-battery noise-floor work.
|
||||
>
|
||||
> ⚠ **Do not assume the direction of the correction.** Training's real MFU is
|
||||
> *unmeasured*, not obviously better. Flex with a BlockMask ought to beat
|
||||
> dense-masked sdpa, but that is a prediction and this investigation has been
|
||||
> unkind to those.
|
||||
>
|
||||
> The root cause was procedural, not technical, and it is written up as
|
||||
> playbook **§4.3**: run 1 recorded no attention backend in its provenance, so
|
||||
> the benchmark/trainer delta was invisible and nobody enumerated it. Run 2
|
||||
> onward records `attn_implementation_requested` **and** `_resolved`, plus the
|
||||
> torch/transformers versions and dynamo's compile counters.
|
||||
|
||||
### 6.1 Where the step time goes
|
||||
|
||||
Real checkpoint, GPU0, `attn_implementation="sdpa"`, PEFT + gradient
|
||||
|
||||
Reference in New Issue
Block a user