perf(gen-seat): record prefill measurements — roughly doubled

Closes the one axis of the original premise left unverified. Measured
cold (cache-busted) on both builds under matching serve configs:

  ~6.7k-token prompt   3,206 -> 6,334 tok/s prefill   (+98%)
  ~27k-token prompt    2,862 -> 5,085 tok/s prefill   (+78%)
  TTFT on a ~27k doc    9.43 -> 5.31 s                (-44%)

Prefill gains far exceed the +18% decode gain, and that ordering is the
expected one: decode at bs=1 is memory-bandwidth-bound and the weights
are 4-bit under either scheme, so little changes; prefill is
compute-bound, which is where native Blackwell FP4 tensor cores replace
the Marlin dequant-to-BF16 path. The summarizer aliases are the
consumers that feel this.

Adds bench/prefill_bench.py plus the raw JSON. The harness deliberately
uses SystemRandom: a seeded nonce regenerates the previous run's prompts
verbatim, prefix caching then serves them, and the first attempt read
~41k tok/s of cache-hit rather than ~5k of actual prefill.
This commit is contained in:
vh
2026-08-15 02:33:50 -07:00
parent fa4f652a39
commit 4a5c3fcccf
5 changed files with 119 additions and 0 deletions
+12
View File
@@ -90,13 +90,25 @@ Speed alone does not justify cutting over a seat backing 7 LiteLLM aliases.
| metric | W4A16 (old) | mixed (new) | delta |
|---|---|---|---|
| decode tok/s, bs=1, cache-busted | 80.12 | **94.53** | **+18.0%** |
| **prefill tok/s, ~6.7k prompt** | 3,206 | **6,334** | **+98%** |
| **prefill tok/s, ~27k prompt** | 2,862 | **5,085** | **+78%** |
| TTFT on a ~27k-token doc | 9.43 s | **5.31 s** | −44% |
| MTP acceptance | 47.8% | 47.7% | unchanged |
| perplexity, 6 held-out passages | 6.941 | 7.059 | +1.7% worse |
| abliteration compliance | 4/4 | 4/4 | preserved |
| weights on disk | 27.7 GB | 22.5 GB | −19% |
**Prefill roughly doubled** — the bigger practical win, and exactly what theory predicts:
decode at bs=1 is memory-bandwidth-bound (weights are 4-bit either way, so little changes),
while prefill is compute-bound and is where Blackwell's native FP4 tensor cores replace the
Marlin dequant-to-BF16 path. This is what the `summarizer` / `summarizer-large` aliases feel
on long documents.
- `quickbench.py` — cache-busted bs=1 decode + MTP acceptance. **Bust the cache:** with
a fixed prompt, prefix caching returns byte-identical timings and you measure nothing.
- `prefill_bench.py` — TTFT on long prompts. Same trap, worse: a *seeded* nonce reproduces
the previous run's prompts verbatim, so prefix caching serves them and you read ~41k tok/s
of cache-hit instead of ~5k of real prefill. Uses `SystemRandom`; never seed it.
- `eval_quality.py` — perplexity, deterministic generations, abliteration survival.
**PPL must be measured with `--speculative-config` OFF**: under MTP, vLLM's
`prompt_logprobs` come back ~uniform over the vocab (median rank ~10^5, logprob