flash-next-seat: full 262K context, KV pinned at a measured 14 GiB, gen-large on the gateway

Operator-directed: raise context to the model's native maximum and take as much KV
as the card safely allows, and expose the seat through LiteLLM as `gen-large`.

  max_model_len     131,072  ->  262,144
  KV cache             8.76  ->  14.00 GiB  (332,721 -> 560,654 tokens)
  concurrency      2.54x@128K ->  2.14x@262K

⚠ 16.00 GiB WAS TRIED FIRST AND IS TOO AGGRESSIVE. A 155,497-token non-repeating
prefill drove GPU 2 to 97,074 of 97,887 MiB and the caching allocator logged "OOM on
device 0 while trying to allocate 488636416 bytes (free: 422117376)" -- 466 MiB wanted
against 403 MiB free. The request completed, so nothing failed visibly; that is one
step before the shape that crashed stacks/mog-sec twice on 2026-09-10 (~1.04 GiB
wanted, ~600 MB free). Backed off to 14.00 GiB, which re-probes clean: zero allocator
warnings, a 155,557-token prefill in 14.2 s, and 2,085 MiB still free at peak.

The reason the first estimate was wrong is worth keeping, because it is not obvious
and it inverts the usual advice: --kv-cache-memory makes vLLM SKIP MEMORY PROFILING
ENTIRELY and ignore --gpu-memory-utilization. The profiler was the thing accounting
for deep-prefill activation, so pinning bytes switched off the protection that the
pin was supposed to formalise. vLLM's own "--kv-cache-memory=18745235968 (17.46 GiB)
to fully utilize gpu memory" line is computed from a profile measured at
max-num-batched-tokens depth and sits 3.5 GiB above what a 150K-token request
survives; open #54764 compounds it, since PLE short-conv prefill pads every request
in a batch to the batch-MAX query length.

max-num-batched-tokens stays at 8192 -- it is what bounds the activation peak, and
doubling max_model_len left the profiled peak unchanged at 1.65 GiB precisely because
the peak tracks chunk size, not context length.

Gateway: `gen-large` added to the LiteLLM model_list, pointing at fv-ml1:8022. One
alias on purpose -- a single alias cannot trip the shared-config enable_thinking
mutation footgun, which needs two over the same (model, api_base). Sampling is the
checkpoint's own declared set (temp 1.0 / top_p 0.95 / top_k 20); presence_penalty,
min_p and repetition_penalty are left unset because the checkpoint declares no
canonical value for them. Verified registered for both the infra-ops admin key and
the shared all-agents key, since a new model behind a scoped allowlist 403s silently.

Also adds services/flash-next-mtp-bench/ -- the MTP measurement campaign and its
rationale. MTP stays off, but on "not yet measured here" rather than on vLLM's
4xH100 recipe number, which is a cross-harness comparison and not evidence about a
TP=1 Blackwell seat.
This commit is contained in:
vh
2026-09-12 23:49:06 -07:00
parent 3132a16ca0
commit 7e62a07341
5 changed files with 387 additions and 15 deletions
+106
View File
@@ -0,0 +1,106 @@
# Does MTP actually hurt Qwen3.8-Flash-Next on one Blackwell card?
Status: **campaign built, running.** Results and a verdict land here with the raw
JSON alongside, so the conclusion can be re-derived rather than taken on faith.
## Why this exists
The `flash-next-seat` stack shipped with MTP speculative decoding **off**, citing
vLLM's published recipe: on 4×H100 it measured MTP as worse at every concurrency
tested — 8–36% lower request throughput, 32–173% higher per-token latency, ~36%
acceptance — and says don't default it on.
**That was not valid evidence about our seat, and defaulting on it was the wrong
call.** Our own measurement rule says the harness is part of the number and that
cross-harness comparisons are invalid, not merely noisy. The recipe's harness
differs from ours on nearly every axis that could plausibly drive the result:
| | vLLM recipe | this seat |
|---|---|---|
| GPUs | 4× H100 80 GB (Hopper) | 1× RTX PRO 6000 96 GB (Blackwell, sm_120) |
| Parallelism | TP=4 — experts sharded 4 ways, all-reduce per layer | **TP=1 — every expert local, no collective** |
| PLE table | offloaded (forced: 80 GB cards can't hold it) | offloaded (chosen) |
| k tested | **3 only** | 1, 2, 3 |
## Two mechanisms that could make MTP lose here, and one that makes k the real question
Spec decoding's usual win is that verifying k+1 tokens costs about the same as
decoding 1, because decode is memory-bandwidth-bound: you read the weights once
either way. **Two things about this model break that assumption.**
1. **Sparse-MoE expert-read amplification.** 512 experts, 10 active per token. At
batch 1, decoding one token touches ~10 experts. Verifying k+1 tokens routes
each position to *its own* 10, largely disjoint — so the weights read scale
with the token count instead of staying flat. The "free verification" premise
does not hold for an ultra-sparse MoE. This is worst at low concurrency and
shrinks as the batch already touches many experts anyway, which is exactly why
the sweep has to cross concurrency and not just report one number.
2. **PLE/UVA fetch amplification.** The 51B n-gram table lives in host RAM and is
read over PCIe. Every token position needs its own row lookups, so k+1 tokens
means k+1× the host round-trips — multiplying traffic on the single slowest
link in the system, the one we deliberately moved off the card.
**And the mechanism that makes k the actual experiment:** the checkpoint's
`mtp_num_hidden_layers` is **1**. The draft head is a *single module run
autoregressively* for k>1 — the pattern our own quant playbook §5.1 flags, where
deeper k improves acceptance and destroys throughput. If throughput falls
monotonically in k while acceptance rises, then **k=1 may win and the recipe's
k=3 number says nothing about it.** Nobody has published k=1 for this model.
So there are coherent reasons MTP could genuinely lose here — and an equally
coherent reason the published number is the wrong number to decide on.
## Design
Five boots. Arms in order: `off_A`, `k1`, `k2`, `k3`, `off_B`.
- **Repeats** — 3 per (arm, concurrency) cell, concurrency ∈ {1, 4, 8}. Median
reported with spread, never a single run.
- **Noise floor** — `off_A` and `off_B` are the *same configuration*, booted first
and last. Their difference is the floor, and because they are separate boots it
includes boot-to-boot variance that three reps inside one boot cannot see.
Running `off_B` last also catches monotonic drift across the campaign.
- **Positive control (a)** — MTP acceptance must be **> 0** on every MTP arm, read
from `/metrics` as a delta. A head that loads uninitialised serves fine and
reports ~0% accept; that is the failure our `gen` seat's `re:^mtp.*`
ignore-list footgun produces. **An arm with ~0% acceptance is void** — it
measured a broken head, not MTP. (Pre-checked: this checkpoint excludes
`mtp.*` and `model.mtp.*` from quantization in both quant configs, and its own
audit reports all 31 MTP tensors byte-identical to source.)
- **Positive control (b)** — aggregate throughput must **rise with concurrency**
inside every arm. Known-true for a batching server; if the harness can't see
it, the harness is blind and its negatives are worthless.
- **Null control** — `off_A` vs `off_B` must show no effect beyond the floor.
- **Stated sensitivity floor** — the report prints "cannot resolve effects smaller
than X" from the observed `off_A`/`off_B` spread. A delta under it is not a
finding.
**Harness parity.** Every arm's argv is *derived from the live compose file* via
`docker compose config`, not retyped, so the arms are provably identical except
for `--speculative-config`. That is what stops a stray flag from quietly becoming
the real independent variable.
## Instruments
- `concbench.py` — the **house** concurrency harness, reused unchanged from
`services/gen-seat-mixed-quant/bench/`. It already avoids the three traps that
have produced confident wrong answers here before: unseeded nonces so prefix
caching can't fake prefill, MTP acceptance read as a **delta** so a long-lived
seat's history doesn't swamp it, and aggregate throughput from wall clock rather
than a sum of per-request rates.
- `run-campaign.sh` — the driver. Runs on fv-ml1 against `127.0.0.1` so the
NH3↔FV mesh hop is not in the measurement.
## Known gaps in this design
- **No TTFT.** `concbench.py` doesn't stream, so prefill latency isn't measured
separately. The MTP mechanisms above are decode-side, so this is an acceptable
omission — but it means the campaign cannot speak to first-token latency.
- **One prompt shape.** ~400 output tokens from a short prompt. MTP acceptance is
workload-dependent; a long-context or code workload could differ and this says
nothing about them.
- **The offload's contribution is not isolated.** If MTP loses, this design cannot
separate mechanism 1 from mechanism 2. The clean discriminator is available and
cheap — re-run the best MTP arm at TP=2 across GPU 2+3 with `cpu_offload: false`
so the table is resident, and compare the MTP delta with and without the PCIe
path. GPU 3 is idle, so this is a follow-up worth doing if the answer matters.