Files
esh-pfi-infrastructure/stacks/flash-next-seat
vh 7e62a07341 flash-next-seat: full 262K context, KV pinned at a measured 14 GiB, gen-large on the gateway
Operator-directed: raise context to the model's native maximum and take as much KV
as the card safely allows, and expose the seat through LiteLLM as `gen-large`.

  max_model_len     131,072  ->  262,144
  KV cache             8.76  ->  14.00 GiB  (332,721 -> 560,654 tokens)
  concurrency      2.54x@128K ->  2.14x@262K

⚠ 16.00 GiB WAS TRIED FIRST AND IS TOO AGGRESSIVE. A 155,497-token non-repeating
prefill drove GPU 2 to 97,074 of 97,887 MiB and the caching allocator logged "OOM on
device 0 while trying to allocate 488636416 bytes (free: 422117376)" -- 466 MiB wanted
against 403 MiB free. The request completed, so nothing failed visibly; that is one
step before the shape that crashed stacks/mog-sec twice on 2026-09-10 (~1.04 GiB
wanted, ~600 MB free). Backed off to 14.00 GiB, which re-probes clean: zero allocator
warnings, a 155,557-token prefill in 14.2 s, and 2,085 MiB still free at peak.

The reason the first estimate was wrong is worth keeping, because it is not obvious
and it inverts the usual advice: --kv-cache-memory makes vLLM SKIP MEMORY PROFILING
ENTIRELY and ignore --gpu-memory-utilization. The profiler was the thing accounting
for deep-prefill activation, so pinning bytes switched off the protection that the
pin was supposed to formalise. vLLM's own "--kv-cache-memory=18745235968 (17.46 GiB)
to fully utilize gpu memory" line is computed from a profile measured at
max-num-batched-tokens depth and sits 3.5 GiB above what a 150K-token request
survives; open #54764 compounds it, since PLE short-conv prefill pads every request
in a batch to the batch-MAX query length.

max-num-batched-tokens stays at 8192 -- it is what bounds the activation peak, and
doubling max_model_len left the profiled peak unchanged at 1.65 GiB precisely because
the peak tracks chunk size, not context length.

Gateway: `gen-large` added to the LiteLLM model_list, pointing at fv-ml1:8022. One
alias on purpose -- a single alias cannot trip the shared-config enable_thinking
mutation footgun, which needs two over the same (model, api_base). Sampling is the
checkpoint's own declared set (temp 1.0 / top_p 0.95 / top_k 20); presence_penalty,
min_p and repetition_penalty are left unset because the checkpoint declares no
canonical value for them. Verified registered for both the infra-ops admin key and
the shared all-agents key, since a new model behind a scoped allowlist 403s silently.

Also adds services/flash-next-mtp-bench/ -- the MTP measurement campaign and its
rationale. MTP stays off, but on "not yet measured here" rather than on vLLM's
4xH100 recipe number, which is a cross-harness comparison and not evidence about a
TP=1 Blackwell seat.
2026-09-12 23:49:06 -07:00
..

flash-next-seat — Qwen3.8-Flash-Next (abliterated), fv-ml1 GPU 2, :8022

The first seat on the fleet whose weights do not fit its card and run anyway.

Qwen3.8-Flash-Next is 176B total — a 125B main model plus a 51B n-gram (PLE) lookup table — activating ~6B parameters per token. The n-gram table is a pure embedding lookup with almost no compute per token, so it lives in pinned host RAM and the GPU reads the rows it needs directly over CUDA UVA on a dedicated stream with async prefetch.

Checkpoint dealignai/Qwen3.8-Flash-Next-ABLITERATED-NVFP4 @ be794b99… (126.0 GiB)
On the card ~78 GiB of 95.6 GiB — routed experts NVFP4 W4A4, rest at source precision
In host RAM 47.7 GiB pinned, FP8 E4M3, 10 model-plefp8-* shards + per-table scalar scale
Context 131,072 to start (native ceiling 262,144) — see Raising context
Speculative decoding none — see Why MTP is off
Gateway wiring none yet — this seat is not in LiteLLM; gen is untouched

Deploy

scripts/deploy-stack.sh fv-ml1 flash-next-seat     # diffs vs live, prompts y/N
# then on the host, first boot only:
ssh infra-ops@10.251.50.54 'cd /opt/docker/compose/flash-next-seat && docker compose config >/dev/null && docker compose up -d'

The .env lives on the host and is never committed. Copy .env.example, set API_KEY, and read the FIRST-BOOT annotations before changing anything else.

Architecture, briefly

Four ideas, and three of them shape the serving config:

  • GDN + QSA. 36 of 48 layers use Gated DeltaNet (linear attention) to compress history; every fourth layer uses Qwen Sparse Attention for long-range retrieval. This is why KV is cheap at depth and why --mamba-cache-dtype float32 matters.
  • N-gram embedding. The 51B lookup table that this seat offloads. Qwen's own framing: capacity with almost no per-token compute.
  • Gated residual / hyper-connections. Four residual branches; excluded from quantization in this checkpoint.
  • MTP head. Present and preserved byte-identically. Deliberately unused.

Why this checkpoint, and the trap that disqualifies most others

vLLM selects the PLE table's weight format from text_config.ple_embedding_dtype, as the first branch of Qwen4ExpPLEEmbeddingMethod.from_quant_config. This checkpoint declares "float8_e4m3fn".

A build that ships an FP8 PLE table without that declaration resolves instead through ModelOpt's *.ple.* exclude to the unquantized method, never registers the weight_scale parameter, and dies on load with no module or parameter named 'ngram_embedding.weight_scale'. gorbatjovy/qwen3.8-flash-next-abliterated-NVFP4-plefp8 is exactly this case. Check that field before trying another build.

Chosen over better-liked alternatives because its provenance states protocols and repeat counts rather than adjectives. From its own qualification-notes.md and metrics files, kept in the model directory:

  • HarmBench, 240 genuinely-harmful behaviours, greedy: 100% compliance at reasoning off / low / xhigh.
  • MMLU 82.11% → 81.93% (−0.18 pp) on an identical harness, 2,280 questions.
  • GSM8K 97.27% (1283/1319), full set, single-shot, temp 0.6 — inside the stated BF16 reference band 97.12–97.50.
  • AIME26 pass@1 98.75% (237/240, SEM 0.61 pp, 30 problems × 8 repeats), majority@8 100%, max_tokens 130,000, 4.9M completion tokens, stop_rate 99.17%.
  • Byte-equality audit of unchanged tensors: 1,562 tensors / 118.4 GB compared, all passed, including all 31 MTP tensors.

⚠ Two honest gaps in that evidence. The routed experts are NVFP4 W4A4, and nobody — including the publisher — has measured this checkpoint at the full 262K context; AIME26's 130K-token generations are the deepest evidence that exists. Separately, validate_checkpoint_report.json in the repo describes the earlier BF16-PLE revision (204 shards / 173.6 GiB), not the published FP8-PLE one.

Rejected alternatives, for the record: orcarouter/…-Uncensored-NVFP4 is gated (access request pending nothing — not requested); nvidia/…-NVFP4 is the cleanest ModelOpt MIXED_PRECISION build but is not abliterated; lovedheart/…-Pruned-RTXPRO-6000 prunes to 448 of 512 experts.

Why MTP is off

Against our house graft-MTP habit, and on purpose.

vLLM's own recipe for this model measured MTP on 4×H100 as worse at every concurrency tested — 8–36% lower request throughput, 32–173% higher per-token latency, driven by ~36% acceptance — and says do not enable it by default. Open issue #55357 reports episodic 0% draft acceptance with repetition collapse inside thinking blocks. Open #55496 reports ModelOpt MIXED_PRECISION failing to load FP8_BLOCK_SCALES MTP experts.

Turning it on is two lines in compose.yaml (documented in place). If you do, measure it here, with repeats, against this seat's own baseline — the numbers above are someone else's hardware.

The upstream situation, as of 2026-09-13

  • #53896 — model support. Merged 2026-08-31. In v0.29.0.
  • #54371 — UVA PLE-offload and Engram tensor parallelism. Merged 2026-09-09T14:32Z. This is the offload this seat uses. Not in v0.29.0, which was cut ~6 h earlier; present in v0.29.1rc0 and in any nightly from 2026-09-10 onward.
  • #53899 — the older, worker-based PLE offload. Open and explicitly paused in favour of #54371. Do not go back to it. Its whole bug family — the TP=1 startup deadlock (#53960), the pidfd_getfd / kernel.yama.ptrace_scope gate, the shared-CUDA-event race under async scheduling, and silently one-step-stale PLE outputs under CUDA graphs — came from the separate worker process and the CUDA-IPC row transfer that the UVA path does not have.

Open issues worth knowing about on SM120, none of them blocking:

Issue What it does Our exposure
#54173 CUBLAS internal error / illegal memory access in the GDN path with prefix caching We enable prefix caching. FN_PREFIX_CACHING= is the one-line rollback.
#54764 PLE short-conv batched prefill pads every request to the batch-max query length Why --max-num-batched-tokens is 8192, not 16384
#54919 Long prefill starves active decode for 3–7 minutes Why context starts at 128K
#54521 Greedy decoding non-deterministic from persistent_topk in prefill Affects any A/B on this seat — establish a noise floor before comparing
#54426 fp8_e4m3 KV on the QSA path is an unmerged RFC Why --kv-cache-dtype is not set to fp8 here

Raising context

128K is a starting value, not a measured one. Before raising it, bisect with a non-repeating prompt — a repeated one hashes to cached blocks and never prefills deep, so it proves nothing. The stacks/mog-sec README records this the hard way: three successive context cuts all sized the KV pool while the crashes were governed by processing depth, which is a different number.

The point of a ceiling is the refusal. Below it the seat serves; above it vLLM returns a clean 400 naming the limit, instead of the engine dying and taking every in-flight request with it.

Not done yet

  • Pin --kv-cache-memory in bytes from the first boot's budget line, replacing the 0.90 ratio. Same discipline as stacks/mog-sec and stacks/erp-seat.
  • Gateway wiring is deliberately absent. Pointing any LiteLLM alias at this seat — in particular displacing gen / summarizer / classifier, which is the long-term intent recorded in henge item 49 — changes what every existing caller receives and is the operator's call, not a deploy-time default.