Operator ruling: the multi-turn degeneration lives in the un-fixed vLLM, not in the weights. The hypothesis that sec's stock-graft MTP head causes it is withdrawn. Two failures produced it. First, a false dichotomy treated as a deduction: having verified gen and sec run an identical engine, I concluded config was eliminated and therefore the weights were responsible. That does not follow. An engine bug present in both seats is not exonerated by the seats being identical; it only means the engine cannot explain a difference between them. It can still explain the failure. Second, and more instructive, the difference being explained may not exist. The premise was a single operator observation made during a session with many concurrent changes. That cannot carry a causal claim, and it became the load-bearing support for a root-cause narrative it could not hold. The same caveat now attaches to the coherent-to-10k observation on the new build: same n, same uncontrolled conditions, opposite direction. The comparison is weak at both ends, so the file no longer presents either sighting as a result. What survives as measured fact is unchanged and still recorded: sec's MTP head is byte-identical to the uncensored base across all 15 tensors, gen's was abliterated in-band, and acceptance differs slightly. None of that is shown to cause degeneration. Adds the generalisable lesson: an observation made while many things are changing cannot support a causal conclusion. It is the inverse of the warning already in the gen-seat compose file, which guards against trusting a negative result from a synthetic probe; this guards against trusting a positive sighting from an uncontrolled session.
10 KiB
DFlash2 speculative decoding — measured on our own stack (2026-08-22)
Operator-driven session. Read the epistemic labels. During the chase we generalised from observations that later proved wrong; this file separates what was measured from what remains hypothesis, and records the wrong turns so nobody re-derives them.
What DFlash2 is
A 2B draft model (3.85 GB bf16) for speculative decoding against Qwen3.8-27B —
incoai/Qwen3.8-27B-DFlash2, Apache-2.0, blog inco.ai/blog/dflash2, upstream z-lab/dflash.
Block diffusion: drafts a whole 8-token block in one pass, with a candidate selector tracing a
path through per-slot top-K. Lossless (greedy matches the target).
vLLM support merged 2026-08-21 05:27 UTC as PR #52816 (b389ac29). Method string is
"dflash", not dflash2.
✅ MEASURED — throughput and acceptance
Single instrument (specbench.py, 8 fixed prompts, temp 0, max_tokens 256), delta against
vLLM's own spec_decode counters. The MTP k=3 numbers reproduce our recorded 58.4% / 55.3%
figures exactly, which is what validates the instrument.
| seat | config | accepted tok/forward | throughput |
|---|---|---|---|
| gen (orcarouter) | MTP k=3 (production) | 2.753 | 114.9 tok/s |
| gen | MTP k=7 (control) | 3.041 | 74.0 tok/s |
| gen | DFlash2 k=7 | 3.254 | 131.9 tok/s |
| sec (M.O.G.-SEC) | MTP k=3 (production) | 2.676 | 110.5 tok/s |
| sec | DFlash2 k=7 | 3.252 | 130.0 tok/s |
⭐ The k=7 MTP control was essential and inverted the obvious read. Going deeper on MTP
improves acceptance (2.753 → 3.041) while destroying throughput (114.9 → 74.0). Our MTP
head is a single module (mtp_num_hidden_layers=1, only mtp.layers.0, 15 tensors) run
autoregressively, so k draft tokens cost k sequential forward passes. "Just raise
num_speculative_tokens" is a trap — without the control I would have recommended it.
DFlash2's win is therefore not better per-token acceptance — our MTP is actually better at position 0 (79.6% vs 75.4%). It is that block drafting makes depth nearly free.
⭐ The drafter is model-agnostic across finetunes — 3.254 (gen) vs 3.252 (sec), a 0.06%
difference, with superimposable per-position curves. One drafter file on /tank serves both.
✅ MEASURED — how DFlash2 runs (answers "can one drafter serve both seats?")
EAGLE3-style coupled, not standalone. In vLLM: load_model(self, target_model) binds it to a
specific target object; pass_hidden_states_to_model=True; gpu_model_runner reads
dflash_config.target_layer_ids → [i+1 …] to register auxiliary hidden-state capture on the
target at layers 5, 19, 33, 47, 61. It even reads the target's RoPE style at load.
Consequences:
- Weights file is shareable (one download, both seats mount it) — gen and sec are architecturally identical on every dimension the drafter needs: 64 layers (deepest tap 61), hidden 5120, intermediate 17408, vocab 248,320 > mask token 248,070.
- VRAM is NOT shareable — 3.85 GB per seat. The drafter lives inside the target's engine process, consuming hidden states mid-forward. Two seats are two processes; there is no cross-process sharing mechanism and there could not be.
✅ MEASURED — it works on our stack, which the card does not claim
The card tests stock BF16 on an H200 with FlashAttention 3. Verified here instead:
abliterated + NVFP4 compressed-tensors target ✓, Blackwell sm_120 ✓, DFlash2 CUDA graphs
captured ✓. None of that was documented anywhere.
🔶 HYPOTHESIS — why our acceptance trails the published numbers
Both our targets land at ~3.25 accepted length against the card's 4.10–5.46 on stock BF16. Finetune drift is ruled out — two different finetunes gave identical results to three decimals. The shared variable is NVFP4 quantization of the target, which is mechanically plausible (the drafter reads quantized hidden states at its five taps). Second candidate: prompt distribution (ours general-purpose, theirs GSM8K/MATH/HumanEval/MBPP/MT-Bench). Neither is confirmed. Settling it needs a BF16 target seat (~56 GB) — a real GPU window.
❌ RETRACTED — the "MTP head mismatch causes the degeneration" hypothesis
Operator ruling, 2026-08-22: this hypothesis is WRONG. The degeneration lives in the un-fixed vLLM, not in the weights. Recorded here rather than deleted, because it was reasoned to confidently enough that a future session could re-derive it.
Two independent failures produced it, and the second is the instructive one:
- I treated a false dichotomy as a deduction. Having verified gen and sec run an identical
engine (same image ID
sha256:bd3236cff208…, same live version0.27.2rc1.dev150+g311b3513aread from inside both processes, same flags bargpu-memory-utilization0.43 vs 0.44), I concluded "config is eliminated, therefore it is the weights." That does not follow. An engine bug present in BOTH seats is not exonerated by the two seats being identical — it just means the engine cannot explain a difference. It can still explain the failure. - The difference I was explaining may not exist. The premise was a single operator observation of sec degenerating at ~2k, made during a session with many concurrent changes. n=1 under heavy concurrent modification is not evidence — see the meta-lesson below.
What survives as fact (measured, still true, just not causal): sec's MTP head is
byte-identical to qwen38-27b-uncensored-bf16 across all 15 tensors — a stock head on a
security-finetuned body, because the Qwen3_5ForConditionalGeneration wrapper never loads the
head, so the finetuning could not reach it. gen's orcarouter head was abliterated in-band by
its author. Acceptance differs slightly (gen 58.4%, sec 55.9%). All true. None of it shown to
cause multi-turn degeneration.
Current standing explanation: the degeneration is an engine bug in the un-fixed vLLM. Both
production seats run 311b3513, which is 172 commits behind GDN spec-decode fix #53077
(merged 2026-08-20). #51113 is present in that build and is therefore necessary but
insufficient on its own.
⭐⭐ META-LESSON — n=1 during a busy session is not evidence
The operator's own framing, and it generalises past this incident: an observation made while many things are being changed at once cannot carry a causal claim, no matter how confidently it is reported. Tonight that single observation became the load-bearing premise for a weights-side hypothesis, a root-cause narrative, and very nearly a recommendation.
This is the same failure the gen-seat compose file already warns about in different words — "a passing probe is NOT sufficient evidence" — inverted. That note guards against trusting a negative result from a synthetic test. This one guards against trusting a positive sighting from an uncontrolled session. Both reduce to: hold the system still, or do not draw causal conclusions from it.
Applies equally to the "coherent to 10k" observation below — same n, same conditions, opposite direction. Neither observation is worth more than the other.
⚠️ CONFOUNDED — and the "before" state is itself unreliable
sec now runs DFlash2 on a newer build and the operator reports coherent to 10k tokens with adversarial nonsense prompts. ⚠ Treat this the same way as the 2k sighting it is being compared against: n=1, uncontrolled session, not evidence. The comparison is weak on both ends.
Two variables changed at once:
- Engine:
311b3513→e9d1398d, +259 commits,behind_by=0(a strict superset), including GDN spec-decode fix #53077 (merged 2026-08-20) that production is 172 commits behind. - Drafter: frozen MTP head → DFlash2 reading live hidden states.
Isolating it = run MTP k=3 on the same new build. Not yet done.
#51113 is present in BOTH builds (verified by ancestry, behind_by=0 each) — so the
"proper upstream fix" our compose comment credits is necessary but insufficient; sec ran it
and still degenerated. Related open upstream: #53180 (quantized Qwen3.8-27B hybrid GDN + MTP
producing silent degenerate output, no fix), #41884 (DFlash + prefix caching on hybrid,
IndexError, workaround is disabling one).
❌ WRONG TURNS — do not repeat
- Version strings are not lineage. The DFlash2 build reports
0.26.1rc1.dev1048and our production nightly0.27.2rc1.dev150, which looks like a regression. It is a setuptools_scm tag-reachability artifact. Use the GitHub compare API and checkbehind_by. - Docker Hub push timestamps lie about source freshness.
nightly-ba07e4a4was pushed 06:12 UTC, comfortably after the 05:27 merge — but cut from a 03:46 commit that predates it. Grep the image for the symbols you need. Believing the timestamp would have cost an RP-seat outage to serve a model the engine could not instantiate. --max-num-batched-tokenswas not the image truncation. Raising it 16,384 → 32,768 on that theory changed nothing and cost ~3 GiB of peak activation, which came straight out of the KV pool. The cap was the tokenizer (§3.14 of the playbook).- "1M needs YaRN, absent from config" is FALSE for the sec quant. It is fully present:
rope_type: yarn,factor: 4.0,original_max_position_embeddings: 262144,max_position_embeddings: 1000000. Context is a KV-memory choice, not a model limit.
Live state — sec is NOT running from its compose stack
vllm-sec-dflash2, a standalone container on sec's port with sec's served names, so the
sec / sec-reasoning gateway aliases work unchanged. /opt/docker/compose/mog-sec is
stopped but unmodified.
| production sec | current | |
|---|---|---|
| image | nightly-311b3513 |
nightly-e9d1398d |
| speculation | MTP k=3 | DFlash2 k=7, drafter /tank/aimodels/qwen38-27b-dflash2-drafter |
| max-model-len | 262,144 | 480,000 |
| KV pool | 418,218 (1.60×) | 526,617 (1.10×) |
| images | 4096² → 16,384 tok | 2048² → ~5,125 tok (--mm-processor-kwargs size cap) |
ROLLBACK is two commands: docker rm -f vllm-sec-dflash2 then docker compose up -d in
/opt/docker/compose/mog-sec.
⚠ --gpu-memory-utilization 0.55 is the stable ceiling while GPU1's other tenants are up.
0.58 sized KV at 594,172 then OOM'd during CUDA graph capture — the process reached 57.49 GiB
against ~57.6 free. Real 1M context needs ~49 GiB of KV and therefore evicting most of GPU1.
Launcher: /tmp/run_sec_dflash2.sh on ana-ml2 (ephemeral — re-derive from this table if lost).