flash-next-seat: full 262K context, KV pinned at a measured 14 GiB, gen-large on the gateway

Operator-directed: raise context to the model's native maximum and take as much KV
as the card safely allows, and expose the seat through LiteLLM as `gen-large`.

  max_model_len     131,072  ->  262,144
  KV cache             8.76  ->  14.00 GiB  (332,721 -> 560,654 tokens)
  concurrency      2.54x@128K ->  2.14x@262K

⚠ 16.00 GiB WAS TRIED FIRST AND IS TOO AGGRESSIVE. A 155,497-token non-repeating
prefill drove GPU 2 to 97,074 of 97,887 MiB and the caching allocator logged "OOM on
device 0 while trying to allocate 488636416 bytes (free: 422117376)" -- 466 MiB wanted
against 403 MiB free. The request completed, so nothing failed visibly; that is one
step before the shape that crashed stacks/mog-sec twice on 2026-09-10 (~1.04 GiB
wanted, ~600 MB free). Backed off to 14.00 GiB, which re-probes clean: zero allocator
warnings, a 155,557-token prefill in 14.2 s, and 2,085 MiB still free at peak.

The reason the first estimate was wrong is worth keeping, because it is not obvious
and it inverts the usual advice: --kv-cache-memory makes vLLM SKIP MEMORY PROFILING
ENTIRELY and ignore --gpu-memory-utilization. The profiler was the thing accounting
for deep-prefill activation, so pinning bytes switched off the protection that the
pin was supposed to formalise. vLLM's own "--kv-cache-memory=18745235968 (17.46 GiB)
to fully utilize gpu memory" line is computed from a profile measured at
max-num-batched-tokens depth and sits 3.5 GiB above what a 150K-token request
survives; open #54764 compounds it, since PLE short-conv prefill pads every request
in a batch to the batch-MAX query length.

max-num-batched-tokens stays at 8192 -- it is what bounds the activation peak, and
doubling max_model_len left the profiled peak unchanged at 1.65 GiB precisely because
the peak tracks chunk size, not context length.

Gateway: `gen-large` added to the LiteLLM model_list, pointing at fv-ml1:8022. One
alias on purpose -- a single alias cannot trip the shared-config enable_thinking
mutation footgun, which needs two over the same (model, api_base). Sampling is the
checkpoint's own declared set (temp 1.0 / top_p 0.95 / top_k 20); presence_penalty,
min_p and repetition_penalty are left unset because the checkpoint declares no
canonical value for them. Verified registered for both the infra-ops admin key and
the shared all-agents key, since a new model behind a scoped allowlist 403s silently.

Also adds services/flash-next-mtp-bench/ -- the MTP measurement campaign and its
rationale. MTP stays off, but on "not yet measured here" rather than on vLLM's
4xH100 recipe number, which is a cross-harness comparison and not evidence about a
TP=1 Blackwell seat.
This commit is contained in:
vh
2026-09-12 23:49:06 -07:00
parent 3132a16ca0
commit 7e62a07341
5 changed files with 387 additions and 15 deletions
+165
View File
@@ -0,0 +1,165 @@
#!/usr/bin/env bash
# run-campaign.sh — does MTP speculative decoding help or hurt Qwen3.8-Flash-Next
# on ONE RTX PRO 6000 with the n-gram table offloaded to host RAM?
#
# WHY THIS EXISTS. vLLM's published recipe measured MTP on 4xH100 as worse at
# every concurrency (8-36% less throughput, 32-173% more per-token latency, ~36%
# acceptance) and we initially defaulted MTP off on that basis. That was a
# cross-harness comparison and therefore not valid evidence about THIS seat:
# 4xH100 is TP=4 Hopper with experts sharded four ways and an all-reduce per
# layer; this is TP=1 Blackwell with every expert local. Our own rule says the
# harness is part of the number and cross-harness comparisons are invalid, not
# merely noisy. So we measure it here.
#
# AND THE RECIPE ONLY TESTED k=3. The checkpoint's mtp_num_hidden_layers is 1, so
# the draft head is a SINGLE module run autoregressively for k>1 (quant playbook
# §5.1: deeper k improves acceptance and destroys throughput). If throughput
# falls monotonically in k while acceptance rises, k=1 may well WIN and the
# recipe's k=3 number says nothing about it. That is the hypothesis this sweeps.
#
# ── MEASUREMENT DISCIPLINE ──────────────────────────────────────────────────
# Every number here carries repeats, a noise floor, a positive control and a null
# control, or it does not get to carry a conclusion.
#
# REPEATS 3 reps per (arm, concurrency) cell; median reported with spread.
# NOISE FLOOR The `off` arm is booted TWICE -- off_A first and off_B last. Both
# arms are the same configuration, so their difference IS the
# floor, and it includes boot-to-boot variance, which three reps
# inside one boot cannot see. Running off_B last also catches any
# monotonic drift across the campaign (thermals, page cache).
# POSITIVE CTL (a) MTP acceptance must be > 0 on every MTP arm. A head that
# loads uninitialised reports ~0% accept and still serves --
# the exact failure our gen seat's `re:^mtp.*` ignore-list
# footgun produces. If acceptance is ~0 the throughput number
# is measuring a broken head, not MTP, and the arm is void.
# (b) Aggregate throughput must RISE with concurrency inside every
# arm. That is known-true for a batching server; if the
# harness cannot see it, the harness is blind and its nulls
# are worthless.
# NULL CONTROL off_A vs off_B must show no effect beyond the floor.
# FLOOR STATED The report prints "cannot resolve effects smaller than X" from
# the observed off_A/off_B spread. A delta under it is not a
# finding.
#
# HARNESS PARITY. Every arm's argv is DERIVED from the live compose file, so the
# arms are provably identical except for --speculative-config. Nothing is
# retyped, which is what stops a stray flag from becoming the real independent
# variable.
#
# The production container is stopped for the duration: same GPU, same port. The
# seat has no consumers yet, so this costs nothing. It is restored at the end.
set -uo pipefail
STACK_DIR=${STACK_DIR:-/opt/docker/compose/flash-next-seat}
OUT=${OUT:-/tank/aimodels/flash-next-mtp-bench}
BENCH=${BENCH:-$OUT/concbench.py}
PORT=${PORT:-8022}
GPU=${GPU:-2}
REPS=${REPS:-3}
CONCS=${CONCS:-"1 4 8"}
MAXTOK=${MAXTOK:-400}
RPS=${RPS:-4} # requests per stream, per concbench
METHOD=${METHOD:-qwen4_exp_mtp}
MODEL=${MODEL:-qwen3.8-flash-next-uncensored}
NAME=fn-mtp-bench
BOOT_TIMEOUT=${BOOT_TIMEOUT:-2400}
mkdir -p "$OUT"
log(){ echo "[$(date -u +%H:%M:%S)] $*" | tee -a "$OUT/campaign.log"; }
# --- derive the production argv + run opts straight from the compose file -----
read_compose() {
sudo -n docker compose -f "$STACK_DIR/compose.yaml" --env-file "$STACK_DIR/.env" config --format json
}
log "deriving argv from $STACK_DIR/compose.yaml"
read_compose > "$OUT/resolved-compose.json" || { log "FATAL: could not resolve compose"; exit 1; }
python3 - "$OUT/resolved-compose.json" "$OUT/argv.txt" "$OUT/envs.txt" <<'PY'
import json,sys
d=json.load(open(sys.argv[1]))
svc=d["services"]["vllm-flash-next"]
open(sys.argv[2],"w").write("\n".join(str(x) for x in svc["command"])+"\n")
env=svc.get("environment") or {}
if isinstance(env,dict): items=[(k,v) for k,v in env.items()]
else: items=[e.split("=",1) for e in env]
open(sys.argv[3],"w").write("\n".join(f"{k}={'' if v is None else v}" for k,v in items)+"\n")
PY
mapfile -t ARGV < "$OUT/argv.txt"
ENVARGS=(); while IFS= read -r l; do [ -n "$l" ] && ENVARGS+=(-e "$l"); done < "$OUT/envs.txt"
log "argv has ${#ARGV[@]} entries; ${#ENVARGS[@]} env flags"
printf '%s\n' "${ARGV[@]}" | sed 's/^/ /' >> "$OUT/campaign.log"
MODEL_DIR=$(python3 -c "
import json,sys
d=json.load(open('$OUT/resolved-compose.json'))
for v in d['services']['vllm-flash-next']['volumes']:
t=v['target'] if isinstance(v,dict) else v.split(':')[1]
s=v['source'] if isinstance(v,dict) else v.split(':')[0]
if t=='/model': print(s)
")
HFCACHE=$(python3 -c "
import json
d=json.load(open('$OUT/resolved-compose.json'))
for v in d['services']['vllm-flash-next']['volumes']:
t=v['target'] if isinstance(v,dict) else v.split(':')[1]
s=v['source'] if isinstance(v,dict) else v.split(':')[0]
if t=='/hfcache': print(s)
")
IMAGE=$(python3 -c "
import json; print(json.load(open('$OUT/resolved-compose.json'))['services']['vllm-flash-next']['image'])")
log "image=$IMAGE model=$MODEL_DIR"
stop_bench(){ sudo -n docker rm -f "$NAME" >/dev/null 2>&1 || true; }
trap 'stop_bench' EXIT
boot(){ # boot <arm-label> [spec-json]
local arm="$1"; shift
local extra=(); [ $# -gt 0 ] && [ -n "${1:-}" ] && extra=(--speculative-config "$1")
stop_bench
log "=== boot arm=$arm ${extra[*]:-(no spec-config)}"
sudo -n docker run -d --name "$NAME" --ipc host --ulimit memlock=-1 \
--gpus "\"device=$GPU\"" -p "$PORT:8000" \
-v "$HFCACHE:/hfcache" -v "$MODEL_DIR:/model:ro" \
"${ENVARGS[@]}" "$IMAGE" "${ARGV[@]}" "${extra[@]}" \
> "$OUT/$arm.cid" 2>"$OUT/$arm.runerr" || { log " docker run FAILED: $(cat "$OUT/$arm.runerr")"; return 1; }
local t0=$SECONDS
while [ $((SECONDS-t0)) -lt "$BOOT_TIMEOUT" ]; do
if curl -sf "http://127.0.0.1:$PORT/health" >/dev/null 2>&1; then
log " healthy after $((SECONDS-t0))s"
sudo -n docker logs "$NAME" > "$OUT/$arm.boot.log" 2>&1
return 0
fi
if ! sudo -n docker ps --format '{{.Names}}' | grep -qx "$NAME"; then
log " CONTAINER DIED during boot"; sudo -n docker logs "$NAME" > "$OUT/$arm.boot.log" 2>&1
tail -25 "$OUT/$arm.boot.log" | sed 's/^/ /' | tee -a "$OUT/campaign.log"; return 1
fi
sleep 15
done
log " BOOT TIMEOUT after ${BOOT_TIMEOUT}s"; sudo -n docker logs "$NAME" > "$OUT/$arm.boot.log" 2>&1; return 1
}
bench_arm(){ # bench_arm <arm-label>
local arm="$1" cargs=()
for c in $CONCS; do cargs+=(--concurrency "$c"); done
for rep in $(seq 1 "$REPS"); do
log " bench $arm rep=$rep"
python3 "$BENCH" --base "http://127.0.0.1:$PORT" --model "$MODEL" \
"${cargs[@]}" --requests-per-stream "$RPS" --max-tokens "$MAXTOK" \
--tag "$arm-rep$rep" --out "$OUT/res-$arm-rep$rep.json" 2>&1 | tee -a "$OUT/campaign.log"
done
}
# --- the campaign -----------------------------------------------------------
# off_A first, the MTP arms in ascending k, off_B LAST so the floor brackets the
# whole run rather than sitting at one end of it.
run_arm(){ boot "$1" "${2:-}" && bench_arm "$1" || log " arm $1 SKIPPED (boot failed)"; }
log "########## CAMPAIGN START (reps=$REPS concs='$CONCS' max_tokens=$MAXTOK) ##########"
run_arm off_A ""
run_arm k1 "{\"method\":\"$METHOD\",\"num_speculative_tokens\":1}"
run_arm k2 "{\"method\":\"$METHOD\",\"num_speculative_tokens\":2}"
run_arm k3 "{\"method\":\"$METHOD\",\"num_speculative_tokens\":3}"
run_arm off_B ""
stop_bench
log "########## CAMPAIGN DONE -- results in $OUT ##########"
log "restoring the production container"
sudo -n docker compose -f "$STACK_DIR/compose.yaml" --env-file "$STACK_DIR/.env" up -d 2>&1 | tail -3 | tee -a "$OUT/campaign.log"