Operator-directed: raise context to the model's native maximum and take as much KV as the card safely allows, and expose the seat through LiteLLM as `gen-large`. max_model_len 131,072 -> 262,144 KV cache 8.76 -> 14.00 GiB (332,721 -> 560,654 tokens) concurrency 2.54x@128K -> 2.14x@262K ⚠ 16.00 GiB WAS TRIED FIRST AND IS TOO AGGRESSIVE. A 155,497-token non-repeating prefill drove GPU 2 to 97,074 of 97,887 MiB and the caching allocator logged "OOM on device 0 while trying to allocate 488636416 bytes (free: 422117376)" -- 466 MiB wanted against 403 MiB free. The request completed, so nothing failed visibly; that is one step before the shape that crashed stacks/mog-sec twice on 2026-09-10 (~1.04 GiB wanted, ~600 MB free). Backed off to 14.00 GiB, which re-probes clean: zero allocator warnings, a 155,557-token prefill in 14.2 s, and 2,085 MiB still free at peak. The reason the first estimate was wrong is worth keeping, because it is not obvious and it inverts the usual advice: --kv-cache-memory makes vLLM SKIP MEMORY PROFILING ENTIRELY and ignore --gpu-memory-utilization. The profiler was the thing accounting for deep-prefill activation, so pinning bytes switched off the protection that the pin was supposed to formalise. vLLM's own "--kv-cache-memory=18745235968 (17.46 GiB) to fully utilize gpu memory" line is computed from a profile measured at max-num-batched-tokens depth and sits 3.5 GiB above what a 150K-token request survives; open #54764 compounds it, since PLE short-conv prefill pads every request in a batch to the batch-MAX query length. max-num-batched-tokens stays at 8192 -- it is what bounds the activation peak, and doubling max_model_len left the profiled peak unchanged at 1.65 GiB precisely because the peak tracks chunk size, not context length. Gateway: `gen-large` added to the LiteLLM model_list, pointing at fv-ml1:8022. One alias on purpose -- a single alias cannot trip the shared-config enable_thinking mutation footgun, which needs two over the same (model, api_base). Sampling is the checkpoint's own declared set (temp 1.0 / top_p 0.95 / top_k 20); presence_penalty, min_p and repetition_penalty are left unset because the checkpoint declares no canonical value for them. Verified registered for both the infra-ops admin key and the shared all-agents key, since a new model behind a scoped allowlist 403s silently. Also adds services/flash-next-mtp-bench/ -- the MTP measurement campaign and its rationale. MTP stays off, but on "not yet measured here" rather than on vLLM's 4xH100 recipe number, which is a cross-harness comparison and not evidence about a TP=1 Blackwell seat.
100 lines
6.7 KiB
Bash
100 lines
6.7 KiB
Bash
# flash-next-seat — Qwen3.8-Flash-Next (abliterated) on fv-ml1 GPU 2, :8022.
|
|
# Copy to .env on the host at /opt/docker/compose/flash-next-seat/.env.
|
|
#
|
|
# This is the INITIAL configuration, stood up 2026-09-13. Values marked FIRST-BOOT
|
|
# are deliberately conservative and expected to be revised once the seat has
|
|
# reported its own memory budget and been bisected for depth. Do not treat them as
|
|
# measured — they are not yet.
|
|
|
|
# ── Image ───────────────────────────────────────────────────────────────────
|
|
# ⚠ MUST contain vLLM #54371 (UVA PLE-offload), merged 2026-09-09T14:32Z.
|
|
# Verified by ancestry rather than version string: this commit is +150 / behind_by=0
|
|
# from merge commit 3116c5d06bfe76501b3dd6b5434bfc7f3274f5e7. v0.29.0 does NOT
|
|
# contain it (cut ~6h before the merge) and neither does any nightly-<sha> tag
|
|
# dated 2026-09-09 or earlier — the nightly build runs ~06:16 UTC.
|
|
FN_IMAGE=vllm/vllm-openai:nightly-eed1f3d0c6043bd494424a22443ee198dd56f657
|
|
API_KEY=replace-me
|
|
|
|
# ── Placement ───────────────────────────────────────────────────────────────
|
|
# GPU 2 was completely idle (2 MiB) before this seat; GPU 3 still is. Every other
|
|
# compose GPU pin on fv-ml1 is 0 or 1, so this seat displaced nothing.
|
|
FN_GPU_ID=2
|
|
FN_PORT=8022
|
|
FN_CONTAINER_NAME=vllm-flash-next
|
|
|
|
# ── Model ───────────────────────────────────────────────────────────────────
|
|
# dealignai/Qwen3.8-Flash-Next-ABLITERATED-NVFP4 @ be794b990578ef3031eccf9f28e675a289a09ee9
|
|
FN_MODEL=/tank/aimodels/qwen38-flash-next-abliterated-nvfp4
|
|
FN_QUANT=modelopt_fp4
|
|
FN_SERVED_NAME=qwen3.8-flash-next-uncensored
|
|
FN_SERVED_NAME_THINK=qwen3.8-flash-next-uncensored-thinking
|
|
|
|
# ── The offload ─────────────────────────────────────────────────────────────
|
|
# The 51B n-gram table (47.7 GiB FP8) lives in pinned host RAM; the GPU reads rows
|
|
# over CUDA UVA. Without this the checkpoint needs ~126 GiB of VRAM and will not
|
|
# start on one 96 GiB card. fv-ml1 has 566 GB RAM / ~388 GB available, so the
|
|
# host side is not a constraint here — unlike every DGX-Spark report upstream,
|
|
# where host and device share one unified pool and "offload" frees nothing.
|
|
FN_ENGRAM_CONFIG={"cpu_offload": true}
|
|
|
|
# Card is dedicated -- one tenant, nothing to compete with. Measured at 262K:
|
|
# weights 74.36 GiB resident, 560,654 KV tokens, 2.14x concurrency.
|
|
# ⚠ This ratio is now ADVISORY ONLY -- see the KV pin below, which overrides it.
|
|
FN_GPU_MEM_UTIL=0.96
|
|
# ⚠⚠ 14.00 GiB, PINNED IN BYTES AND MEASURED THE HARD WAY (2026-09-13).
|
|
# 16.00 GiB was tried first and nearly OOM'd: a 155,497-token prefill drove GPU 2 to
|
|
# 97,074 of 97,887 MiB and the allocator logged "OOM on device 0 while trying to
|
|
# allocate 488636416 bytes (free: 422117376)" -- 466 MiB wanted, 403 MiB free. The
|
|
# request survived but that is one step before the mog-sec crash shape.
|
|
# ⚠ WHY THE ESTIMATE WAS WRONG: --kv-cache-memory makes vLLM SKIP MEMORY PROFILING
|
|
# and ignore --gpu-memory-utilization entirely. The profiler was the thing accounting
|
|
# for deep-prefill activation; pinning bytes turns it off. Do NOT take vLLM's
|
|
# "17.46 GiB to fully utilize" suggestion -- it is computed from a profile taken at
|
|
# max-num-batched-tokens depth and is 3.5 GiB above what a 150K request survives.
|
|
# Re-raising requires re-running the deep probe and reading the allocator log.
|
|
# VERIFIED at 14.00 GiB: 0 OOM warnings, 155,557-token prefill in 14.2 s, 2,085 MiB
|
|
# still free on the card at peak.
|
|
FN_KV_CACHE_MEMORY=15032385536
|
|
|
|
# FULL NATIVE 262,144 (operator-directed 2026-09-13). The KV pool holds ~641K
|
|
# tokens, so a single max-length request fits with ~2.4x concurrency to spare.
|
|
# ⚠ STARTUP IS NOT A DEPTH TEST. Two open upstream issues make depth the risky
|
|
# axis -- #54764 (PLE short-conv prefill pads every request in a batch to the
|
|
# batch-MAX query length) and #54919 (long prefill starving active decode for 3-7
|
|
# minutes) -- and vLLM's own recipe admits a single 262K request was never tested.
|
|
# If deep requests misbehave, --max-num-batched-tokens is the lever, not this.
|
|
FN_MAX_MODEL_LEN=262144
|
|
FN_MAX_NUM_SEQS=16
|
|
FN_MAX_NUM_BATCHED_TOKENS=8192
|
|
FN_MAMBA_CACHE_MODE=align
|
|
|
|
# ── Prefix caching ──────────────────────────────────────────────────────────
|
|
# ⚠ THE ROLLBACK LEVER for open #54173 (CUBLAS_STATUS_INTERNAL_ERROR / illegal
|
|
# memory access in the GDN path, WITH prefix caching). Set to the empty string to
|
|
# disable. Leave FN_MAMBA_CACHE_MODE=align either way — Qwen4Exp raises on "all".
|
|
FN_PREFIX_CACHING=--enable-prefix-caching
|
|
|
|
# ── Vision ──────────────────────────────────────────────────────────────────
|
|
# 4194304 px = 2048x2048 -> ~5,125 image tokens. The checkpoint's own preprocessor
|
|
# declares 16777216 (4096x4096) -> ~16,384 tokens, which is both wasteful and fatal
|
|
# on builds enforcing the image-token count check. Same trap as stacks/mog-sec.
|
|
FN_MM_PROCESSOR_KWARGS={"size": {"longest_edge": 4194304, "shortest_edge": 65536}}
|
|
FN_LIMIT_MM={"image": 4}
|
|
|
|
# ── Misc ────────────────────────────────────────────────────────────────────
|
|
FN_REASONING_PARSER=qwen3
|
|
FN_REASONING_EFFORT=medium
|
|
FN_TOOL_CALL_PARSER=qwen3_xml
|
|
# Intentionally EMPTY. expandable_segments has corrupted retained tensors on this
|
|
# box before (quant playbook §3.10) and has never been tested against a pinned
|
|
# host allocation handed to UVA.
|
|
FN_ALLOC_CONF=
|
|
|
|
# ── NOT SET, on purpose ─────────────────────────────────────────────────────
|
|
# --speculative-config : MTP is off. vLLM's own recipe measured it WORSE at every
|
|
# concurrency on 4xH100 (8-36% less throughput, 32-173%
|
|
# more latency, ~36% acceptance) and open #55357 reports
|
|
# episodic 0% acceptance with repetition collapse.
|
|
# --kv-cache-dtype fp8 : fp8_e4m3 KV on this model's QSA path is an unmerged RFC
|
|
# (#54426). Do not copy it over from gen/mog-sec.
|