Files
esh-pfi-infrastructure/stacks/flash-next-seat/.env.example
T
vh 1b5d6ba23a docs(flash-next-seat): dealignai weights deleted — record that no local rollback exists
Operator instruction: delete the displaced dealignai checkpoint. 125 GiB
reclaimed from /tank (59% -> 57% used). Verified before removing: not mounted
by any running or exited container, no symlinks, no inodes shared with the
converted orcarouter directory.

Every "rollback is two .env keys" statement across the stack README, the
.env.example, persistent-memory and its detail file was true when written and
is false now -- the .env backup still names paths that no longer exist.
Corrected in place rather than left as false reassurance, since a stale
rollback instruction is discovered precisely when it is needed.

Reverting this seat now costs a 126 GiB re-download. The quality A/B against
dealignai is likewise no longer runnable locally: its reference arm is gone.

The pristine 170 GiB orcarouter download is retained deliberately -- it is what
makes the PLE bf16->FP8 conversion reproducible without re-fetching -- and that
is now recorded so a future session does not reclaim it as an obvious duplicate.

Also notes that ~75 GiB of non-PLE shards are duplicated between the pristine
and converted orca directories (the convert's hardlinks hit EXDEV across two
container bind mounts); both now sit directly on /tank, so relinking would
reclaim it if /tank ever tightens.
2026-09-14 02:51:54 -07:00

115 lines
7.8 KiB
Bash

# flash-next-seat — Qwen3.8-Flash-Next (abliterated) on fv-ml1 GPU 2, :8022.
# Copy to .env on the host at /opt/docker/compose/flash-next-seat/.env.
#
# This is the INITIAL configuration, stood up 2026-09-13. Values marked FIRST-BOOT
# are deliberately conservative and expected to be revised once the seat has
# reported its own memory budget and been bisected for depth. Do not treat them as
# measured — they are not yet.
# ── Image ───────────────────────────────────────────────────────────────────
# ⚠ MUST contain vLLM #54371 (UVA PLE-offload), merged 2026-09-09T14:32Z.
# Verified by ancestry rather than version string: this commit is +150 / behind_by=0
# from merge commit 3116c5d06bfe76501b3dd6b5434bfc7f3274f5e7. v0.29.0 does NOT
# contain it (cut ~6h before the merge) and neither does any nightly-<sha> tag
# dated 2026-09-09 or earlier — the nightly build runs ~06:16 UTC.
FN_IMAGE=vllm/vllm-openai:nightly-eed1f3d0c6043bd494424a22443ee198dd56f657
API_KEY=replace-me
# ── Placement ───────────────────────────────────────────────────────────────
# GPU 2 was completely idle (2 MiB) before this seat; GPU 3 still is. Every other
# compose GPU pin on fv-ml1 is 0 or 1, so this seat displaced nothing.
FN_GPU_ID=2
FN_PORT=8022
FN_CONTAINER_NAME=vllm-flash-next
# ── Model ───────────────────────────────────────────────────────────────────
# orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4, with its bf16 PLE converted to FP8
# in-house (2026-09-14) so vLLM's from_quant_config branch 1 selects the FP8 PLE method.
# ⚠ compressed-tensors, NOT modelopt_fp4 — this checkpoint is weight-only on BOTH axes
# (W8 float attn / W4 float experts, input_activations: null), which is why it was chosen
# over the dealignai W4A4 build.
# ⚠⚠ NO LOCAL ROLLBACK: the dealignai weights were DELETED 2026-09-14 (operator instruction).
# Reverting means re-downloading 126 GiB of dealignai/Qwen3.8-Flash-Next-ABLITERATED-NVFP4.
# The pristine 170 GiB orcarouter source is kept at qwen38-flash-next-orcarouter-nvfp4 so the
# PLE bf16→FP8 conversion can be redone without re-fetching.
FN_MODEL=/tank/aimodels/qwen38-flash-next-orcarouter-nvfp4-plefp8
FN_QUANT=compressed-tensors
FN_SERVED_NAME=qwen3.8-flash-next-uncensored
FN_SERVED_NAME_THINK=qwen3.8-flash-next-uncensored-thinking
# ── The offload ─────────────────────────────────────────────────────────────
# The 51B n-gram table (47.7 GiB FP8) lives in pinned host RAM; the GPU reads rows
# over CUDA UVA. Without this the checkpoint needs ~126 GiB of VRAM and will not
# start on one 96 GiB card. fv-ml1 has 566 GB RAM / ~388 GB available, so the
# host side is not a constraint here — unlike every DGX-Spark report upstream,
# where host and device share one unified pool and "offload" frees nothing.
FN_ENGRAM_CONFIG={"cpu_offload": true}
# Card is dedicated -- one tenant, nothing to compete with. Measured at 262K:
# weights 74.36 GiB resident, 560,654 KV tokens, 2.14x concurrency.
# ⚠ This ratio is now ADVISORY ONLY -- see the KV pin below, which overrides it.
FN_GPU_MEM_UTIL=0.96
# ⚠⚠ 14.00 GiB, PINNED IN BYTES AND MEASURED THE HARD WAY (2026-09-13).
# 16.00 GiB was tried first and nearly OOM'd: a 155,497-token prefill drove GPU 2 to
# 97,074 of 97,887 MiB and the allocator logged "OOM on device 0 while trying to
# allocate 488636416 bytes (free: 422117376)" -- 466 MiB wanted, 403 MiB free. The
# request survived but that is one step before the mog-sec crash shape.
# ⚠ WHY THE ESTIMATE WAS WRONG: --kv-cache-memory makes vLLM SKIP MEMORY PROFILING
# and ignore --gpu-memory-utilization entirely. The profiler was the thing accounting
# for deep-prefill activation; pinning bytes turns it off. Do NOT take vLLM's
# "17.46 GiB to fully utilize" suggestion -- it is computed from a profile taken at
# max-num-batched-tokens depth and is 3.5 GiB above what a 150K request survives.
# Re-raising requires re-running the deep probe and reading the allocator log.
# VERIFIED at 14.00 GiB: 0 OOM warnings, 155,557-token prefill in 14.2 s, 2,085 MiB
# still free on the card at peak.
# ⚠ MTP ON (k=3) REQUIRES A SMALLER KV POOL THAN THE NO-SPEC SEAT.
# Measured 2026-09-13: the draft head adds ~5.08 GiB of weights and raises
# per-token KV cost ~16%. At the old 15032385536 (14 GiB) the engine OOMs at
# init with MTP enabled -- 79.44 GiB weights + 14 GiB KV = 93.44 against a
# 91.17 GiB budget. 8589934592 (8 GiB) is PROVEN to boot; 10737418240 (10 GiB)
# leaves ~1.73 GiB for activations and is the current setting. If the seat
# OOMs at init, drop to 8589934592.
FN_KV_CACHE_MEMORY=10737418240
# FULL NATIVE 262,144 (operator-directed 2026-09-13). The KV pool holds ~641K
# tokens, so a single max-length request fits with ~2.4x concurrency to spare.
# ⚠ STARTUP IS NOT A DEPTH TEST. Two open upstream issues make depth the risky
# axis -- #54764 (PLE short-conv prefill pads every request in a batch to the
# batch-MAX query length) and #54919 (long prefill starving active decode for 3-7
# minutes) -- and vLLM's own recipe admits a single 262K request was never tested.
# If deep requests misbehave, --max-num-batched-tokens is the lever, not this.
FN_MAX_MODEL_LEN=262144
FN_MAX_NUM_SEQS=16
FN_MAX_NUM_BATCHED_TOKENS=8192
FN_MAMBA_CACHE_MODE=align
# ── Prefix caching ──────────────────────────────────────────────────────────
# ⚠ THE ROLLBACK LEVER for open #54173 (CUBLAS_STATUS_INTERNAL_ERROR / illegal
# memory access in the GDN path, WITH prefix caching). Set to the empty string to
# disable. Leave FN_MAMBA_CACHE_MODE=align either way — Qwen4Exp raises on "all".
FN_PREFIX_CACHING=--enable-prefix-caching
# ── Vision ──────────────────────────────────────────────────────────────────
# 4194304 px = 2048x2048 -> ~5,125 image tokens. The checkpoint's own preprocessor
# declares 16777216 (4096x4096) -> ~16,384 tokens, which is both wasteful and fatal
# on builds enforcing the image-token count check. Same trap as stacks/mog-sec.
FN_MM_PROCESSOR_KWARGS={"size": {"longest_edge": 4194304, "shortest_edge": 65536}}
FN_LIMIT_MM={"image": 4}
# ── Misc ────────────────────────────────────────────────────────────────────
FN_REASONING_PARSER=qwen3
FN_REASONING_EFFORT=medium
FN_TOOL_CALL_PARSER=qwen3_xml
# Intentionally EMPTY. expandable_segments has corrupted retained tensors on this
# box before (quant playbook §3.10) and has never been tested against a pinned
# host allocation handed to UVA.
FN_ALLOC_CONF=
# ── NOT SET, on purpose ─────────────────────────────────────────────────────
# --speculative-config : MTP is off. vLLM's own recipe measured it WORSE at every
# concurrency on 4xH100 (8-36% less throughput, 32-173%
# more latency, ~36% acceptance) and open #55357 reports
# episodic 0% acceptance with repetition collapse.
# --kv-cache-dtype fp8 : fp8_e4m3 KV on this model's QSA path is an unmerged RFC
# (#54426). Do not copy it over from gen/mog-sec.