Files
esh-pfi-infrastructure/services/flash-next-mtp-bench/run-campaign.sh
T
vh a91b841d86 feat(fv-ml1): generate the seat inventory from the live box instead of maintaining it by hand
The seat documentation must stay current, and a hand-written document cannot.
The LiteLLM config described char-rp as a 31B model on a host and GPU it had not
been on since 2026-08-24 -- three weeks of silent drift in a file that read as
authoritative, and the reason a seat spent that period serving a model nobody
intended. Anything typed here drifts the same way; anything read off the running
containers cannot.

scripts/seat-inventory.py derives the whole document from the host:

- placement and VRAM from nvidia-smi compute-apps, mapped to containers through
  /proc/<pid>/cgroup -- nvidia-smi reports the vLLM engine child while docker
  reports the container pid, so matching them directly silently yields nothing
- weights and KV tokens parsed from each engine's own startup log, not derived
  arithmetically, with concurrency computed as KV tokens over context
- architecture, layer and expert counts, and the exact quantization group scheme
  (W4A4 vs W4A16 distinguished) from each model's config.json
- speculative-decoding method and k from the container argv, which is how the
  three incompatible methods on this box became visible
- lineage from the .PROVENANCE.txt SIBLING files -- they sit beside the model
  directory, not inside it, which is why an earlier pass wrongly reported two
  fully-documented seats as having no provenance
- gateway aliases resolved from the LiteLLM config on ana-docker

--check compares the committed document against the live box and exits non-zero
when they diverge, ignoring only the generation timestamp. Suitable for CI or a
scheduled drift alarm; read-only throughout, safe against production.

Also commits the KV_CACHE_BYTES override added to the MTP campaign runner, which
asserts the flag exists in the derived argv and aborts rather than running a
campaign that silently ignored it.
2026-09-13 23:01:44 -07:00

260 lines
13 KiB
Bash
Executable File

#!/usr/bin/env bash
# run-campaign.sh — does MTP speculative decoding help or hurt Qwen3.8-Flash-Next
# on ONE RTX PRO 6000 with the n-gram table offloaded to host RAM?
#
# WHY THIS EXISTS. vLLM's published recipe measured MTP on 4xH100 as worse at
# every concurrency (8-36% less throughput, 32-173% more per-token latency, ~36%
# acceptance) and we initially defaulted MTP off on that basis. That was a
# cross-harness comparison and therefore not valid evidence about THIS seat:
# 4xH100 is TP=4 Hopper with experts sharded four ways and an all-reduce per
# layer; this is TP=1 Blackwell with every expert local. Our own rule says the
# harness is part of the number and cross-harness comparisons are invalid, not
# merely noisy. So we measure it here.
#
# AND THE RECIPE ONLY TESTED k=3. The checkpoint's mtp_num_hidden_layers is 1, so
# the draft head is a SINGLE module run autoregressively for k>1 (quant playbook
# §5.1: deeper k improves acceptance and destroys throughput). If throughput
# falls monotonically in k while acceptance rises, k=1 may well WIN and the
# recipe's k=3 number says nothing about it. That is the hypothesis this sweeps.
#
# ── MEASUREMENT DISCIPLINE ──────────────────────────────────────────────────
# Every number here carries repeats, a noise floor, a positive control and a null
# control, or it does not get to carry a conclusion.
#
# REPEATS 3 reps per (arm, concurrency) cell; median reported with spread.
# NOISE FLOOR The `off` arm is booted TWICE -- off_A first and off_B last. Both
# arms are the same configuration, so their difference IS the
# floor, and it includes boot-to-boot variance, which three reps
# inside one boot cannot see. Running off_B last also catches any
# monotonic drift across the campaign (thermals, page cache).
# POSITIVE CTL (a) MTP acceptance must be > 0 on every MTP arm. A head that
# loads uninitialised reports ~0% accept and still serves --
# the exact failure our gen seat's `re:^mtp.*` ignore-list
# footgun produces. If acceptance is ~0 the throughput number
# is measuring a broken head, not MTP, and the arm is void.
# (b) Aggregate throughput must RISE with concurrency inside every
# arm. That is known-true for a batching server; if the
# harness cannot see it, the harness is blind and its nulls
# are worthless.
# NULL CONTROL off_A vs off_B must show no effect beyond the floor.
# FLOOR STATED The report prints "cannot resolve effects smaller than X" from
# the observed off_A/off_B spread. A delta under it is not a
# finding.
#
# HARNESS PARITY. Every arm's argv is DERIVED from the live compose file, so the
# arms are provably identical except for --speculative-config. Nothing is
# retyped, which is what stops a stray flag from becoming the real independent
# variable.
#
# ⚠ RUNS ON GPU 3, PORT 8023 -- operator-directed 2026-09-13 so the campaign runs in
# PARALLEL with live testing on the production seat (GPU 2, :8022). This script must
# therefore NEVER touch the production container: there is no compose down, no compose
# up, and nothing addressed by the production container name. The only container it
# creates or removes is `fn-mtp-bench`.
#
# ⚠ POWER. This puts TWO of the four cards under load at once, which is the condition
# that tripped the Anaheim rack breaker on 2026-08-26 and 2026-09-11, and the Fountain
# Valley circuit was specced while every record still said the box had two GPUs rather
# than four. The operator accepted that risk explicitly. Per-GPU power draw is sampled
# to power.log throughout so there is a record either way.
set -uo pipefail
STACK_DIR=${STACK_DIR:-/opt/docker/compose/flash-next-seat}
OUT=${OUT:-/tank/aimodels/flash-next-mtp-bench}
BENCH=${BENCH:-$OUT/concbench.py}
PORT=${PORT:-8023}
GPU=${GPU:-3}
REPS=${REPS:-3}
CONCS=${CONCS:-"1 4 8"}
MAXTOK=${MAXTOK:-400}
RPS=${RPS:-4} # requests per stream, per concbench
METHOD=${METHOD:-qwen4_exp_mtp}
MODEL=${MODEL:-qwen3.8-flash-next-uncensored}
NAME_BASE=fn-mtp-bench
NAME="" # set per-arm by boot(); never reused, so the name can never collide
BOOT_TIMEOUT=${BOOT_TIMEOUT:-2400}
mkdir -p "$OUT"
log(){ echo "[$(date -u +%H:%M:%S)] $*" | tee -a "$OUT/campaign.log"; }
# Per-GPU power + memory every 10 s for the life of the campaign. Two cards under
# load is the risk the operator accepted; this is the record of what it actually drew.
( while :; do
echo "$(date -u +%H:%M:%S) $(nvidia-smi --query-gpu=index,power.draw,memory.used,utilization.gpu --format=csv,noheader | tr '\n' '|')" >> "$OUT/power.log"
sleep 10
done ) & POWER_PID=$!
# --- derive the production argv + run opts straight from the compose file -----
read_compose() {
sudo -n docker compose -f "$STACK_DIR/compose.yaml" --env-file "$STACK_DIR/.env" config --format json
}
log "deriving argv from $STACK_DIR/compose.yaml"
read_compose > "$OUT/resolved-compose.json" || { log "FATAL: could not resolve compose"; exit 1; }
python3 - "$OUT/resolved-compose.json" "$OUT/argv.txt" "$OUT/envs.txt" <<'PY'
import json,sys
d=json.load(open(sys.argv[1]))
svc=d["services"]["vllm-flash-next"]
open(sys.argv[2],"w").write("\n".join(str(x) for x in svc["command"])+"\n")
env=svc.get("environment") or {}
if isinstance(env,dict): items=[(k,v) for k,v in env.items()]
else: items=[e.split("=",1) for e in env]
open(sys.argv[3],"w").write("\n".join(f"{k}={'' if v is None else v}" for k,v in items)+"\n")
PY
mapfile -t ARGV < "$OUT/argv.txt"
ENVARGS=(); while IFS= read -r l; do [ -n "$l" ] && ENVARGS+=(-e "$l"); done < "$OUT/envs.txt"
log "argv has ${#ARGV[@]} entries; ${#ENVARGS[@]} env flags"
# --- optional single-flag override of the production argv --------------------
# MTP's draft head costs ~5.08 GiB of weights (74.36 -> 79.44 GiB, measured
# 2026-09-13 and FLAT in k: identical for k=1, k=2 and k=3). That does not fit
# alongside the seat's 14 GiB pinned --kv-cache-memory under
# --gpu-memory-utilization 0.96, so every MTP arm OOMs at engine init while the
# no-spec arms boot fine. Pinning a smaller KV budget for EVERY arm makes room
# without confounding the comparison -- and it is free for this benchmark, which
# at conc<=8 with 400-token completions never touches more than a few thousand
# KV tokens against a cache sized in the hundreds of thousands.
if [ -n "${KV_CACHE_BYTES:-}" ]; then
kv_found=0
for i in "${!ARGV[@]}"; do
if [ "${ARGV[$i]}" = "--kv-cache-memory" ]; then
log "override: --kv-cache-memory ${ARGV[$((i+1))]} -> $KV_CACHE_BYTES"
ARGV[$((i+1))]="$KV_CACHE_BYTES"
kv_found=1
break
fi
done
# Fail loudly rather than run a campaign that silently ignored the override --
# a run whose knob did nothing is worse than a run that refused to start.
[ "$kv_found" -eq 1 ] || { log "FATAL: KV_CACHE_BYTES set but --kv-cache-memory absent from derived argv"; exit 1; }
fi
printf '%s\n' "${ARGV[@]}" | sed 's/^/ /' >> "$OUT/campaign.log"
MODEL_DIR=$(python3 -c "
import json,sys
d=json.load(open('$OUT/resolved-compose.json'))
for v in d['services']['vllm-flash-next']['volumes']:
t=v['target'] if isinstance(v,dict) else v.split(':')[1]
s=v['source'] if isinstance(v,dict) else v.split(':')[0]
if t=='/model': print(s)
")
HFCACHE=$(python3 -c "
import json
d=json.load(open('$OUT/resolved-compose.json'))
for v in d['services']['vllm-flash-next']['volumes']:
t=v['target'] if isinstance(v,dict) else v.split(':')[1]
s=v['source'] if isinstance(v,dict) else v.split(':')[0]
if t=='/hfcache': print(s)
")
IMAGE=$(python3 -c "
import json; print(json.load(open('$OUT/resolved-compose.json'))['services']['vllm-flash-next']['image'])")
log "image=$IMAGE model=$MODEL_DIR"
stop_bench(){
# Remove the PREVIOUS arm's container (if any). Deliberately does NOT wait for
# the container record to disappear: measured 2026-09-13, tearing down a
# container holding ~92 GB of GPU memory plus the offloaded PLE mapping leaves
# an Exited record holding its name for anywhere between 0s and ~60s. That
# latency is unpredictable, so arms get unique names (see boot) and the record
# is simply left to be reaped by cleanup at the end.
[ -n "${NAME:-}" ] && sudo -n docker rm -f "$NAME" >/dev/null 2>&1 || true
# Do wait on GPU memory: that IS what the next container needs, and unlike the
# container record it clears promptly and reports truthfully.
local t0=$SECONDS
while [ "$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits -i "$GPU" 2>/dev/null || echo 0)" -gt 1000 ]; do
if [ $((SECONDS-t0)) -ge 180 ]; then
log " WARN: GPU $GPU still holding memory after 180s; continuing anyway"
break
fi
sleep 2
done
}
cleanup(){
stop_bench
# reap every per-arm container this campaign created
for n in $(sudo -n docker ps -a --format '{{.Names}}' 2>/dev/null | grep "^${NAME_BASE}-" || true); do
sudo -n docker rm -f "$n" >/dev/null 2>&1 || true
done
[ -n "${POWER_PID:-}" ] && kill "$POWER_PID" 2>/dev/null
}
trap 'cleanup' EXIT
boot(){ # boot <arm-label> [spec-json]
local arm="$1"; shift
local extra=(); [ $# -gt 0 ] && [ -n "${1:-}" ] && extra=(--speculative-config "$1")
stop_bench
NAME="${NAME_BASE}-${arm}" # unique per arm: collision is impossible by construction
log "=== boot arm=$arm ${extra[*]:-(no spec-config)}"
# The daemon can still hold the name after the container is gone from
# `docker ps -a`, so a Conflict here is transient -- retry rather than
# forfeiting the whole arm (a skipped arm is silent data loss).
local attempt=1
while :; do
if sudo -n docker run -d --name "$NAME" --ipc host --ulimit memlock=-1 \
--gpus "\"device=$GPU\"" -p "$PORT:8000" \
-v "$HFCACHE:/hfcache" -v "$MODEL_DIR:/model:ro" \
"${ENVARGS[@]}" "$IMAGE" "${ARGV[@]}" "${extra[@]}" \
> "$OUT/$arm.cid" 2>"$OUT/$arm.runerr"; then
break
fi
if grep -q "already in use" "$OUT/$arm.runerr" && [ "$attempt" -lt 12 ]; then
log " name still held by daemon (attempt $attempt) -- retrying in 5s"
attempt=$((attempt+1)); sleep 5; continue
fi
log " docker run FAILED: $(cat "$OUT/$arm.runerr")"; return 1
done
local t0=$SECONDS
while [ $((SECONDS-t0)) -lt "$BOOT_TIMEOUT" ]; do
if curl -sf "http://127.0.0.1:$PORT/health" >/dev/null 2>&1; then
log " healthy after $((SECONDS-t0))s"
sudo -n docker logs "$NAME" > "$OUT/$arm.boot.log" 2>&1
return 0
fi
if ! sudo -n docker ps --format '{{.Names}}' | grep -qx "$NAME"; then
log " CONTAINER DIED during boot"; sudo -n docker logs "$NAME" > "$OUT/$arm.boot.log" 2>&1
tail -25 "$OUT/$arm.boot.log" | sed 's/^/ /' | tee -a "$OUT/campaign.log"; return 1
fi
sleep 15
done
log " BOOT TIMEOUT after ${BOOT_TIMEOUT}s"; sudo -n docker logs "$NAME" > "$OUT/$arm.boot.log" 2>&1; return 1
}
bench_arm(){ # bench_arm <arm-label>
local arm="$1" cargs=()
for c in $CONCS; do cargs+=(--concurrency "$c"); done
for rep in $(seq 1 "$REPS"); do
log " bench $arm rep=$rep"
python3 "$BENCH" --base "http://127.0.0.1:$PORT" --model "$MODEL" \
"${cargs[@]}" --requests-per-stream "$RPS" --max-tokens "$MAXTOK" \
--tag "$arm-rep$rep" --out "$OUT/res-$arm-rep$rep.json" 2>&1 | tee -a "$OUT/campaign.log"
done
}
# --- the campaign -----------------------------------------------------------
# off_A first, the MTP arms in ascending k, off_B LAST so the floor brackets the
# whole run rather than sitting at one end of it.
run_arm(){ boot "$1" "${2:-}" && bench_arm "$1" || log " arm $1 SKIPPED (boot failed)"; }
log "########## CAMPAIGN START (reps=$REPS concs='$CONCS' max_tokens=$MAXTOK) ##########"
run_arm off_A ""
run_arm k1 "{\"method\":\"$METHOD\",\"num_speculative_tokens\":1}"
run_arm k2 "{\"method\":\"$METHOD\",\"num_speculative_tokens\":2}"
run_arm k3 "{\"method\":\"$METHOD\",\"num_speculative_tokens\":3}"
run_arm off_B ""
stop_bench
missing=0
for a in off_A k1 k2 k3 off_B; do
for r in $(seq 1 "$REPS"); do
[ -s "$OUT/res-$a-rep$r.json" ] || { log " MISSING RESULT: $a rep$r"; missing=$((missing+1)); }
done
done
if [ "$missing" -gt 0 ]; then
log "########## CAMPAIGN INCOMPLETE -- $missing missing result(s); DO NOT treat as a finished run ##########"
else
log "########## CAMPAIGN DONE -- results in $OUT ##########"
fi
# Deliberately does NOT touch the production seat. It is live on GPU 2 / :8022 and
# serving the operator's `gen-large` traffic throughout; this campaign only ever
# created and removed `fn-mtp-bench` on GPU 3.
log "peak power draw seen per card:"
awk -F'|' '{for(i=1;i<=NF;i++) if($i ~ /W/){split($i,a,", "); n=split(a[1],b," "); idx=b[n]; gsub(/ W/,"",a[2]); if(a[2]+0>m[idx]) m[idx]=a[2]+0}} END{for(g in m) printf " GPU %s peak %.0f W\n", g, m[g]}' "$OUT/power.log" | sort | tee -a "$OUT/campaign.log"