feat(fv-ml1): generate the seat inventory from the live box instead of maintaining it by hand

The seat documentation must stay current, and a hand-written document cannot.
The LiteLLM config described char-rp as a 31B model on a host and GPU it had not
been on since 2026-08-24 -- three weeks of silent drift in a file that read as
authoritative, and the reason a seat spent that period serving a model nobody
intended. Anything typed here drifts the same way; anything read off the running
containers cannot.

scripts/seat-inventory.py derives the whole document from the host:

- placement and VRAM from nvidia-smi compute-apps, mapped to containers through
  /proc/<pid>/cgroup -- nvidia-smi reports the vLLM engine child while docker
  reports the container pid, so matching them directly silently yields nothing
- weights and KV tokens parsed from each engine's own startup log, not derived
  arithmetically, with concurrency computed as KV tokens over context
- architecture, layer and expert counts, and the exact quantization group scheme
  (W4A4 vs W4A16 distinguished) from each model's config.json
- speculative-decoding method and k from the container argv, which is how the
  three incompatible methods on this box became visible
- lineage from the .PROVENANCE.txt SIBLING files -- they sit beside the model
  directory, not inside it, which is why an earlier pass wrongly reported two
  fully-documented seats as having no provenance
- gateway aliases resolved from the LiteLLM config on ana-docker

--check compares the committed document against the live box and exits non-zero
when they diverge, ignoring only the generation timestamp. Suitable for CI or a
scheduled drift alarm; read-only throughout, safe against production.

Also commits the KV_CACHE_BYTES override added to the MTP campaign runner, which
asserts the flag exists in the derived argv and aborts rather than running a
campaign that silently ignored it.
This commit is contained in:
vh
2026-09-13 23:01:44 -07:00
parent 2d83a895c1
commit a91b841d86
3 changed files with 393 additions and 129 deletions
@@ -103,6 +103,30 @@ PY
mapfile -t ARGV < "$OUT/argv.txt"
ENVARGS=(); while IFS= read -r l; do [ -n "$l" ] && ENVARGS+=(-e "$l"); done < "$OUT/envs.txt"
log "argv has ${#ARGV[@]} entries; ${#ENVARGS[@]} env flags"
# --- optional single-flag override of the production argv --------------------
# MTP's draft head costs ~5.08 GiB of weights (74.36 -> 79.44 GiB, measured
# 2026-09-13 and FLAT in k: identical for k=1, k=2 and k=3). That does not fit
# alongside the seat's 14 GiB pinned --kv-cache-memory under
# --gpu-memory-utilization 0.96, so every MTP arm OOMs at engine init while the
# no-spec arms boot fine. Pinning a smaller KV budget for EVERY arm makes room
# without confounding the comparison -- and it is free for this benchmark, which
# at conc<=8 with 400-token completions never touches more than a few thousand
# KV tokens against a cache sized in the hundreds of thousands.
if [ -n "${KV_CACHE_BYTES:-}" ]; then
kv_found=0
for i in "${!ARGV[@]}"; do
if [ "${ARGV[$i]}" = "--kv-cache-memory" ]; then
log "override: --kv-cache-memory ${ARGV[$((i+1))]} -> $KV_CACHE_BYTES"
ARGV[$((i+1))]="$KV_CACHE_BYTES"
kv_found=1
break
fi
done
# Fail loudly rather than run a campaign that silently ignored the override --
# a run whose knob did nothing is worse than a run that refused to start.
[ "$kv_found" -eq 1 ] || { log "FATAL: KV_CACHE_BYTES set but --kv-cache-memory absent from derived argv"; exit 1; }
fi
printf '%s\n' "${ARGV[@]}" | sed 's/^/ /' >> "$OUT/campaign.log"
MODEL_DIR=$(python3 -c "