Files
esh-pfi-infrastructure/docs/pfi/fv-ml1-gpu-seat-inventory.md
T
vh a91b841d86 feat(fv-ml1): generate the seat inventory from the live box instead of maintaining it by hand
The seat documentation must stay current, and a hand-written document cannot.
The LiteLLM config described char-rp as a 31B model on a host and GPU it had not
been on since 2026-08-24 -- three weeks of silent drift in a file that read as
authoritative, and the reason a seat spent that period serving a model nobody
intended. Anything typed here drifts the same way; anything read off the running
containers cannot.

scripts/seat-inventory.py derives the whole document from the host:

- placement and VRAM from nvidia-smi compute-apps, mapped to containers through
  /proc/<pid>/cgroup -- nvidia-smi reports the vLLM engine child while docker
  reports the container pid, so matching them directly silently yields nothing
- weights and KV tokens parsed from each engine's own startup log, not derived
  arithmetically, with concurrency computed as KV tokens over context
- architecture, layer and expert counts, and the exact quantization group scheme
  (W4A4 vs W4A16 distinguished) from each model's config.json
- speculative-decoding method and k from the container argv, which is how the
  three incompatible methods on this box became visible
- lineage from the .PROVENANCE.txt SIBLING files -- they sit beside the model
  directory, not inside it, which is why an earlier pass wrongly reported two
  fully-documented seats as having no provenance
- gateway aliases resolved from the LiteLLM config on ana-docker

--check compares the committed document against the live box and exits non-zero
when they diverge, ignoring only the generation timestamp. Suitable for CI or a
scheduled drift alarm; read-only throughout, safe against production.

Also commits the KV_CACHE_BYTES override added to the MTP campaign runner, which
asserts the flag exists in the derived argv and aborts rather than running a
campaign that silently ignored it.
2026-09-13 23:01:44 -07:00

6.0 KiB
Raw Blame History

fv-ml1 — GPU seat inventory and model lineage

Generated 2026-09-14 06:00 UTC by scripts/seat-inventory.py, read from the running containers on 100.64.0.7 — docker inspect, nvidia-smi, each model's own config.json, and the .PROVENANCE.txt siblings on /tank.

⚠ .PROVENANCE.txt lives beside the model directory, not inside it: /tank/aimodels/<model>.PROVENANCE.txt. ls <model>/ will not show it.

Placement, KV cache and concurrency

GPU seat VRAM weights KV tokens ctx concurrency util
0 vllm-mog-sec 46.6 GiB 25.47 GiB 342,920 163840 2.09× 0.50
0 vllm-gen 38.3 GiB 21.97 GiB 268,205 262144 1.02× 0.38
1 vllm-meromero-rp 37.9 GiB 19.51 GiB 266,334 262144 1.02× 0.40
1 vllm-erp-seat 26.3 GiB 15.9 GiB 534,649 262144 2.04× 0.30
1 vllm-reward 9.0 GiB 4.41 GiB 26,224 16384 1.60× 0.10
1 vllm-coder 6.0 GiB 2.98 GiB 112,624 8192 13.75× 0.06
1 vllm-embed 3.4 GiB 1.12 GiB 10,272 8192 1.25× 0.03
1 vllm-rerank-a3 2.1 GiB 1.06 GiB — 8192 — 0.03
2 vllm-flash-next 94.7 GiB 79.44 GiB 344,155 262144 1.31× 0.96

Concurrency = KV tokens ÷ context: how many full-length requests fit at once. Below ~1.0× the seat cannot hold even one conversation at its declared context.

Lineage and quantization

vllm-gen — GPU 0

  • serves: qwen3.8-27b-uncensored, qwen3.8-27b-uncensored-thinking
  • model: /tank/aimodels/qwen38-27b-orcarouter-nvfp4-mixed
  • architecture: Qwen3_5ForConditionalGeneration (qwen3_5), 64 layers
  • quantization: compressed-tensors / mixed-precision — W8A8 (float-quantized), W4A4 (nvfp4-pack-quantized)
  • speculative decoding: {"method": "qwen3_5_mtp", "num_speculative_tokens": 3}
  • image: vllm/vllm-openai:nightly-311b3513af33bc29b4acb2fde2e9313e5e9966a0

vllm-mog-sec — GPU 0

  • serves: mog-sec-27b, mog-sec-27b-thinking
  • model: /tank/aimodels/mog-sec-27b-nvfp4-mixed
  • architecture: Qwen3_5ForConditionalGeneration (qwen3_5), 64 layers
  • quantization: compressed-tensors / mixed-precision — W8A8 (float-quantized), W4A4 (nvfp4-pack-quantized)
  • speculative decoding: {"method": "dflash", "model": "/drafter", "num_speculative_tokens": 7}
  • image: vllm/vllm-openai:nightly-e9d1398d9edfd90fcc1cf783805240e3effec013
  • provenance:
    model: Blackfrost-Research/M.O.G.-SEC-27B-1M-CTX-BF16  (bf16, pen-test seat source)
    source_url: https://huggingface.co/Blackfrost-Research/M.O.G.-SEC-27B-1M-CTX-BF16
    revision_pinned: deede67794b4eaaf31f016d02a0aaf71f1a303b9
    pulled_by: infra-ops (as llmuser)
    pulled_at_utc: 2026-08-21T09:10Z
    size_on_disk: 52 GB (18 shards, index total_size 55.6 GB)
    

vllm-coder — GPU 1

  • serves: qwen2.5-coder-1.5b
  • model: ?
  • image: vllm/vllm-openai:latest ⚠ floating tag

vllm-embed — GPU 1

  • serves: Qwen/Qwen3-Embedding-0.6B
  • model: ?
  • image: vllm/vllm-openai:latest ⚠ floating tag

vllm-erp-seat — GPU 1

  • serves: G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16
  • model: /tank/aimodels/G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16
  • architecture: Gemma4ForConditionalGeneration (gemma4), 30 layers, 128 experts
  • quantization: compressed-tensors / nvfp4-pack-quantized — W4A16 (nvfp4-pack-quantized)
  • image: vllm/vllm-openai:nightly-311b3513af33bc29b4acb2fde2e9313e5e9966a0

vllm-meromero-rp — GPU 1

  • serves: char-rp, char-rp-thinking
  • model: /tank/aimodels/meromero-v2-nvfp4-work/G4-MeroMero-v2-31B-NVFP4A16
  • architecture: Gemma4ForConditionalGeneration (gemma4), 60 layers
  • quantization: compressed-tensors / nvfp4-pack-quantized — W4A16 (nvfp4-pack-quantized)
  • image: vllm/vllm-openai:v0.26.0

vllm-rerank-a3 — GPU 1

  • serves: BAAI/bge-reranker-v2-m3
  • model: ?
  • image: vllm/vllm-openai:v0.24.0

vllm-reward — GPU 1

  • serves: Skywork/Skywork-Reward-V2-Llama-3.1-8B-AWQ
  • model: /tank/aimodels/llm/Skywork-Reward-V2-Llama-3.1-8B-AWQ
  • architecture: LlamaForSequenceClassification (llama), 32 layers
  • quantization: compressed-tensors / pack-quantized — W4A16 (pack-quantized)
  • image: vllm/vllm-openai:latest ⚠ floating tag

vllm-flash-next — GPU 2

  • serves: qwen3.8-flash-next-uncensored, qwen3.8-flash-next-uncensored-thinking
  • model: /tank/aimodels/qwen38-flash-next-abliterated-nvfp4
  • architecture: Qwen4ExpForConditionalGeneration (qwen4_exp), 48 layers, 512 experts
  • quantization: modelopt / None — W4A4 (None)
  • speculative decoding: {"method": "mtp", "num_speculative_tokens": 3}
  • image: vllm/vllm-openai:nightly-eed1f3d0c6043bd494424a22443ee198dd56f657

Gateway aliases resolving to this host

19 aliases. Ports with no listening seat are marked dead.

alias port
char-rp 8016
char-rp-fast 8021
char-rp-reasoning 8016
chat-judge 8015
classifier 8015
coder-fast 8020
erp-tune-v2 8098
gemma4-26b-a4b-it-base 8099
gen 8015
gen-large 8022
gen-reasoning 8015
image-judge 8015
qwen-image-bench 8015
qwen3-embedding 8001
reranker 8013
sec 8019
sec-reasoning 8019
summarizer 8015
summarizer-large 8015

Regenerate with scripts/seat-inventory.py after ANY seat change — model swap, quant change, context or utilization edit, or speculative-decoding change. Run --check in CI to catch a stale document.