Adds docs/pfi/llm-seat-catalog.md, the durable hand-curated record of what each seat IS -- lineage, provenance, model-card facts, quantization, speculative decoding, licenses, and measured warm tok/s + deep-prefill depth results with their harness and date. It complements the auto-generated fv-ml1-gpu-seat-inventory.md (live placement/KV/concurrency): the two split by volatility, and the catalog defers to the inventory for any live number. Adds scripts/seat-bench.py so the catalog's speed/depth numbers are reproducible (warm decode tok/s, n=3, conc=1, median; deep prefill at ~0.97x max-model-len with an allocator-log OOM scan). Serial by design -- concurrent deep prefills would confound both OOM and tok/s. Captures the 2026-09-14 measurements: all six generative seats prefill to ~255K (coder ~16K) with zero OOM/CUBLAS/illegal-memory; warm decode 62.7-337.3 tok/s; per-seat VRAM. seat-inventory.py now cross-links the catalog in its footer.
5.8 KiB
fv-ml1 — GPU seat inventory and model lineage
Generated 2026-09-14 17:45 UTC by scripts/seat-inventory.py, read from the running
containers on 100.64.0.7 — docker inspect, nvidia-smi, each model's own
config.json, and the .PROVENANCE.txt siblings on /tank.
⚠ .PROVENANCE.txt lives beside the model directory, not inside it:
/tank/aimodels/<model>.PROVENANCE.txt. ls <model>/ will not show it.
Placement, KV cache and concurrency
| GPU | seat | VRAM | weights | KV tokens | ctx | concurrency | util |
|---|---|---|---|---|---|---|---|
| 0 | vllm-cyberprev |
46.0 GiB | 25.11 GiB | 359,372 | 262144 | 1.37× | 0.40 |
| 0 | vllm-gen-small |
35.8 GiB | 23.98 GiB | 670,142 | 262144 | 2.56× | 0.48 |
| 1 | vllm-meromero-rp |
43.0 GiB | 19.51 GiB | 320,774 | 262144 | 1.22× | 0.52 |
| 1 | vllm-erp-seat |
27.1 GiB | 15.9 GiB | 534,649 | 262144 | 2.04× | 0.24 |
| 1 | vllm-reward |
9.0 GiB | 4.41 GiB | 26,224 | 16384 | 1.60× | 0.10 |
| 1 | vllm-coder |
6.1 GiB | 2.98 GiB | 77,056 | 16384 | 4.70× | 0.055 |
| 1 | vllm-embed |
3.4 GiB | 1.12 GiB | 10,272 | 8192 | 1.25× | 0.03 |
| 1 | vllm-rerank-a3 |
2.1 GiB | 1.06 GiB | — | 8192 | — | 0.03 |
| 2 | vllm-flash-next |
93.1 GiB | 76.82 GiB | 344,155 | 262144 | 1.31× | 0.96 |
Concurrency = KV tokens ÷ context: how many full-length requests fit at once. Below ~1.0× the seat cannot hold even one conversation at its declared context.
Lineage and quantization
vllm-cyberprev — GPU 0
- serves:
cyberprev-27b,cyberprev-27b-thinking - model:
/tank/aimodels/cyberprev-nvfp4-mixed-v2 - architecture: Qwen3_5ForConditionalGeneration (qwen3_5), 64 layers
- quantization: compressed-tensors / mixed-precision — W8A8 (float-quantized), W4A4 (nvfp4-pack-quantized)
- speculative decoding:
{"method": "dflash", "model": "/drafter", "num_speculative_tokens": 7} - image:
vllm/vllm-openai:nightly-e9d1398d9edfd90fcc1cf783805240e3effec013
vllm-gen-small — GPU 0
- serves:
gen-small,gen-small-thinking - model:
/tank/aimodels/qwen36-35b-a3b-heretic-nvfp4 - architecture: Qwen3_5MoeForConditionalGeneration (qwen3_5_moe), 40 layers, 256 experts
- quantization: modelopt / None — W4A4 (None)
- speculative decoding:
{"method": "qwen3_5_mtp", "num_speculative_tokens": 3} - image:
vllm/vllm-openai:nightly-e9d1398d9edfd90fcc1cf783805240e3effec013
vllm-coder — GPU 1
- serves:
qwen2.5-coder-1.5b - model:
? - image:
vllm/vllm-openai:v0.24.0
vllm-embed — GPU 1
- serves:
Qwen/Qwen3-Embedding-0.6B - model:
? - image:
vllm/vllm-openai:latest⚠ floating tag
vllm-erp-seat — GPU 1
- serves:
G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16 - model:
/tank/aimodels/G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16 - architecture: Gemma4ForConditionalGeneration (gemma4), 30 layers, 128 experts
- quantization: compressed-tensors / nvfp4-pack-quantized — W4A16 (nvfp4-pack-quantized)
- image:
vllm/vllm-openai:nightly-311b3513af33bc29b4acb2fde2e9313e5e9966a0
vllm-meromero-rp — GPU 1
- serves:
char-rp,char-rp-thinking - model:
/tank/aimodels/meromero-v2-nvfp4-work/G4-MeroMero-v2-31B-NVFP4A16 - architecture: Gemma4ForConditionalGeneration (gemma4), 60 layers
- quantization: compressed-tensors / nvfp4-pack-quantized — W4A16 (nvfp4-pack-quantized)
- image:
vllm/vllm-openai:v0.26.0
vllm-rerank-a3 — GPU 1
- serves:
BAAI/bge-reranker-v2-m3 - model:
? - image:
vllm/vllm-openai:v0.24.0
vllm-reward — GPU 1
- serves:
Skywork/Skywork-Reward-V2-Llama-3.1-8B-AWQ - model:
/tank/aimodels/llm/Skywork-Reward-V2-Llama-3.1-8B-AWQ - architecture: LlamaForSequenceClassification (llama), 32 layers
- quantization: compressed-tensors / pack-quantized — W4A16 (pack-quantized)
- image:
vllm/vllm-openai:latest⚠ floating tag
vllm-flash-next — GPU 2
- serves:
qwen3.8-flash-next-uncensored,qwen3.8-flash-next-uncensored-thinking - model:
/tank/aimodels/qwen38-flash-next-orcarouter-nvfp4-plefp8 - architecture: Qwen4ExpForConditionalGeneration (qwen4_exp), 48 layers, 512 experts
- quantization: compressed-tensors / mixed-precision — W8A16 (naive-quantized), W4A16 (nvfp4-pack-quantized)
- speculative decoding:
{"method": "mtp", "num_speculative_tokens": 3} - image:
vllm/vllm-openai:nightly-eed1f3d0c6043bd494424a22443ee198dd56f657
Gateway aliases resolving to this host
22 aliases. Ports with no listening seat are marked dead.
| alias | port |
|---|---|
char-rp |
8016 |
char-rp-fast |
8021 |
char-rp-reasoning |
8016 |
chat-judge |
8022 |
classifier |
8026 |
classifier-large |
8022 |
coder-fast |
8020 |
erp-tune-v2 |
8098 |
gemma4-26b-a4b-it-base |
8099 |
gen |
8022 |
gen-large |
8022 |
gen-reasoning |
8022 |
gen-small |
8026 |
gen-small-reasoning |
8026 |
image-judge |
8022 |
qwen-image-bench |
8022 |
qwen3-embedding |
8001 |
reranker |
8013 |
sec |
8025 |
sec-reasoning |
8025 |
summarizer |
8026 |
summarizer-large |
8022 |
Lineage, provenance, model cards, measured tok/s and depth results live in the hand-curated companion llm-seat-catalog.md.
Regenerate with scripts/seat-inventory.py after ANY seat change —
model swap, quant change, context or utilization edit, or speculative-decoding
change. Run --check in CI to catch a stale document.