feat(sglang): stage vLLM-vs-SGLang bench stack on ana-ml2

SGLang 0.5.13 confirmed to support our formats on Blackwell sm_120
(compressed-tensors NVFP4 W4A4, fp8, modelopt_fp4, petit_nvfp4, fp4_e2m1 KV),
so the bench can be a real NVFP4 head-to-head. Parameterized compose (model/
quant/GPU via .env) + a common streaming load generator (bench.py: agg tok/s,
TTFT p50/p99, TPOT) so both engines are driven identically on an exclusive GPU.
Bench-oriented; promote to a real stack only if SGLang wins. Launch deferred
until the NVFP4 eval frees a GPU.
This commit is contained in:
vh
2026-06-12 22:40:26 -07:00
parent 19a07b96ab
commit 5f049cb4ad
4 changed files with 243 additions and 0 deletions
+60
View File
@@ -0,0 +1,60 @@
# sglang — vLLM-vs-SGLang bench on ana-ml2
Stood up to benchmark **SGLang against vLLM** on the same model + hardware, to
see whether SGLang's throughput/latency wins justify it as a serving option
(or a replacement) for the granite path on the Blackwells.
**Bench-oriented, not a permanent service** (yet). If SGLang wins decisively →
promote to a real stack + add a gateway entry. Otherwise tear it down after.
## Capability (checked 2026-06-13)
SGLang 0.5.13 (torch 2.11+cu130) supports our formats on Blackwell sm_120:
`compressed-tensors` (the llm-compressor NVFP4 W4A4 output), `fp8`,
`modelopt_fp4`, `petit_nvfp4`, `mxfp4`, and `fp4_e2m1` KV. So the bench can be a
real **NVFP4 head-to-head**, not just FP8.
## The one rule for a fair bench
**Exclusive GPU, same everything.** Both engines must run on a card with NO
co-tenants (no eval endpoints, no llama-swap hot-load), same model, same context
length, same prompt profile, same concurrency sweep, driven by the SAME load
generator (`bench.py`) — not each engine's self-flattering built-in benchmark.
The contention that skewed the earlier vLLM throughput probe is exactly what to
avoid here.
## Run
```bash
# 1. On ana-ml2, after the eval frees a GPU: cp .env.example .env, set
# SGLANG_MODEL / SGLANG_QUANT to match the vLLM config under test, and
# SGLANG_GPU_ID to an EXCLUSIVE card.
scripts/deploy-stack.sh ana-ml2 sglang
# (or docker compose up -d on the host)
# 2. Bench SGLang:
python3 stacks/sglang/bench.py --url http://10.250.50.54:30000/v1 \
--model granite-4.1-8b-nvfp4 --concurrency 1 10 50 100 200 --in-tokens 2048 --out-tokens 256
# 3. Stop SGLang, bring up vLLM on the SAME GPU + model, bench identically:
python3 stacks/sglang/bench.py --url http://10.250.50.54:8006/v1 \
--model granite-4.1-8b-nvfp4 --concurrency 1 10 50 100 200 --in-tokens 2048 --out-tokens 256
# 4. Repeat the sweep at --in-tokens 30000 (the prefill-heavy agent-memory
# regime, where the engines can diverge sharply).
```
## Metrics (`bench.py` reports)
- **agg_tok/s** — aggregate output throughput at concurrency N (the headline)
- **ttft_p50 / p99** — time-to-first-token (prefill latency; matters most at high in-tokens)
- **tpot_ms** — time-per-output-token (decode latency; the per-stream UX number)
`mem-fraction-static` is SGLang's `gpu-memory-utilization` analog; set it high
(0.85–0.90) on an exclusive 96 GB card.
## Bench target
Bench whichever format wins Brokkr's quality eval (the production-relevant one):
8B-NVFP4-W4A4 if that's the path, else FP8. Benching a format we won't ship is
academic. Optionally run both formats to see if the engine ranking flips.