5f049cb4ad
SGLang 0.5.13 confirmed to support our formats on Blackwell sm_120 (compressed-tensors NVFP4 W4A4, fp8, modelopt_fp4, petit_nvfp4, fp4_e2m1 KV), so the bench can be a real NVFP4 head-to-head. Parameterized compose (model/ quant/GPU via .env) + a common streaming load generator (bench.py: agg tok/s, TTFT p50/p99, TPOT) so both engines are driven identically on an exclusive GPU. Bench-oriented; promote to a real stack only if SGLang wins. Launch deferred until the NVFP4 eval frees a GPU.
61 lines
2.6 KiB
Markdown
61 lines
2.6 KiB
Markdown
# sglang — vLLM-vs-SGLang bench on ana-ml2
|
||
|
||
Stood up to benchmark **SGLang against vLLM** on the same model + hardware, to
|
||
see whether SGLang's throughput/latency wins justify it as a serving option
|
||
(or a replacement) for the granite path on the Blackwells.
|
||
|
||
**Bench-oriented, not a permanent service** (yet). If SGLang wins decisively →
|
||
promote to a real stack + add a gateway entry. Otherwise tear it down after.
|
||
|
||
## Capability (checked 2026-06-13)
|
||
|
||
SGLang 0.5.13 (torch 2.11+cu130) supports our formats on Blackwell sm_120:
|
||
`compressed-tensors` (the llm-compressor NVFP4 W4A4 output), `fp8`,
|
||
`modelopt_fp4`, `petit_nvfp4`, `mxfp4`, and `fp4_e2m1` KV. So the bench can be a
|
||
real **NVFP4 head-to-head**, not just FP8.
|
||
|
||
## The one rule for a fair bench
|
||
|
||
**Exclusive GPU, same everything.** Both engines must run on a card with NO
|
||
co-tenants (no eval endpoints, no llama-swap hot-load), same model, same context
|
||
length, same prompt profile, same concurrency sweep, driven by the SAME load
|
||
generator (`bench.py`) — not each engine's self-flattering built-in benchmark.
|
||
The contention that skewed the earlier vLLM throughput probe is exactly what to
|
||
avoid here.
|
||
|
||
## Run
|
||
|
||
```bash
|
||
# 1. On ana-ml2, after the eval frees a GPU: cp .env.example .env, set
|
||
# SGLANG_MODEL / SGLANG_QUANT to match the vLLM config under test, and
|
||
# SGLANG_GPU_ID to an EXCLUSIVE card.
|
||
scripts/deploy-stack.sh ana-ml2 sglang
|
||
# (or docker compose up -d on the host)
|
||
|
||
# 2. Bench SGLang:
|
||
python3 stacks/sglang/bench.py --url http://10.250.50.54:30000/v1 \
|
||
--model granite-4.1-8b-nvfp4 --concurrency 1 10 50 100 200 --in-tokens 2048 --out-tokens 256
|
||
|
||
# 3. Stop SGLang, bring up vLLM on the SAME GPU + model, bench identically:
|
||
python3 stacks/sglang/bench.py --url http://10.250.50.54:8006/v1 \
|
||
--model granite-4.1-8b-nvfp4 --concurrency 1 10 50 100 200 --in-tokens 2048 --out-tokens 256
|
||
|
||
# 4. Repeat the sweep at --in-tokens 30000 (the prefill-heavy agent-memory
|
||
# regime, where the engines can diverge sharply).
|
||
```
|
||
|
||
## Metrics (`bench.py` reports)
|
||
|
||
- **agg_tok/s** — aggregate output throughput at concurrency N (the headline)
|
||
- **ttft_p50 / p99** — time-to-first-token (prefill latency; matters most at high in-tokens)
|
||
- **tpot_ms** — time-per-output-token (decode latency; the per-stream UX number)
|
||
|
||
`mem-fraction-static` is SGLang's `gpu-memory-utilization` analog; set it high
|
||
(0.85–0.90) on an exclusive 96 GB card.
|
||
|
||
## Bench target
|
||
|
||
Bench whichever format wins Brokkr's quality eval (the production-relevant one):
|
||
8B-NVFP4-W4A4 if that's the path, else FP8. Benching a format we won't ship is
|
||
academic. Optionally run both formats to see if the engine ranking flips.
|