SGLang 0.5.13 confirmed to support our formats on Blackwell sm_120 (compressed-tensors NVFP4 W4A4, fp8, modelopt_fp4, petit_nvfp4, fp4_e2m1 KV), so the bench can be a real NVFP4 head-to-head. Parameterized compose (model/ quant/GPU via .env) + a common streaming load generator (bench.py: agg tok/s, TTFT p50/p99, TPOT) so both engines are driven identically on an exclusive GPU. Bench-oriented; promote to a real stack only if SGLang wins. Launch deferred until the NVFP4 eval frees a GPU.
sglang — vLLM-vs-SGLang bench on ana-ml2
Stood up to benchmark SGLang against vLLM on the same model + hardware, to see whether SGLang's throughput/latency wins justify it as a serving option (or a replacement) for the granite path on the Blackwells.
Bench-oriented, not a permanent service (yet). If SGLang wins decisively → promote to a real stack + add a gateway entry. Otherwise tear it down after.
Capability (checked 2026-06-13)
SGLang 0.5.13 (torch 2.11+cu130) supports our formats on Blackwell sm_120:
compressed-tensors (the llm-compressor NVFP4 W4A4 output), fp8,
modelopt_fp4, petit_nvfp4, mxfp4, and fp4_e2m1 KV. So the bench can be a
real NVFP4 head-to-head, not just FP8.
The one rule for a fair bench
Exclusive GPU, same everything. Both engines must run on a card with NO
co-tenants (no eval endpoints, no llama-swap hot-load), same model, same context
length, same prompt profile, same concurrency sweep, driven by the SAME load
generator (bench.py) — not each engine's self-flattering built-in benchmark.
The contention that skewed the earlier vLLM throughput probe is exactly what to
avoid here.
Run
# 1. On ana-ml2, after the eval frees a GPU: cp .env.example .env, set
# SGLANG_MODEL / SGLANG_QUANT to match the vLLM config under test, and
# SGLANG_GPU_ID to an EXCLUSIVE card.
scripts/deploy-stack.sh ana-ml2 sglang
# (or docker compose up -d on the host)
# 2. Bench SGLang:
python3 stacks/sglang/bench.py --url http://10.250.50.54:30000/v1 \
--model granite-4.1-8b-nvfp4 --concurrency 1 10 50 100 200 --in-tokens 2048 --out-tokens 256
# 3. Stop SGLang, bring up vLLM on the SAME GPU + model, bench identically:
python3 stacks/sglang/bench.py --url http://10.250.50.54:8006/v1 \
--model granite-4.1-8b-nvfp4 --concurrency 1 10 50 100 200 --in-tokens 2048 --out-tokens 256
# 4. Repeat the sweep at --in-tokens 30000 (the prefill-heavy agent-memory
# regime, where the engines can diverge sharply).
Metrics (bench.py reports)
- agg_tok/s — aggregate output throughput at concurrency N (the headline)
- ttft_p50 / p99 — time-to-first-token (prefill latency; matters most at high in-tokens)
- tpot_ms — time-per-output-token (decode latency; the per-stream UX number)
mem-fraction-static is SGLang's gpu-memory-utilization analog; set it high
(0.85–0.90) on an exclusive 96 GB card.
Bench target
Bench whichever format wins Brokkr's quality eval (the production-relevant one): 8B-NVFP4-W4A4 if that's the path, else FP8. Benching a format we won't ship is academic. Optionally run both formats to see if the engine ranking flips.