# sglang — vLLM-vs-SGLang bench on ana-ml2 Stood up to benchmark **SGLang against vLLM** on the same model + hardware, to see whether SGLang's throughput/latency wins justify it as a serving option (or a replacement) for the granite path on the Blackwells. **Bench-oriented, not a permanent service** (yet). If SGLang wins decisively → promote to a real stack + add a gateway entry. Otherwise tear it down after. ## Capability (checked 2026-06-13) SGLang 0.5.13 (torch 2.11+cu130) supports our formats on Blackwell sm_120: `compressed-tensors` (the llm-compressor NVFP4 W4A4 output), `fp8`, `modelopt_fp4`, `petit_nvfp4`, `mxfp4`, and `fp4_e2m1` KV. So the bench can be a real **NVFP4 head-to-head**, not just FP8. ## The one rule for a fair bench **Exclusive GPU, same everything.** Both engines must run on a card with NO co-tenants (no eval endpoints, no llama-swap hot-load), same model, same context length, same prompt profile, same concurrency sweep, driven by the SAME load generator (`bench.py`) — not each engine's self-flattering built-in benchmark. The contention that skewed the earlier vLLM throughput probe is exactly what to avoid here. ## Run ```bash # 1. On ana-ml2, after the eval frees a GPU: cp .env.example .env, set # SGLANG_MODEL / SGLANG_QUANT to match the vLLM config under test, and # SGLANG_GPU_ID to an EXCLUSIVE card. scripts/deploy-stack.sh ana-ml2 sglang # (or docker compose up -d on the host) # 2. Bench SGLang: python3 stacks/sglang/bench.py --url http://10.250.50.54:30000/v1 \ --model granite-4.1-8b-nvfp4 --concurrency 1 10 50 100 200 --in-tokens 2048 --out-tokens 256 # 3. Stop SGLang, bring up vLLM on the SAME GPU + model, bench identically: python3 stacks/sglang/bench.py --url http://10.250.50.54:8006/v1 \ --model granite-4.1-8b-nvfp4 --concurrency 1 10 50 100 200 --in-tokens 2048 --out-tokens 256 # 4. Repeat the sweep at --in-tokens 30000 (the prefill-heavy agent-memory # regime, where the engines can diverge sharply). ``` ## Metrics (`bench.py` reports) - **agg_tok/s** — aggregate output throughput at concurrency N (the headline) - **ttft_p50 / p99** — time-to-first-token (prefill latency; matters most at high in-tokens) - **tpot_ms** — time-per-output-token (decode latency; the per-stream UX number) `mem-fraction-static` is SGLang's `gpu-memory-utilization` analog; set it high (0.85–0.90) on an exclusive 96 GB card. ## Bench target Bench whichever format wins Brokkr's quality eval (the production-relevant one): 8B-NVFP4-W4A4 if that's the path, else FP8. Benching a format we won't ship is academic. Optionally run both formats to see if the engine ranking flips.