Files
vh 5f049cb4ad feat(sglang): stage vLLM-vs-SGLang bench stack on ana-ml2
SGLang 0.5.13 confirmed to support our formats on Blackwell sm_120
(compressed-tensors NVFP4 W4A4, fp8, modelopt_fp4, petit_nvfp4, fp4_e2m1 KV),
so the bench can be a real NVFP4 head-to-head. Parameterized compose (model/
quant/GPU via .env) + a common streaming load generator (bench.py: agg tok/s,
TTFT p50/p99, TPOT) so both engines are driven identically on an exclusive GPU.
Bench-oriented; promote to a real stack only if SGLang wins. Launch deferred
until the NVFP4 eval frees a GPU.
2026-06-12 22:40:26 -07:00
..

sglang — vLLM-vs-SGLang bench on ana-ml2

Stood up to benchmark SGLang against vLLM on the same model + hardware, to see whether SGLang's throughput/latency wins justify it as a serving option (or a replacement) for the granite path on the Blackwells.

Bench-oriented, not a permanent service (yet). If SGLang wins decisively → promote to a real stack + add a gateway entry. Otherwise tear it down after.

Capability (checked 2026-06-13)

SGLang 0.5.13 (torch 2.11+cu130) supports our formats on Blackwell sm_120: compressed-tensors (the llm-compressor NVFP4 W4A4 output), fp8, modelopt_fp4, petit_nvfp4, mxfp4, and fp4_e2m1 KV. So the bench can be a real NVFP4 head-to-head, not just FP8.

The one rule for a fair bench

Exclusive GPU, same everything. Both engines must run on a card with NO co-tenants (no eval endpoints, no llama-swap hot-load), same model, same context length, same prompt profile, same concurrency sweep, driven by the SAME load generator (bench.py) — not each engine's self-flattering built-in benchmark. The contention that skewed the earlier vLLM throughput probe is exactly what to avoid here.

Run

# 1. On ana-ml2, after the eval frees a GPU: cp .env.example .env, set
#    SGLANG_MODEL / SGLANG_QUANT to match the vLLM config under test, and
#    SGLANG_GPU_ID to an EXCLUSIVE card.
scripts/deploy-stack.sh ana-ml2 sglang
# (or docker compose up -d on the host)

# 2. Bench SGLang:
python3 stacks/sglang/bench.py --url http://10.250.50.54:30000/v1 \
    --model granite-4.1-8b-nvfp4 --concurrency 1 10 50 100 200 --in-tokens 2048 --out-tokens 256

# 3. Stop SGLang, bring up vLLM on the SAME GPU + model, bench identically:
python3 stacks/sglang/bench.py --url http://10.250.50.54:8006/v1 \
    --model granite-4.1-8b-nvfp4 --concurrency 1 10 50 100 200 --in-tokens 2048 --out-tokens 256

# 4. Repeat the sweep at --in-tokens 30000 (the prefill-heavy agent-memory
#    regime, where the engines can diverge sharply).

Metrics (bench.py reports)

  • agg_tok/s — aggregate output throughput at concurrency N (the headline)
  • ttft_p50 / p99 — time-to-first-token (prefill latency; matters most at high in-tokens)
  • tpot_ms — time-per-output-token (decode latency; the per-stream UX number)

mem-fraction-static is SGLang's gpu-memory-utilization analog; set it high (0.850.90) on an exclusive 96 GB card.

Bench target

Bench whichever format wins Brokkr's quality eval (the production-relevant one): 8B-NVFP4-W4A4 if that's the path, else FP8. Benching a format we won't ship is academic. Optionally run both formats to see if the engine ranking flips.