Files
esh-pfi-infrastructure/stacks/sglang/README.md
T
vh 5f049cb4ad feat(sglang): stage vLLM-vs-SGLang bench stack on ana-ml2
SGLang 0.5.13 confirmed to support our formats on Blackwell sm_120
(compressed-tensors NVFP4 W4A4, fp8, modelopt_fp4, petit_nvfp4, fp4_e2m1 KV),
so the bench can be a real NVFP4 head-to-head. Parameterized compose (model/
quant/GPU via .env) + a common streaming load generator (bench.py: agg tok/s,
TTFT p50/p99, TPOT) so both engines are driven identically on an exclusive GPU.
Bench-oriented; promote to a real stack only if SGLang wins. Launch deferred
until the NVFP4 eval frees a GPU.
2026-06-12 22:40:26 -07:00

61 lines
2.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# sglang — vLLM-vs-SGLang bench on ana-ml2
Stood up to benchmark **SGLang against vLLM** on the same model + hardware, to
see whether SGLang's throughput/latency wins justify it as a serving option
(or a replacement) for the granite path on the Blackwells.
**Bench-oriented, not a permanent service** (yet). If SGLang wins decisively →
promote to a real stack + add a gateway entry. Otherwise tear it down after.
## Capability (checked 2026-06-13)
SGLang 0.5.13 (torch 2.11+cu130) supports our formats on Blackwell sm_120:
`compressed-tensors` (the llm-compressor NVFP4 W4A4 output), `fp8`,
`modelopt_fp4`, `petit_nvfp4`, `mxfp4`, and `fp4_e2m1` KV. So the bench can be a
real **NVFP4 head-to-head**, not just FP8.
## The one rule for a fair bench
**Exclusive GPU, same everything.** Both engines must run on a card with NO
co-tenants (no eval endpoints, no llama-swap hot-load), same model, same context
length, same prompt profile, same concurrency sweep, driven by the SAME load
generator (`bench.py`) — not each engine's self-flattering built-in benchmark.
The contention that skewed the earlier vLLM throughput probe is exactly what to
avoid here.
## Run
```bash
# 1. On ana-ml2, after the eval frees a GPU: cp .env.example .env, set
# SGLANG_MODEL / SGLANG_QUANT to match the vLLM config under test, and
# SGLANG_GPU_ID to an EXCLUSIVE card.
scripts/deploy-stack.sh ana-ml2 sglang
# (or docker compose up -d on the host)
# 2. Bench SGLang:
python3 stacks/sglang/bench.py --url http://10.250.50.54:30000/v1 \
--model granite-4.1-8b-nvfp4 --concurrency 1 10 50 100 200 --in-tokens 2048 --out-tokens 256
# 3. Stop SGLang, bring up vLLM on the SAME GPU + model, bench identically:
python3 stacks/sglang/bench.py --url http://10.250.50.54:8006/v1 \
--model granite-4.1-8b-nvfp4 --concurrency 1 10 50 100 200 --in-tokens 2048 --out-tokens 256
# 4. Repeat the sweep at --in-tokens 30000 (the prefill-heavy agent-memory
# regime, where the engines can diverge sharply).
```
## Metrics (`bench.py` reports)
- **agg_tok/s** — aggregate output throughput at concurrency N (the headline)
- **ttft_p50 / p99** — time-to-first-token (prefill latency; matters most at high in-tokens)
- **tpot_ms** — time-per-output-token (decode latency; the per-stream UX number)
`mem-fraction-static` is SGLang's `gpu-memory-utilization` analog; set it high
(0.850.90) on an exclusive 96 GB card.
## Bench target
Bench whichever format wins Brokkr's quality eval (the production-relevant one):
8B-NVFP4-W4A4 if that's the path, else FP8. Benching a format we won't ship is
academic. Optionally run both formats to see if the engine ranking flips.