Reclaimed Qwen3.5-9B's over-provisioned KV (20x conc @ 32k) and handed it to granite. granite: 51200->131072 ctx (305k-token pool, 2.33x worst-case; PagedAttention => ~2.2x more short-request concurrency from the bigger pool), util 0.36->0.35. qwen: 32768->65536 ctx (8.13x), util 0.40->0.35. Trio unchanged (chunked inputs, 8k plenty). Leaves ~3.7GB free on the shared card. Start-order matters (trim qwen first, then grow granite) — vLLM requires free>=util*total at startup.
29 lines
1.2 KiB
Bash
29 lines
1.2 KiB
Bash
# Qwen3.5-9B VL (FP8) on ana-ml2 — copy to .env on the host and fill.
|
|
# Real .env lives on ana-ml2 at /opt/docker/compose/qwen35-vl/.env (gitignored).
|
|
|
|
# Pinned nightly digest — carries the Qwen3.5-VL vision-FP8 exclusion fix that
|
|
# :latest (v0.19.1) lacks. Re-pin to :latest once the fix reaches a stable
|
|
# release (see README + compose header).
|
|
QWEN_IMAGE=vllm/vllm-openai@sha256:49211ab2155b21a2dc35f3583f5b545f5e55e77daf8f86df49977c71d5f2f528
|
|
|
|
QWEN_CONTAINER_NAME=vllm-qwen35
|
|
QWEN_MODEL=Qwen/Qwen3.5-9B
|
|
QWEN_SERVED_NAME=qwen3.5-9b-fp8
|
|
QWEN_PORT=8007
|
|
|
|
# GPU 1 = shared with the granite summarizer + embed/rerank/reward trio.
|
|
# GPU 0 is kept free for hot-reloading large models.
|
|
QWEN_GPU_ID=1
|
|
|
|
# util 0.35 (~33.6 GB) / max-len 65536 — GPU-1 rebalance 2026-06-13. Qwen was
|
|
# wildly over-provisioned (20x conc @ 32k); trimmed to free room for granite while
|
|
# DOUBLING qwen's own context (32k->65k, still ~8x conc). On this shared card vLLM
|
|
# needs free >= util*total (cap ~0.51 here); floor to start at 65k is ~0.34.
|
|
# Raising max-len is free for short requests (PagedAttention = KV per actual token).
|
|
QWEN_GPU_MEM_UTIL=0.35
|
|
QWEN_MAX_MODEL_LEN=65536
|
|
|
|
# Optional
|
|
HF_TOKEN=
|
|
API_KEY=
|