tune(vllm): GPU-1 rebalance — granite 131k ctx, qwen 65k ctx, ~3.5GB free

Reclaimed Qwen3.5-9B's over-provisioned KV (20x conc @ 32k) and handed it
to granite. granite: 51200->131072 ctx (305k-token pool, 2.33x worst-case;
PagedAttention => ~2.2x more short-request concurrency from the bigger pool),
util 0.36->0.35. qwen: 32768->65536 ctx (8.13x), util 0.40->0.35. Trio
unchanged (chunked inputs, 8k plenty). Leaves ~3.7GB free on the shared
card. Start-order matters (trim qwen first, then grow granite) — vLLM
requires free>=util*total at startup.
This commit is contained in:
vh
2026-06-13 08:46:46 -07:00
parent 38186be1a7
commit 1e2a3a13b5
2 changed files with 19 additions and 12 deletions
+7 -4
View File
@@ -15,10 +15,13 @@ QWEN_PORT=8007
# GPU 0 is kept free for hot-reloading large models.
QWEN_GPU_ID=1
# 0.40 (~38 GB) — above the ~34 GB start floor, ~10 GB card headroom over prod.
# On this shared card vLLM needs free >= util*total, so util is capped ~0.51.
QWEN_GPU_MEM_UTIL=0.40
QWEN_MAX_MODEL_LEN=32768
# util 0.35 (~33.6 GB) / max-len 65536 — GPU-1 rebalance 2026-06-13. Qwen was
# wildly over-provisioned (20x conc @ 32k); trimmed to free room for granite while
# DOUBLING qwen's own context (32k->65k, still ~8x conc). On this shared card vLLM
# needs free >= util*total (cap ~0.51 here); floor to start at 65k is ~0.34.
# Raising max-len is free for short requests (PagedAttention = KV per actual token).
QWEN_GPU_MEM_UTIL=0.35
QWEN_MAX_MODEL_LEN=65536
# Optional
HF_TOKEN=