tune(gpu1): grow selene 0.13→0.17 + qwen36 0.32→0.34 into the buffer
Put GPU1's idle ~11 GB buffer to work on the two KV-bound models that gained live consumers from the worldtree migration (granite + the pooling models under-use their util, so growing them is wasted): - selene 0.13→0.17: KV 2.53→6.33 GiB, concurrency 1.27x→3.16x @32K (Domari judge) - qwen36 0.32→0.34: KV 7.73→9.63 GiB, concurrency 2.92x→3.64x @131K (arbo judge + worldtree actor/echo + gateway) GPU1 free now ~5.6 GB (safe floor for single-service recreates).
This commit is contained in:
@@ -22,9 +22,11 @@ QWEN_GPU_ID=1
|
||||
# fp16 KV + CUDA-graph. fp16 KV (compose drops --kv-cache-dtype fp8): the NVFP4
|
||||
# swap freed enough room to run full-precision KV. Hybrid attn (10/40 full-attn)
|
||||
# keeps even fp16 KV cheap. GPU-1 budget (2026-06-15 rebalance, pinned): qwen36
|
||||
# 0.32 + granite 0.34/131072 (RESTORED from the FP8-era 0.24/64K) + trio 0.16 =
|
||||
# ~0.82, ~24 GB free headroom. Recreate ONE service at a time (profiling race).
|
||||
QWEN_GPU_MEM_UTIL=0.32
|
||||
# GPU-1 budget (2026-06-16): qwen36 0.34 (grown from 0.32 for the arbo-judge +
|
||||
# worldtree actor/echo + gateway load) + granite 0.34/131072 + selene 0.17 + trio
|
||||
# 0.16. Nominal sum >1.0 but the pooling models + granite under-use their util, so
|
||||
# it fits with ~5-6 GB physical free. Recreate ONE service at a time (profiling race).
|
||||
QWEN_GPU_MEM_UTIL=0.34
|
||||
QWEN_MAX_MODEL_LEN=131072
|
||||
# Sampler-warmup OOM guard on the shared GPU (248K vocab × default 1024 seqs is
|
||||
# a huge transient). 32 is plenty for a vision endpoint.
|
||||
|
||||
Reference in New Issue
Block a user