feat(vllm): replace phi4-mini with Granite 4.1 8B summarizer + retune GPU 1
Granite 4.1 8B beat phi4-mini on precision in brokkr's R15 P03 eval, so it's the new production summarizer/dreamer for nevermore. - vllm-phi4 -> vllm-granite: official IBM FP8 (ibm-granite/granite-4.1-8b-fp8, compressed-tensors), GPU 1, 50K ctx, FP8-KV, CUDA graphs. Same :8004 slot. - GPU 1 retune: the embed/rerank/reward trio was over-provisioned (embed ran a 5.89x KV pool, reward 3.90x). Trimmed utils 0.20/0.20/0.30 -> 0.07/0.07/0.18, freeing ~10 GB so granite runs with CUDA graphs (not --enforce-eager) and keeps ~10 GB free as a hedge for future Granite text-LoRAs (--enable-lora). - LiteLLM: phi4-mini model_list entry -> granite-4.1-8b (hosted_vllm @ :8004); explicit entry shadows the '*' wildcard's llama-swap route. - nevermore repointed (LLAMA_SWAP_MODEL=granite-4.1-8b via the gateway) live. Verified end-to-end: vLLM :8004 generates, gateway routes (gateway-granite-ok), KV 86,768 tokens/1.69x at 50K, 0 restarts, GPU 1 10.3 GB free.
This commit is contained in:
@@ -17,11 +17,14 @@
|
||||
# route — left direct; see README.
|
||||
|
||||
model_list:
|
||||
# --- Phi-4-mini (generative chat) — production summarizer + dreaming
|
||||
# agent. Full prompt + completion captured per call. ---
|
||||
- model_name: phi4-mini
|
||||
# --- Granite 4.1 8B (generative chat) — production summarizer + dreaming
|
||||
# agent. Replaced phi4-mini 2026-06-05 (beat it on precision in brokkr's
|
||||
# R15 P03 eval). vLLM on ana-ml2 GPU 1, official FP8, 50K ctx. Explicit
|
||||
# entry shadows the "*" wildcard's llama-swap route for this name. Full
|
||||
# prompt + completion captured per call. ---
|
||||
- model_name: granite-4.1-8b
|
||||
litellm_params:
|
||||
model: hosted_vllm/phi4-mini
|
||||
model: hosted_vllm/granite-4.1-8b
|
||||
api_base: http://10.250.50.54:8004/v1
|
||||
api_key: os.environ/VLLM_API_KEY
|
||||
model_info:
|
||||
|
||||
Reference in New Issue
Block a user