feat(vllm): add phi4-mini FP8 summarizer/dreamer on ana-ml2; retire granite from llama-swap
phi4-mini supersedes the granite-4-small llama-swap pin as the summarizer + dreaming agent. New vllm-phi4 service: Phi-4-mini-instruct, vLLM-native FP8 (near-lossless on RTX 6000 Ada cc 8.9), 50K ctx, FP8 KV cache, GPU 1, :8004. llama-swap: removed granite-4-small (depinned) + granite-4-micro config — both superseded. CONFIG ONLY; the GGUFs stay on disk. Frees granite-4-small's ~24 GB (it was pinned at 120K ctx). Placement: phi4 on GPU 1 with the embed/rerank/reward trio (~1.2 GB margin at 50K); keeps GPU 0 clear for llama-swap heavy models. docs/roadmap.md captures the deferred vLLM observability (Langfuse req/resp tracing + Prometheus/Grafana). Deploy order: llama-swap config (free granite) -> vllm-phi4 -> repoint nevermore.
This commit is contained in:
@@ -445,36 +445,18 @@ models:
|
||||
# GRANITE MODELS (IBM)
|
||||
# ==========================================================================
|
||||
|
||||
"granite-4-small":
|
||||
name: "Granite 4.0 Small Q4_K_M"
|
||||
description: "IBM Granite 4.0 Small. Deterministic utility model for structured tasks."
|
||||
ttl: 0 # pinned — member of the `pinned` group, never unloads
|
||||
cmd: |
|
||||
/app/llama-server
|
||||
--context-shift
|
||||
--model /models/unsloth_granite-4.0-h-small-GGUF/granite-4.0-h-small-Q4_K_M.gguf
|
||||
--port ${PORT}
|
||||
--n-gpu-layers 999
|
||||
--ctx-size 120000
|
||||
--flash-attn on
|
||||
--top-p 1.0
|
||||
--temp 0.0
|
||||
--top-k 0
|
||||
# ── "granite-4-small" REMOVED 2026-06-04 ──────────────────────────────────
|
||||
# Superseded by phi4-mini (summarizer + dreaming agent), now served via vLLM
|
||||
# FP8 on ana-ml2 (stacks/vllm → vllm-phi4, :8004). It was pinned at 120K ctx
|
||||
# (~24 GB resident: ~6 GB weights + ~18 GB KV); removing it reclaims that VRAM.
|
||||
# DOWNSTREAM: repoint the news-digest curator from this llama-swap endpoint to
|
||||
# the phi4-mini vLLM endpoint at/before deploy.
|
||||
|
||||
"granite-4-micro":
|
||||
name: "Granite 4.0 Micro Q4_K_M"
|
||||
description: "IBM Granite 4.0 Micro. Ultra-lightweight for fast structured responses."
|
||||
ttl: 600
|
||||
cmd: |
|
||||
/app/llama-server
|
||||
--context-shift
|
||||
--model /models/ibm-granite_granite-4.0-micro-GGUF/granite-4.0-micro-Q4_K_M.gguf
|
||||
--port ${PORT}
|
||||
--n-gpu-layers 999
|
||||
--ctx-size 32768
|
||||
--flash-attn on
|
||||
--temp 0.0
|
||||
--top-p 1.0
|
||||
# ── "granite-4-micro" config REMOVED 2026-06-04 ───────────────────────────
|
||||
# Retired from llama-swap alongside granite-4-small (both superseded by
|
||||
# phi4-mini). Per operator: CONFIG ONLY — the GGUF stays on disk at
|
||||
# /models/ibm-granite_granite-4.0-micro-GGUF/granite-4.0-micro-Q4_K_M.gguf
|
||||
# (NOT deleted), so this entry can be restored later if needed.
|
||||
|
||||
# ==========================================================================
|
||||
# JUDGE / EVAL MODELS
|
||||
@@ -605,13 +587,12 @@ groups:
|
||||
#
|
||||
# Current pins:
|
||||
# qwen3.5-9b — ~6 GB at Q4 + KV. General-purpose chat baseline.
|
||||
# granite-4-small — ~5-6 GB at Q4_K_M + 120K KV. Used by news-digest
|
||||
# curator twice daily; pinning avoids the cold-load
|
||||
# latency and prevents qwen3.6-27b (and similar)
|
||||
# from evicting it when both are needed concurrently.
|
||||
# VRAM budget: ~12 GB persistent in the pin slot. Single RTX 6000 Ada
|
||||
# is 48 GB, so this leaves ~36 GB for whichever non-pinned model the
|
||||
# user invokes alongside (qwen3.6-27b at ~30 GB fits cleanly).
|
||||
# VRAM budget: ~6 GB persistent in the pin slot. Single RTX 6000 Ada
|
||||
# is 48 GB, so this leaves ~40 GB for whichever non-pinned model the
|
||||
# user invokes alongside.
|
||||
#
|
||||
# granite-4-small WAS pinned here; removed 2026-06-04 — superseded by
|
||||
# phi4-mini (vLLM FP8, stacks/vllm → vllm-phi4). Freed ~24 GB (120K KV).
|
||||
#
|
||||
# qwen3.6-35-a3b WAS in this group; removed 2026-04-27 because its
|
||||
# ~29 GB at Q6_K_XL pushed concurrent loads OOM. Now lives outside
|
||||
@@ -622,4 +603,3 @@ groups:
|
||||
persistent: true
|
||||
members:
|
||||
- "qwen3.5-9b"
|
||||
- "granite-4-small"
|
||||
|
||||
Reference in New Issue
Block a user