feat(vllm): add phi4-mini FP8 summarizer/dreamer on ana-ml2; retire granite from llama-swap

phi4-mini supersedes the granite-4-small llama-swap pin as the summarizer +
dreaming agent. New vllm-phi4 service: Phi-4-mini-instruct, vLLM-native FP8
(near-lossless on RTX 6000 Ada cc 8.9), 50K ctx, FP8 KV cache, GPU 1, :8004.

llama-swap: removed granite-4-small (depinned) + granite-4-micro config —
both superseded. CONFIG ONLY; the GGUFs stay on disk. Frees granite-4-small's
~24 GB (it was pinned at 120K ctx).

Placement: phi4 on GPU 1 with the embed/rerank/reward trio (~1.2 GB margin at
50K); keeps GPU 0 clear for llama-swap heavy models. docs/roadmap.md captures
the deferred vLLM observability (Langfuse req/resp tracing + Prometheus/Grafana).

Deploy order: llama-swap config (free granite) -> vllm-phi4 -> repoint nevermore.
This commit is contained in:
vh
2026-06-03 23:21:45 -07:00
parent 5cd4678d9f
commit 40a374b809
4 changed files with 151 additions and 37 deletions
+17 -37
View File
@@ -445,36 +445,18 @@ models:
# GRANITE MODELS (IBM)
# ==========================================================================
"granite-4-small":
name: "Granite 4.0 Small Q4_K_M"
description: "IBM Granite 4.0 Small. Deterministic utility model for structured tasks."
ttl: 0 # pinned — member of the `pinned` group, never unloads
cmd: |
/app/llama-server
--context-shift
--model /models/unsloth_granite-4.0-h-small-GGUF/granite-4.0-h-small-Q4_K_M.gguf
--port ${PORT}
--n-gpu-layers 999
--ctx-size 120000
--flash-attn on
--top-p 1.0
--temp 0.0
--top-k 0
# ── "granite-4-small" REMOVED 2026-06-04 ──────────────────────────────────
# Superseded by phi4-mini (summarizer + dreaming agent), now served via vLLM
# FP8 on ana-ml2 (stacks/vllm → vllm-phi4, :8004). It was pinned at 120K ctx
# (~24 GB resident: ~6 GB weights + ~18 GB KV); removing it reclaims that VRAM.
# DOWNSTREAM: repoint the news-digest curator from this llama-swap endpoint to
# the phi4-mini vLLM endpoint at/before deploy.
"granite-4-micro":
name: "Granite 4.0 Micro Q4_K_M"
description: "IBM Granite 4.0 Micro. Ultra-lightweight for fast structured responses."
ttl: 600
cmd: |
/app/llama-server
--context-shift
--model /models/ibm-granite_granite-4.0-micro-GGUF/granite-4.0-micro-Q4_K_M.gguf
--port ${PORT}
--n-gpu-layers 999
--ctx-size 32768
--flash-attn on
--temp 0.0
--top-p 1.0
# ── "granite-4-micro" config REMOVED 2026-06-04 ───────────────────────────
# Retired from llama-swap alongside granite-4-small (both superseded by
# phi4-mini). Per operator: CONFIG ONLY — the GGUF stays on disk at
# /models/ibm-granite_granite-4.0-micro-GGUF/granite-4.0-micro-Q4_K_M.gguf
# (NOT deleted), so this entry can be restored later if needed.
# ==========================================================================
# JUDGE / EVAL MODELS
@@ -605,13 +587,12 @@ groups:
#
# Current pins:
# qwen3.5-9b — ~6 GB at Q4 + KV. General-purpose chat baseline.
# granite-4-small — ~5-6 GB at Q4_K_M + 120K KV. Used by news-digest
# curator twice daily; pinning avoids the cold-load
# latency and prevents qwen3.6-27b (and similar)
# from evicting it when both are needed concurrently.
# VRAM budget: ~12 GB persistent in the pin slot. Single RTX 6000 Ada
# is 48 GB, so this leaves ~36 GB for whichever non-pinned model the
# user invokes alongside (qwen3.6-27b at ~30 GB fits cleanly).
# VRAM budget: ~6 GB persistent in the pin slot. Single RTX 6000 Ada
# is 48 GB, so this leaves ~40 GB for whichever non-pinned model the
# user invokes alongside.
#
# granite-4-small WAS pinned here; removed 2026-06-04 — superseded by
# phi4-mini (vLLM FP8, stacks/vllm → vllm-phi4). Freed ~24 GB (120K KV).
#
# qwen3.6-35-a3b WAS in this group; removed 2026-04-27 because its
# ~29 GB at Q6_K_XL pushed concurrent loads OOM. Now lives outside
@@ -622,4 +603,3 @@ groups:
persistent: true
members:
- "qwen3.5-9b"
- "granite-4-small"