c985ede07b
selene: AtlaAI Selene-1-Mini-Llama-3.1-8B judge restored on vLLM after the llama-swap teardown took its Q6_K GGUF offline. FP8 (dynamic --quantization fp8; FP8 >= the validated Q6_K fidelity, and text-only Llama so no vision- tower-noise risk; NVFP4's W4A4 too aggressive for a precision judge). GPU1 util 0.13 (8.51 GiB weights + 2.53 GiB KV, 32K ctx, 1.27x concurrency), ~11 GB GPU1 buffer left. Gateway selene-1-mini-8b → :8011 (shadows the * wildcard that used to reach it via llama-swap). Judge smoke: scored an unfaithful claim 1/5 correctly. mistral-small-4: max-model-len 131072 → 262144 (full native 256K) for novel-length consistency-checking. KV pool is util-bound (~862K tokens), so 256K costs no extra VRAM — max concurrency just drops to 3.29x at full length. max-num-seqs 64 → 32 keeps the warmup transient flat (scales with seqs × len), so it fits the tight GPU0 (free unchanged at 5.2 GB). Verified loaded + healthy.
206 lines
8.9 KiB
YAML
206 lines
8.9 KiB
YAML
# LiteLLM gateway config — fronts the vLLM services on ana-ml2
|
|
# (10.250.50.54) and logs every request + response so they're
|
|
# inspectable in the Logs UI at http://10.250.50.70:4000/ui.
|
|
#
|
|
# Deploys to /opt/docker/conf/litellm/config.yaml (mounted read-only
|
|
# into the container at /app/config.yaml).
|
|
#
|
|
# Model-name → upstream mapping:
|
|
# phi4-mini → vLLM :8004 (generative chat)
|
|
# qwen3-embedding → vLLM :8001 (/v1/embeddings)
|
|
# qwen3-reranker → vLLM :8002 (/rerank)
|
|
# * (wildcard) → llama-swap :9292 (the swappable generative zoo)
|
|
#
|
|
# The wildcard fronts llama-swap so its whole model zoo logs through the
|
|
# gateway without per-model registration. The vllm-reward classifier
|
|
# (:8003) is a pooling /classify endpoint with no first-class LiteLLM
|
|
# route — left direct; see README.
|
|
|
|
model_list:
|
|
# --- Granite 4.1 8B (generative chat) — production summarizer + dreaming
|
|
# agent. Replaced phi4-mini 2026-06-05 (beat it on precision in brokkr's
|
|
# R15 P03 eval). vLLM on ana-ml2 GPU 1, official FP8, 50K ctx. Explicit
|
|
# entry shadows the "*" wildcard's llama-swap route for this name. Full
|
|
# prompt + completion captured per call. ---
|
|
- model_name: granite-4.1-8b
|
|
litellm_params:
|
|
model: hosted_vllm/granite-4.1-8b
|
|
api_base: http://10.250.50.54:8004/v1
|
|
api_key: os.environ/VLLM_API_KEY
|
|
model_info:
|
|
mode: chat
|
|
|
|
# --- Qwen3.6-35B-A3B vision-language MoE (official FP8) — vision + chat. vLLM
|
|
# on ana-ml2 GPU 1, :8007. Explicit entry shadows the "*" wildcard llama-swap
|
|
# route. REPLACED qwen3.5-9b-fp8 2026-06-14 (the 9B is retired; this is a
|
|
# 35B-A3B MoE — served under its TRUE name, never aliased under the old one).
|
|
#
|
|
# THINKING SPLIT (2026-06-15, operator call — mirrors the glm-5.1 pattern
|
|
# above). One hybrid checkpoint; the per-request `enable_thinking` switch
|
|
# picks the mode. LiteLLM forwards extra_body verbatim to vLLM, where
|
|
# chat_template_kwargs lands in the chat template. vLLM runs
|
|
# --reasoning-parser qwen3 so reasoning surfaces as reasoning_content. ---
|
|
# qwen3.6-35b-a3b: thinking DISABLED by default. The checkpoint defaults
|
|
# thinking ON; enable_thinking=false forces the empty <think></think> block.
|
|
# Reasoning is opt-in via qwen3.6-35b-a3b-thinking below.
|
|
- model_name: qwen3.6-35b-a3b
|
|
litellm_params:
|
|
model: hosted_vllm/qwen3.6-35b-a3b
|
|
api_base: http://10.250.50.54:8007/v1
|
|
api_key: os.environ/VLLM_API_KEY
|
|
extra_body:
|
|
chat_template_kwargs:
|
|
enable_thinking: false
|
|
model_info:
|
|
mode: chat
|
|
# qwen3.6-35b-a3b-thinking: identical upstream checkpoint, thinking ENABLED
|
|
# (opt-in reasoning). The qwen3 reasoning-parser splits <think>…</think> into
|
|
# reasoning_content; content holds just the answer.
|
|
- model_name: qwen3.6-35b-a3b-thinking
|
|
litellm_params:
|
|
model: hosted_vllm/qwen3.6-35b-a3b
|
|
api_base: http://10.250.50.54:8007/v1
|
|
api_key: os.environ/VLLM_API_KEY
|
|
extra_body:
|
|
chat_template_kwargs:
|
|
enable_thinking: true
|
|
model_info:
|
|
mode: chat
|
|
|
|
# --- Mistral Small 4 (official NVFP4) — creative-writing / general text. 119B
|
|
# MoE (6.5B active), vLLM on ana-ml2 GPU 0 (dedicated 96 GB Blackwell), :8010.
|
|
# Explicit entry shadows the "*" wildcard. TEXT-ONLY for now — vLLM 0.23.0's
|
|
# Mistral multimodal processor crashes at startup (loaded with image/video
|
|
# limit 0); vision returns when vLLM patches it. Deployed 2026-06-15. ---
|
|
- model_name: mistral-small-4
|
|
litellm_params:
|
|
model: hosted_vllm/mistral-small-4
|
|
api_base: http://10.250.50.54:8010/v1
|
|
api_key: os.environ/VLLM_API_KEY
|
|
model_info:
|
|
mode: chat
|
|
|
|
# mistral-small-4-reasoning: same upstream checkpoint, reasoning ON (2026-06-15,
|
|
# operator wanted "medium" — but Mistral's reasoning_effort is BINARY, only
|
|
# 'none' or 'high' (a medium/low request 400s). 'high' is the sole reasoning-ON
|
|
# level, so it carries the -reasoning intent. LiteLLM forwards extra_body to
|
|
# vLLM; the mistral reasoning-parser splits [THINK]…[/THINK] into reasoning_content.
|
|
- model_name: mistral-small-4-reasoning
|
|
litellm_params:
|
|
model: hosted_vllm/mistral-small-4
|
|
api_base: http://10.250.50.54:8010/v1
|
|
api_key: os.environ/VLLM_API_KEY
|
|
extra_body:
|
|
reasoning_effort: high
|
|
model_info:
|
|
mode: chat
|
|
|
|
# --- Selene 1 Mini 8B (AtlaAI judge, FP8) — restored on GPU1 after the
|
|
# llama-swap teardown (was the Q6_K GGUF in the swap zoo). vLLM dynamic fp8,
|
|
# :8011. Explicit entry shadows the "*" wildcard (which used to reach it via
|
|
# llama-swap). Hallucination/RAG-faithfulness judge; callers set temp ~0.01. ---
|
|
- model_name: selene-1-mini-8b
|
|
litellm_params:
|
|
model: hosted_vllm/selene-1-mini-8b
|
|
api_base: http://10.250.50.54:8011/v1
|
|
api_key: os.environ/VLLM_API_KEY
|
|
model_info:
|
|
mode: chat
|
|
|
|
# --- Qwen3 embeddings ---
|
|
- model_name: qwen3-embedding
|
|
litellm_params:
|
|
model: hosted_vllm/Qwen/Qwen3-Embedding-0.6B
|
|
api_base: http://10.250.50.54:8001/v1
|
|
api_key: os.environ/VLLM_API_KEY
|
|
model_info:
|
|
mode: embedding
|
|
|
|
# --- Qwen3 reranker (proxy /rerank route) ---
|
|
- model_name: qwen3-reranker
|
|
litellm_params:
|
|
model: hosted_vllm/Qwen/Qwen3-Reranker-0.6B
|
|
api_base: http://10.250.50.54:8002/v1
|
|
api_key: os.environ/VLLM_API_KEY
|
|
model_info:
|
|
mode: rerank
|
|
|
|
# --- z.ai GLM (cloud API) — fronted for unified logging across local
|
|
# + cloud inference. Explicit entries, so they win over the "*"
|
|
# wildcard below (no collision with llama-swap's glm4.7-flash etc.
|
|
# — different model IDs). NOTE: paid API; only gateway-keyed callers
|
|
# can reach these, but they DO spend z.ai credits. Key in .env. ---
|
|
# glm-5.1: thinking DISABLED by default (2026-06-11, operator call). LiteLLM
|
|
# strips a top-level `thinking` param (drop_params), but forwards `extra_body`
|
|
# verbatim to z.ai, where the native thinking:{type:disabled} control lands —
|
|
# verified reasoning_tokens→0. Reasoning is opt-in via glm-5.1-reasoning below.
|
|
- model_name: glm-5.1
|
|
litellm_params:
|
|
model: openai/glm-5.1
|
|
api_base: https://api.z.ai/api/coding/paas/v4
|
|
api_key: os.environ/Z_AI_API_KEY
|
|
extra_body:
|
|
thinking:
|
|
type: disabled
|
|
# glm-5.1-reasoning: identical upstream, thinking ENABLED (opt-in reasoning).
|
|
- model_name: glm-5.1-reasoning
|
|
litellm_params:
|
|
model: openai/glm-5.1
|
|
api_base: https://api.z.ai/api/coding/paas/v4
|
|
api_key: os.environ/Z_AI_API_KEY
|
|
extra_body:
|
|
thinking:
|
|
type: enabled
|
|
- model_name: glm-5-turbo
|
|
litellm_params:
|
|
model: openai/glm-5-turbo
|
|
api_base: https://api.z.ai/api/coding/paas/v4
|
|
api_key: os.environ/Z_AI_API_KEY
|
|
- model_name: glm-4.7
|
|
litellm_params:
|
|
model: openai/glm-4.7
|
|
api_base: https://api.z.ai/api/coding/paas/v4
|
|
api_key: os.environ/Z_AI_API_KEY
|
|
- model_name: glm-4.5-air
|
|
litellm_params:
|
|
model: openai/glm-4.5-air
|
|
api_base: https://api.z.ai/api/coding/paas/v4
|
|
api_key: os.environ/Z_AI_API_KEY
|
|
|
|
# --- llama-swap passthrough (the swappable generative LLM zoo on
|
|
# ana-ml2:9292) ---
|
|
# Wildcard: any model name NOT matched by an exact entry above routes to
|
|
# llama-swap, which swaps the requested model into GPU on demand. This lets
|
|
# the gateway front the WHOLE swappable zoo (artemis / selene / qwen3.x / …)
|
|
# for logging + auth WITHOUT registering each model here — keep adding and
|
|
# swapping models in llama-swap freely; litellm logs them all. litellm does
|
|
# no inference; llama-swap still does all the model loading + serving.
|
|
# Exact matches above (phi4-mini / qwen3-embedding / qwen3-reranker) win;
|
|
# this only catches everything else. `openai/*` forwards the requested model
|
|
# name verbatim to llama-swap's OpenAI-compatible endpoint.
|
|
- model_name: "*"
|
|
litellm_params:
|
|
model: openai/*
|
|
api_base: http://10.250.50.54:9292/v1
|
|
api_key: "noauth" # llama-swap takes no auth; placeholder bearer
|
|
|
|
general_settings:
|
|
master_key: os.environ/LITELLM_MASTER_KEY
|
|
database_url: os.environ/DATABASE_URL
|
|
store_model_in_db: true
|
|
# THE log switch: persists full request messages + response bodies into
|
|
# SpendLogs so they render in the Logs UI. Without this you get metadata
|
|
# (tokens, latency, model) but not the prompt/completion text.
|
|
store_prompts_in_spend_logs: true
|
|
|
|
litellm_settings:
|
|
# vLLM rejects some OpenAI params other backends accept; drop silently
|
|
# rather than 400 the caller.
|
|
drop_params: true
|
|
# --- Langfuse trace export (live 2026-06-05). Full prompt/completion +
|
|
# reasoning + tok-derivable latency traces ship to the Langfuse stack on
|
|
# ana-docker (project "gateway"). Keys + host in .env. The gateway and
|
|
# every consumer stay pointed here — this callback is the whole upgrade. ---
|
|
success_callback: ["langfuse"]
|
|
failure_callback: ["langfuse"]
|