Files
esh-pfi-infrastructure/stacks/litellm
vh 569e1af9ca feat(homepage): split AI fleet into role-based groups on a dedicated AI tab
Move the ~22-service flat "AI Systems" group off the Main tab into a new
four-tab layout (Main / AI / Infrastructure / Toolchain). The AI tab sorts
the inference fleet by function into seven groups:

  AI - Inference        gen, char-rp, char-rp-reasoning, Granite summarizer
  AI - Eval & Retrieval Selene, Skywork Reward, Qwen3 rerank/embed, image-bench
  AI - Gateways & Chat  LiteLLM, Asset Engine, Gateway Chat, Open WebUI, ...
  AI - Speech (TTS)     Chatterbox Fast, Kokoro, mOrpheus
  AI - Audio Tools      Parakeet ASR, YT Voice Clipper
  AI - Image & Media    ComfyUI, Arbo
  AI - Dormant          stopped rollback seats + retired auditions

Relabel each stack's homepage.group so canonical stacks/ matches the live
containers on ana-ml2, ana-docker, and irv-ml1. Dormant stacks were refreshed
with `docker compose up --no-start` so they carry the new label while staying
stopped (compose-start rollback preserved). settings.yaml drives tab/order/
columns; services.yaml and README updated to the new scheme.
2026-07-14 20:05:50 -07:00
..

litellm

OpenAI-compatible gateway in front of the vLLM services on ana-ml2, standing in the request path so every request + response is logged and inspectable in a browser. This is the thing vLLM does not give us: Dozzle shows vLLM's stdout (connection/request metadata) but not the full prompt/completion bodies. LiteLLM captures both, per call, with a Logs UI.

Server: ana-docker (10.250.50.70) Port: 4000 (proxy API + admin/Logs UI at /ui) — configurable in .env Backs: the vllm stack on ana-ml2 (10.250.50.54)

Why it exists

phi4-mini is becoming a production summarizer + "dreaming" agent. Being able to read exactly what it was asked and what it answered is the difference between debuggable and opaque. See docs/roadmap.md → "Observability for the vLLM stack". This is the lean first cut of that roadmap item — see Langfuse-ready below for the upgrade path.

What routes through it

Consumers point their OpenAI base_url at http://10.250.50.70:4000 and pick a model by name; the gateway forwards to the right vLLM port and logs the round-trip.

model name (here) upstream vLLM port logged
phi4-mini generative chat :8004 full prompt + completion
qwen3-embedding /v1/embeddings :8001 input + vector metadata
qwen3-reranker /rerank :8002 query + docs + scores

Not routed: the vllm-reward Skywork classifier (:8003) is a pooling /classify endpoint with no first-class LiteLLM route — callers hit it directly for now. The generative model is the high-value target for req/resp visibility and it routes cleanly here. (If reward logging is wanted later, LiteLLM pass_through_endpoints can cover it.)

The log switch

Full prompt/response text shows in the Logs UI because of store_prompts_in_spend_logs: true in conf/config.yaml. Without it you'd get metadata only (tokens, latency, model name) — not the text. The Postgres sidecar (litellm-db) is the store.

Langfuse-ready

This deliberately does not stand up Langfuse's heavy v3 stack (ClickHouse + Redis + MinIO + Postgres + app containers). To graduate to full Langfuse traces later:

  1. Stand up (or point at) a Langfuse instance.
  2. Set LANGFUSE_PUBLIC_KEY / LANGFUSE_SECRET_KEY / LANGFUSE_HOST in .env.
  3. Uncomment success_callback / failure_callback in conf/config.yaml.
  4. docker compose up -d to restart.

No re-architecture: the gateway and every consumer stay pointed here.

Deploy

# 1. Sync canonical → ana-docker (compose + conf/config.yaml)
scripts/deploy-stack.sh ana-docker litellm

# 2. On the server: create .env from the template and fill secrets
ssh ana-docker 'cd /opt/docker/compose/litellm && cp -n .env.example .env'
#   generate the keys:
#     openssl rand -hex 24 | sed 's/^/sk-/'   # LITELLM_MASTER_KEY
#     openssl rand -hex 32                     # LITELLM_SALT_KEY
#     openssl rand -hex 24                     # POSTGRES_PASSWORD
$EDITOR  # fill .env on the server

# 3. Sanity-parse then launch
ssh ana-docker 'cd /opt/docker/compose/litellm && docker compose config >/dev/null && docker compose up -d && docker compose ps'

.env.example is the only env file in git. The real .env (master key, salt, Postgres password) lives on the server and is gitignored.

Smoke test

# liveness (no auth)
curl -fsS http://10.250.50.70:4000/health/liveliness   # -> "I'm alive!"

# a chat round-trip (uses the master key), then look for it in the Logs UI
curl -s http://10.250.50.70:4000/v1/chat/completions \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"phi4-mini","messages":[{"role":"user","content":"say hi"}]}'

# embeddings
curl -s http://10.250.50.70:4000/v1/embeddings \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3-embedding","input":"hello"}'

Then open http://10.250.50.70:4000/ui (log in with the master key) → Logs tab → the calls appear with full request + response.

Notes

  • Both boxes are Anaheim (10.250.0.0/16) so the ana-docker → ana-ml2 hop is LAN-local; negligible added latency.
  • VLLM_API_KEY is blank by default because the vllm stack ships API_KEY= empty. Set it here only if you set it there.
  • LITELLM_SALT_KEY must be set once and never changed — rotating it makes any keys stored in Postgres undecryptable.
  • Empty tools: [] strippingconf/strip_empty_tools.py is a pre-call hook (registered via litellm_settings.callbacks) that drops an empty/None tools field (and any orphaned tool_choice) before forwarding. vLLM 400s on tools: [] ("tools must not be an empty array"); drop_params doesn't catch empty values, only unsupported params. It runs on every request, so all vLLM-backed models are covered, and only fires when tools is present-and-empty (real tools pass through untouched). The file mounts at /app/strip_empty_tools.py beside config.yaml because LiteLLM resolves callbacks relative to the config dir. Note: real tool-calls additionally need the upstream vLLM server launched with --enable-auto-tool-choice — a vLLM-side flag, separate from this gateway.