Files
esh-pfi-infrastructure/stacks/litellm/README.md
T
vh d1bea13994 fix(litellm): strip empty tools:[] before forwarding to vLLM
vLLM's OpenAI server 400s on an empty tools array ("tools must not be an
empty array"), which broke every gateway call carrying tools:[] (clients
that send it to mean "no tools" -- OpenAI tolerates it, vLLM does not).
drop_params doesn't help: it drops unsupported PARAMS, not empty VALUES.

Add a CustomLogger async_pre_call_hook (conf/strip_empty_tools.py) that
pops an empty/None tools field (+ orphaned tool_choice) before forwarding,
registered globally via litellm_settings.callbacks so it covers every
vLLM-backed model, not just mistral-small-4. Mounted at
/app/strip_empty_tools.py beside config.yaml (LiteLLM resolves callbacks
relative to the config dir). Surgical: only fires when tools is present
and empty; real tools pass through untouched.

Verified on live gateway (1.87.0): mistral-small-4 and granite-4.1-8b
with tools:[] now 200 (were 400); no-tools baseline unchanged; a real
tool still passes through.
2026-06-16 00:43:57 -07:00

120 lines
5.2 KiB
Markdown

# litellm
OpenAI-compatible **gateway** in front of the vLLM services on ana-ml2,
standing in the request path so every request + response is **logged and
inspectable in a browser**. This is the thing vLLM does not give us:
Dozzle shows vLLM's stdout (connection/request metadata) but not the full
prompt/completion bodies. LiteLLM captures both, per call, with a Logs UI.
**Server:** ana-docker (`10.250.50.70`)
**Port:** `4000` (proxy API + admin/Logs UI at `/ui`) — configurable in `.env`
**Backs:** the `vllm` stack on ana-ml2 (`10.250.50.54`)
## Why it exists
phi4-mini is becoming a production summarizer + "dreaming" agent. Being
able to read exactly what it was asked and what it answered is the
difference between debuggable and opaque. See `docs/roadmap.md` →
"Observability for the vLLM stack". This is the **lean first cut** of that
roadmap item — see *Langfuse-ready* below for the upgrade path.
## What routes through it
Consumers point their OpenAI `base_url` at `http://10.250.50.70:4000` and
pick a model **by name**; the gateway forwards to the right vLLM port and
logs the round-trip.
| model name (here) | upstream | vLLM port | logged |
|---|---|---|---|
| `phi4-mini` | generative chat | `:8004` | **full prompt + completion** |
| `qwen3-embedding` | `/v1/embeddings` | `:8001` | input + vector metadata |
| `qwen3-reranker` | `/rerank` | `:8002` | query + docs + scores |
**Not routed:** the `vllm-reward` Skywork classifier (`:8003`) is a pooling
`/classify` endpoint with no first-class LiteLLM route — callers hit it
directly for now. The generative model is the high-value target for
req/resp visibility and it routes cleanly here. (If reward logging is
wanted later, LiteLLM `pass_through_endpoints` can cover it.)
## The log switch
Full prompt/response text shows in the Logs UI because of
`store_prompts_in_spend_logs: true` in `conf/config.yaml`. Without it you'd
get metadata only (tokens, latency, model name) — not the text. The
Postgres sidecar (`litellm-db`) is the store.
## Langfuse-ready
This deliberately does **not** stand up Langfuse's heavy v3 stack
(ClickHouse + Redis + MinIO + Postgres + app containers). To graduate to
full Langfuse traces later:
1. Stand up (or point at) a Langfuse instance.
2. Set `LANGFUSE_PUBLIC_KEY` / `LANGFUSE_SECRET_KEY` / `LANGFUSE_HOST` in `.env`.
3. Uncomment `success_callback` / `failure_callback` in `conf/config.yaml`.
4. `docker compose up -d` to restart.
No re-architecture: the gateway and every consumer stay pointed here.
## Deploy
```bash
# 1. Sync canonical → ana-docker (compose + conf/config.yaml)
scripts/deploy-stack.sh ana-docker litellm
# 2. On the server: create .env from the template and fill secrets
ssh ana-docker 'cd /opt/docker/compose/litellm && cp -n .env.example .env'
# generate the keys:
# openssl rand -hex 24 | sed 's/^/sk-/' # LITELLM_MASTER_KEY
# openssl rand -hex 32 # LITELLM_SALT_KEY
# openssl rand -hex 24 # POSTGRES_PASSWORD
$EDITOR # fill .env on the server
# 3. Sanity-parse then launch
ssh ana-docker 'cd /opt/docker/compose/litellm && docker compose config >/dev/null && docker compose up -d && docker compose ps'
```
> `.env.example` is the only env file in git. The real `.env` (master key,
> salt, Postgres password) lives on the server and is gitignored.
## Smoke test
```bash
# liveness (no auth)
curl -fsS http://10.250.50.70:4000/health/liveliness # -> "I'm alive!"
# a chat round-trip (uses the master key), then look for it in the Logs UI
curl -s http://10.250.50.70:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"phi4-mini","messages":[{"role":"user","content":"say hi"}]}'
# embeddings
curl -s http://10.250.50.70:4000/v1/embeddings \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"qwen3-embedding","input":"hello"}'
```
Then open `http://10.250.50.70:4000/ui` (log in with the master key) →
**Logs** tab → the calls appear with full request + response.
## Notes
- Both boxes are Anaheim (`10.250.0.0/16`) so the ana-docker → ana-ml2 hop
is LAN-local; negligible added latency.
- `VLLM_API_KEY` is blank by default because the `vllm` stack ships
`API_KEY=` empty. Set it here only if you set it there.
- `LITELLM_SALT_KEY` must be set **once** and never changed — rotating it
makes any keys stored in Postgres undecryptable.
- **Empty `tools: []` stripping** — `conf/strip_empty_tools.py` is a pre-call
hook (registered via `litellm_settings.callbacks`) that drops an empty/None
`tools` field (and any orphaned `tool_choice`) before forwarding. vLLM 400s on
`tools: []` ("tools must not be an empty array"); `drop_params` doesn't catch
empty *values*, only unsupported params. It runs on **every** request, so all
vLLM-backed models are covered, and only fires when `tools` is present-and-empty
(real tools pass through untouched). The file mounts at `/app/strip_empty_tools.py`
beside `config.yaml` because LiteLLM resolves callbacks relative to the config
dir. Note: real tool-calls additionally need the upstream vLLM server launched
with `--enable-auto-tool-choice` — a vLLM-side flag, separate from this gateway.