# litellm OpenAI-compatible **gateway** in front of the vLLM services on ana-ml2, standing in the request path so every request + response is **logged and inspectable in a browser**. This is the thing vLLM does not give us: Dozzle shows vLLM's stdout (connection/request metadata) but not the full prompt/completion bodies. LiteLLM captures both, per call, with a Logs UI. **Server:** ana-docker (`10.250.50.70`) **Port:** `4000` (proxy API + admin/Logs UI at `/ui`) — configurable in `.env` **Backs:** the `vllm` stack on ana-ml2 (`10.250.50.54`) ## Why it exists phi4-mini is becoming a production summarizer + "dreaming" agent. Being able to read exactly what it was asked and what it answered is the difference between debuggable and opaque. See `docs/roadmap.md` → "Observability for the vLLM stack". This is the **lean first cut** of that roadmap item — see *Langfuse-ready* below for the upgrade path. ## What routes through it Consumers point their OpenAI `base_url` at `http://10.250.50.70:4000` and pick a model **by name**; the gateway forwards to the right vLLM port and logs the round-trip. | model name (here) | upstream | vLLM port | logged | |---|---|---|---| | `phi4-mini` | generative chat | `:8004` | **full prompt + completion** | | `qwen3-embedding` | `/v1/embeddings` | `:8001` | input + vector metadata | | `qwen3-reranker` | `/rerank` | `:8002` | query + docs + scores | **Not routed:** the `vllm-reward` Skywork classifier (`:8003`) is a pooling `/classify` endpoint with no first-class LiteLLM route — callers hit it directly for now. The generative model is the high-value target for req/resp visibility and it routes cleanly here. (If reward logging is wanted later, LiteLLM `pass_through_endpoints` can cover it.) ## The log switch Full prompt/response text shows in the Logs UI because of `store_prompts_in_spend_logs: true` in `conf/config.yaml`. Without it you'd get metadata only (tokens, latency, model name) — not the text. The Postgres sidecar (`litellm-db`) is the store. ## Langfuse-ready This deliberately does **not** stand up Langfuse's heavy v3 stack (ClickHouse + Redis + MinIO + Postgres + app containers). To graduate to full Langfuse traces later: 1. Stand up (or point at) a Langfuse instance. 2. Set `LANGFUSE_PUBLIC_KEY` / `LANGFUSE_SECRET_KEY` / `LANGFUSE_HOST` in `.env`. 3. Uncomment `success_callback` / `failure_callback` in `conf/config.yaml`. 4. `docker compose up -d` to restart. No re-architecture: the gateway and every consumer stay pointed here. ## ⚠ `reasoning_effort` is not a universal vocabulary `gen-reasoning` accepts **only** `xhigh` (its default), `medium` and `low`, and returns HTTP 400 on anything else: Unexpected reasoning effort high. Supported types are xhigh (default), medium, and low. That is the *default* value of several clients, so the seat presents as broken rather than as one enum value out of step. `conf/reasoning_effort_map.py` is a pre-call hook that maps `high` and `max` onto `xhigh` for that model group only. Measured 2026-09-02 across every local seat before scoping it: | model | `reasoning_effort: high` | |---|---| | `gen-reasoning` | **rejected** → mapped | | `gen`, `sec`, `char-rp-reasoning`, `summarizer` | accepted → untouched | Paid passthroughs (`gen-frontier*`, `glm*`, `kimi*`) were deliberately **not** probed — they spend vendor credits — and are not mapped. **Add a model to `EFFORT_MAP` only after measuring that it actually rejects the value.** ⚠ **A hook file needs a compose change, not just a conf push.** Callbacks are bind-mounted per-file beside `config.yaml`, so a new hook requires a new volume line and `docker compose up -d litellm` (a `restart` will not pick it up — the volume only attaches at container creation). Target the service by name; a bare `up -d` bounces the DB too. ## Deploy ```bash # 1. Sync canonical → ana-docker (compose + conf/config.yaml) scripts/deploy-stack.sh ana-docker litellm # 2. On the server: create .env from the template and fill secrets ssh ana-docker 'cd /opt/docker/compose/litellm && cp -n .env.example .env' # generate the keys: # openssl rand -hex 24 | sed 's/^/sk-/' # LITELLM_MASTER_KEY # openssl rand -hex 32 # LITELLM_SALT_KEY # openssl rand -hex 24 # POSTGRES_PASSWORD $EDITOR # fill .env on the server # 3. Sanity-parse then launch ssh ana-docker 'cd /opt/docker/compose/litellm && docker compose config >/dev/null && docker compose up -d && docker compose ps' ``` > `.env.example` is the only env file in git. The real `.env` (master key, > salt, Postgres password) lives on the server and is gitignored. ## Smoke test ```bash # liveness (no auth) curl -fsS http://10.250.50.70:4000/health/liveliness # -> "I'm alive!" # a chat round-trip (uses the master key), then look for it in the Logs UI curl -s http://10.250.50.70:4000/v1/chat/completions \ -H "Authorization: Bearer $LITELLM_MASTER_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"phi4-mini","messages":[{"role":"user","content":"say hi"}]}' # embeddings curl -s http://10.250.50.70:4000/v1/embeddings \ -H "Authorization: Bearer $LITELLM_MASTER_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"qwen3-embedding","input":"hello"}' ``` Then open `http://10.250.50.70:4000/ui` (log in with the master key) → **Logs** tab → the calls appear with full request + response. ## Notes - Both boxes are Anaheim (`10.250.0.0/16`) so the ana-docker → ana-ml2 hop is LAN-local; negligible added latency. - `VLLM_API_KEY` is blank by default because the `vllm` stack ships `API_KEY=` empty. Set it here only if you set it there. - `LITELLM_SALT_KEY` must be set **once** and never changed — rotating it makes any keys stored in Postgres undecryptable. - **Empty `tools: []` stripping** — `conf/strip_empty_tools.py` is a pre-call hook (registered via `litellm_settings.callbacks`) that drops an empty/None `tools` field (and any orphaned `tool_choice`) before forwarding. vLLM 400s on `tools: []` ("tools must not be an empty array"); `drop_params` doesn't catch empty *values*, only unsupported params. It runs on **every** request, so all vLLM-backed models are covered, and only fires when `tools` is present-and-empty (real tools pass through untouched). The file mounts at `/app/strip_empty_tools.py` beside `config.yaml` because LiteLLM resolves callbacks relative to the config dir. Note: real tool-calls additionally need the upstream vLLM server launched with `--enable-auto-tool-choice` — a vLLM-side flag, separate from this gateway.