LiteLLM proxy fronting the vLLM services on ana-ml2 so every request + response is captured and inspectable in a browser Logs UI — the visibility vLLM itself lacks (Dozzle shows only connection metadata). - compose: litellm (proxy + /ui Logs) + litellm-db (Postgres store) - conf/config.yaml: routes phi4-mini (chat, :8004), qwen3-embedding (:8001), qwen3-reranker (:8002); store_prompts_in_spend_logs persists full prompt/completion text. reward classifier (:8003) stays direct (no first-class LiteLLM route). - Langfuse-ready: lean first cut intentionally skips Langfuse's heavy v3 stack; graduating is one env-var + callback step, no re-architecture. - roadmap: mark the vLLM-observability item's first cut as shipped. Lean first cut of docs/roadmap.md "Observability for the vLLM stack".
110 lines
4.4 KiB
Markdown
110 lines
4.4 KiB
Markdown
# litellm
|
|
|
|
OpenAI-compatible **gateway** in front of the vLLM services on ana-ml2,
|
|
standing in the request path so every request + response is **logged and
|
|
inspectable in a browser**. This is the thing vLLM does not give us:
|
|
Dozzle shows vLLM's stdout (connection/request metadata) but not the full
|
|
prompt/completion bodies. LiteLLM captures both, per call, with a Logs UI.
|
|
|
|
**Server:** ana-docker (`10.250.50.70`)
|
|
**Port:** `4000` (proxy API + admin/Logs UI at `/ui`) — configurable in `.env`
|
|
**Backs:** the `vllm` stack on ana-ml2 (`10.250.50.54`)
|
|
|
|
## Why it exists
|
|
|
|
phi4-mini is becoming a production summarizer + "dreaming" agent. Being
|
|
able to read exactly what it was asked and what it answered is the
|
|
difference between debuggable and opaque. See `docs/roadmap.md` →
|
|
"Observability for the vLLM stack". This is the **lean first cut** of that
|
|
roadmap item — see *Langfuse-ready* below for the upgrade path.
|
|
|
|
## What routes through it
|
|
|
|
Consumers point their OpenAI `base_url` at `http://10.250.50.70:4000` and
|
|
pick a model **by name**; the gateway forwards to the right vLLM port and
|
|
logs the round-trip.
|
|
|
|
| model name (here) | upstream | vLLM port | logged |
|
|
|---|---|---|---|
|
|
| `phi4-mini` | generative chat | `:8004` | **full prompt + completion** |
|
|
| `qwen3-embedding` | `/v1/embeddings` | `:8001` | input + vector metadata |
|
|
| `qwen3-reranker` | `/rerank` | `:8002` | query + docs + scores |
|
|
|
|
**Not routed:** the `vllm-reward` Skywork classifier (`:8003`) is a pooling
|
|
`/classify` endpoint with no first-class LiteLLM route — callers hit it
|
|
directly for now. The generative model is the high-value target for
|
|
req/resp visibility and it routes cleanly here. (If reward logging is
|
|
wanted later, LiteLLM `pass_through_endpoints` can cover it.)
|
|
|
|
## The log switch
|
|
|
|
Full prompt/response text shows in the Logs UI because of
|
|
`store_prompts_in_spend_logs: true` in `conf/config.yaml`. Without it you'd
|
|
get metadata only (tokens, latency, model name) — not the text. The
|
|
Postgres sidecar (`litellm-db`) is the store.
|
|
|
|
## Langfuse-ready
|
|
|
|
This deliberately does **not** stand up Langfuse's heavy v3 stack
|
|
(ClickHouse + Redis + MinIO + Postgres + app containers). To graduate to
|
|
full Langfuse traces later:
|
|
|
|
1. Stand up (or point at) a Langfuse instance.
|
|
2. Set `LANGFUSE_PUBLIC_KEY` / `LANGFUSE_SECRET_KEY` / `LANGFUSE_HOST` in `.env`.
|
|
3. Uncomment `success_callback` / `failure_callback` in `conf/config.yaml`.
|
|
4. `docker compose up -d` to restart.
|
|
|
|
No re-architecture: the gateway and every consumer stay pointed here.
|
|
|
|
## Deploy
|
|
|
|
```bash
|
|
# 1. Sync canonical → ana-docker (compose + conf/config.yaml)
|
|
scripts/deploy-stack.sh ana-docker litellm
|
|
|
|
# 2. On the server: create .env from the template and fill secrets
|
|
ssh ana-docker 'cd /opt/docker/compose/litellm && cp -n .env.example .env'
|
|
# generate the keys:
|
|
# openssl rand -hex 24 | sed 's/^/sk-/' # LITELLM_MASTER_KEY
|
|
# openssl rand -hex 32 # LITELLM_SALT_KEY
|
|
# openssl rand -hex 24 # POSTGRES_PASSWORD
|
|
$EDITOR # fill .env on the server
|
|
|
|
# 3. Sanity-parse then launch
|
|
ssh ana-docker 'cd /opt/docker/compose/litellm && docker compose config >/dev/null && docker compose up -d && docker compose ps'
|
|
```
|
|
|
|
> `.env.example` is the only env file in git. The real `.env` (master key,
|
|
> salt, Postgres password) lives on the server and is gitignored.
|
|
|
|
## Smoke test
|
|
|
|
```bash
|
|
# liveness (no auth)
|
|
curl -fsS http://10.250.50.70:4000/health/liveliness # -> "I'm alive!"
|
|
|
|
# a chat round-trip (uses the master key), then look for it in the Logs UI
|
|
curl -s http://10.250.50.70:4000/v1/chat/completions \
|
|
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"model":"phi4-mini","messages":[{"role":"user","content":"say hi"}]}'
|
|
|
|
# embeddings
|
|
curl -s http://10.250.50.70:4000/v1/embeddings \
|
|
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"model":"qwen3-embedding","input":"hello"}'
|
|
```
|
|
|
|
Then open `http://10.250.50.70:4000/ui` (log in with the master key) →
|
|
**Logs** tab → the calls appear with full request + response.
|
|
|
|
## Notes
|
|
|
|
- Both boxes are Anaheim (`10.250.0.0/16`) so the ana-docker → ana-ml2 hop
|
|
is LAN-local; negligible added latency.
|
|
- `VLLM_API_KEY` is blank by default because the `vllm` stack ships
|
|
`API_KEY=` empty. Set it here only if you set it there.
|
|
- `LITELLM_SALT_KEY` must be set **once** and never changed — rotating it
|
|
makes any keys stored in Postgres undecryptable.
|