Files
esh-pfi-infrastructure/stacks/litellm/README.md
T
vh 83b2ec1a8a feat(litellm): add vLLM request/response logging gateway on ana-docker
LiteLLM proxy fronting the vLLM services on ana-ml2 so every request +
response is captured and inspectable in a browser Logs UI — the
visibility vLLM itself lacks (Dozzle shows only connection metadata).

- compose: litellm (proxy + /ui Logs) + litellm-db (Postgres store)
- conf/config.yaml: routes phi4-mini (chat, :8004), qwen3-embedding
  (:8001), qwen3-reranker (:8002); store_prompts_in_spend_logs persists
  full prompt/completion text. reward classifier (:8003) stays direct
  (no first-class LiteLLM route).
- Langfuse-ready: lean first cut intentionally skips Langfuse's heavy v3
  stack; graduating is one env-var + callback step, no re-architecture.
- roadmap: mark the vLLM-observability item's first cut as shipped.

Lean first cut of docs/roadmap.md "Observability for the vLLM stack".
2026-06-04 01:17:50 -07:00

110 lines
4.4 KiB
Markdown

# litellm
OpenAI-compatible **gateway** in front of the vLLM services on ana-ml2,
standing in the request path so every request + response is **logged and
inspectable in a browser**. This is the thing vLLM does not give us:
Dozzle shows vLLM's stdout (connection/request metadata) but not the full
prompt/completion bodies. LiteLLM captures both, per call, with a Logs UI.
**Server:** ana-docker (`10.250.50.70`)
**Port:** `4000` (proxy API + admin/Logs UI at `/ui`) — configurable in `.env`
**Backs:** the `vllm` stack on ana-ml2 (`10.250.50.54`)
## Why it exists
phi4-mini is becoming a production summarizer + "dreaming" agent. Being
able to read exactly what it was asked and what it answered is the
difference between debuggable and opaque. See `docs/roadmap.md` →
"Observability for the vLLM stack". This is the **lean first cut** of that
roadmap item — see *Langfuse-ready* below for the upgrade path.
## What routes through it
Consumers point their OpenAI `base_url` at `http://10.250.50.70:4000` and
pick a model **by name**; the gateway forwards to the right vLLM port and
logs the round-trip.
| model name (here) | upstream | vLLM port | logged |
|---|---|---|---|
| `phi4-mini` | generative chat | `:8004` | **full prompt + completion** |
| `qwen3-embedding` | `/v1/embeddings` | `:8001` | input + vector metadata |
| `qwen3-reranker` | `/rerank` | `:8002` | query + docs + scores |
**Not routed:** the `vllm-reward` Skywork classifier (`:8003`) is a pooling
`/classify` endpoint with no first-class LiteLLM route — callers hit it
directly for now. The generative model is the high-value target for
req/resp visibility and it routes cleanly here. (If reward logging is
wanted later, LiteLLM `pass_through_endpoints` can cover it.)
## The log switch
Full prompt/response text shows in the Logs UI because of
`store_prompts_in_spend_logs: true` in `conf/config.yaml`. Without it you'd
get metadata only (tokens, latency, model name) — not the text. The
Postgres sidecar (`litellm-db`) is the store.
## Langfuse-ready
This deliberately does **not** stand up Langfuse's heavy v3 stack
(ClickHouse + Redis + MinIO + Postgres + app containers). To graduate to
full Langfuse traces later:
1. Stand up (or point at) a Langfuse instance.
2. Set `LANGFUSE_PUBLIC_KEY` / `LANGFUSE_SECRET_KEY` / `LANGFUSE_HOST` in `.env`.
3. Uncomment `success_callback` / `failure_callback` in `conf/config.yaml`.
4. `docker compose up -d` to restart.
No re-architecture: the gateway and every consumer stay pointed here.
## Deploy
```bash
# 1. Sync canonical → ana-docker (compose + conf/config.yaml)
scripts/deploy-stack.sh ana-docker litellm
# 2. On the server: create .env from the template and fill secrets
ssh ana-docker 'cd /opt/docker/compose/litellm && cp -n .env.example .env'
# generate the keys:
# openssl rand -hex 24 | sed 's/^/sk-/' # LITELLM_MASTER_KEY
# openssl rand -hex 32 # LITELLM_SALT_KEY
# openssl rand -hex 24 # POSTGRES_PASSWORD
$EDITOR # fill .env on the server
# 3. Sanity-parse then launch
ssh ana-docker 'cd /opt/docker/compose/litellm && docker compose config >/dev/null && docker compose up -d && docker compose ps'
```
> `.env.example` is the only env file in git. The real `.env` (master key,
> salt, Postgres password) lives on the server and is gitignored.
## Smoke test
```bash
# liveness (no auth)
curl -fsS http://10.250.50.70:4000/health/liveliness # -> "I'm alive!"
# a chat round-trip (uses the master key), then look for it in the Logs UI
curl -s http://10.250.50.70:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"phi4-mini","messages":[{"role":"user","content":"say hi"}]}'
# embeddings
curl -s http://10.250.50.70:4000/v1/embeddings \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"qwen3-embedding","input":"hello"}'
```
Then open `http://10.250.50.70:4000/ui` (log in with the master key) →
**Logs** tab → the calls appear with full request + response.
## Notes
- Both boxes are Anaheim (`10.250.0.0/16`) so the ana-docker → ana-ml2 hop
is LAN-local; negligible added latency.
- `VLLM_API_KEY` is blank by default because the `vllm` stack ships
`API_KEY=` empty. Set it here only if you set it there.
- `LITELLM_SALT_KEY` must be set **once** and never changed — rotating it
makes any keys stored in Postgres undecryptable.