LLM observability for the fleet — pretty trace UI over the gateway: prompts, completions, reasoning, latency, token counts. The pretty layer LiteLLM's spend_logs lacked. - stacks/langfuse: v3 self-host stack (web/worker/postgres/clickhouse/redis/ minio) on ana-docker, adapted from upstream. UI on :3001 (gitea owns :3000). Project + API keys auto-provisioned via LANGFUSE_INIT_*. HOSTNAME=0.0.0.0 on langfuse-web so it's reachable via the published port while also on tnet. - litellm: enabled success_callback/failure_callback: ["langfuse"] (the passthrough env was already wired); keys + host go in the litellm .env. Verified: stack healthy, project keys authenticate, and a real gateway call landed a litellm-acompletion trace in Langfuse within ~6s. Secrets live only in the server .env (never committed).
litellm
OpenAI-compatible gateway in front of the vLLM services on ana-ml2, standing in the request path so every request + response is logged and inspectable in a browser. This is the thing vLLM does not give us: Dozzle shows vLLM's stdout (connection/request metadata) but not the full prompt/completion bodies. LiteLLM captures both, per call, with a Logs UI.
Server: ana-docker (10.250.50.70)
Port: 4000 (proxy API + admin/Logs UI at /ui) — configurable in .env
Backs: the vllm stack on ana-ml2 (10.250.50.54)
Why it exists
phi4-mini is becoming a production summarizer + "dreaming" agent. Being
able to read exactly what it was asked and what it answered is the
difference between debuggable and opaque. See docs/roadmap.md →
"Observability for the vLLM stack". This is the lean first cut of that
roadmap item — see Langfuse-ready below for the upgrade path.
What routes through it
Consumers point their OpenAI base_url at http://10.250.50.70:4000 and
pick a model by name; the gateway forwards to the right vLLM port and
logs the round-trip.
| model name (here) | upstream | vLLM port | logged |
|---|---|---|---|
phi4-mini |
generative chat | :8004 |
full prompt + completion |
qwen3-embedding |
/v1/embeddings |
:8001 |
input + vector metadata |
qwen3-reranker |
/rerank |
:8002 |
query + docs + scores |
Not routed: the vllm-reward Skywork classifier (:8003) is a pooling
/classify endpoint with no first-class LiteLLM route — callers hit it
directly for now. The generative model is the high-value target for
req/resp visibility and it routes cleanly here. (If reward logging is
wanted later, LiteLLM pass_through_endpoints can cover it.)
The log switch
Full prompt/response text shows in the Logs UI because of
store_prompts_in_spend_logs: true in conf/config.yaml. Without it you'd
get metadata only (tokens, latency, model name) — not the text. The
Postgres sidecar (litellm-db) is the store.
Langfuse-ready
This deliberately does not stand up Langfuse's heavy v3 stack (ClickHouse + Redis + MinIO + Postgres + app containers). To graduate to full Langfuse traces later:
- Stand up (or point at) a Langfuse instance.
- Set
LANGFUSE_PUBLIC_KEY/LANGFUSE_SECRET_KEY/LANGFUSE_HOSTin.env. - Uncomment
success_callback/failure_callbackinconf/config.yaml. docker compose up -dto restart.
No re-architecture: the gateway and every consumer stay pointed here.
Deploy
# 1. Sync canonical → ana-docker (compose + conf/config.yaml)
scripts/deploy-stack.sh ana-docker litellm
# 2. On the server: create .env from the template and fill secrets
ssh ana-docker 'cd /opt/docker/compose/litellm && cp -n .env.example .env'
# generate the keys:
# openssl rand -hex 24 | sed 's/^/sk-/' # LITELLM_MASTER_KEY
# openssl rand -hex 32 # LITELLM_SALT_KEY
# openssl rand -hex 24 # POSTGRES_PASSWORD
$EDITOR # fill .env on the server
# 3. Sanity-parse then launch
ssh ana-docker 'cd /opt/docker/compose/litellm && docker compose config >/dev/null && docker compose up -d && docker compose ps'
.env.exampleis the only env file in git. The real.env(master key, salt, Postgres password) lives on the server and is gitignored.
Smoke test
# liveness (no auth)
curl -fsS http://10.250.50.70:4000/health/liveliness # -> "I'm alive!"
# a chat round-trip (uses the master key), then look for it in the Logs UI
curl -s http://10.250.50.70:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"phi4-mini","messages":[{"role":"user","content":"say hi"}]}'
# embeddings
curl -s http://10.250.50.70:4000/v1/embeddings \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"qwen3-embedding","input":"hello"}'
Then open http://10.250.50.70:4000/ui (log in with the master key) →
Logs tab → the calls appear with full request + response.
Notes
- Both boxes are Anaheim (
10.250.0.0/16) so the ana-docker → ana-ml2 hop is LAN-local; negligible added latency. VLLM_API_KEYis blank by default because thevllmstack shipsAPI_KEY=empty. Set it here only if you set it there.LITELLM_SALT_KEYmust be set once and never changed — rotating it makes any keys stored in Postgres undecryptable.