Files
vh 9171e6a20f feat(langfuse): stand up Langfuse v3 + wire the LiteLLM trace callback
LLM observability for the fleet — pretty trace UI over the gateway: prompts,
completions, reasoning, latency, token counts. The pretty layer LiteLLM's
spend_logs lacked.

- stacks/langfuse: v3 self-host stack (web/worker/postgres/clickhouse/redis/
  minio) on ana-docker, adapted from upstream. UI on :3001 (gitea owns :3000).
  Project + API keys auto-provisioned via LANGFUSE_INIT_*. HOSTNAME=0.0.0.0 on
  langfuse-web so it's reachable via the published port while also on tnet.
- litellm: enabled success_callback/failure_callback: ["langfuse"] (the
  passthrough env was already wired); keys + host go in the litellm .env.

Verified: stack healthy, project keys authenticate, and a real gateway call
landed a litellm-acompletion trace in Langfuse within ~6s. Secrets live only in
the server .env (never committed).
2026-06-05 11:35:01 -07:00
..

litellm

OpenAI-compatible gateway in front of the vLLM services on ana-ml2, standing in the request path so every request + response is logged and inspectable in a browser. This is the thing vLLM does not give us: Dozzle shows vLLM's stdout (connection/request metadata) but not the full prompt/completion bodies. LiteLLM captures both, per call, with a Logs UI.

Server: ana-docker (10.250.50.70) Port: 4000 (proxy API + admin/Logs UI at /ui) — configurable in .env Backs: the vllm stack on ana-ml2 (10.250.50.54)

Why it exists

phi4-mini is becoming a production summarizer + "dreaming" agent. Being able to read exactly what it was asked and what it answered is the difference between debuggable and opaque. See docs/roadmap.md → "Observability for the vLLM stack". This is the lean first cut of that roadmap item — see Langfuse-ready below for the upgrade path.

What routes through it

Consumers point their OpenAI base_url at http://10.250.50.70:4000 and pick a model by name; the gateway forwards to the right vLLM port and logs the round-trip.

model name (here) upstream vLLM port logged
phi4-mini generative chat :8004 full prompt + completion
qwen3-embedding /v1/embeddings :8001 input + vector metadata
qwen3-reranker /rerank :8002 query + docs + scores

Not routed: the vllm-reward Skywork classifier (:8003) is a pooling /classify endpoint with no first-class LiteLLM route — callers hit it directly for now. The generative model is the high-value target for req/resp visibility and it routes cleanly here. (If reward logging is wanted later, LiteLLM pass_through_endpoints can cover it.)

The log switch

Full prompt/response text shows in the Logs UI because of store_prompts_in_spend_logs: true in conf/config.yaml. Without it you'd get metadata only (tokens, latency, model name) — not the text. The Postgres sidecar (litellm-db) is the store.

Langfuse-ready

This deliberately does not stand up Langfuse's heavy v3 stack (ClickHouse + Redis + MinIO + Postgres + app containers). To graduate to full Langfuse traces later:

  1. Stand up (or point at) a Langfuse instance.
  2. Set LANGFUSE_PUBLIC_KEY / LANGFUSE_SECRET_KEY / LANGFUSE_HOST in .env.
  3. Uncomment success_callback / failure_callback in conf/config.yaml.
  4. docker compose up -d to restart.

No re-architecture: the gateway and every consumer stay pointed here.

Deploy

# 1. Sync canonical → ana-docker (compose + conf/config.yaml)
scripts/deploy-stack.sh ana-docker litellm

# 2. On the server: create .env from the template and fill secrets
ssh ana-docker 'cd /opt/docker/compose/litellm && cp -n .env.example .env'
#   generate the keys:
#     openssl rand -hex 24 | sed 's/^/sk-/'   # LITELLM_MASTER_KEY
#     openssl rand -hex 32                     # LITELLM_SALT_KEY
#     openssl rand -hex 24                     # POSTGRES_PASSWORD
$EDITOR  # fill .env on the server

# 3. Sanity-parse then launch
ssh ana-docker 'cd /opt/docker/compose/litellm && docker compose config >/dev/null && docker compose up -d && docker compose ps'

.env.example is the only env file in git. The real .env (master key, salt, Postgres password) lives on the server and is gitignored.

Smoke test

# liveness (no auth)
curl -fsS http://10.250.50.70:4000/health/liveliness   # -> "I'm alive!"

# a chat round-trip (uses the master key), then look for it in the Logs UI
curl -s http://10.250.50.70:4000/v1/chat/completions \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"phi4-mini","messages":[{"role":"user","content":"say hi"}]}'

# embeddings
curl -s http://10.250.50.70:4000/v1/embeddings \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3-embedding","input":"hello"}'

Then open http://10.250.50.70:4000/ui (log in with the master key) → Logs tab → the calls appear with full request + response.

Notes

  • Both boxes are Anaheim (10.250.0.0/16) so the ana-docker → ana-ml2 hop is LAN-local; negligible added latency.
  • VLLM_API_KEY is blank by default because the vllm stack ships API_KEY= empty. Set it here only if you set it there.
  • LITELLM_SALT_KEY must be set once and never changed — rotating it makes any keys stored in Postgres undecryptable.