The box is physically at Fountain Valley, renamed, renumbered onto 10.251/16, and serving inference again. This lands the repo half of that. Host: hostname ana-ml2 -> fv-ml1, pinned to 10.251.50.54 by a dnsmasq reservation so the address the runbook, DNS and LiteLLM all assume is the address it actually has. Its headscale node is renamed too. The sweep ran from scripts/fv-ml1-rename-sweep.sh, whose allowlist is the reason this diff touches current-state files and not the record. Dated persistent-memory entries, archival-memory and incident notes still say ana-ml2 in 31 and 62 places respectively, because that is what the box was when those things happened. Rewriting them would make the history lie. LiteLLM was the load-bearing piece and needed more than the api_base sed the runbook describes. Twenty api_base entries repointed, but a grep-and-verify pass also caught a LIVE pass_through_endpoints target for the scalar-judge reward route still on the old address -- an api_base-only substitution would have left it dead. Four prose references describing current state were repointed as well; one historical note recording where a hand-test was run is deliberately left pointing at 10.250.50.54. Two facts in the server tables were wrong and are corrected here. The site is Fountain Valley, not Anaheim. And the box has FOUR RTX PRO 6000 Blackwell Max-Q, not two -- verified by nvidia-smi -L and independently by PCI enumeration of four GB202GL devices. That is 391 GB of VRAM rather than 196, which changes what fits on it. DNS: fv-ml1, fv-ml1-bmc and fv-gw added under the fv site via the piggyback approach, scriberr re-homed, and the ana-ml2 records removed. Applied to all three resolvers. The BMC record carries a warning that its 802.1q VLAN tag must stay disabled -- it shipped tagging VLAN 250 into an untagged port, which made it invisible to every network-side diagnostic and is the reason it appeared dead through several cable changes. Verified end to end: summarizer and sec both answer through the Anaheim gateway across the mesh to FV seats on different ports.
litellm
OpenAI-compatible gateway in front of the vLLM services on ana-ml2, standing in the request path so every request + response is logged and inspectable in a browser. This is the thing vLLM does not give us: Dozzle shows vLLM's stdout (connection/request metadata) but not the full prompt/completion bodies. LiteLLM captures both, per call, with a Logs UI.
Server: ana-docker (10.250.50.70)
Port: 4000 (proxy API + admin/Logs UI at /ui) — configurable in .env
Backs: the vllm stack on ana-ml2 (10.250.50.54)
Why it exists
phi4-mini is becoming a production summarizer + "dreaming" agent. Being
able to read exactly what it was asked and what it answered is the
difference between debuggable and opaque. See docs/roadmap.md →
"Observability for the vLLM stack". This is the lean first cut of that
roadmap item — see Langfuse-ready below for the upgrade path.
What routes through it
Consumers point their OpenAI base_url at http://10.250.50.70:4000 and
pick a model by name; the gateway forwards to the right vLLM port and
logs the round-trip.
| model name (here) | upstream | vLLM port | logged |
|---|---|---|---|
phi4-mini |
generative chat | :8004 |
full prompt + completion |
qwen3-embedding |
/v1/embeddings |
:8001 |
input + vector metadata |
qwen3-reranker |
/rerank |
:8002 |
query + docs + scores |
Not routed: the vllm-reward Skywork classifier (:8003) is a pooling
/classify endpoint with no first-class LiteLLM route — callers hit it
directly for now. The generative model is the high-value target for
req/resp visibility and it routes cleanly here. (If reward logging is
wanted later, LiteLLM pass_through_endpoints can cover it.)
The log switch
Full prompt/response text shows in the Logs UI because of
store_prompts_in_spend_logs: true in conf/config.yaml. Without it you'd
get metadata only (tokens, latency, model name) — not the text. The
Postgres sidecar (litellm-db) is the store.
Langfuse-ready
This deliberately does not stand up Langfuse's heavy v3 stack (ClickHouse + Redis + MinIO + Postgres + app containers). To graduate to full Langfuse traces later:
- Stand up (or point at) a Langfuse instance.
- Set
LANGFUSE_PUBLIC_KEY/LANGFUSE_SECRET_KEY/LANGFUSE_HOSTin.env. - Uncomment
success_callback/failure_callbackinconf/config.yaml. docker compose up -dto restart.
No re-architecture: the gateway and every consumer stay pointed here.
⚠ reasoning_effort is not a universal vocabulary
gen-reasoning accepts only xhigh (its default), medium and low, and
returns HTTP 400 on anything else:
Unexpected reasoning effort high. Supported types are xhigh (default),
medium, and low.
That is the default value of several clients, so the seat presents as broken
rather than as one enum value out of step. conf/reasoning_effort_map.py is a
pre-call hook that maps high and max onto xhigh for that model group only.
Measured 2026-09-02 across every local seat before scoping it:
| model | reasoning_effort: high |
|---|---|
gen-reasoning |
rejected → mapped |
gen, sec, char-rp-reasoning, summarizer |
accepted → untouched |
Paid passthroughs (gen-frontier*, glm*, kimi*) were deliberately not
probed — they spend vendor credits — and are not mapped. Add a model to
EFFORT_MAP only after measuring that it actually rejects the value.
⚠ A hook file needs a compose change, not just a conf push. Callbacks are
bind-mounted per-file beside config.yaml, so a new hook requires a new volume
line and docker compose up -d litellm (a restart will not pick it up — the
volume only attaches at container creation). Target the service by name; a bare
up -d bounces the DB too.
Deploy
# 1. Sync canonical → ana-docker (compose + conf/config.yaml)
scripts/deploy-stack.sh ana-docker litellm
# 2. On the server: create .env from the template and fill secrets
ssh ana-docker 'cd /opt/docker/compose/litellm && cp -n .env.example .env'
# generate the keys:
# openssl rand -hex 24 | sed 's/^/sk-/' # LITELLM_MASTER_KEY
# openssl rand -hex 32 # LITELLM_SALT_KEY
# openssl rand -hex 24 # POSTGRES_PASSWORD
$EDITOR # fill .env on the server
# 3. Sanity-parse then launch
ssh ana-docker 'cd /opt/docker/compose/litellm && docker compose config >/dev/null && docker compose up -d && docker compose ps'
.env.exampleis the only env file in git. The real.env(master key, salt, Postgres password) lives on the server and is gitignored.
Smoke test
# liveness (no auth)
curl -fsS http://10.250.50.70:4000/health/liveliness # -> "I'm alive!"
# a chat round-trip (uses the master key), then look for it in the Logs UI
curl -s http://10.250.50.70:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"phi4-mini","messages":[{"role":"user","content":"say hi"}]}'
# embeddings
curl -s http://10.250.50.70:4000/v1/embeddings \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"qwen3-embedding","input":"hello"}'
Then open http://10.250.50.70:4000/ui (log in with the master key) →
Logs tab → the calls appear with full request + response.
Notes
- Both boxes are Anaheim (
10.250.0.0/16) so the ana-docker → ana-ml2 hop is LAN-local; negligible added latency. VLLM_API_KEYis blank by default because thevllmstack shipsAPI_KEY=empty. Set it here only if you set it there.LITELLM_SALT_KEYmust be set once and never changed — rotating it makes any keys stored in Postgres undecryptable.- Empty
tools: []stripping —conf/strip_empty_tools.pyis a pre-call hook (registered vialitellm_settings.callbacks) that drops an empty/Nonetoolsfield (and any orphanedtool_choice) before forwarding. vLLM 400s ontools: []("tools must not be an empty array");drop_paramsdoesn't catch empty values, only unsupported params. It runs on every request, so all vLLM-backed models are covered, and only fires whentoolsis present-and-empty (real tools pass through untouched). The file mounts at/app/strip_empty_tools.pybesideconfig.yamlbecause LiteLLM resolves callbacks relative to the config dir. Note: real tool-calls additionally need the upstream vLLM server launched with--enable-auto-tool-choice— a vLLM-side flag, separate from this gateway.