Operator instruction: take down the existing sec seat (mog-sec) and promote hotdogs (cyberprev) into the sec and sec-reasoning gateway seats. - mog-sec container (vllm-mog-sec, :8019, fv-ml1 GPU0) taken down; ~48 GB freed on GPU0 (cyberprev, already co-resident there, is now the sole GPU0 chat seat). - Gateway sec -> hosted_vllm/cyberprev-27b @ :8025; sec-reasoning -> hosted_vllm/cyberprev-27b-thinking @ :8025. sec/sec-reasoning are ROLE aliases, so this is a promotion, not silent substitution (samplers were already identical between the sec blocks and cyberprev, so only model+api_base changed). - Removed the standalone cyberprev-27b / cyberprev-reasoning gateway aliases added in the prior commit -- now redundant with sec/sec-reasoning, and the fleet convention is a role alias on the gateway with the model's served-name only at the vLLM layer (as mog-sec had). cyberprev's vLLM served-names are unchanged. - Verified e2e through the gateway: sec answers (nmap -sV version detection), sec-reasoning answers with a thinking split (127 reasoning tokens); retired mog-sec-27b now 400s. Note: mog-sec was the fleet's only offense+defense/blue-team seat; the sec role is now offense-only (cyberprev tool-calling). Operator-directed after reviewing the capability comparison. mog-sec stack files retained for a future restore.
litellm
OpenAI-compatible gateway in front of the vLLM services on fv-ml1, standing in the request path so every request + response is logged and inspectable in a browser. This is the thing vLLM does not give us: Dozzle shows vLLM's stdout (connection/request metadata) but not the full prompt/completion bodies. LiteLLM captures both, per call, with a Logs UI.
Server: ana-docker (10.250.50.70)
Port: 4000 (proxy API + admin/Logs UI at /ui) — configurable in .env
Backs: the vllm stack on fv-ml1 (10.251.50.54)
Why it exists
phi4-mini is becoming a production summarizer + "dreaming" agent. Being
able to read exactly what it was asked and what it answered is the
difference between debuggable and opaque. See docs/roadmap.md →
"Observability for the vLLM stack". This is the lean first cut of that
roadmap item — see Langfuse-ready below for the upgrade path.
What routes through it
Consumers point their OpenAI base_url at http://10.250.50.70:4000 and
pick a model by name; the gateway forwards to the right vLLM port and
logs the round-trip.
| model name (here) | upstream | vLLM port | logged |
|---|---|---|---|
phi4-mini |
generative chat | :8004 |
full prompt + completion |
qwen3-embedding |
/v1/embeddings |
:8001 |
input + vector metadata |
qwen3-reranker |
/rerank |
:8002 |
query + docs + scores |
Not routed: the vllm-reward Skywork classifier (:8003) is a pooling
/classify endpoint with no first-class LiteLLM route — callers hit it
directly for now. The generative model is the high-value target for
req/resp visibility and it routes cleanly here. (If reward logging is
wanted later, LiteLLM pass_through_endpoints can cover it.)
The log switch
Full prompt/response text shows in the Logs UI because of
store_prompts_in_spend_logs: true in conf/config.yaml. Without it you'd
get metadata only (tokens, latency, model name) — not the text. The
Postgres sidecar (litellm-db) is the store.
Langfuse-ready
This deliberately does not stand up Langfuse's heavy v3 stack (ClickHouse + Redis + MinIO + Postgres + app containers). To graduate to full Langfuse traces later:
- Stand up (or point at) a Langfuse instance.
- Set
LANGFUSE_PUBLIC_KEY/LANGFUSE_SECRET_KEY/LANGFUSE_HOSTin.env. - Uncomment
success_callback/failure_callbackinconf/config.yaml. docker compose up -dto restart.
No re-architecture: the gateway and every consumer stay pointed here.
⚠ reasoning_effort is not a universal vocabulary
gen-reasoning accepts only xhigh (its default), medium and low, and
returns HTTP 400 on anything else:
Unexpected reasoning effort high. Supported types are xhigh (default),
medium, and low.
That is the default value of several clients, so the seat presents as broken
rather than as one enum value out of step. conf/reasoning_effort_map.py is a
pre-call hook that maps high and max onto xhigh for that model group only.
Measured 2026-09-02 across every local seat before scoping it:
| model | reasoning_effort: high |
|---|---|
gen-reasoning |
rejected → mapped |
gen, sec, char-rp-reasoning, summarizer |
accepted → untouched |
Paid passthroughs (gen-frontier*, glm*, kimi*) were deliberately not
probed — they spend vendor credits — and are not mapped. Add a model to
EFFORT_MAP only after measuring that it actually rejects the value.
⚠ A hook file needs a compose change, not just a conf push. Callbacks are
bind-mounted per-file beside config.yaml, so a new hook requires a new volume
line and docker compose up -d litellm (a restart will not pick it up — the
volume only attaches at container creation). Target the service by name; a bare
up -d bounces the DB too.
Deploy
# 1. Sync canonical → ana-docker (compose + conf/config.yaml)
scripts/deploy-stack.sh ana-docker litellm
# 2. On the server: create .env from the template and fill secrets
ssh ana-docker 'cd /opt/docker/compose/litellm && cp -n .env.example .env'
# generate the keys:
# openssl rand -hex 24 | sed 's/^/sk-/' # LITELLM_MASTER_KEY
# openssl rand -hex 32 # LITELLM_SALT_KEY
# openssl rand -hex 24 # POSTGRES_PASSWORD
$EDITOR # fill .env on the server
# 3. Sanity-parse then launch
ssh ana-docker 'cd /opt/docker/compose/litellm && docker compose config >/dev/null && docker compose up -d && docker compose ps'
.env.exampleis the only env file in git. The real.env(master key, salt, Postgres password) lives on the server and is gitignored.
Smoke test
# liveness (no auth)
curl -fsS http://10.250.50.70:4000/health/liveliness # -> "I'm alive!"
# a chat round-trip (uses the master key), then look for it in the Logs UI
curl -s http://10.250.50.70:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"phi4-mini","messages":[{"role":"user","content":"say hi"}]}'
# embeddings
curl -s http://10.250.50.70:4000/v1/embeddings \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"qwen3-embedding","input":"hello"}'
Then open http://10.250.50.70:4000/ui (log in with the master key) →
Logs tab → the calls appear with full request + response.
Notes
- Both boxes are Anaheim (
10.250.0.0/16) so the ana-docker → fv-ml1 hop is LAN-local; negligible added latency. VLLM_API_KEYis blank by default because thevllmstack shipsAPI_KEY=empty. Set it here only if you set it there.LITELLM_SALT_KEYmust be set once and never changed — rotating it makes any keys stored in Postgres undecryptable.- Empty
tools: []stripping —conf/strip_empty_tools.pyis a pre-call hook (registered vialitellm_settings.callbacks) that drops an empty/Nonetoolsfield (and any orphanedtool_choice) before forwarding. vLLM 400s ontools: []("tools must not be an empty array");drop_paramsdoesn't catch empty values, only unsupported params. It runs on every request, so all vLLM-backed models are covered, and only fires whentoolsis present-and-empty (real tools pass through untouched). The file mounts at/app/strip_empty_tools.pybesideconfig.yamlbecause LiteLLM resolves callbacks relative to the config dir. Note: real tool-calls additionally need the upstream vLLM server launched with--enable-auto-tool-choice— a vLLM-side flag, separate from this gateway.