Files
esh-pfi-infrastructure/stacks/litellm
vh 3462b5336c config(litellm): apply canonical Qwen3.8 sampling; fix presence_penalty on the thinking alias
Sourced from upstream rather than tuned by hand. Qwen/Qwen3.8-27B card
'Best Practices' 1 and unsloth/Qwen3.8-27B 1 are byte-identical:

  Thinking: temperature=1.0 top_p=0.95 top_k=20 min_p=0.0
            presence_penalty=0.0 repetition_penalty=1.0
  Instruct: temperature=0.7 top_p=0.80 top_k=20 min_p=0.0
            presence_penalty=1.5 repetition_penalty=1.0

REAL BUG FIXED: gen-reasoning carried presence_penalty=1.5 -- the
INSTRUCT-mode value applied to a THINKING deployment, where canonical is
0.0. Corrected.

gen was already canonical; added the missing explicit min_p and
repetition_penalty so the full set is visible at the call site rather than
relying on backend defaults that happen to agree.

DELIBERATELY NOT canonicalised: summarizer, classifier, image-judge and
qwen-image-bench run temperature=0 (and the judges top_k=1,
repetition_penalty=1.05) because determinism is the point of those seats.
Forcing temperature=0.7 on a classifier to match a chat preset would break
their contract, so canonical is applied only where the alias is actually
doing open-ended generation.

Recorded against presence_penalty=1.5, which upstream itself hedges on
verbatim: 'you can adjust the presence_penalty parameter between 0 and 2
to reduce endless repetition. However, using a higher value may
occasionally result in language mixing and a slight decrease in model
performance.' 1.5 is high in that band and is the operator's suspected
trigger for the multi-turn degradation. Left at canonical so the baseline
is defensible, with the caveat and the 0.0-0.5 fallback documented inline
as the first dial to move if it recurs.
2026-08-16 16:01:22 -07:00
..

litellm

OpenAI-compatible gateway in front of the vLLM services on ana-ml2, standing in the request path so every request + response is logged and inspectable in a browser. This is the thing vLLM does not give us: Dozzle shows vLLM's stdout (connection/request metadata) but not the full prompt/completion bodies. LiteLLM captures both, per call, with a Logs UI.

Server: ana-docker (10.250.50.70) Port: 4000 (proxy API + admin/Logs UI at /ui) — configurable in .env Backs: the vllm stack on ana-ml2 (10.250.50.54)

Why it exists

phi4-mini is becoming a production summarizer + "dreaming" agent. Being able to read exactly what it was asked and what it answered is the difference between debuggable and opaque. See docs/roadmap.md → "Observability for the vLLM stack". This is the lean first cut of that roadmap item — see Langfuse-ready below for the upgrade path.

What routes through it

Consumers point their OpenAI base_url at http://10.250.50.70:4000 and pick a model by name; the gateway forwards to the right vLLM port and logs the round-trip.

model name (here) upstream vLLM port logged
phi4-mini generative chat :8004 full prompt + completion
qwen3-embedding /v1/embeddings :8001 input + vector metadata
qwen3-reranker /rerank :8002 query + docs + scores

Not routed: the vllm-reward Skywork classifier (:8003) is a pooling /classify endpoint with no first-class LiteLLM route — callers hit it directly for now. The generative model is the high-value target for req/resp visibility and it routes cleanly here. (If reward logging is wanted later, LiteLLM pass_through_endpoints can cover it.)

The log switch

Full prompt/response text shows in the Logs UI because of store_prompts_in_spend_logs: true in conf/config.yaml. Without it you'd get metadata only (tokens, latency, model name) — not the text. The Postgres sidecar (litellm-db) is the store.

Langfuse-ready

This deliberately does not stand up Langfuse's heavy v3 stack (ClickHouse + Redis + MinIO + Postgres + app containers). To graduate to full Langfuse traces later:

  1. Stand up (or point at) a Langfuse instance.
  2. Set LANGFUSE_PUBLIC_KEY / LANGFUSE_SECRET_KEY / LANGFUSE_HOST in .env.
  3. Uncomment success_callback / failure_callback in conf/config.yaml.
  4. docker compose up -d to restart.

No re-architecture: the gateway and every consumer stay pointed here.

Deploy

# 1. Sync canonical → ana-docker (compose + conf/config.yaml)
scripts/deploy-stack.sh ana-docker litellm

# 2. On the server: create .env from the template and fill secrets
ssh ana-docker 'cd /opt/docker/compose/litellm && cp -n .env.example .env'
#   generate the keys:
#     openssl rand -hex 24 | sed 's/^/sk-/'   # LITELLM_MASTER_KEY
#     openssl rand -hex 32                     # LITELLM_SALT_KEY
#     openssl rand -hex 24                     # POSTGRES_PASSWORD
$EDITOR  # fill .env on the server

# 3. Sanity-parse then launch
ssh ana-docker 'cd /opt/docker/compose/litellm && docker compose config >/dev/null && docker compose up -d && docker compose ps'

.env.example is the only env file in git. The real .env (master key, salt, Postgres password) lives on the server and is gitignored.

Smoke test

# liveness (no auth)
curl -fsS http://10.250.50.70:4000/health/liveliness   # -> "I'm alive!"

# a chat round-trip (uses the master key), then look for it in the Logs UI
curl -s http://10.250.50.70:4000/v1/chat/completions \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"phi4-mini","messages":[{"role":"user","content":"say hi"}]}'

# embeddings
curl -s http://10.250.50.70:4000/v1/embeddings \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3-embedding","input":"hello"}'

Then open http://10.250.50.70:4000/ui (log in with the master key) → Logs tab → the calls appear with full request + response.

Notes

  • Both boxes are Anaheim (10.250.0.0/16) so the ana-docker → ana-ml2 hop is LAN-local; negligible added latency.
  • VLLM_API_KEY is blank by default because the vllm stack ships API_KEY= empty. Set it here only if you set it there.
  • LITELLM_SALT_KEY must be set once and never changed — rotating it makes any keys stored in Postgres undecryptable.
  • Empty tools: [] strippingconf/strip_empty_tools.py is a pre-call hook (registered via litellm_settings.callbacks) that drops an empty/None tools field (and any orphaned tool_choice) before forwarding. vLLM 400s on tools: [] ("tools must not be an empty array"); drop_params doesn't catch empty values, only unsupported params. It runs on every request, so all vLLM-backed models are covered, and only fires when tools is present-and-empty (real tools pass through untouched). The file mounts at /app/strip_empty_tools.py beside config.yaml because LiteLLM resolves callbacks relative to the config dir. Note: real tool-calls additionally need the upstream vLLM server launched with --enable-auto-tool-choice — a vLLM-side flag, separate from this gateway.