Every agent on nh3-dev — Claude, Codex, Grok, Aider — needs the same answers: what runs here, how do I call it, what will bite me. Until now that lived in ~/.claude/CLAUDE.md, which only Claude sessions load, and it was interleaved with operator preferences that other families have no use for. Two-tier by design, matching the persistent-memory split: FLEETTOOLS.md is a 135-line index an agent reads whole, and each entry links to a detail file it opens only when it actually needs that tool. Reading the index costs about a fifth of reading the tree. Detail paths are absolute so they resolve from any working directory, since a non-Claude agent will cat the path rather than follow a markdown link. ~/FLEETTOOLS.md symlinks to the index for discovery. Rule zero is that live inventories get queried, not transcribed: Homepage /api/services, asset-engine /api/v1/services, LiteLLM /v1/models, and every FastAPI seat's /openapi.json. A copied service table would be stale within a month and this repo already has a standing rule against second copies that drift. Contents verified against the running fleet rather than copied from existing docs: binaries resolved on PATH, seven endpoints probed live, the LiteLLM roster counted at 40 models where the old note said ~30. No credentials are included; the vault and its CLI are pointed at instead.
2.6 KiB
LiteLLM gateway — the fleet's OpenAI-compatible front door
Base URL: http://10.250.50.70:4000/v1 (ana-docker) · Logs UI: :4000/ui
Fronts the ana-ml2 / fv-ml1 / irv-ml1 vLLM seats plus paid vendor passthroughs.
⚠ Never hardcode the model roster — query it
curl -s -H "Authorization: Bearer $KEY" http://10.250.50.70:4000/v1/models \
| python3 -c "import sys,json;[print(m['id']) for m in json.load(sys.stdin)['data']]"
The roster drifts constantly (40 entries as of 2026-09-18; a doc that names them goes stale within weeks). Any list written down here would be a second copy that rots.
Keys
The gateway 401s without a virtual API key. Three tiers:
| key | where | reach |
|---|---|---|
shared all-agents (all-agents-local) |
~/.claude/CLAUDE.md § Global tools |
every model, paid passthroughs included |
| infra-ops admin | ~/.config/litellm/infra-ops-key (0600) |
admin operations |
| project-scoped | request from infra-ops via althing | isolated, auditable spend |
⚠ The shared key is NOT free-local-only. It reaches paid GLM/z.ai and Kimi passthroughs, and a paid call on it spends real vendor credits shared across every session and project holding it. There is no isolation. For anything where "shows up on someone else's tab" matters, get a project-scoped key.
Stable aliases worth knowing
summarizer / summarizer-large / classifier — the always-available local seat for
summarization, naming/labeling, classification and triage. Essentially free compute;
reach for it before spending premium API tokens on work a small local model handles.
Embeddings and rerankers (qwen3-embedding, qwen3-reranker, reranker) are local
and free too. Anything named *-frontier, glm-* or kimi-* is PAID.
Traps
- LiteLLM mutates shared deployment config in-place. Per-request sampler params
bleed into the SHARED
litellm_paramsfor that deployment. Fix: give each variant a distinct served-model-name rather than reusing one with different params. - Retiring a model orphans scoped keys. A stale allowlist produces a silent
per-endpoint 403. Audit
/key/list+/key/infoon every repoint. - The
:4000/uiplayground cannot test vLLM backends — it sendstools: [], which vLLM rejects with a 400. Use a chat frontend (Open WebUI, Gateway Chat) instead. - "Random model calls" in the logs are usually AUTH FAILURES — empty
api_keyor a missingBearerprefix, showing as zero-spend 401s. - Kimi models 400 on any
temperatureother than 1 — omit the field entirely rather than sending 0. They answer incontentwith chain-of-thought inreasoning_content.