Files
esh-pfi-infrastructure/stacks/lobe-chat/README.md
T
vh 303fb7a5aa feat(lobe-chat): add the sec seats to the picker, drop two retired ones
Two independent gates kept the new `sec` family out of Lobe, and only one of
them was visible from the symptom.

The picker never auto-discovers. `OPENAI_MODEL_LIST=-all,+<names>` clears
Lobe's built-in OpenAI catalogue and re-adds one model per `+name`, so anything
added to LiteLLM stays invisible until this list is edited and the container
bounced. That pin is deliberate — an unpinned picker offers models that fail on
click — but it means the list rots in both directions, and it had:

- `sec` / `sec-reasoning` missing (hosted_vllm/mog-sec-27b{,-thinking} on
  ana-ml2:8019, added to config.yaml earlier today), and
- `char-rp-reasoning` / `char-rp-fable` still listed after being retired
  upstream, i.e. two picker entries that 400 on click. Verified: a call to
  char-rp-fable now returns 400 Bad Request.

The list is now curated to live, chat-capable, free-local seats — eleven, each
round-tripped through the container after the bounce. The paid family stays out
deliberately; that is now a picker decision rather than a key one.

Which is the other half of this commit: the `lobe-chat-esh` key is no longer
scoped to free local models. On the operator's instruction infra-ops swapped its
explicit array for the `all-proxy-models` access group, so it now reaches the
paid passthroughs with `max_budget: None`. The README documented the old posture
as current, which made it a security claim that was no longer true; it now
carries the change, what it costs, and the fact that the picker is the only
remaining gate.
2026-08-21 08:04:10 -07:00

247 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# lobe-chat — chat frontend over the LiteLLM gateway (esh-docker-vm)
Evaluation replacement for the hand-rolled `gateway-chat` single-file HTML
surface, which the operator does not want to keep improving — it has already
produced two defects (a 1024 `max_tokens` default that read as model degeneracy,
and a `NaN`→`null` `max_tokens` bug).
- **Host:** esh-docker-vm (10.0.50.45) · **Port:** 3210 · **URL:** http://10.0.50.45:3210
- **Backend:** LiteLLM gateway at `10.250.50.70:4000/v1` (reachable from ESH, ~30 ms)
## Why Lobe over Open WebUI
Weight, measured from the registries rather than from marketing:
| | compressed | layers |
|---|---|---|
| Lobe Chat | **143 MB** | 1 |
| Open WebUI | 1,825 MB | 19 |
12.8×. Open WebUI was declined by the operator in June 2026 on weight grounds and
that objection still holds. (Its recorded *secondary* objection — the empty-`tools`
400 against vLLM — is now moot: the `strip_empty_tools` callback covers the normal
API path and only failed to protect LiteLLM's own built-in playground.)
## ANSWERED — Lobe TTS is a SPLIT: endpoint env-driven, voice/model/format UI-only
Resolved 2026-08-16 against the running image (`.next` server bundle + route probe),
not docs:
- **There is a server-side route** `(backend)/webapi/tts/openai/route.js` — TTS goes
browser → Lobe server → the **OpenAI provider**, whose server-side base URL is
`OPENAI_PROXY_URL`. So the TTS *endpoint* inherits the gateway. (Probing the route
bare returns 401 — it exists and wants the browser's provider payload — not 404.)
- **No `process.env.*TTS*` / `*AUDIO*` / `*SPEECH*` vars exist at all.** Voice, model,
and the enable-toggle live in a **client-side settings store**
(bundle key `TTS_SETTING_KEY = 'tts'`), configured in Settings, per browser. There is
no `AUDIO_TTS_*` equivalent to Open WebUI's.
**Verdict against the operator's "manageable OR scriptable" criterion:** the
load-bearing part (endpoint → gateway) IS env-scriptable and already wired; voice and
model are a **one-time UI setup per browser**, not a maintenance surface. That is
"manageable", which was acceptable.
> ### ⚠ CORRECTION 2026-08-17 (tts-dev) — this section's two footnotes were wrong,
> and one of them meant TTS did not work at all
>
> Both claims below were reasoned from the wrong hop. Re-measured end-to-end from
> esh-docker-vm with this stack's own `.env` key:
>
> 1. **"An unknown `model` routes to the gateway default" — false on this path.**
> True direct-to-`:8198`, which accepts-and-ignores `model`. But requests go
> through **LiteLLM**, which resolves the model name *first*. Lobe's default is
> `tts-1`, and `POST $OPENAI_PROXY_URL/audio/speech {"model":"tts-1",...}` returned
> **403 `key not allowed to access model`** (the scoped key's allow-list held
> `ext-tts` alone; on an unscoped key it is a 400 `Invalid model name`). **So TTS
> was dead on arrival from deploy until 2026-08-17** — the endpoint being wired is
> necessary, not sufficient. **Fixed the same day**: infra-ops aliased `tts-1`,
> `tts-1-hd` and `gpt-4o-mini-tts` to ext-tts's upstream and extended the
> `lobe-chat-esh` key allow-list 20 → 23 models, so the stock payload now works
> with **zero client-side settings** (verified from this host with this `.env`:
> 200, `audio/mpeg`, real MP3). Setting the UI model field to `ext-tts` remains
> valid and harmless.
> 2. **"`response_format: mp3` … set it in the UI" — not possible.** Lobe's OpenAI
> TTS client sends exactly three fields (server bundle `chunks/29685.js`):
> `{ input, model: options?.model || 'tts-1', voice }`. `response_format` is never
> sent and has no UI setting, so **format cannot be chosen from this side at all.**
> You get the fleet gateway's default, which today is WAV (~123 KB for three words;
> 23.5 MB for a 245 s turn), and LiteLLM labels it `content-type: audio/mpeg`
> regardless of the bytes. Chrome sniffs and plays it. Fixing the size is
> tts-dev's side of the fence, not a setting here.
>
> **Nothing to do in the browser any more.** The alias + allow-list pair above
> removed the manual step entirely; a fresh browser, a phone, or any stock OpenAI
> client works out of the box.
>
> ⚠️ **Coupling worth knowing:** `tts-1`, `tts-1-hd` and `gpt-4o-mini-tts` are
> independent LiteLLM DB rows pointing at ext-tts's upstream, not references to
> `ext-tts`. If that upstream ever repoints, all four move together or stock
> clients quietly land on a dead engine.
Trade recorded for the record: Open WebUI has dedicated `AUDIO_TTS_*` env vars (fully
scriptable) at 1,825 MB; Lobe is endpoint-env + UI-cosmetic at 143 MB. Operator chose
to proceed with Lobe.
## TTS — verified with tts-dev, 2026-08-16
**Route: the `ext-tts` LiteLLM alias. Do NOT go direct to the seat.** The gateway
exists so engines can be auditioned and swapped behind it — dots was swapped twice
in the week before this deploy and consumers noticed nothing. Direct-to-seat means
eating every engine change.
**Auth:** the gateway virtual key, nothing more. The seat itself has no auth
(LAN/WireGuard-internal).
**Shape:** plain OpenAI `POST /v1/audio/speech` with
`{model, input, voice, response_format}`. No deviations; a stock OpenAI client works
drop-in. mp3/opus/aac/flac transcode; wav/pcm pass through byte-verbatim. `speed`
honoured 0.25–4.0.
⚠️ **`model` is the one field that is NOT free-form on this route.** That tolerance
belongs to `:8198`, which accepts-and-ignores `model`; **LiteLLM in front of it does
not** — it resolves the name first, so `tts-1` never reaches the gateway (403 on the
scoped key, 400 on an unscoped one). Set Lobe's TTS model to **`ext-tts`**. See the
2026-08-17 correction above.
### ⚠️ CORRECTION — my earlier voice foot-gun warning was wrong
This file previously warned that any non-fleet OpenAI voice name 404s and could trip
the router cooldown, and told you to pin the voice. **That was stale and mostly
unfounded.** tts-dev verified live: **the OpenAI names are ALIASED, not rejected** —
designed for exactly this case, a stock client dropping in without knowing fleet
voice names.
```
alloy, echo, onyx, ash -> donut nova -> miranda
shimmer, coral -> emmie fable, sage -> glados
```
Authoritative list (13): `computer`, `computer-soft`, `computer-urgent`, `donut`,
`emmie`, `glados`, `miranda`, `sindra`, `sindra-excited`, `sindra-sad`,
`sindra-soft`, `sindra-sultry`, `sindra-whisper`.
**Voice surface is now fully safe (tts-dev, 2026-08-16, tts-stack `c55bc3c`).**
`ballad`→emmie and `verse`→donut were the last two unaliased OpenAI names; they are
now aliased and live. Verified: the entire modern OpenAI voice set returns 200, and
only a genuinely-unknown string (`wharrgarbl`) 404s. **A stock Lobe voice picker
cannot produce a 404 from any voice it would plausibly offer, so it cannot trip the
router cooldown.** The env-vs-UI voice question is therefore moot for safety — the
voice string does not need pinning. (Format is a different matter; see below — and
it is not pinnable from here at all.)
### The real constraint is SIZE, not length
The 71.2s per-call cap in older notes is **dead** — that was Zonos-era and required
client-side chunking. dots chunks server-side. tts-dev threw a 592-word single call
at it: HTTP 200, 66.4 s wall, **245 s of audio**, byte-identical ending under ASR
diff. Long assistant turns are handled, not clipped.
But: **245 s of WAV is 23.5 MB**, and WAV is what this deploy gets — Lobe never sends
`response_format`, so the fleet gateway's default (native lane → engine → wav) is what
lands. **That is not fixable from this stack**: no env var, no UI field, not in the
payload. It is a fleet-gateway default, and tts-dev owns it. Until it changes, expect
a few MB per spoken turn on the LAN — wasteful, not fatal.
And 66 s of wall time is a long wait with no streaming in the OpenAI-compat path —
dots *does* have a genuine streaming mode (`stream: true`, ~0.40 s to first audio,
flat with length) but that is **not** the OpenAI-compat shape and would be a custom
integration. Lobe sends no `speed` either, so the gateway's atempo path is unreachable
from here.
### ⚠️ Capacity — tell tts-dev if this ramps
**The seat serializes generation — one render at a time, by design, no continuous
batching.** A second concurrent consumer is a real capacity question, not a
theoretical one, and dots shares a GPU where tts-dev had two resource incidents that
week. Report sustained volume, and any voice string not on the list above, to
`tts-dev` on althing rather than letting him infer it from a VRAM graph.
## Credential posture
> ### ⚠️ CHANGED 2026-08-21 — this key is NO LONGER free-local-only
>
> On the operator's explicit instruction, infra-ops replaced the key's explicit
> model array with the **`all-proxy-models` access group**. `lobe-chat-esh` now
> reaches **everything behind the gateway, paid passthroughs included** —
> `glm-*`, `kimi-k3`, `gen-frontier*`. The scoping described below is history,
> kept because the reasoning still explains what the guard rail was for.
>
> **What this means in practice:** a paid call through this key spends real
> vendor credits, and the key has **`max_budget: None`** — no ceiling. The UI
> in front of it is LAN-exposed with an access code and nothing else. A spend
> cap is available on request from infra-ops; the operator has not asked for
> one, so it is uncapped by instruction, not by oversight.
>
> Two side-effects of the access group worth knowing: it **auto-includes future
> models**, so no more per-seat key edits, and it only ever resolves **live**
> models — which is how five retired names (`char-rp-fable`,
> `char-rp-reasoning`, `lfm2.5-2.6b`, `qwen3-reranker`,
> `reranker-a4-gte-modernbert`) dropped off it silently.
>
> **The picker is now the only gate.** `OPENAI_MODEL_LIST` decides what a user
> can select; the key no longer restricts anything. As of 2026-08-21 that list
> deliberately excludes the paid family — so exposing paid models to this UI is
> one env edit away, and should stay a conscious decision rather than a drift.
The original posture, for the record:
Deliberately **not** the shared all-agents key — that reaches the paid passthroughs
(GLM, Kimi), and a LAN-exposed chat UI holding it would let anyone who can reach the
port spend vendor credits from a pool shared across every project.
This stack used a purpose-minted LiteLLM virtual key, `key_alias: lobe-chat-esh`,
scoped to the 20 free **local** models. Scoping was verified at mint time, both
directions:
- `gen` → answers
- `glm-5.2`, `kimi-k3`, `gen-frontier` → `key not allowed to access model`
Secrets are in the vault, never in git. `.env` on the host is `0600`:
```bash
secret get esh-docker-vm/lobe-chat-litellm-key # -> OPENAI_API_KEY
secret get esh-docker-vm/lobe-chat-access-code # -> ACCESS_CODE (UI gate)
secret get esh-docker-vm/lobe-chat-key-vaults-secret # -> KEY_VAULTS_SECRET
```
`ACCESS_CODE` matters: this is a home-lab LAN segment with nothing in front of it.
## Verified on deploy (2026-08-16)
- container healthy; `http://10.0.50.45:3210/` → 307 → `/chat` → 200
- from **inside** the container: `GET /v1/models` returns the fleet seats, and a
`gen` chat round-trip returns `"ok"` — so the app's own network path and key work,
not merely the host's
- image on disk 617 MB (143 MB compressed)
## Deploy
```bash
scripts/deploy-stack.sh esh-docker-vm lobe-chat --compose
# on host: populate .env from the vault (see above), chmod 600, then
ssh lkraven@10.0.50.45 'cd /opt/docker/compose/lobe-chat && docker compose up -d'
```
`lkraven` owns `/opt/docker` and is in the `docker` group on this host, so no sudo
is needed. Note ESH is outside the infra-ops NOPASSWD grant.
## System Agent — why `gpt-5-mini` was being called
Lobe has a **System Agent**: a background model, separate from your chat model,
that auto-names conversations, summarizes history, rewrites RAG queries, and
generates assistant metadata. **Its default is `openai/gpt-5-mini`.** Because our
OpenAI provider is the gateway, that model name is forwarded verbatim to LiteLLM,
and the scoped key (local models only) 403s it — filling the gateway log with
"Tried to access gpt-5-mini" and silently breaking auto-naming.
Fixed by env (`SYSTEM_AGENT`), so this is scriptable, not a UI setting — a point
IN Lobe's favour on the manageable-by-agent axis, unlike the output-token cap:
```
SYSTEM_AGENT=topic=openai/summarizer,translation=openai/gen,agentMeta=openai/gen,queryRewrite=openai/summarizer,historyCompress=openai/summarizer,thread=openai/summarizer
```
All six documented keys are set explicitly — any omitted key falls back to the
`gpt-5-mini` default. `summarizer`/`classifier` are the same seat as `gen` at
temp 0, the right fit for naming/summarizing; `gen` where output quality matters.