Two independent gates kept the new `sec` family out of Lobe, and only one of
them was visible from the symptom.
The picker never auto-discovers. `OPENAI_MODEL_LIST=-all,+<names>` clears
Lobe's built-in OpenAI catalogue and re-adds one model per `+name`, so anything
added to LiteLLM stays invisible until this list is edited and the container
bounced. That pin is deliberate — an unpinned picker offers models that fail on
click — but it means the list rots in both directions, and it had:
- `sec` / `sec-reasoning` missing (hosted_vllm/mog-sec-27b{,-thinking} on
ana-ml2:8019, added to config.yaml earlier today), and
- `char-rp-reasoning` / `char-rp-fable` still listed after being retired
upstream, i.e. two picker entries that 400 on click. Verified: a call to
char-rp-fable now returns 400 Bad Request.
The list is now curated to live, chat-capable, free-local seats — eleven, each
round-tripped through the container after the bounce. The paid family stays out
deliberately; that is now a picker decision rather than a key one.
Which is the other half of this commit: the `lobe-chat-esh` key is no longer
scoped to free local models. On the operator's instruction infra-ops swapped its
explicit array for the `all-proxy-models` access group, so it now reaches the
paid passthroughs with `max_budget: None`. The README documented the old posture
as current, which made it a security claim that was no longer true; it now
carries the change, what it costs, and the fact that the picker is the only
remaining gate.
247 lines
13 KiB
Markdown
247 lines
13 KiB
Markdown
# lobe-chat — chat frontend over the LiteLLM gateway (esh-docker-vm)
|
||
|
||
Evaluation replacement for the hand-rolled `gateway-chat` single-file HTML
|
||
surface, which the operator does not want to keep improving — it has already
|
||
produced two defects (a 1024 `max_tokens` default that read as model degeneracy,
|
||
and a `NaN`→`null` `max_tokens` bug).
|
||
|
||
- **Host:** esh-docker-vm (10.0.50.45) · **Port:** 3210 · **URL:** http://10.0.50.45:3210
|
||
- **Backend:** LiteLLM gateway at `10.250.50.70:4000/v1` (reachable from ESH, ~30 ms)
|
||
|
||
## Why Lobe over Open WebUI
|
||
|
||
Weight, measured from the registries rather than from marketing:
|
||
|
||
| | compressed | layers |
|
||
|---|---|---|
|
||
| Lobe Chat | **143 MB** | 1 |
|
||
| Open WebUI | 1,825 MB | 19 |
|
||
|
||
12.8×. Open WebUI was declined by the operator in June 2026 on weight grounds and
|
||
that objection still holds. (Its recorded *secondary* objection — the empty-`tools`
|
||
400 against vLLM — is now moot: the `strip_empty_tools` callback covers the normal
|
||
API path and only failed to protect LiteLLM's own built-in playground.)
|
||
|
||
## ANSWERED — Lobe TTS is a SPLIT: endpoint env-driven, voice/model/format UI-only
|
||
|
||
Resolved 2026-08-16 against the running image (`.next` server bundle + route probe),
|
||
not docs:
|
||
|
||
- **There is a server-side route** `(backend)/webapi/tts/openai/route.js` — TTS goes
|
||
browser → Lobe server → the **OpenAI provider**, whose server-side base URL is
|
||
`OPENAI_PROXY_URL`. So the TTS *endpoint* inherits the gateway. (Probing the route
|
||
bare returns 401 — it exists and wants the browser's provider payload — not 404.)
|
||
- **No `process.env.*TTS*` / `*AUDIO*` / `*SPEECH*` vars exist at all.** Voice, model,
|
||
and the enable-toggle live in a **client-side settings store**
|
||
(bundle key `TTS_SETTING_KEY = 'tts'`), configured in Settings, per browser. There is
|
||
no `AUDIO_TTS_*` equivalent to Open WebUI's.
|
||
|
||
**Verdict against the operator's "manageable OR scriptable" criterion:** the
|
||
load-bearing part (endpoint → gateway) IS env-scriptable and already wired; voice and
|
||
model are a **one-time UI setup per browser**, not a maintenance surface. That is
|
||
"manageable", which was acceptable.
|
||
|
||
> ### ⚠ CORRECTION 2026-08-17 (tts-dev) — this section's two footnotes were wrong,
|
||
> and one of them meant TTS did not work at all
|
||
>
|
||
> Both claims below were reasoned from the wrong hop. Re-measured end-to-end from
|
||
> esh-docker-vm with this stack's own `.env` key:
|
||
>
|
||
> 1. **"An unknown `model` routes to the gateway default" — false on this path.**
|
||
> True direct-to-`:8198`, which accepts-and-ignores `model`. But requests go
|
||
> through **LiteLLM**, which resolves the model name *first*. Lobe's default is
|
||
> `tts-1`, and `POST $OPENAI_PROXY_URL/audio/speech {"model":"tts-1",...}` returned
|
||
> **403 `key not allowed to access model`** (the scoped key's allow-list held
|
||
> `ext-tts` alone; on an unscoped key it is a 400 `Invalid model name`). **So TTS
|
||
> was dead on arrival from deploy until 2026-08-17** — the endpoint being wired is
|
||
> necessary, not sufficient. **Fixed the same day**: infra-ops aliased `tts-1`,
|
||
> `tts-1-hd` and `gpt-4o-mini-tts` to ext-tts's upstream and extended the
|
||
> `lobe-chat-esh` key allow-list 20 → 23 models, so the stock payload now works
|
||
> with **zero client-side settings** (verified from this host with this `.env`:
|
||
> 200, `audio/mpeg`, real MP3). Setting the UI model field to `ext-tts` remains
|
||
> valid and harmless.
|
||
> 2. **"`response_format: mp3` … set it in the UI" — not possible.** Lobe's OpenAI
|
||
> TTS client sends exactly three fields (server bundle `chunks/29685.js`):
|
||
> `{ input, model: options?.model || 'tts-1', voice }`. `response_format` is never
|
||
> sent and has no UI setting, so **format cannot be chosen from this side at all.**
|
||
> You get the fleet gateway's default, which today is WAV (~123 KB for three words;
|
||
> 23.5 MB for a 245 s turn), and LiteLLM labels it `content-type: audio/mpeg`
|
||
> regardless of the bytes. Chrome sniffs and plays it. Fixing the size is
|
||
> tts-dev's side of the fence, not a setting here.
|
||
>
|
||
> **Nothing to do in the browser any more.** The alias + allow-list pair above
|
||
> removed the manual step entirely; a fresh browser, a phone, or any stock OpenAI
|
||
> client works out of the box.
|
||
>
|
||
> ⚠️ **Coupling worth knowing:** `tts-1`, `tts-1-hd` and `gpt-4o-mini-tts` are
|
||
> independent LiteLLM DB rows pointing at ext-tts's upstream, not references to
|
||
> `ext-tts`. If that upstream ever repoints, all four move together or stock
|
||
> clients quietly land on a dead engine.
|
||
|
||
Trade recorded for the record: Open WebUI has dedicated `AUDIO_TTS_*` env vars (fully
|
||
scriptable) at 1,825 MB; Lobe is endpoint-env + UI-cosmetic at 143 MB. Operator chose
|
||
to proceed with Lobe.
|
||
|
||
## TTS — verified with tts-dev, 2026-08-16
|
||
|
||
**Route: the `ext-tts` LiteLLM alias. Do NOT go direct to the seat.** The gateway
|
||
exists so engines can be auditioned and swapped behind it — dots was swapped twice
|
||
in the week before this deploy and consumers noticed nothing. Direct-to-seat means
|
||
eating every engine change.
|
||
|
||
**Auth:** the gateway virtual key, nothing more. The seat itself has no auth
|
||
(LAN/WireGuard-internal).
|
||
|
||
**Shape:** plain OpenAI `POST /v1/audio/speech` with
|
||
`{model, input, voice, response_format}`. No deviations; a stock OpenAI client works
|
||
drop-in. mp3/opus/aac/flac transcode; wav/pcm pass through byte-verbatim. `speed`
|
||
honoured 0.25–4.0.
|
||
|
||
⚠️ **`model` is the one field that is NOT free-form on this route.** That tolerance
|
||
belongs to `:8198`, which accepts-and-ignores `model`; **LiteLLM in front of it does
|
||
not** — it resolves the name first, so `tts-1` never reaches the gateway (403 on the
|
||
scoped key, 400 on an unscoped one). Set Lobe's TTS model to **`ext-tts`**. See the
|
||
2026-08-17 correction above.
|
||
|
||
### ⚠️ CORRECTION — my earlier voice foot-gun warning was wrong
|
||
|
||
This file previously warned that any non-fleet OpenAI voice name 404s and could trip
|
||
the router cooldown, and told you to pin the voice. **That was stale and mostly
|
||
unfounded.** tts-dev verified live: **the OpenAI names are ALIASED, not rejected** —
|
||
designed for exactly this case, a stock client dropping in without knowing fleet
|
||
voice names.
|
||
|
||
```
|
||
alloy, echo, onyx, ash -> donut nova -> miranda
|
||
shimmer, coral -> emmie fable, sage -> glados
|
||
```
|
||
|
||
Authoritative list (13): `computer`, `computer-soft`, `computer-urgent`, `donut`,
|
||
`emmie`, `glados`, `miranda`, `sindra`, `sindra-excited`, `sindra-sad`,
|
||
`sindra-soft`, `sindra-sultry`, `sindra-whisper`.
|
||
|
||
**Voice surface is now fully safe (tts-dev, 2026-08-16, tts-stack `c55bc3c`).**
|
||
`ballad`→emmie and `verse`→donut were the last two unaliased OpenAI names; they are
|
||
now aliased and live. Verified: the entire modern OpenAI voice set returns 200, and
|
||
only a genuinely-unknown string (`wharrgarbl`) 404s. **A stock Lobe voice picker
|
||
cannot produce a 404 from any voice it would plausibly offer, so it cannot trip the
|
||
router cooldown.** The env-vs-UI voice question is therefore moot for safety — the
|
||
voice string does not need pinning. (Format is a different matter; see below — and
|
||
it is not pinnable from here at all.)
|
||
|
||
### The real constraint is SIZE, not length
|
||
|
||
The 71.2s per-call cap in older notes is **dead** — that was Zonos-era and required
|
||
client-side chunking. dots chunks server-side. tts-dev threw a 592-word single call
|
||
at it: HTTP 200, 66.4 s wall, **245 s of audio**, byte-identical ending under ASR
|
||
diff. Long assistant turns are handled, not clipped.
|
||
|
||
But: **245 s of WAV is 23.5 MB**, and WAV is what this deploy gets — Lobe never sends
|
||
`response_format`, so the fleet gateway's default (native lane → engine → wav) is what
|
||
lands. **That is not fixable from this stack**: no env var, no UI field, not in the
|
||
payload. It is a fleet-gateway default, and tts-dev owns it. Until it changes, expect
|
||
a few MB per spoken turn on the LAN — wasteful, not fatal.
|
||
|
||
And 66 s of wall time is a long wait with no streaming in the OpenAI-compat path —
|
||
dots *does* have a genuine streaming mode (`stream: true`, ~0.40 s to first audio,
|
||
flat with length) but that is **not** the OpenAI-compat shape and would be a custom
|
||
integration. Lobe sends no `speed` either, so the gateway's atempo path is unreachable
|
||
from here.
|
||
|
||
### ⚠️ Capacity — tell tts-dev if this ramps
|
||
|
||
**The seat serializes generation — one render at a time, by design, no continuous
|
||
batching.** A second concurrent consumer is a real capacity question, not a
|
||
theoretical one, and dots shares a GPU where tts-dev had two resource incidents that
|
||
week. Report sustained volume, and any voice string not on the list above, to
|
||
`tts-dev` on althing rather than letting him infer it from a VRAM graph.
|
||
|
||
## Credential posture
|
||
|
||
> ### ⚠️ CHANGED 2026-08-21 — this key is NO LONGER free-local-only
|
||
>
|
||
> On the operator's explicit instruction, infra-ops replaced the key's explicit
|
||
> model array with the **`all-proxy-models` access group**. `lobe-chat-esh` now
|
||
> reaches **everything behind the gateway, paid passthroughs included** —
|
||
> `glm-*`, `kimi-k3`, `gen-frontier*`. The scoping described below is history,
|
||
> kept because the reasoning still explains what the guard rail was for.
|
||
>
|
||
> **What this means in practice:** a paid call through this key spends real
|
||
> vendor credits, and the key has **`max_budget: None`** — no ceiling. The UI
|
||
> in front of it is LAN-exposed with an access code and nothing else. A spend
|
||
> cap is available on request from infra-ops; the operator has not asked for
|
||
> one, so it is uncapped by instruction, not by oversight.
|
||
>
|
||
> Two side-effects of the access group worth knowing: it **auto-includes future
|
||
> models**, so no more per-seat key edits, and it only ever resolves **live**
|
||
> models — which is how five retired names (`char-rp-fable`,
|
||
> `char-rp-reasoning`, `lfm2.5-2.6b`, `qwen3-reranker`,
|
||
> `reranker-a4-gte-modernbert`) dropped off it silently.
|
||
>
|
||
> **The picker is now the only gate.** `OPENAI_MODEL_LIST` decides what a user
|
||
> can select; the key no longer restricts anything. As of 2026-08-21 that list
|
||
> deliberately excludes the paid family — so exposing paid models to this UI is
|
||
> one env edit away, and should stay a conscious decision rather than a drift.
|
||
|
||
The original posture, for the record:
|
||
|
||
Deliberately **not** the shared all-agents key — that reaches the paid passthroughs
|
||
(GLM, Kimi), and a LAN-exposed chat UI holding it would let anyone who can reach the
|
||
port spend vendor credits from a pool shared across every project.
|
||
|
||
This stack used a purpose-minted LiteLLM virtual key, `key_alias: lobe-chat-esh`,
|
||
scoped to the 20 free **local** models. Scoping was verified at mint time, both
|
||
directions:
|
||
|
||
- `gen` → answers
|
||
- `glm-5.2`, `kimi-k3`, `gen-frontier` → `key not allowed to access model`
|
||
|
||
Secrets are in the vault, never in git. `.env` on the host is `0600`:
|
||
|
||
```bash
|
||
secret get esh-docker-vm/lobe-chat-litellm-key # -> OPENAI_API_KEY
|
||
secret get esh-docker-vm/lobe-chat-access-code # -> ACCESS_CODE (UI gate)
|
||
secret get esh-docker-vm/lobe-chat-key-vaults-secret # -> KEY_VAULTS_SECRET
|
||
```
|
||
|
||
`ACCESS_CODE` matters: this is a home-lab LAN segment with nothing in front of it.
|
||
|
||
## Verified on deploy (2026-08-16)
|
||
|
||
- container healthy; `http://10.0.50.45:3210/` → 307 → `/chat` → 200
|
||
- from **inside** the container: `GET /v1/models` returns the fleet seats, and a
|
||
`gen` chat round-trip returns `"ok"` — so the app's own network path and key work,
|
||
not merely the host's
|
||
- image on disk 617 MB (143 MB compressed)
|
||
|
||
## Deploy
|
||
|
||
```bash
|
||
scripts/deploy-stack.sh esh-docker-vm lobe-chat --compose
|
||
# on host: populate .env from the vault (see above), chmod 600, then
|
||
ssh lkraven@10.0.50.45 'cd /opt/docker/compose/lobe-chat && docker compose up -d'
|
||
```
|
||
|
||
`lkraven` owns `/opt/docker` and is in the `docker` group on this host, so no sudo
|
||
is needed. Note ESH is outside the infra-ops NOPASSWD grant.
|
||
|
||
## System Agent — why `gpt-5-mini` was being called
|
||
|
||
Lobe has a **System Agent**: a background model, separate from your chat model,
|
||
that auto-names conversations, summarizes history, rewrites RAG queries, and
|
||
generates assistant metadata. **Its default is `openai/gpt-5-mini`.** Because our
|
||
OpenAI provider is the gateway, that model name is forwarded verbatim to LiteLLM,
|
||
and the scoped key (local models only) 403s it — filling the gateway log with
|
||
"Tried to access gpt-5-mini" and silently breaking auto-naming.
|
||
|
||
Fixed by env (`SYSTEM_AGENT`), so this is scriptable, not a UI setting — a point
|
||
IN Lobe's favour on the manageable-by-agent axis, unlike the output-token cap:
|
||
|
||
```
|
||
SYSTEM_AGENT=topic=openai/summarizer,translation=openai/gen,agentMeta=openai/gen,queryRewrite=openai/summarizer,historyCompress=openai/summarizer,thread=openai/summarizer
|
||
```
|
||
|
||
All six documented keys are set explicitly — any omitted key falls back to the
|
||
`gpt-5-mini` default. `summarizer`/`classifier` are the same seat as `gen` at
|
||
temp 0, the right fit for naming/summarizing; `gen` where output quality matters.
|