feat(lobe-chat): stand up Lobe Chat on esh-docker-vm over the LiteLLM gateway
Replacement candidate for the hand-rolled gateway-chat HTML surface, which the operator does not want to keep improving -- it has already produced two defects tonight. Chosen over Open WebUI on weight, measured from the registries rather than recalled: Lobe 143 MB compressed / 1 layer vs Open WebUI 1,825 MB / 19 layers, a 12.8x difference. Open WebUI was declined in June 2026 on weight and that still holds; its secondary recorded objection (empty-tools 400 against vLLM) is now moot since strip_empty_tools covers the normal API path and only missed LiteLLM's built-in playground. CREDENTIAL POSTURE: deliberately NOT the shared all-agents key, which reaches the paid GLM/Kimi passthroughs -- a LAN-exposed chat UI holding it would let anyone reaching the port spend vendor credits from a pool shared across every project. Minted a scoped LiteLLM virtual key (key_alias lobe-chat-esh) limited to the 20 free local models, and verified the scoping BOTH ways: gen answers, glm-5.2 / kimi-k3 / gen-frontier all return 'key not allowed to access model'. Secrets vaulted, host .env 0600. Verified from INSIDE the container, not just from the host: /v1/models returns the fleet seats and a gen round-trip returns 'ok', so the app's own network path and key both work. Container healthy, / -> 307 -> /chat -> 200. Documents the open question this deploy exists to answer: whether Lobe's TTS is ENV-configurable or UI-only. That is the operator's deciding criterion and is NOT yet established -- Open WebUI has dedicated AUDIO_TTS_* vars, Lobe documents a shared OPENAI_PROXY_URL which should carry TTS since LiteLLM serves audio/speech on the same base, but that is inference. Also records the ext-tts voice foot-gun: unknown voices 404 and can trip the router cooldown, so the voice must be pinned rather than left at a UI default.
This commit is contained in:
@@ -0,0 +1,84 @@
|
||||
# lobe-chat — chat frontend over the LiteLLM gateway (esh-docker-vm)
|
||||
|
||||
Evaluation replacement for the hand-rolled `gateway-chat` single-file HTML
|
||||
surface, which the operator does not want to keep improving — it has already
|
||||
produced two defects (a 1024 `max_tokens` default that read as model degeneracy,
|
||||
and a `NaN`→`null` `max_tokens` bug).
|
||||
|
||||
- **Host:** esh-docker-vm (10.0.50.45) · **Port:** 3210 · **URL:** http://10.0.50.45:3210
|
||||
- **Backend:** LiteLLM gateway at `10.250.50.70:4000/v1` (reachable from ESH, ~30 ms)
|
||||
|
||||
## Why Lobe over Open WebUI
|
||||
|
||||
Weight, measured from the registries rather than from marketing:
|
||||
|
||||
| | compressed | layers |
|
||||
|---|---|---|
|
||||
| Lobe Chat | **143 MB** | 1 |
|
||||
| Open WebUI | 1,825 MB | 19 |
|
||||
|
||||
12.8×. Open WebUI was declined by the operator in June 2026 on weight grounds and
|
||||
that objection still holds. (Its recorded *secondary* objection — the empty-`tools`
|
||||
400 against vLLM — is now moot: the `strip_empty_tools` callback covers the normal
|
||||
API path and only failed to protect LiteLLM's own built-in playground.)
|
||||
|
||||
## ⚠️ The open question this deploy exists to answer
|
||||
|
||||
**Is Lobe's TTS configurable by ENV, or only through the settings UI?** That is the
|
||||
operator's deciding criterion — manageable/scriptable by an agent — and it is
|
||||
unresolved. Open WebUI has dedicated `AUDIO_TTS_ENGINE` / `AUDIO_TTS_OPENAI_API_BASE_URL`
|
||||
/ `AUDIO_TTS_MODEL` / `AUDIO_TTS_VOICE`. Lobe documents a **shared** `OPENAI_PROXY_URL`,
|
||||
which *should* carry TTS because LiteLLM serves `/v1/chat/completions` and
|
||||
`/v1/audio/speech` on the same base — but that is inference, not verification.
|
||||
|
||||
If it turns out UI-only, the real choice is: 1.8 GB with genuine scriptability, or
|
||||
143 MB with click-ops. That is an operator call, not an agent one.
|
||||
|
||||
## ⚠️ ext-tts voice foot-gun
|
||||
|
||||
`ext-tts` accepts `donut` / `emmie` / `glados` / `miranda` (+ emotion variants) and
|
||||
the OpenAI aliases `nova` / `alloy`. **Any other OpenAI voice name (`echo`, `fable`,
|
||||
`onyx`, `shimmer`) 404s and can trip the LiteLLM router cooldown.** Pin the voice
|
||||
explicitly rather than accepting whatever the UI defaults to.
|
||||
|
||||
## Credential posture
|
||||
|
||||
Deliberately **not** the shared all-agents key — that reaches the paid passthroughs
|
||||
(GLM, Kimi), and a LAN-exposed chat UI holding it would let anyone who can reach the
|
||||
port spend vendor credits from a pool shared across every project.
|
||||
|
||||
This stack uses a purpose-minted LiteLLM virtual key, `key_alias: lobe-chat-esh`,
|
||||
scoped to the 20 free **local** models. Scoping was verified at mint time, both
|
||||
directions:
|
||||
|
||||
- `gen` → answers
|
||||
- `glm-5.2`, `kimi-k3`, `gen-frontier` → `key not allowed to access model`
|
||||
|
||||
Secrets are in the vault, never in git. `.env` on the host is `0600`:
|
||||
|
||||
```bash
|
||||
secret get esh-docker-vm/lobe-chat-litellm-key # -> OPENAI_API_KEY
|
||||
secret get esh-docker-vm/lobe-chat-access-code # -> ACCESS_CODE (UI gate)
|
||||
secret get esh-docker-vm/lobe-chat-key-vaults-secret # -> KEY_VAULTS_SECRET
|
||||
```
|
||||
|
||||
`ACCESS_CODE` matters: this is a home-lab LAN segment with nothing in front of it.
|
||||
|
||||
## Verified on deploy (2026-08-16)
|
||||
|
||||
- container healthy; `http://10.0.50.45:3210/` → 307 → `/chat` → 200
|
||||
- from **inside** the container: `GET /v1/models` returns the fleet seats, and a
|
||||
`gen` chat round-trip returns `"ok"` — so the app's own network path and key work,
|
||||
not merely the host's
|
||||
- image on disk 617 MB (143 MB compressed)
|
||||
|
||||
## Deploy
|
||||
|
||||
```bash
|
||||
scripts/deploy-stack.sh esh-docker-vm lobe-chat --compose
|
||||
# on host: populate .env from the vault (see above), chmod 600, then
|
||||
ssh lkraven@10.0.50.45 'cd /opt/docker/compose/lobe-chat && docker compose up -d'
|
||||
```
|
||||
|
||||
`lkraven` owns `/opt/docker` and is in the `docker` group on this host, so no sudo
|
||||
is needed. Note ESH is outside the infra-ops NOPASSWD grant.
|
||||
Reference in New Issue
Block a user