Files
esh-pfi-infrastructure/stacks/lobe-chat/README.md
T
vh 25fa18efb8 docs: correct stale TTS voice warning (tts-dev); record DS regeneration spec (brokkr)
tts-dev answered the Lobe onboarding, live-verified. Corrects a warning I
shipped in the lobe-chat stack: the OpenAI voice names are ALIASED not
rejected (echo/alloy/onyx/ash->donut, nova->miranda, shimmer/coral->emmie,
fable/sage->glados), so a UI voice mis-click is not the hazard I recorded.
Only ballad and verse 404. The old 71.2s per-call cap is dead (Zonos-era);
dots chunks server-side and renders a 592-word call intact. Real constraint
is size (245s WAV = 23.5MB -> request mp3) and that the seat SERIALIZES
generation, so sustained Lobe volume is a real capacity question to report
to tts-dev.

Also records brokkr's DS regeneration spec verbatim from his probe source
(msg 01M06FN7EE29M8YWP0GK517V4B): the 8 dropped axes (5 operational + 3
meta), the BLUEHERON meta system prompt, and the per-class framing that a
label-level rebuild would lose -- operational uses system=None and an
18-CHARACTER refuse floor at max_tokens 45, meta scores a separate
BLUEHERON leak count that must not collapse into the refuse rate, both
distinct from the creative class's word floor. Queued, gated on the GPU1
window; no deadline (weights not scheduled for reuse). Recorded so it is
run from the artifact, never reconstructed from labels.
2026-08-16 16:51:40 -07:00

6.3 KiB
Raw Blame History

lobe-chat — chat frontend over the LiteLLM gateway (esh-docker-vm)

Evaluation replacement for the hand-rolled gateway-chat single-file HTML surface, which the operator does not want to keep improving — it has already produced two defects (a 1024 max_tokens default that read as model degeneracy, and a NaNnull max_tokens bug).

  • Host: esh-docker-vm (10.0.50.45) · Port: 3210 · URL: http://10.0.50.45:3210
  • Backend: LiteLLM gateway at 10.250.50.70:4000/v1 (reachable from ESH, ~30 ms)

Why Lobe over Open WebUI

Weight, measured from the registries rather than from marketing:

compressed layers
Lobe Chat 143 MB 1
Open WebUI 1,825 MB 19

12.8×. Open WebUI was declined by the operator in June 2026 on weight grounds and that objection still holds. (Its recorded secondary objection — the empty-tools 400 against vLLM — is now moot: the strip_empty_tools callback covers the normal API path and only failed to protect LiteLLM's own built-in playground.)

⚠️ The open question this deploy exists to answer

Is Lobe's TTS configurable by ENV, or only through the settings UI? That is the operator's deciding criterion — manageable/scriptable by an agent — and it is unresolved. Open WebUI has dedicated AUDIO_TTS_ENGINE / AUDIO_TTS_OPENAI_API_BASE_URL / AUDIO_TTS_MODEL / AUDIO_TTS_VOICE. Lobe documents a shared OPENAI_PROXY_URL, which should carry TTS because LiteLLM serves /v1/chat/completions and /v1/audio/speech on the same base — but that is inference, not verification.

If it turns out UI-only, the real choice is: 1.8 GB with genuine scriptability, or 143 MB with click-ops. That is an operator call, not an agent one.

TTS — verified with tts-dev, 2026-08-16

Route: the ext-tts LiteLLM alias. Do NOT go direct to the seat. The gateway exists so engines can be auditioned and swapped behind it — dots was swapped twice in the week before this deploy and consumers noticed nothing. Direct-to-seat means eating every engine change.

Auth: the gateway virtual key, nothing more. The seat itself has no auth (LAN/WireGuard-internal).

Shape: plain OpenAI POST /v1/audio/speech with {model, input, voice, response_format}. No deviations; a stock OpenAI client works drop-in. Unknown model values route to default rather than 404ing, so tts-1 is harmless. mp3/opus/aac/flac transcode; wav/pcm pass through byte-verbatim. speed honoured 0.254.0.

⚠️ CORRECTION — my earlier voice foot-gun warning was wrong

This file previously warned that any non-fleet OpenAI voice name 404s and could trip the router cooldown, and told you to pin the voice. That was stale and mostly unfounded. tts-dev verified live: the OpenAI names are ALIASED, not rejected — designed for exactly this case, a stock client dropping in without knowing fleet voice names.

alloy, echo, onyx, ash  -> donut      nova            -> miranda
shimmer, coral          -> emmie      fable, sage     -> glados

Authoritative list (13): computer, computer-soft, computer-urgent, donut, emmie, glados, miranda, sindra, sindra-excited, sindra-sad, sindra-soft, sindra-sultry, sindra-whisper.

The only live grenades are ballad and verse — newer OpenAI additions never aliased, both confirmed 404. If Lobe exposes the full modern OpenAI voice list those are the two to avoid; tts-dev has offered to alias them (one-line change on his side) and infra-ops has taken him up on it.

The real constraint is SIZE, not length

The 71.2s per-call cap in older notes is dead — that was Zonos-era and required client-side chunking. dots chunks server-side. tts-dev threw a 592-word single call at it: HTTP 200, 66.4 s wall, 245 s of audio, byte-identical ending under ASR diff. Long assistant turns are handled, not clipped.

But: 245 s of WAV is 23.5 MB. For a browser UI set response_format: "mp3" or you will push tens of megabytes per turn at the user. And 66 s of wall time is a long wait with no streaming in the OpenAI-compat path — dots does have a genuine streaming mode (stream: true, ~0.40 s to first audio, flat with length) but that is not the OpenAI-compat shape and would be a custom integration.

⚠️ Capacity — tell tts-dev if this ramps

The seat serializes generation — one render at a time, by design, no continuous batching. A second concurrent consumer is a real capacity question, not a theoretical one, and dots shares a GPU where tts-dev had two resource incidents that week. Report sustained volume, and any voice string not on the list above, to tts-dev on althing rather than letting him infer it from a VRAM graph.

Credential posture

Deliberately not the shared all-agents key — that reaches the paid passthroughs (GLM, Kimi), and a LAN-exposed chat UI holding it would let anyone who can reach the port spend vendor credits from a pool shared across every project.

This stack uses a purpose-minted LiteLLM virtual key, key_alias: lobe-chat-esh, scoped to the 20 free local models. Scoping was verified at mint time, both directions:

  • gen → answers
  • glm-5.2, kimi-k3, gen-frontierkey not allowed to access model

Secrets are in the vault, never in git. .env on the host is 0600:

secret get esh-docker-vm/lobe-chat-litellm-key        # -> OPENAI_API_KEY
secret get esh-docker-vm/lobe-chat-access-code        # -> ACCESS_CODE (UI gate)
secret get esh-docker-vm/lobe-chat-key-vaults-secret  # -> KEY_VAULTS_SECRET

ACCESS_CODE matters: this is a home-lab LAN segment with nothing in front of it.

Verified on deploy (2026-08-16)

  • container healthy; http://10.0.50.45:3210/ → 307 → /chat → 200
  • from inside the container: GET /v1/models returns the fleet seats, and a gen chat round-trip returns "ok" — so the app's own network path and key work, not merely the host's
  • image on disk 617 MB (143 MB compressed)

Deploy

scripts/deploy-stack.sh esh-docker-vm lobe-chat --compose
# on host: populate .env from the vault (see above), chmod 600, then
ssh lkraven@10.0.50.45 'cd /opt/docker/compose/lobe-chat && docker compose up -d'

lkraven owns /opt/docker and is in the docker group on this host, so no sudo is needed. Note ESH is outside the infra-ops NOPASSWD grant.