docs(lobe-chat): correct the TTS notes — the deploy's TTS never worked

Two claims in this stack's docs were reasoned from the wrong hop, and one of
them hid a dead feature since deploy. Re-measured from esh-docker-vm against
the live .env:

1. "An unknown `model` routes to the gateway default" — true of :8198, false of
   the path Lobe takes. LiteLLM resolves the model name first, so Lobe's default
   `tts-1` returns 403 (`key not allowed to access model`) and never reaches the
   gateway. `ext-tts` returns 200 + audio. The endpoint inheriting
   OPENAI_PROXY_URL is necessary but not sufficient: Settings -> TTS -> OpenAI
   TTS model -> `ext-tts` is a required one-time step per browser, and removing
   it needs a LiteLLM alias plus a key allow-list entry (both master-key, so
   infra-ops).

2. "`response_format: mp3` ... set it in the UI" — not possible. Lobe's OpenAI
   TTS client sends `{input, model, voice}` and nothing else (server bundle
   chunks/29685.js), so format is not selectable from this stack at any level.
   The deploy gets the fleet gateway's default (WAV, ~23.5 MB for a 245 s turn),
   relabelled `audio/mpeg` by LiteLLM. That is tts-dev's fence, not this one's.

Comment/doc only — no functional change, so the host copy needs no redeploy.
This commit is contained in:
vh
2026-08-17 21:28:59 -07:00
parent d1f4f1cb96
commit ca8c0a318e
2 changed files with 69 additions and 24 deletions
+13 -4
View File
@@ -34,10 +34,19 @@
# CORRECTION to an earlier note in this file: the OpenAI voice names are
# ALIASED, not rejected (echo/alloy/onyx/ash->donut, nova->miranda,
# shimmer/coral->emmie, fable/sage->glados), so a UI mis-click is NOT the
# hazard I first recorded. Only `ballad` and `verse` 404.
# The real constraint is SIZE: 245s of WAV is 23.5 MB, so a browser UI should
# request `response_format: "mp3"`. The old 71.2s per-call cap is dead (that
# was Zonos-era); dots chunks server-side and a 592-word call renders intact.
# hazard I first recorded. Only `ballad` and `verse` 404 -- both aliased since.
# ⚠️ REQUIRED UI STEP (tts-dev 2026-08-17): Settings -> TTS -> OpenAI TTS model
# -> `ext-tts`. There is no env var for it. Lobe's default is `tts-1`, and while
# :8198 ignores `model`, LITELLM IN FRONT OF IT DOES NOT -- it resolves the name
# first, so `tts-1` returns 403 (`key not allowed to access model`; 400 on an
# unscoped key) and TTS does nothing. Verified from this host with this .env.
# Per browser -- the TTS settings store is client-side.
# SIZE: Lobe sends only {input, model, voice} -- no `response_format`, no
# `speed`, and no UI field for either -- so you get the gateway's default (WAV,
# ~23.5 MB for a 245 s turn), relabelled `audio/mpeg` by LiteLLM. Not fixable
# from this stack: it is a fleet-gateway default and tts-dev owns it.
# The old 71.2s per-call cap is dead (that was Zonos-era); dots chunks
# server-side and a 592-word call renders intact.
# ⚠️ The seat SERIALIZES generation -- one render at a time, no continuous
# batching -- so sustained volume from here is a real capacity question for a
# shared GPU. Report ramp to tts-dev on althing.