docs(lobe-chat): correct the TTS notes — the deploy's TTS never worked
Two claims in this stack's docs were reasoned from the wrong hop, and one of
them hid a dead feature since deploy. Re-measured from esh-docker-vm against
the live .env:
1. "An unknown `model` routes to the gateway default" — true of :8198, false of
the path Lobe takes. LiteLLM resolves the model name first, so Lobe's default
`tts-1` returns 403 (`key not allowed to access model`) and never reaches the
gateway. `ext-tts` returns 200 + audio. The endpoint inheriting
OPENAI_PROXY_URL is necessary but not sufficient: Settings -> TTS -> OpenAI
TTS model -> `ext-tts` is a required one-time step per browser, and removing
it needs a LiteLLM alias plus a key allow-list entry (both master-key, so
infra-ops).
2. "`response_format: mp3` ... set it in the UI" — not possible. Lobe's OpenAI
TTS client sends `{input, model, voice}` and nothing else (server bundle
chunks/29685.js), so format is not selectable from this stack at any level.
The deploy gets the fleet gateway's default (WAV, ~23.5 MB for a 245 s turn),
relabelled `audio/mpeg` by LiteLLM. That is tts-dev's fence, not this one's.
Comment/doc only — no functional change, so the host copy needs no redeploy.
This commit is contained in:
@@ -34,10 +34,19 @@
|
||||
# CORRECTION to an earlier note in this file: the OpenAI voice names are
|
||||
# ALIASED, not rejected (echo/alloy/onyx/ash->donut, nova->miranda,
|
||||
# shimmer/coral->emmie, fable/sage->glados), so a UI mis-click is NOT the
|
||||
# hazard I first recorded. Only `ballad` and `verse` 404.
|
||||
# The real constraint is SIZE: 245s of WAV is 23.5 MB, so a browser UI should
|
||||
# request `response_format: "mp3"`. The old 71.2s per-call cap is dead (that
|
||||
# was Zonos-era); dots chunks server-side and a 592-word call renders intact.
|
||||
# hazard I first recorded. Only `ballad` and `verse` 404 -- both aliased since.
|
||||
# ⚠️ REQUIRED UI STEP (tts-dev 2026-08-17): Settings -> TTS -> OpenAI TTS model
|
||||
# -> `ext-tts`. There is no env var for it. Lobe's default is `tts-1`, and while
|
||||
# :8198 ignores `model`, LITELLM IN FRONT OF IT DOES NOT -- it resolves the name
|
||||
# first, so `tts-1` returns 403 (`key not allowed to access model`; 400 on an
|
||||
# unscoped key) and TTS does nothing. Verified from this host with this .env.
|
||||
# Per browser -- the TTS settings store is client-side.
|
||||
# SIZE: Lobe sends only {input, model, voice} -- no `response_format`, no
|
||||
# `speed`, and no UI field for either -- so you get the gateway's default (WAV,
|
||||
# ~23.5 MB for a 245 s turn), relabelled `audio/mpeg` by LiteLLM. Not fixable
|
||||
# from this stack: it is a fleet-gateway default and tts-dev owns it.
|
||||
# The old 71.2s per-call cap is dead (that was Zonos-era); dots chunks
|
||||
# server-side and a 592-word call renders intact.
|
||||
# ⚠️ The seat SERIALIZES generation -- one render at a time, no continuous
|
||||
# batching -- so sustained volume from here is a real capacity question for a
|
||||
# shared GPU. Report ramp to tts-dev on althing.
|
||||
|
||||
Reference in New Issue
Block a user