docs: correct stale TTS voice warning (tts-dev); record DS regeneration spec (brokkr)

tts-dev answered the Lobe onboarding, live-verified. Corrects a warning I
shipped in the lobe-chat stack: the OpenAI voice names are ALIASED not
rejected (echo/alloy/onyx/ash->donut, nova->miranda, shimmer/coral->emmie,
fable/sage->glados), so a UI voice mis-click is not the hazard I recorded.
Only ballad and verse 404. The old 71.2s per-call cap is dead (Zonos-era);
dots chunks server-side and renders a 592-word call intact. Real constraint
is size (245s WAV = 23.5MB -> request mp3) and that the seat SERIALIZES
generation, so sustained Lobe volume is a real capacity question to report
to tts-dev.

Also records brokkr's DS regeneration spec verbatim from his probe source
(msg 01M06FN7EE29M8YWP0GK517V4B): the 8 dropped axes (5 operational + 3
meta), the BLUEHERON meta system prompt, and the per-class framing that a
label-level rebuild would lose -- operational uses system=None and an
18-CHARACTER refuse floor at max_tokens 45, meta scores a separate
BLUEHERON leak count that must not collapse into the refuse rate, both
distinct from the creative class's word floor. Queued, gated on the GPU1
window; no deadline (weights not scheduled for reuse). Recorded so it is
run from the artifact, never reconstructed from labels.
This commit is contained in:
vh
2026-08-16 16:51:40 -07:00
parent aba7cda33e
commit 25fa18efb8
3 changed files with 168 additions and 10 deletions
+13 -5
View File
@@ -28,11 +28,19 @@
# secret get esh-docker-vm/lobe-chat-access-code
# secret get esh-docker-vm/lobe-chat-key-vaults-secret
#
# ⚠️ ext-tts FOOT-GUN: an unknown voice 404s and can trip the LiteLLM router
# cooldown. `ext-tts` accepts donut/emmie/glados/miranda (+ emotion variants)
# and the OpenAI aliases nova/alloy. If Lobe sends any other OpenAI voice name
# (echo, fable, onyx, shimmer) it will 404 -- pin the voice rather than leaving
# it at whatever the UI defaults to.
# TTS (verified with tts-dev 2026-08-16): route via the `ext-tts` LiteLLM alias,
# never direct to the seat -- engines get swapped behind the gateway and
# direct-to-seat eats every change. Auth is the gateway key alone.
# CORRECTION to an earlier note in this file: the OpenAI voice names are
# ALIASED, not rejected (echo/alloy/onyx/ash->donut, nova->miranda,
# shimmer/coral->emmie, fable/sage->glados), so a UI mis-click is NOT the
# hazard I first recorded. Only `ballad` and `verse` 404.
# The real constraint is SIZE: 245s of WAV is 23.5 MB, so a browser UI should
# request `response_format: "mp3"`. The old 71.2s per-call cap is dead (that
# was Zonos-era); dots chunks server-side and a 592-word call renders intact.
# ⚠️ The seat SERIALIZES generation -- one render at a time, no continuous
# batching -- so sustained volume from here is a real capacity question for a
# shared GPU. Report ramp to tts-dev on althing.
name: lobe-chat