feat(morpheus): staged clone voices + max_tokens 3500 (context-clamped)

- max_tokens default 2400->3500 (~42s) in wrapper + gateway-chat client, with a _cap()
  clamp so prompt+gen never exceeds MAX_CTX (4096) — a cloning ref block is ~1100 tokens,
  so an unclamped 3500 would overflow context on the clone path.
- Staged clone voices: /voices dir of <name>.wav + <name>.txt, each encoded to its Orpheus
  reference block at startup; voice="<name>" zero-shot clones it. Beatrice (a chatterbox
  reference) staged as the first normal-voice clone. GET /voices lists baddy + clones.
- compose: mount voices dir + pass MORPHEUS_MAX_LEN to the wrapper (clamp must match engine).

vLLM concurrency (measured, --max-num-seqs 8, 250-tok reqs): near-linear batching — 8
concurrent finish in the same ~2.8s as 1 (707 tok/s, 8.1x single, flat per-req latency).
Chunked-sentence production can fan out for ~8x throughput; CPU SNAC decode is the scale
bottleneck, not generation.
This commit is contained in:
vh
2026-07-09 01:35:14 -07:00
parent f363fe6c84
commit 0655a37bf6
5 changed files with 41 additions and 8 deletions
+7 -1
View File
@@ -24,9 +24,15 @@ lower time-to-first-audio is a future enhancement.)
- `POST /tts` → `audio/wav`. Body: `{"text": "...", "voice": "baddy", "temperature": 0.6,
"max_tokens": 1200, "repetition_penalty": 1.1}`.
- **Zero-shot clone:** add `"reference_audio_b64": "<base64 WAV>"` + `"reference_text":
- **Zero-shot clone (ad-hoc):** add `"reference_audio_b64": "<base64 WAV>"` + `"reference_text":
"<its transcript>"`. Keep `repetition_penalty <= 1.1` for cloning (higher penalizes the
in-context reference audio tokens and breaks generation).
- **Staged clone voices:** drop `<name>.wav` + `<name>.txt` (its transcript) into the voices
dir (`/home/lkraven/morpheus/voices/`); each is encoded to its reference block once at
startup, so `voice: "<name>"` zero-shot clones it (e.g. `beatrice`). `GET /voices` lists them.
- `max_tokens` defaults to 3500 (~42 s), auto-clamped so prompt + gen never exceeds the
4096 context (a cloning reference block is ~1,100 tokens). `repetition_penalty` 1.1 is
load-bearing — at 1.0 the model never emits end-of-speech and rambles to the cap.
- `GET /voices`, `GET /health`, `GET /docs` (OpenAPI UI).
**Expressive tags** (baddy is trained for these): `<sigh> <gasp> <laugh> <chuckle> <pant>