feat(morpheus): staged clone voices + max_tokens 3500 (context-clamped)
- max_tokens default 2400->3500 (~42s) in wrapper + gateway-chat client, with a _cap() clamp so prompt+gen never exceeds MAX_CTX (4096) — a cloning ref block is ~1100 tokens, so an unclamped 3500 would overflow context on the clone path. - Staged clone voices: /voices dir of <name>.wav + <name>.txt, each encoded to its Orpheus reference block at startup; voice="<name>" zero-shot clones it. Beatrice (a chatterbox reference) staged as the first normal-voice clone. GET /voices lists baddy + clones. - compose: mount voices dir + pass MORPHEUS_MAX_LEN to the wrapper (clamp must match engine). vLLM concurrency (measured, --max-num-seqs 8, 250-tok reqs): near-linear batching — 8 concurrent finish in the same ~2.8s as 1 (707 tok/s, 8.1x single, flat per-req latency). Chunked-sentence production can fan out for ~8x throughput; CPU SNAC decode is the scale bottleneck, not generation.
This commit is contained in:
@@ -24,9 +24,15 @@ lower time-to-first-audio is a future enhancement.)
|
||||
|
||||
- `POST /tts` → `audio/wav`. Body: `{"text": "...", "voice": "baddy", "temperature": 0.6,
|
||||
"max_tokens": 1200, "repetition_penalty": 1.1}`.
|
||||
- **Zero-shot clone:** add `"reference_audio_b64": "<base64 WAV>"` + `"reference_text":
|
||||
- **Zero-shot clone (ad-hoc):** add `"reference_audio_b64": "<base64 WAV>"` + `"reference_text":
|
||||
"<its transcript>"`. Keep `repetition_penalty <= 1.1` for cloning (higher penalizes the
|
||||
in-context reference audio tokens and breaks generation).
|
||||
- **Staged clone voices:** drop `<name>.wav` + `<name>.txt` (its transcript) into the voices
|
||||
dir (`/home/lkraven/morpheus/voices/`); each is encoded to its reference block once at
|
||||
startup, so `voice: "<name>"` zero-shot clones it (e.g. `beatrice`). `GET /voices` lists them.
|
||||
- `max_tokens` defaults to 3500 (~42 s), auto-clamped so prompt + gen never exceeds the
|
||||
4096 context (a cloning reference block is ~1,100 tokens). `repetition_penalty` 1.1 is
|
||||
load-bearing — at 1.0 the model never emits end-of-speech and rambles to the cap.
|
||||
- `GET /voices`, `GET /health`, `GET /docs` (OpenAPI UI).
|
||||
|
||||
**Expressive tags** (baddy is trained for these): `<sigh> <gasp> <laugh> <chuckle> <pant>
|
||||
|
||||
Reference in New Issue
Block a user