feat(morpheus): permanent mOrpheus TTS stack (vLLM bf16 + SNAC/FastAPI wrapper) on irv-ml1

Two-container stack serving MrDragonFox/mOrpheus (uncensored Orpheus TTS, Llama-3.2-3B
-> SNAC 24kHz). vllm-morpheus (GPU/3090) emits Orpheus audio tokens; morpheus-tts (CPU)
SNAC-decodes them to WAV and exposes POST /tts (baddy voice + zero-shot cloning). Deployed
+ tested end-to-end (28/28 valid frames, valid WAV, reachable over WG).

Hard-won config, all encoded in compose/README:
- bf16 REQUIRED: --quantization fp8 destroys audio-token generation (0 valid SNAC frames
  even at greedy). Footprint ~7.9GB.
- Image PINNED to v0.23.0: 'latest' ships Blackwell oink/aiter kernels that crash on Ampere
  import.
- 3090 (not the comfy-contended A6000); --enforce-eager to fit the shared card.
- RTF ~1.0 end-to-end (gen ~98 tok/s / RTF 0.84 + CPU decode + HTTP).

INTERNAL RESEARCH ONLY (CC-BY-NC-4.0); do not expose externally.
This commit is contained in:
vh
2026-07-09 00:49:25 -07:00
parent 99a4a1721f
commit 01eedd8d27
6 changed files with 315 additions and 0 deletions
+54
View File
@@ -0,0 +1,54 @@
# mOrpheus — uncensored Orpheus TTS (irv-ml1)
Permanent serving stack for **`MrDragonFox/mOrpheus_3B-1Base_early_preview-v1-25000`** — an
uncensored Orpheus TTS finetune (Llama-3.2-3B LLM → SNAC 24 kHz audio). Trained speaker
**"baddy"**; supports **zero-shot voice cloning** from a reference clip.
> **INTERNAL RESEARCH ONLY.** License is **CC-BY-NC-4.0** (non-commercial). Do **not** expose
> this endpoint externally or use it in any commercial-facing product.
## Shape
Two containers (see `compose.yaml`):
| service | where | role |
|---|---|---|
| `vllm-morpheus` | GPU (3090), FP8 | serves the mOrpheus LLM; emits Orpheus audio tokens |
| `morpheus-tts` | CPU | SNAC-decodes tokens → 24 kHz WAV; the public `/tts` endpoint |
**Real-time:** ~165 tok/s single-stream on the 3090 (FP8) ⇒ **RTF ≈ 0.50 (2× real-time)**,
measured. A ~4 s clip generates in ~2 s. (Whole-clip decode in v1; chunked streaming for
lower time-to-first-audio is a future enhancement.)
## Endpoints (`http://10.100.79.3:8299`)
- `POST /tts` → `audio/wav`. Body: `{"text": "...", "voice": "baddy", "temperature": 0.6,
"max_tokens": 1200, "repetition_penalty": 1.1}`.
- **Zero-shot clone:** add `"reference_audio_b64": "<base64 WAV>"` + `"reference_text":
"<its transcript>"`. Keep `repetition_penalty <= 1.1` for cloning (higher penalizes the
in-context reference audio tokens and breaks generation).
- `GET /voices`, `GET /health`, `GET /docs` (OpenAPI UI).
**Expressive tags** (baddy is trained for these): `<sigh> <gasp> <laugh> <chuckle> <pant>
<groan> <moan>` etc. Use **real carrier sentences with sparse, sentence-boundary tags** —
stacking many tags with little text sends this early checkpoint into a repeat-loop.
## Deploy (irv-ml1, as `lkraven` — docker-group, no sudo)
```bash
# one-time: stage weights (from the audition dir or a fresh pull-hf-repo) + copy the stack
mkdir -p /home/lkraven/morpheus/models
mv /home/lkraven/orpheus-audition/models/mOrpheus /home/lkraven/morpheus/models/
mv /home/lkraven/orpheus-audition/models/snac_24khz /home/lkraven/morpheus/models/
# copy compose.yaml + tts/ to /home/lkraven/morpheus/, cp .env.example .env
cd /home/lkraven/morpheus && docker compose build && docker compose up -d
```
## Gotchas
- **Pin `vllm/vllm-openai:v0.23.0`** — `latest` ships Blackwell-only kernels (oink/aiter)
that crash on Ampere *import*. Do not bump to `latest` on this box.
- **GPU = 3090, not the A6000** — the A6000 is comfy's and spikes to ~41 GB without warning
(OOM'd two launches). FP8's ~5 GB footprint coexists with the 3090 audio zoo.
- **FP8** on Ampere is a VRAM save (upcast), no compute speedup — real-time comes from vLLM.
- Canonical copy lives here; deployed copy is `/home/lkraven/morpheus/` on irv-ml1.