stacks/index-tts: add streaming WAV endpoint (wrapper 0.2.0)

IndexTTS-2's tts.infer(stream_return=True) is a generator that yields
audio chunks per text segment as they finish, plus inter-segment
silence. Expose this via the existing POST /v1/audio/speech with a new
"stream": true field on the request body.

Wire-up:
  - 44-byte WAV header emitted up front with placeholder data length
    (0xFFFFFFFF) so chunks can be written before total samples are
    known. Players that read until EOF (mpv, ffplay, aplay, sox,
    browsers via <audio>) handle this fine.
  - Each yielded chunk goes through _chunk_to_pcm_bytes(), which
    handles torch tensors / numpy arrays in either int16 or float
    (-1..1) form.
  - 22050 Hz mono int16 — IndexTTS-2's hardcoded output shape.

Time-to-first-audio drops from full-file latency to ~one-segment
latency. Single-sentence inputs barely benefit; long passages /
multi-paragraph reads benefit a lot. Strict metadata parsers may
balk at the placeholder size — request without stream for a
closed-length WAV in that case.

INDEX_TTS_TAG bumped to v2 to force a rebuild.
This commit is contained in:
vh
2026-04-25 14:50:29 -07:00
parent ab696ecbd1
commit 54fef0e9d8
3 changed files with 115 additions and 12 deletions
+31 -1
View File
@@ -72,9 +72,39 @@ curl -X POST http://10.100.79.3:8192/v1/audio/speech \
curl http://10.100.79.3:8192/healthz
```
Output is always WAV (PCM_16, 22050 Hz — IndexTTS-2's native rate).
Output is always WAV (PCM_16, 22050 Hz mono — IndexTTS-2's native rate).
`response_format` other than `wav` is rejected.
### Streaming (since 0.2.0)
Add `"stream": true` to any request to stream the WAV as it generates.
IndexTTS-2 emits one chunk per text segment (~120 tokens) as soon as it
finishes synthesizing it; long inputs start playing while the rest is
still being generated.
```bash
# Pipe straight into a player. Time-to-first-audio drops from
# whole-file latency to ~one-segment latency.
curl -fsS -X POST http://10.100.79.3:8192/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"input":"long passage of text...","voice":"glados","stream":true}' \
| mpv --no-cache -
# Or save while playing (tee).
curl -fsS -X POST http://10.100.79.3:8192/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"input":"...","voice":"glados","stream":true}' \
| tee out.wav | mpv -
```
The streaming WAV uses a placeholder data-length in the header
(`0xFFFFFFFF`) so chunks can be written before the total is known.
Players that read until EOF (mpv, ffplay, aplay, sox, browsers via
`<audio>`) handle this fine. Strict parsers (some metadata extractors,
foobar2000 default settings) may complain about the size. If that
matters, request without `stream` and you get a normal closed-length
WAV.
## Voice library
Flat dirs on the host (bind-mounted; survives container recreates):