stacks/index-tts: add streaming WAV endpoint (wrapper 0.2.0)
IndexTTS-2's tts.infer(stream_return=True) is a generator that yields
audio chunks per text segment as they finish, plus inter-segment
silence. Expose this via the existing POST /v1/audio/speech with a new
"stream": true field on the request body.
Wire-up:
- 44-byte WAV header emitted up front with placeholder data length
(0xFFFFFFFF) so chunks can be written before total samples are
known. Players that read until EOF (mpv, ffplay, aplay, sox,
browsers via <audio>) handle this fine.
- Each yielded chunk goes through _chunk_to_pcm_bytes(), which
handles torch tensors / numpy arrays in either int16 or float
(-1..1) form.
- 22050 Hz mono int16 — IndexTTS-2's hardcoded output shape.
Time-to-first-audio drops from full-file latency to ~one-segment
latency. Single-sentence inputs barely benefit; long passages /
multi-paragraph reads benefit a lot. Strict metadata parsers may
balk at the placeholder size — request without stream for a
closed-length WAV in that case.
INDEX_TTS_TAG bumped to v2 to force a rebuild.
This commit is contained in:
@@ -72,9 +72,39 @@ curl -X POST http://10.100.79.3:8192/v1/audio/speech \
|
||||
curl http://10.100.79.3:8192/healthz
|
||||
```
|
||||
|
||||
Output is always WAV (PCM_16, 22050 Hz — IndexTTS-2's native rate).
|
||||
Output is always WAV (PCM_16, 22050 Hz mono — IndexTTS-2's native rate).
|
||||
`response_format` other than `wav` is rejected.
|
||||
|
||||
### Streaming (since 0.2.0)
|
||||
|
||||
Add `"stream": true` to any request to stream the WAV as it generates.
|
||||
IndexTTS-2 emits one chunk per text segment (~120 tokens) as soon as it
|
||||
finishes synthesizing it; long inputs start playing while the rest is
|
||||
still being generated.
|
||||
|
||||
```bash
|
||||
# Pipe straight into a player. Time-to-first-audio drops from
|
||||
# whole-file latency to ~one-segment latency.
|
||||
curl -fsS -X POST http://10.100.79.3:8192/v1/audio/speech \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"input":"long passage of text...","voice":"glados","stream":true}' \
|
||||
| mpv --no-cache -
|
||||
|
||||
# Or save while playing (tee).
|
||||
curl -fsS -X POST http://10.100.79.3:8192/v1/audio/speech \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"input":"...","voice":"glados","stream":true}' \
|
||||
| tee out.wav | mpv -
|
||||
```
|
||||
|
||||
The streaming WAV uses a placeholder data-length in the header
|
||||
(`0xFFFFFFFF`) so chunks can be written before the total is known.
|
||||
Players that read until EOF (mpv, ffplay, aplay, sox, browsers via
|
||||
`<audio>`) handle this fine. Strict parsers (some metadata extractors,
|
||||
foobar2000 default settings) may complain about the size. If that
|
||||
matters, request without `stream` and you get a normal closed-length
|
||||
WAV.
|
||||
|
||||
## Voice library
|
||||
|
||||
Flat dirs on the host (bind-mounted; survives container recreates):
|
||||
|
||||
Reference in New Issue
Block a user