ace-step: stream audio bytes inline; catalog v3 → v4
The pre-fix wrapper at stacks/ace-step/infer-api.py returned a JSON
{output_path: "..."} reference to a file written inside the
container at /app/outputs/. That path was unreachable from outside
the container — every consumer got 134 bytes of JSON-pretending-to-
be-WAV instead of audio. Surfaced by the asset_engine consumer's
end-to-end smoke (althing thread 01KRCJF7NGMXYE9F62Q1A6KFD4 msg 5);
my own earlier smoke missed it because I checked HTTP=200 and stopped
reading instead of inspecting the response body.
Wrapper now reads back the file the pipeline writes and streams the
bytes via fastapi.responses.Response with media_type set from the
audio_format request field (audio/wav | audio/mpeg | audio/flac).
The in-container path is exposed via X-Output-Path header for log
correlation but is no longer load-bearing.
Verified end-to-end against live ace-step on irv-ml1:
POST /generate -> HTTP 200 in 80s
content-type: audio/wav
content-length: 945226
x-output-path: /app/outputs/output_cfe87d1d....wav
$ file response.wav
RIFF (little-endian) data, WAVE audio, Microsoft PCM, 16 bit,
stereo 48000 Hz
Catalog: ace-step bumped version 3 -> 4. Dropped
response.output_field (no longer applicable). reproducibility.notes
expanded to record both the v2 18-arg-tuple fix and this v4
inline-streaming change so the history is auditable from the
catalog itself.
Stale ACEStepOutput Pydantic model left in infer-api.py for now —
unused but small; future cleanup.
This commit is contained in:
@@ -752,7 +752,7 @@ services:
|
||||
Apache-2.0 hybrid diffusion+LLM music generation. Multi-minute lyric-aware
|
||||
songs with vocals + instrumentation.
|
||||
category: music
|
||||
version: 3
|
||||
version: 4
|
||||
host: irv-ml1
|
||||
endpoint: http://10.100.79.3:8210/generate
|
||||
method: POST
|
||||
@@ -948,16 +948,23 @@ services:
|
||||
Wrapper-side cleanup queued — once the upstream model defaults this,
|
||||
the catalog field will become optional or be dropped entirely.
|
||||
response:
|
||||
# As of wrapper version that ships with image local/ace-step:v1
|
||||
# post 2026-05-11, /generate streams audio bytes inline with
|
||||
# Content-Type set from the audio_format request field. The
|
||||
# in-container output_path is exposed via X-Output-Path header
|
||||
# for log correlation but is no longer load-bearing.
|
||||
type: audio
|
||||
mime_from_field: audio_format
|
||||
output_field: output_path
|
||||
reproducibility:
|
||||
seedable: true
|
||||
deterministic: true
|
||||
notes: >
|
||||
actual_seeds parameter exposed; identical seeds + params = identical audio.
|
||||
Local infer-api.py patches upstream's broken 24-arg pipeline signature
|
||||
(was 18 in upstream — caused crashes with audio_duration in `format` slot).
|
||||
(was 18 in upstream — caused crashes with audio_duration in `format` slot)
|
||||
AND inline-streams the generated audio bytes (was returning a JSON
|
||||
path reference to a file inside the container, which was unreachable
|
||||
from outside).
|
||||
estimated_latency:
|
||||
cold_start_s: 30
|
||||
warm_per_unit: "~10–60s depending on audio_duration + infer_step"
|
||||
|
||||
Reference in New Issue
Block a user