ace-step + stable-audio-open: deploy music + SFX generation to irv-ml1

Two new audio-generation stacks alongside the TTS slate:

ace-step :8210 — Apache 2.0 music generation foundation model
(hybrid diffusion + LLM). Lyric-aware multi-minute songs. ~10-12 GB
VRAM during inference, A6000-pinned. Custom Dockerfile patches
upstream's torch/cu126 resolution bug (--extra-index-url cu126 was
falling back to pypi-default cu13 wheels, mismatching torchvision).

stable-audio-open :8211 — Stability AI 1.21B latent-diffusion SFX +
ambience. Up to 47s clips at 44.1 kHz. ~6 GB VRAM in fp16,
A6000-pinned. Custom FastAPI shim around diffusers' StableAudioPipeline
(no upstream HTTP server). Dockerfile pins torchsde explicitly —
diffusers doesn't pull it as a hard dep but
CosineDPMSolverMultistepScheduler needs it.
This commit is contained in:
vh
2026-04-28 09:11:23 -07:00
parent 0ba41e02ea
commit 4a4c09177f
11 changed files with 732 additions and 0 deletions
+48
View File
@@ -0,0 +1,48 @@
# Stable Audio Open 1.0 stack tunables. Copy to `.env` on irv-ml1
# before deploying.
# ── image ────────────────────────────────────────────────────────────
# Local image tag — bump when you change Dockerfile or server.py to
# force a fresh build.
SAO_TAG=v1
# Which Stable Audio model to load. As of 2026-04 the only released
# checkpoint is 1.0; future revisions can swap here without touching
# compose.yaml or server.py.
SAO_MODEL=stabilityai/stable-audio-open-1.0
# ── network ──────────────────────────────────────────────────────────
# Host port (container listens on 8000 internally).
# Reservations on irv-ml1: 8188 ComfyUI, 8190 CosyVoice, 8191 Qwen3-TTS,
# 8192 IndexTTS-2, 8193 Kokoro, 8194 VibeVoice, 8195 Fish, 8196
# Chatterbox, 8197 Voxtral, 8210 ACE-Step (music), 8765 Parakeet ASR.
SAO_PORT=8211
SAO_BIND=0.0.0.0
# ── runtime / GPU ────────────────────────────────────────────────────
# GPU pinning. "0" = RTX 3090 (24 GB), "1" = RTX A6000 (48 GB).
# A6000 (1) recommended — Fish s2-pro lives there at ~17 GB; SAO adds
# ~6 GB practical (model fp16 + small VAE working set), and ACE-Step
# adds another ~12 GB during inference. Total ~35 GB / 48 GB still
# leaves headroom. The 3090 is full with the TTS slate.
SAO_GPU_DEVICES=1
# ── HuggingFace auth ─────────────────────────────────────────────────
# HF token — REQUIRED. Stable Audio Open is gated; you must:
# 1. Visit https://huggingface.co/stabilityai/stable-audio-open-1.0
# and accept the Stability AI Community License (one click).
# 2. Generate a read token at
# https://huggingface.co/settings/tokens.
# 3. Paste it here.
# Without this, the first model download 401s and the container
# crashloops.
SAO_HF_TOKEN=
# ── persistent storage on the host ───────────────────────────────────
# HF cache — first start pulls the model (~6 GB) into this dir.
# Persistent across container recreates so we don't re-pull.
SAO_CACHE_DIR=/worktank/stable-audio-open/hf_cache
# Generated audio output — clients can pull from here for any flow
# that wants a file path instead of a streamed WAV body.
SAO_OUTPUTS_DIR=/worktank/stable-audio-open/outputs
+43
View File
@@ -0,0 +1,43 @@
# Stable Audio Open 1.0 inference image.
# pytorch/pytorch base ships torch + cuda + cudnn already linked, so
# we only layer the diffusers stack + a libsndfile for soundfile + the
# fastapi shim. Smaller and faster to build than starting from
# nvidia/cuda and pip-installing torch ourselves.
FROM pytorch/pytorch:2.5.1-cuda12.4-cudnn9-runtime AS base
ENV PYTHONUNBUFFERED=1 \
PYTHONDONTWRITEBYTECODE=1 \
PIP_NO_CACHE_DIR=1 \
PIP_DISABLE_PIP_VERSION_CHECK=1 \
HF_HOME=/app/hf_cache
# libsndfile1 is the C lib soundfile binds to. Without it the pip
# install of soundfile succeeds but `import soundfile` fails at
# runtime with OSError: cannot find libsndfile.
RUN apt-get update && apt-get install -y --no-install-recommends \
libsndfile1 \
&& rm -rf /var/lib/apt/lists/*
# protobuf + sentencepiece are pulled in by the T5 text encoder
# (Stable Audio Open uses google/t5-base-cb under the hood).
# accelerate gates the .to(device) fast path for diffusers.
# torchsde is required by CosineDPMSolverMultistepScheduler — diffusers
# doesn't pull it as a hard dep; without it, pipeline init fails with
# "CosineDPMSolverMultistepScheduler requires the torchsde library".
RUN pip install \
"diffusers>=0.27.0" \
"transformers>=4.40.0" \
accelerate \
protobuf \
sentencepiece \
soundfile \
torchsde \
fastapi \
"uvicorn[standard]" \
pydantic
WORKDIR /app
COPY server.py /app/server.py
EXPOSE 8000
CMD ["uvicorn", "server:app", "--host", "0.0.0.0", "--port", "8000"]
+55
View File
@@ -0,0 +1,55 @@
# stable-audio-open
Stability AI's Stable Audio Open 1.0 — text-to-audio latent diffusion.
Strong on SFX, foley, ambience, short loops. Not a music model — it
does not generate intelligible vocals or structured songs (use
`ace-step` for that).
| | |
|---|---|
| host | `irv-ml1` |
| port | `8211` |
| GPU | A6000 (`device_ids: ["1"]`) |
| VRAM | ~6 GB in fp16 |
| max clip | 47 s at 44.1 kHz |
| upstream | https://github.com/Stability-AI/stable-audio-tools |
| model | `stabilityai/stable-audio-open-1.0` (gated) |
| license | Stability AI Community (non-commercial / personal / research) |
## API surface
`server.py` (custom FastAPI shim) exposes:
- `GET /health` — returns 200 once the model is loaded.
- `POST /v1/audio/sfx` — returns a `audio/wav` blob.
```jsonc
{
"prompt": "a vintage typewriter clacking in a quiet room",
"negative_prompt": "Low quality.", // optional, default "Low quality."
"duration": 10.0, // seconds, 0.5 – 47
"steps": 100, // 10 – 300, more = better quality
"seed": 42, // optional
"cfg_scale": 7.0 // 0 – 20
}
```
Why a custom shim: there's no upstream Docker image and no upstream
HTTP server for Stable Audio Open. Diffusers exposes
`StableAudioPipeline` cleanly — the shim is ~70 lines.
## Deploy
```bash
scripts/elway irv-ml1 --playbook playbooks/deploy-stable-audio-open.yaml
```
Pre-deploy: visit https://huggingface.co/stabilityai/stable-audio-open-1.0
once and accept the Community License (HF token alone is not enough —
the gate is per-model). Then put the token in `SAO_HF_TOKEN` in `.env`
on the host.
## Tunables
See `.env.example` — copy to `.env` on the host (lives at
`/opt/docker/compose/stable-audio-open/.env`, gitignored).
+61
View File
@@ -0,0 +1,61 @@
# Stable Audio Open 1.0 — Stability AI's open-weight latent-diffusion
# SFX/ambience generator. 1.21B params, ~4-6 GB VRAM in fp16, up to
# 47 s clips at 44.1 kHz. Strong on text-aligned sound effects, foley,
# field-recording-style ambience. NOT a music model — it does not
# generate intelligible vocals or structured songs (use ACE-Step for
# that).
#
# LICENSE: Stability AI Community License. Personal / research use is
# free; commercial use requires a separate license from Stability
# (https://stability.ai/license). Same posture we already accepted
# for Voxtral.
#
# No upstream Docker image — we ship a custom Dockerfile + a small
# FastAPI shim (server.py) that wraps diffusers' StableAudioPipeline
# and exposes POST /v1/audio/sfx.
#
# All tunables live in .env — edit that, not this file.
services:
stable-audio-open:
image: local/stable-audio-open:${SAO_TAG}
build:
# Build context is the compose dir on the host — the playbook
# uploads server.py + Dockerfile alongside this compose.yaml.
context: .
dockerfile: Dockerfile
container_name: stable-audio-open
restart: unless-stopped
runtime: nvidia
ports:
- "${SAO_BIND:-0.0.0.0}:${SAO_PORT}:8000"
environment:
- NVIDIA_VISIBLE_DEVICES=${SAO_GPU_DEVICES:-1}
- SAO_MODEL=${SAO_MODEL:-stabilityai/stable-audio-open-1.0}
- HF_HOME=/app/hf_cache
# Model is gated on HuggingFace (you must accept the Stability
# Community License once on the model page before the token can
# download it). Set SAO_HF_TOKEN in .env. Without this, the
# first model download 401s and the container crashloops.
- HF_TOKEN=${SAO_HF_TOKEN}
volumes:
- ${SAO_CACHE_DIR}:/app/hf_cache
- ${SAO_OUTPUTS_DIR}:/app/outputs
healthcheck:
# /health is set by server.py — returns 200 once FastAPI is up
# AND the pipeline finished loading (lifespan blocks startup
# until the model is in VRAM).
test: ["CMD-SHELL", "python -c \"import urllib.request,sys; sys.exit(0 if urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=5).status==200 else 1)\""]
interval: 30s
timeout: 10s
retries: 3
# First boot pulls the model (~6 GB) into HF cache + loads to
# VRAM. Cold start ~3-5 min on a fast pipe; subsequent starts
# are ~30 s.
start_period: 600s
labels:
- homepage.group=AI Systems
- homepage.name=Stable Audio Open
- homepage.icon=mdi-waveform
- homepage.description=Diffusion SFX/ambience generator — up to 47s at 44.1 kHz (irv-ml1)
- homepage.href=http://10.100.79.3:${SAO_PORT}
+80
View File
@@ -0,0 +1,80 @@
# FastAPI shim around diffusers' StableAudioPipeline.
# Single endpoint POST /v1/audio/sfx returns a WAV blob.
# Model is loaded once on startup and held in process memory.
import io
import os
import time
from contextlib import asynccontextmanager
from typing import Optional
import soundfile as sf
import torch
from diffusers import StableAudioPipeline
from fastapi import FastAPI, HTTPException, Response
from pydantic import BaseModel, Field
MODEL_ID = os.environ.get("SAO_MODEL", "stabilityai/stable-audio-open-1.0")
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
DTYPE = torch.float16 if DEVICE == "cuda" else torch.float32
state: dict = {}
@asynccontextmanager
async def lifespan(app: FastAPI):
print(f"[sao] loading {MODEL_ID} on {DEVICE} ({DTYPE})", flush=True)
t0 = time.time()
pipe = StableAudioPipeline.from_pretrained(MODEL_ID, torch_dtype=DTYPE)
pipe = pipe.to(DEVICE)
state["pipe"] = pipe
print(f"[sao] loaded in {time.time() - t0:.1f}s", flush=True)
yield
state.clear()
app = FastAPI(lifespan=lifespan)
class SfxRequest(BaseModel):
prompt: str = Field(..., min_length=1)
negative_prompt: Optional[str] = "Low quality."
duration: float = Field(10.0, gt=0.5, le=47.0)
steps: int = Field(100, ge=10, le=300)
seed: Optional[int] = None
cfg_scale: float = Field(7.0, gt=0.0, le=20.0)
@app.get("/health")
def health():
return {
"status": "ok",
"model": MODEL_ID,
"device": DEVICE,
"loaded": "pipe" in state,
}
@app.post("/v1/audio/sfx")
def sfx(req: SfxRequest):
pipe = state.get("pipe")
if pipe is None:
raise HTTPException(503, "model not loaded yet")
generator = None
if req.seed is not None:
generator = torch.Generator(DEVICE).manual_seed(req.seed)
audio = pipe(
req.prompt,
negative_prompt=req.negative_prompt,
num_inference_steps=req.steps,
audio_end_in_s=req.duration,
num_waveforms_per_prompt=1,
generator=generator,
).audios
waveform = audio[0].T.float().cpu().numpy()
buf = io.BytesIO()
sf.write(buf, waveform, pipe.vae.sampling_rate, format="WAV")
buf.seek(0)
return Response(content=buf.read(), media_type="audio/wav")