ace-step + stable-audio-open: deploy music + SFX generation to irv-ml1
Two new audio-generation stacks alongside the TTS slate: ace-step :8210 — Apache 2.0 music generation foundation model (hybrid diffusion + LLM). Lyric-aware multi-minute songs. ~10-12 GB VRAM during inference, A6000-pinned. Custom Dockerfile patches upstream's torch/cu126 resolution bug (--extra-index-url cu126 was falling back to pypi-default cu13 wheels, mismatching torchvision). stable-audio-open :8211 — Stability AI 1.21B latent-diffusion SFX + ambience. Up to 47s clips at 44.1 kHz. ~6 GB VRAM in fp16, A6000-pinned. Custom FastAPI shim around diffusers' StableAudioPipeline (no upstream HTTP server). Dockerfile pins torchsde explicitly — diffusers doesn't pull it as a hard dep but CosineDPMSolverMultistepScheduler needs it.
This commit is contained in:
@@ -0,0 +1,47 @@
|
||||
# ACE-Step 1.5 stack tunables. Copy to `.env` on irv-ml1 before
|
||||
# deploying.
|
||||
|
||||
# ── build pin ────────────────────────────────────────────────────────
|
||||
# SHA of ace-step/ACE-Step to build from. Use the FULL 40-char SHA —
|
||||
# docker buildx git source resolver rejects short hashes. `main` works
|
||||
# at first deploy; pin to a real SHA before any production cutover so
|
||||
# upstream commits don't surprise you on next rebuild.
|
||||
ACE_STEP_SHA=main
|
||||
|
||||
# Local image tag — bump when you change build context to force a
|
||||
# fresh layer build.
|
||||
ACE_STEP_TAG=v1
|
||||
|
||||
# ── network ──────────────────────────────────────────────────────────
|
||||
# Host port (container listens on 8000 internally — infer-api.py
|
||||
# hardcodes uvicorn.run(host=0.0.0.0, port=8000)).
|
||||
# Reservations on irv-ml1: 8188 ComfyUI, 8190 CosyVoice, 8191 Qwen3-TTS,
|
||||
# 8192 IndexTTS-2, 8193 Kokoro, 8194 VibeVoice, 8195 Fish, 8196
|
||||
# Chatterbox, 8197 Voxtral, 8765 Parakeet ASR. 8210 starts the
|
||||
# audio-generation block (music + SFX) so future TTS adds can keep
|
||||
# going from 8198+.
|
||||
ACE_STEP_PORT=8210
|
||||
ACE_STEP_BIND=0.0.0.0
|
||||
|
||||
# ── runtime / GPU ────────────────────────────────────────────────────
|
||||
# GPU pinning. "0" = RTX 3090 (24 GB), "1" = RTX A6000 (48 GB).
|
||||
# A6000 (1) recommended — Fish s2-pro lives there at ~17 GB, and
|
||||
# ACE-Step adds ~10-12 GB during inference, leaving comfortable
|
||||
# headroom on the 48 GB card. The 3090 is full with the TTS slate.
|
||||
ACE_STEP_GPU_DEVICES=1
|
||||
|
||||
# ── persistent storage on the host ───────────────────────────────────
|
||||
# Model checkpoints — primary spot for any manually-staged checkpoints.
|
||||
# ACE-Step's auto-download lands in HF_HOME (cache dir below).
|
||||
ACE_STEP_CHECKPOINTS_DIR=/worktank/ace-step/checkpoints
|
||||
|
||||
# Generated audio output — clients can pull from here via the
|
||||
# returned file path in the /generate response.
|
||||
ACE_STEP_OUTPUTS_DIR=/worktank/ace-step/outputs
|
||||
|
||||
# Application logs.
|
||||
ACE_STEP_LOGS_DIR=/worktank/ace-step/logs
|
||||
|
||||
# HF cache — first start pulls the ACE-Step checkpoint (~5-10 GB)
|
||||
# into this dir. Persistent across container recreates.
|
||||
ACE_STEP_CACHE_DIR=/worktank/ace-step/hf_cache
|
||||
@@ -0,0 +1,66 @@
|
||||
# Custom Dockerfile for ACE-Step 1.5.
|
||||
#
|
||||
# Mirrors upstream's Dockerfile structure, but fixes a CUDA-version
|
||||
# mismatch that crashloops the upstream image as of April 2026:
|
||||
# * upstream's requirements.txt lists `torch torchvision torchaudio`
|
||||
# with no version pins;
|
||||
# * upstream's pip install uses `--extra-index-url cu126`, which is
|
||||
# a FALLBACK only — pypi default wins for resolution;
|
||||
# * pypi-default torch is now cu13, so torch installs cu13 + the
|
||||
# cu126 fallback only kicks in for torchvision/torchaudio →
|
||||
# `RuntimeError: Detected that PyTorch and torchvision were compiled
|
||||
# with different CUDA major versions`.
|
||||
#
|
||||
# Fix: install torch/torchvision/torchaudio FIRST from the cu126 index
|
||||
# (forced via --index-url, not --extra-index-url). Then `pip install
|
||||
# -r requirements.txt` sees they're already satisfied and leaves them
|
||||
# alone.
|
||||
#
|
||||
# Also: command is `python3 infer-api.py` (REST), not `gui.py` (Gradio)
|
||||
# — see compose.yaml command override; CMD here is the same default
|
||||
# so the image works standalone too.
|
||||
FROM nvidia/cuda:12.6.0-runtime-ubuntu22.04 AS base
|
||||
|
||||
ENV PYTHONDONTWRITEBYTECODE=1 \
|
||||
PYTHONUNBUFFERED=1 \
|
||||
HF_HUB_ENABLE_HF_TRANSFER=1 \
|
||||
DEBIAN_FRONTEND=noninteractive
|
||||
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends \
|
||||
python3.10 \
|
||||
python3-pip \
|
||||
python3-venv \
|
||||
python3-dev \
|
||||
build-essential \
|
||||
git \
|
||||
curl \
|
||||
ca-certificates \
|
||||
&& apt-get clean \
|
||||
&& rm -rf /var/lib/apt/lists/* \
|
||||
&& ln -sf /usr/bin/python3 /usr/bin/python
|
||||
|
||||
RUN python -m venv /opt/venv
|
||||
ENV PATH="/opt/venv/bin:$PATH"
|
||||
|
||||
WORKDIR /app
|
||||
|
||||
# Clone upstream. Bake the SHA into a layer-cache key so a different
|
||||
# SHA invalidates everything below.
|
||||
ARG ACE_STEP_REF=main
|
||||
RUN git clone https://github.com/ace-step/ACE-Step.git . \
|
||||
&& git checkout ${ACE_STEP_REF} \
|
||||
&& echo "ace-step ref: $(git rev-parse HEAD)"
|
||||
|
||||
# Pre-install torch/torchvision/torchaudio from the cu126 index — this
|
||||
# satisfies the unpinned entries in requirements.txt so the next pip
|
||||
# install doesn't re-resolve them from pypi default (cu13).
|
||||
RUN pip install --no-cache-dir --upgrade pip \
|
||||
&& pip install --no-cache-dir \
|
||||
torch torchvision torchaudio \
|
||||
--index-url https://download.pytorch.org/whl/cu126 \
|
||||
&& pip install --no-cache-dir hf_transfer peft \
|
||||
&& pip install --no-cache-dir -r requirements.txt \
|
||||
&& pip install --no-cache-dir .
|
||||
|
||||
EXPOSE 8000
|
||||
CMD ["python3", "infer-api.py"]
|
||||
@@ -0,0 +1,53 @@
|
||||
# ace-step
|
||||
|
||||
ACE-Step 1.5 — Apache 2.0 open-source music generation foundation
|
||||
model. Hybrid diffusion + LLM. Generates lyric-aware multi-minute
|
||||
songs (vocals + instrumentation).
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| host | `irv-ml1` |
|
||||
| port | `8210` |
|
||||
| GPU | A6000 (`device_ids: ["1"]`) |
|
||||
| VRAM | ~10-12 GB during inference |
|
||||
| upstream | https://github.com/ace-step/ACE-Step |
|
||||
| license | Apache 2.0 |
|
||||
|
||||
## API surface
|
||||
|
||||
`infer-api.py` (FastAPI) exposes:
|
||||
|
||||
- `GET /health` — liveness, returns 200 once the process is up
|
||||
(model is lazy-loaded on first /generate).
|
||||
- `POST /generate` — body: `ACEStepInput` Pydantic model with
|
||||
~27 params (prompt, lyrics, audio_duration, guidance_scale, etc.).
|
||||
Returns `{status, output_path, message}`.
|
||||
|
||||
The container does NOT expose the Gradio UI — we override the upstream
|
||||
default `python3 acestep/gui.py` with `python3 infer-api.py`. If you
|
||||
want the Gradio UI for ad-hoc experimentation, run a one-off:
|
||||
|
||||
```bash
|
||||
ssh irv-ml1 'docker exec -it ace-step python3 acestep/gui.py --server_name 0.0.0.0 --port 7865'
|
||||
```
|
||||
|
||||
…and port-forward 7865 to your laptop.
|
||||
|
||||
## Deploy
|
||||
|
||||
```bash
|
||||
scripts/elway irv-ml1 --playbook playbooks/deploy-ace-step.yaml
|
||||
```
|
||||
|
||||
Idempotent. Cold build is ~10-15 min (CUDA + torch + transformers +
|
||||
spacy + audio deps). First `/generate` triggers the model download
|
||||
(~5-10 GB) and warmup (~30-60 s).
|
||||
|
||||
## Tunables
|
||||
|
||||
See `.env.example` — copy to `.env` on the host (lives at
|
||||
`/opt/docker/compose/ace-step/.env`, gitignored). Common knobs:
|
||||
|
||||
- `ACE_STEP_SHA` — pin upstream commit
|
||||
- `ACE_STEP_GPU_DEVICES` — GPU index
|
||||
- `ACE_STEP_*_DIR` — bind-mount paths under `/worktank/ace-step/`
|
||||
@@ -0,0 +1,70 @@
|
||||
# ACE-Step 1.5 — open-source music generation foundation model
|
||||
# (April 2026). Hybrid diffusion + LLM architecture, Apache 2.0.
|
||||
# ~50-80 s for a 4-minute song on A6000; under 4 GB VRAM at idle,
|
||||
# ~10-12 GB during inference. Beats YuE / DiffRhythm on the
|
||||
# speed/coherence trade.
|
||||
#
|
||||
# We launch upstream's REST API (`infer-api.py`) instead of the
|
||||
# default `gui.py` (Gradio). The REST surface is what we'll point
|
||||
# clients + automation at; Gradio is dev-time eye candy.
|
||||
#
|
||||
# Image is built locally from upstream's repo via docker buildx git
|
||||
# context, same pattern as fish-s2.
|
||||
#
|
||||
# All tunables live in .env — edit that, not this file.
|
||||
|
||||
services:
|
||||
ace-step:
|
||||
image: local/ace-step:${ACE_STEP_TAG}
|
||||
build:
|
||||
# Build from local Dockerfile (not upstream's git context) — we
|
||||
# ship a patched Dockerfile that fixes upstream's torch/cu126
|
||||
# resolution bug. Playbook uploads Dockerfile alongside this
|
||||
# compose.yaml.
|
||||
context: .
|
||||
dockerfile: Dockerfile
|
||||
args:
|
||||
ACE_STEP_REF: ${ACE_STEP_SHA}
|
||||
container_name: ace-step
|
||||
restart: unless-stopped
|
||||
runtime: nvidia
|
||||
ports:
|
||||
# Container default for infer-api.py is 8000 (hardcoded
|
||||
# uvicorn.run(host=0.0.0.0, port=8000) — no flags). Map host
|
||||
# ACE_STEP_PORT to it.
|
||||
- "${ACE_STEP_BIND:-0.0.0.0}:${ACE_STEP_PORT}:8000"
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${ACE_STEP_GPU_DEVICES:-1}
|
||||
# ACE_OUTPUT_DIR is read by acestep at generation time — keep
|
||||
# in sync with the bind mount below.
|
||||
- ACE_OUTPUT_DIR=/app/outputs
|
||||
# HF_HOME points the HuggingFace cache at the bind mount so the
|
||||
# ~5-10 GB checkpoint download survives container recreates.
|
||||
- HF_HOME=/app/hf_cache
|
||||
volumes:
|
||||
- ${ACE_STEP_CHECKPOINTS_DIR}:/app/checkpoints
|
||||
- ${ACE_STEP_OUTPUTS_DIR}:/app/outputs
|
||||
- ${ACE_STEP_LOGS_DIR}:/app/logs
|
||||
- ${ACE_STEP_CACHE_DIR}:/app/hf_cache
|
||||
# Override upstream's default `python3 acestep/gui.py` with the
|
||||
# REST API entry point. infer-api.py self-binds 0.0.0.0:8000 and
|
||||
# exposes POST /generate + GET /health.
|
||||
command: ["python3", "infer-api.py"]
|
||||
healthcheck:
|
||||
# /health is the cheapest signal infer-api.py exposes — returns
|
||||
# 200 as soon as the FastAPI app is up. The pipeline lazy-loads
|
||||
# on first /generate, so /health says "process alive" not
|
||||
# "model warm". Good enough for a liveness signal; first
|
||||
# /generate has the ~30-60 s warmup baked in.
|
||||
test: ["CMD-SHELL", "python3 -c \"import urllib.request,sys; sys.exit(0 if urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=5).status==200 else 1)\""]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
# First boot pulls ACE-Step checkpoint (~5-10 GB) into HF cache.
|
||||
start_period: 600s
|
||||
labels:
|
||||
- homepage.group=AI Systems
|
||||
- homepage.name=ACE-Step
|
||||
- homepage.icon=mdi-music-note-eighth
|
||||
- homepage.description=Open-source music generation — 4-min song in ~60s, lyrics + style prompts (irv-ml1)
|
||||
- homepage.href=http://10.100.79.3:${ACE_STEP_PORT}
|
||||
Reference in New Issue
Block a user