feat(omnivoice): new TTS stack — k2-fsa/OmniVoice on irv-ml1 3090
Zero-shot, massively-multilingual (600+ language) voice-cloning + voice-design TTS (diffusion-LM, Apache-2.0). No official image, so a thin CUDA container around the pip package running upstream's own Gradio demo (no FastAPI wrapper). Pinned to GPU 0 (3090) — the A6000 is ComfyUI-exclusive — port 8199. Built + verified live on irv-ml1 (Gradio 200, container healthy). Surface is the Gradio UI + Gradio API, NOT OpenAI-compat /v1/audio/speech (wrap later if asset-engine should consume it). deploy-omnivoice.yaml builds local + verifies.
This commit is contained in:
@@ -0,0 +1,54 @@
|
||||
# syntax=docker/dockerfile:1.6
|
||||
#
|
||||
# OmniVoice (k2-fsa/OmniVoice) — zero-shot, massively-multilingual (600+
|
||||
# languages) voice-cloning + voice-design TTS, diffusion-LM architecture,
|
||||
# Apache-2.0. Upstream ships a pip package + its own Gradio demo
|
||||
# (`omnivoice-demo`); there's no official image, so we build a thin CUDA
|
||||
# container around the pip package and run its Gradio server directly.
|
||||
# Unlike index-tts we DON'T need a FastAPI wrapper — OmniVoice serves itself.
|
||||
|
||||
ARG CUDA_BASE=nvidia/cuda:12.8.0-cudnn-runtime-ubuntu22.04
|
||||
FROM ${CUDA_BASE}
|
||||
|
||||
ENV DEBIAN_FRONTEND=noninteractive \
|
||||
PIP_ROOT_USER_ACTION=ignore \
|
||||
PYTHONUNBUFFERED=1 \
|
||||
PATH="/opt/venv/bin:${PATH}"
|
||||
|
||||
RUN apt-get update \
|
||||
&& apt-get install -y --no-install-recommends \
|
||||
python3.10 python3.10-venv python3-pip \
|
||||
git ffmpeg libsndfile1 \
|
||||
ca-certificates wget \
|
||||
&& rm -rf /var/lib/apt/lists/* \
|
||||
&& python3.10 -m venv /opt/venv
|
||||
|
||||
# CUDA 12.8 torch wheels (per OmniVoice's documented install line).
|
||||
RUN pip install --no-cache-dir \
|
||||
torch==2.8.0 torchaudio==2.8.0 \
|
||||
--index-url https://download.pytorch.org/whl/cu128
|
||||
|
||||
# OmniVoice from PyPI (+ huggingface_hub for the entrypoint weight pre-warm).
|
||||
# Optional reproducible pin via the OMNIVOICE_VERSION build arg (empty=latest).
|
||||
ARG OMNIVOICE_VERSION=
|
||||
RUN pip install --no-cache-dir "omnivoice${OMNIVOICE_VERSION:+==${OMNIVOICE_VERSION}}" huggingface_hub
|
||||
|
||||
# Fail the build loudly if the console script name isn't what we expect,
|
||||
# rather than crash-loop at runtime. Logs the actual omni* entrypoints.
|
||||
RUN echo "omni console scripts:" && (ls /opt/venv/bin | grep -i omni || true) \
|
||||
&& command -v omnivoice-demo >/dev/null \
|
||||
|| { echo "ERROR: omnivoice-demo CLI not found after install"; exit 1; }
|
||||
|
||||
WORKDIR /app
|
||||
COPY entrypoint.sh /usr/local/bin/entrypoint.sh
|
||||
RUN chmod +x /usr/local/bin/entrypoint.sh
|
||||
|
||||
EXPOSE 8001
|
||||
|
||||
# Gradio serves HTML at / — 200 once the UI is up (weights load lazily on
|
||||
# first synth; the entrypoint pre-warms them). Generous start period.
|
||||
HEALTHCHECK --interval=30s --timeout=10s --start-period=900s --retries=3 \
|
||||
CMD wget -q -O /dev/null http://127.0.0.1:8001/ || exit 1
|
||||
|
||||
ENTRYPOINT ["/usr/local/bin/entrypoint.sh"]
|
||||
CMD ["omnivoice-demo", "--ip", "0.0.0.0", "--port", "8001"]
|
||||
Reference in New Issue
Block a user