feat(omnivoice): expose full generation surface (voice-design, language, diffusion params)
Wrapper /v1/audio/speech now accepts OmniVoice's whole surface: - voice (clone, now OPTIONAL) and/or instruct (voice DESIGN). instruct is a CONTROLLED vocabulary (gender/age/pitch/accent/whisper tags, comma-separated), not free prose — discoverable at the new /v1/audio/instruct-items endpoint (23 items). - language (Auto + 647, new /v1/audio/languages endpoint), speed, duration. - diffusion controls: num_step, guidance_scale, denoise, preprocess_prompt, postprocess_output; plus a generation_overrides JSON passthrough for expert GenerationConfig knobs (t_shift, layer_penalty_factor, position/class temperature, audio_chunk_*). - at least one of voice/instruct required (else 400). Catalog (services.yaml): omnivoice v1 -> v2, 13 schema-valid fields; instruct as a controlled-vocab text field sourced from the items endpoint. Verified live on irv-ml1: clone, voice-design (instruct-only), and tuned-param synths all -> 24 kHz PCM_16 WAV; 647 languages; 23 instruct items.
This commit is contained in:
@@ -572,7 +572,7 @@ services:
|
|||||||
(diffusion-LM, RTF ~0.025). Apache-2.0. Behind our own FastAPI wrapper
|
(diffusion-LM, RTF ~0.025). Apache-2.0. Behind our own FastAPI wrapper
|
||||||
(stacks/omnivoice/app.py); voices are the reused chatterbox reference clips.
|
(stacks/omnivoice/app.py); voices are the reused chatterbox reference clips.
|
||||||
category: tts
|
category: tts
|
||||||
version: 1
|
version: 2
|
||||||
status: ready
|
status: ready
|
||||||
host: irv-ml1
|
host: irv-ml1
|
||||||
lifecycle:
|
lifecycle:
|
||||||
@@ -592,15 +592,90 @@ services:
|
|||||||
label: Text
|
label: Text
|
||||||
required: true
|
required: true
|
||||||
max_length: 5000
|
max_length: 5000
|
||||||
|
# Voice source — at least one of voice (clone) / instruct (design) is required.
|
||||||
- name: voice
|
- name: voice
|
||||||
type: select
|
type: select
|
||||||
label: Speaker Voice
|
label: Speaker Voice (clone)
|
||||||
|
optional: true
|
||||||
source_url: http://10.100.79.3:8199/v1/audio/voices
|
source_url: http://10.100.79.3:8199/v1/audio/voices
|
||||||
source_jsonpath: $.voices[*]
|
source_jsonpath: $.voices[*]
|
||||||
description: >
|
description: >
|
||||||
Zero-shot clone target — a reference clip staged in
|
Zero-shot clone target — a reference clip in /worktank/omnivoice/voices/
|
||||||
/worktank/omnivoice/voices/ (reused chatterbox voices; 33 at deploy).
|
(reused chatterbox voices; 33 at deploy). Omit to design a voice via
|
||||||
Live list at /v1/audio/voices.
|
instruct instead. Live list at /v1/audio/voices.
|
||||||
|
- name: instruct
|
||||||
|
type: text
|
||||||
|
label: Voice Design (instruct)
|
||||||
|
optional: true
|
||||||
|
source_url: http://10.100.79.3:8199/v1/audio/instruct-items
|
||||||
|
source_jsonpath: $.instruct_items[*]
|
||||||
|
description: >
|
||||||
|
Voice DESIGN — a comma-separated list of CONTROLLED attribute tags (not
|
||||||
|
free prose), e.g. "british accent, elderly, male, low pitch". Valid tags
|
||||||
|
(gender/age/pitch/accent/whisper) at /v1/audio/instruct-items. Use instead
|
||||||
|
of, or together with, a clone voice.
|
||||||
|
- name: language
|
||||||
|
type: select
|
||||||
|
label: Language
|
||||||
|
optional: true
|
||||||
|
default: Auto
|
||||||
|
source_url: http://10.100.79.3:8199/v1/audio/languages
|
||||||
|
source_jsonpath: $.languages[*]
|
||||||
|
description: "Auto-detects when left as Auto; 600+ languages supported."
|
||||||
|
- name: speed
|
||||||
|
type: slider
|
||||||
|
label: Speed
|
||||||
|
optional: true
|
||||||
|
min: 0.5
|
||||||
|
max: 1.5
|
||||||
|
default: 1.0
|
||||||
|
description: "1.0 = normal; >1 faster, <1 slower. Ignored if duration is set."
|
||||||
|
- name: duration
|
||||||
|
type: number
|
||||||
|
label: Duration (seconds)
|
||||||
|
optional: true
|
||||||
|
description: "Fixed output length in seconds; overrides speed when set."
|
||||||
|
- name: num_step
|
||||||
|
type: slider
|
||||||
|
label: Inference Steps
|
||||||
|
optional: true
|
||||||
|
min: 4
|
||||||
|
max: 64
|
||||||
|
default: 32
|
||||||
|
description: "Diffusion steps. Lower = faster, higher = better quality."
|
||||||
|
- name: guidance_scale
|
||||||
|
type: slider
|
||||||
|
label: Guidance Scale (CFG)
|
||||||
|
optional: true
|
||||||
|
min: 0.0
|
||||||
|
max: 4.0
|
||||||
|
default: 2.0
|
||||||
|
- name: denoise
|
||||||
|
type: bool
|
||||||
|
label: Denoise
|
||||||
|
optional: true
|
||||||
|
default: true
|
||||||
|
- name: preprocess_prompt
|
||||||
|
type: bool
|
||||||
|
label: Preprocess Prompt
|
||||||
|
optional: true
|
||||||
|
default: true
|
||||||
|
description: "Silence-trim + punctuate the reference (clone mode)."
|
||||||
|
- name: postprocess_output
|
||||||
|
type: bool
|
||||||
|
label: Postprocess Output
|
||||||
|
optional: true
|
||||||
|
default: true
|
||||||
|
description: "Remove long silences from the generated audio."
|
||||||
|
- name: generation_overrides
|
||||||
|
type: json
|
||||||
|
label: Advanced (GenerationConfig)
|
||||||
|
optional: true
|
||||||
|
description: >
|
||||||
|
Expert OmniVoiceGenerationConfig overrides as a JSON object — keys:
|
||||||
|
t_shift (0.1), layer_penalty_factor (5.0), position_temperature (5.0),
|
||||||
|
class_temperature (0.0), audio_chunk_duration (15.0),
|
||||||
|
audio_chunk_threshold (30.0). Unknown keys ignored.
|
||||||
- name: response_format
|
- name: response_format
|
||||||
type: select
|
type: select
|
||||||
options: [wav]
|
options: [wav]
|
||||||
@@ -621,9 +696,12 @@ services:
|
|||||||
warm_per_unit: "full-utterance (no streaming)"
|
warm_per_unit: "full-utterance (no streaming)"
|
||||||
license: "Apache-2.0"
|
license: "Apache-2.0"
|
||||||
notes: |
|
notes: |
|
||||||
v1 is clone-only (voice = a staged reference clip); OmniVoice's voice-DESIGN
|
Two voice sources, combinable: voice (clone a staged reference clip) and/or
|
||||||
and language/instruct controls are not yet exposed in the wrapper. No streaming.
|
instruct (free-text voice DESIGN); at least one required. Full generation
|
||||||
Voices reused from chatterbox /refs (staged into /worktank/omnivoice/voices/).
|
surface exposed — language (600+), speed, duration, num_step, guidance_scale,
|
||||||
|
denoise, preprocess/postprocess — with expert GenerationConfig knobs (t_shift,
|
||||||
|
layer/position/class temperature, audio_chunk_*) via the generation_overrides
|
||||||
|
JSON field. No streaming. Voices reused from chatterbox /refs.
|
||||||
|
|
||||||
- id: qwen3-tts
|
- id: qwen3-tts
|
||||||
name: Qwen3-TTS 1.7B
|
name: Qwen3-TTS 1.7B
|
||||||
|
|||||||
+93
-14
@@ -7,17 +7,25 @@ ${OMNIVOICE_VOICES_DIR} (the reused chatterbox /refs/*.wav). A voice-clone promp
|
|||||||
is precomputed once per voice at startup (the loaded Whisper ASR auto-transcribes
|
is precomputed once per voice at startup (the loaded Whisper ASR auto-transcribes
|
||||||
each reference) and cached, so per-request latency is just generation.
|
each reference) and cached, so per-request latency is just generation.
|
||||||
|
|
||||||
|
Exposes OmniVoice's full generation surface:
|
||||||
|
- clone (voice=<staged ref>) and/or voice-DESIGN (instruct=<free-text style>)
|
||||||
|
- language (Auto + 600+), speed, duration
|
||||||
|
- diffusion controls: num_step, guidance_scale, denoise, preprocess_prompt,
|
||||||
|
postprocess_output, plus a generation_overrides passthrough for expert knobs
|
||||||
|
(t_shift, layer_penalty_factor, position_temperature, class_temperature, ...).
|
||||||
|
|
||||||
Endpoints:
|
Endpoints:
|
||||||
GET /healthz -> readiness (200 once model + >=1 voice are loaded)
|
GET /healthz -> readiness (200 once model + >=1 voice are loaded)
|
||||||
GET /v1/audio/voices -> {"voices": [<name>, ...]}
|
GET /v1/audio/voices -> {"voices": [<name>, ...]}
|
||||||
POST /v1/audio/speech -> {input, voice, response_format=wav} -> audio/wav
|
GET /v1/audio/languages -> {"languages": ["Auto", <display name>, ...]}
|
||||||
|
POST /v1/audio/speech -> audio/wav
|
||||||
"""
|
"""
|
||||||
import glob
|
import glob
|
||||||
import io
|
import io
|
||||||
import logging
|
import logging
|
||||||
import os
|
import os
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Optional
|
from typing import Any, Dict, Optional
|
||||||
|
|
||||||
import numpy as np
|
import numpy as np
|
||||||
import soundfile as sf
|
import soundfile as sf
|
||||||
@@ -28,14 +36,26 @@ from pydantic import BaseModel
|
|||||||
|
|
||||||
from omnivoice import OmniVoice, OmniVoiceGenerationConfig
|
from omnivoice import OmniVoice, OmniVoiceGenerationConfig
|
||||||
|
|
||||||
|
try:
|
||||||
|
from omnivoice.utils.lang_map import LANG_NAMES, lang_display_name
|
||||||
|
LANGUAGES = ["Auto"] + sorted(lang_display_name(n) for n in LANG_NAMES)
|
||||||
|
except Exception: # noqa: BLE001
|
||||||
|
LANGUAGES = ["Auto"]
|
||||||
|
|
||||||
|
try:
|
||||||
|
# Voice-DESIGN `instruct` is a CONTROLLED vocabulary (gender / age / pitch /
|
||||||
|
# accent / whisper tags), NOT free prose — surfaced so callers can discover it.
|
||||||
|
from omnivoice.utils.voice_design import _INSTRUCT_VALID_EN
|
||||||
|
INSTRUCT_ITEMS = sorted(_INSTRUCT_VALID_EN)
|
||||||
|
except Exception: # noqa: BLE001
|
||||||
|
INSTRUCT_ITEMS = []
|
||||||
|
|
||||||
logging.basicConfig(level=os.environ.get("OMNIVOICE_LOG_LEVEL", "INFO"))
|
logging.basicConfig(level=os.environ.get("OMNIVOICE_LOG_LEVEL", "INFO"))
|
||||||
log = logging.getLogger("omnivoice-api")
|
log = logging.getLogger("omnivoice-api")
|
||||||
|
|
||||||
CKPT = os.environ.get("OMNIVOICE_CKPT", "k2-fsa/OmniVoice")
|
CKPT = os.environ.get("OMNIVOICE_CKPT", "k2-fsa/OmniVoice")
|
||||||
VOICES_DIR = os.environ.get("OMNIVOICE_VOICES_DIR", "/app/voices")
|
VOICES_DIR = os.environ.get("OMNIVOICE_VOICES_DIR", "/app/voices")
|
||||||
ASR_MODEL = os.environ.get("OMNIVOICE_ASR_MODEL", "openai/whisper-large-v3-turbo")
|
ASR_MODEL = os.environ.get("OMNIVOICE_ASR_MODEL", "openai/whisper-large-v3-turbo")
|
||||||
NUM_STEP = int(os.environ.get("OMNIVOICE_NUM_STEP", "32"))
|
|
||||||
GUIDANCE = float(os.environ.get("OMNIVOICE_GUIDANCE_SCALE", "2.0"))
|
|
||||||
|
|
||||||
app = FastAPI(title="OmniVoice TTS (asset-engine wrapper)")
|
app = FastAPI(title="OmniVoice TTS (asset-engine wrapper)")
|
||||||
|
|
||||||
@@ -46,9 +66,24 @@ SR: int = 24000
|
|||||||
|
|
||||||
class SpeechRequest(BaseModel):
|
class SpeechRequest(BaseModel):
|
||||||
input: str
|
input: str
|
||||||
voice: str
|
# Voice source — at least one of voice (clone) / instruct (design) is required.
|
||||||
|
voice: Optional[str] = None # staged reference clip -> clone timbre
|
||||||
|
instruct: Optional[str] = None # free-text voice DESIGN / style
|
||||||
|
# Generation controls (defaults mirror the upstream demo).
|
||||||
|
language: Optional[str] = "Auto" # "Auto" -> auto-detect
|
||||||
|
speed: Optional[float] = None # 0.5–1.5; ignored if duration set
|
||||||
|
duration: Optional[float] = None # fixed seconds; overrides speed
|
||||||
|
num_step: int = 32 # 4–64 diffusion steps
|
||||||
|
guidance_scale: float = 2.0 # 0.0–4.0 CFG
|
||||||
|
denoise: bool = True
|
||||||
|
preprocess_prompt: bool = True
|
||||||
|
postprocess_output: bool = True
|
||||||
|
# Expert passthrough into OmniVoiceGenerationConfig (t_shift,
|
||||||
|
# layer_penalty_factor, position_temperature, class_temperature,
|
||||||
|
# audio_chunk_duration, audio_chunk_threshold). Unknown keys are dropped.
|
||||||
|
generation_overrides: Optional[Dict[str, Any]] = None
|
||||||
response_format: str = "wav"
|
response_format: str = "wav"
|
||||||
model: Optional[str] = None # ignored (single model); OpenAI-compat field
|
model: Optional[str] = None # ignored (single model); OpenAI-compat
|
||||||
|
|
||||||
|
|
||||||
@app.on_event("startup")
|
@app.on_event("startup")
|
||||||
@@ -70,14 +105,15 @@ def _load() -> None:
|
|||||||
log.info("voice ready: %s", name)
|
log.info("voice ready: %s", name)
|
||||||
except Exception as exc: # noqa: BLE001
|
except Exception as exc: # noqa: BLE001
|
||||||
log.warning("voice %s failed to load: %s", name, exc)
|
log.warning("voice %s failed to load: %s", name, exc)
|
||||||
log.info("%d voices loaded: %s", len(PROMPTS), sorted(PROMPTS))
|
log.info("%d voices loaded; %d languages", len(PROMPTS), len(LANGUAGES))
|
||||||
|
|
||||||
|
|
||||||
@app.get("/healthz")
|
@app.get("/healthz")
|
||||||
def healthz():
|
def healthz():
|
||||||
if MODEL is None or not PROMPTS:
|
if MODEL is None or not PROMPTS:
|
||||||
raise HTTPException(status_code=503, detail="not ready")
|
raise HTTPException(status_code=503, detail="not ready")
|
||||||
return {"status": "ok", "voices": len(PROMPTS), "sampling_rate": SR}
|
return {"status": "ok", "voices": len(PROMPTS), "languages": len(LANGUAGES),
|
||||||
|
"sampling_rate": SR}
|
||||||
|
|
||||||
|
|
||||||
@app.get("/v1/audio/voices")
|
@app.get("/v1/audio/voices")
|
||||||
@@ -85,6 +121,17 @@ def voices():
|
|||||||
return {"voices": sorted(PROMPTS.keys())}
|
return {"voices": sorted(PROMPTS.keys())}
|
||||||
|
|
||||||
|
|
||||||
|
@app.get("/v1/audio/languages")
|
||||||
|
def languages():
|
||||||
|
return {"languages": LANGUAGES}
|
||||||
|
|
||||||
|
|
||||||
|
@app.get("/v1/audio/instruct-items")
|
||||||
|
def instruct_items():
|
||||||
|
# Valid comma-separable voice-DESIGN attribute tags (English).
|
||||||
|
return {"instruct_items": INSTRUCT_ITEMS}
|
||||||
|
|
||||||
|
|
||||||
@app.post("/v1/audio/speech")
|
@app.post("/v1/audio/speech")
|
||||||
def speech(req: SpeechRequest):
|
def speech(req: SpeechRequest):
|
||||||
if MODEL is None:
|
if MODEL is None:
|
||||||
@@ -94,15 +141,47 @@ def speech(req: SpeechRequest):
|
|||||||
if req.response_format not in ("wav", "", None):
|
if req.response_format not in ("wav", "", None):
|
||||||
raise HTTPException(status_code=400, detail="only response_format=wav is supported")
|
raise HTTPException(status_code=400, detail="only response_format=wav is supported")
|
||||||
|
|
||||||
prompt = PROMPTS.get(req.voice)
|
has_instruct = bool(req.instruct and req.instruct.strip())
|
||||||
if prompt is None:
|
if not req.voice and not has_instruct:
|
||||||
raise HTTPException(
|
raise HTTPException(
|
||||||
status_code=404,
|
status_code=400,
|
||||||
detail=f"unknown voice '{req.voice}'; have {sorted(PROMPTS)}",
|
detail="provide a 'voice' (clone a staged reference) and/or 'instruct' (design a voice)",
|
||||||
)
|
)
|
||||||
|
|
||||||
gen = OmniVoiceGenerationConfig(num_step=NUM_STEP, guidance_scale=GUIDANCE, denoise=True)
|
cfg = {
|
||||||
out = MODEL.generate(text=req.input.strip(), voice_clone_prompt=prompt, generation_config=gen)
|
"num_step": req.num_step,
|
||||||
|
"guidance_scale": req.guidance_scale,
|
||||||
|
"denoise": req.denoise,
|
||||||
|
"preprocess_prompt": req.preprocess_prompt,
|
||||||
|
"postprocess_output": req.postprocess_output,
|
||||||
|
**(req.generation_overrides or {}),
|
||||||
|
}
|
||||||
|
gen = OmniVoiceGenerationConfig.from_dict(cfg)
|
||||||
|
|
||||||
|
kw: Dict[str, Any] = {"text": req.input.strip(), "generation_config": gen}
|
||||||
|
|
||||||
|
if req.voice:
|
||||||
|
prompt = PROMPTS.get(req.voice)
|
||||||
|
if prompt is None:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=404,
|
||||||
|
detail=f"unknown voice '{req.voice}'; have {sorted(PROMPTS)}",
|
||||||
|
)
|
||||||
|
kw["voice_clone_prompt"] = prompt
|
||||||
|
if has_instruct:
|
||||||
|
kw["instruct"] = req.instruct.strip()
|
||||||
|
if req.language and req.language != "Auto":
|
||||||
|
kw["language"] = req.language
|
||||||
|
if req.speed is not None:
|
||||||
|
kw["speed"] = req.speed
|
||||||
|
if req.duration is not None:
|
||||||
|
kw["duration"] = req.duration
|
||||||
|
|
||||||
|
try:
|
||||||
|
out = MODEL.generate(**kw)
|
||||||
|
except Exception as exc: # noqa: BLE001
|
||||||
|
raise HTTPException(status_code=400, detail=f"{type(exc).__name__}: {exc}")
|
||||||
|
|
||||||
audio = out[0] if isinstance(out, (list, tuple)) else out
|
audio = out[0] if isinstance(out, (list, tuple)) else out
|
||||||
audio = np.asarray(audio, dtype=np.float32)
|
audio = np.asarray(audio, dtype=np.float32)
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user