catalog(asset-engine): add zonos-gateway service (full expressive dials)
New TTS service entry + reproducibility_audit row for the zonos-gateway wrapper (irv-ml1:8890) — the ext-tts-aliased OpenAI facade over Zonos. 23 fields across Text&voice / Expression / Prosody / Quality / Sampling / Output section groups; live voice dropdown from /v1/voices; response format pcm|wav (audition UI forces wav). Distinct from the older down zonos :8203 entry. jsonschema-validated.
This commit is contained in:
@@ -2312,6 +2312,289 @@ services:
|
|||||||
source for defaults/ranges. Adapter not yet deployed/verified — flip to
|
source for defaults/ranges. Adapter not yet deployed/verified — flip to
|
||||||
ready (or experimental) after the first successful generation through 8203.
|
ready (or experimental) after the first successful generation through 8203.
|
||||||
|
|
||||||
|
- id: zonos-gateway
|
||||||
|
name: Zonos Gateway (expressive)
|
||||||
|
description: >
|
||||||
|
OpenAI-compatible streaming facade over the Zonos engine (kept stock),
|
||||||
|
exposing Zonos's full expressive control surface: emotion directions
|
||||||
|
(happy / sad / angry / surprised) plus a valence/arousal axis pair,
|
||||||
|
classifier-free-guidance on emotion, accurate-vs-expressive mode,
|
||||||
|
speaking-rate conditioning, quality-metric targets, and the full
|
||||||
|
sampling stack — all reachable from named presets (neutral / warm /
|
||||||
|
excited / sad / intense / whisper) that seed the dials before explicit
|
||||||
|
overrides win. Streams s16le PCM (or a WAV wrapper) from
|
||||||
|
/v1/audio/speech. The LiteLLM `ext-tts` alias points at this gateway.
|
||||||
|
category: tts
|
||||||
|
version: 1
|
||||||
|
status: experimental
|
||||||
|
host: irv-ml1
|
||||||
|
lifecycle:
|
||||||
|
stack: zonos-gateway
|
||||||
|
vram_gb: 16
|
||||||
|
gpu_device_id: 0
|
||||||
|
endpoint: http://10.100.79.3:8890/v1/audio/speech
|
||||||
|
method: POST
|
||||||
|
content_type: application/json
|
||||||
|
streamable: true
|
||||||
|
model:
|
||||||
|
id: Zyphra/ZONOS2
|
||||||
|
revision: null
|
||||||
|
image: local/zonos-gateway:0.1.0
|
||||||
|
section_groups:
|
||||||
|
- id: basic
|
||||||
|
label: Text & voice
|
||||||
|
- id: expression
|
||||||
|
label: Expression
|
||||||
|
hint: Emotion conditioning. A preset seeds these; explicit dials win.
|
||||||
|
- id: prosody
|
||||||
|
label: Prosody
|
||||||
|
hint: Speaking-rate conditioning. Leave the enable toggles off for the model's native pacing.
|
||||||
|
- id: quality
|
||||||
|
label: Quality target
|
||||||
|
hint: Advanced — raw metric targets (LUFS, silence, bandlimit) Zonos buckets internally.
|
||||||
|
- id: sampling
|
||||||
|
label: Sampling
|
||||||
|
- id: output
|
||||||
|
label: Output
|
||||||
|
fields:
|
||||||
|
- name: input
|
||||||
|
type: textarea
|
||||||
|
label: Text to synthesize
|
||||||
|
section: basic
|
||||||
|
required: true
|
||||||
|
max_length: 5000
|
||||||
|
description: >
|
||||||
|
Text to speak. OpenAI-style `input` field; the gateway streams the
|
||||||
|
synthesized audio back.
|
||||||
|
- name: voice
|
||||||
|
type: select
|
||||||
|
label: Voice
|
||||||
|
section: basic
|
||||||
|
default: Cora
|
||||||
|
source_url: http://10.100.79.3:8890/v1/voices
|
||||||
|
source_jsonpath: $.voices[*].name
|
||||||
|
description: >
|
||||||
|
Predefined Zonos voice. Live-enumerated from /v1/voices so the list
|
||||||
|
auto-syncs with the deployed voice pack (Cora is the default).
|
||||||
|
- name: preset
|
||||||
|
type: select
|
||||||
|
label: Expressive preset
|
||||||
|
section: expression
|
||||||
|
required: false
|
||||||
|
options: [neutral, warm, excited, sad, intense, whisper]
|
||||||
|
default: neutral
|
||||||
|
description: >
|
||||||
|
Named expressive preset applied before explicit dials; any explicit
|
||||||
|
emotion/prosody/quality dial you set overrides the preset's value.
|
||||||
|
- name: emotion_enabled
|
||||||
|
type: bool
|
||||||
|
label: Enable emotion conditioning
|
||||||
|
section: expression
|
||||||
|
required: false
|
||||||
|
default: false
|
||||||
|
description: >
|
||||||
|
Turn emotion conditioning on. Required for the emotion_* dials to
|
||||||
|
bite — a preset that sets emotion turns this on for you.
|
||||||
|
- name: emotion_valence
|
||||||
|
type: slider
|
||||||
|
label: Valence
|
||||||
|
section: expression
|
||||||
|
min: -1.0
|
||||||
|
max: 1.0
|
||||||
|
step: 0.05
|
||||||
|
default: 0.0
|
||||||
|
description: Pleasantness axis. -1 negative, +1 positive.
|
||||||
|
- name: emotion_arousal
|
||||||
|
type: slider
|
||||||
|
label: Arousal
|
||||||
|
section: expression
|
||||||
|
min: -1.0
|
||||||
|
max: 1.0
|
||||||
|
step: 0.05
|
||||||
|
default: 0.0
|
||||||
|
description: Energy/activation axis. -1 calm, +1 excited.
|
||||||
|
- name: emotion_strength
|
||||||
|
type: slider
|
||||||
|
label: Emotion strength
|
||||||
|
section: expression
|
||||||
|
min: 0.0
|
||||||
|
max: 2.0
|
||||||
|
step: 0.05
|
||||||
|
default: 1.0
|
||||||
|
description: Overall scale on the emotion direction. 1.0 = as specified.
|
||||||
|
- name: emotion_cfg_scale
|
||||||
|
type: slider
|
||||||
|
label: Emotion CFG scale
|
||||||
|
section: expression
|
||||||
|
min: 1.0
|
||||||
|
max: 3.0
|
||||||
|
step: 0.1
|
||||||
|
default: 1.0
|
||||||
|
description: >
|
||||||
|
Classifier-free-guidance on emotion. 1.0 = off; >1 amplifies
|
||||||
|
expression.
|
||||||
|
- name: emotion_sliders
|
||||||
|
type: json
|
||||||
|
label: Per-emotion weights (advanced)
|
||||||
|
section: expression
|
||||||
|
optional: true
|
||||||
|
description: >
|
||||||
|
Advanced — per-emotion weight dict {happy|sad|angry|surprised: -1..1};
|
||||||
|
higher = stronger. Overrides the coarse valence/arousal directions
|
||||||
|
with explicit per-emotion control. Omit to use valence/arousal.
|
||||||
|
- name: accurate_mode
|
||||||
|
type: bool
|
||||||
|
label: Accurate mode
|
||||||
|
section: expression
|
||||||
|
required: false
|
||||||
|
default: true
|
||||||
|
description: >
|
||||||
|
true = faithful to the reference voice; false = more
|
||||||
|
expressive/looser.
|
||||||
|
- name: speaking_rate_enabled
|
||||||
|
type: bool
|
||||||
|
label: Enable speaking-rate conditioning
|
||||||
|
section: prosody
|
||||||
|
required: false
|
||||||
|
default: false
|
||||||
|
description: >
|
||||||
|
Turn speaking-rate conditioning on. Required for speed /
|
||||||
|
speaking_rate / speaking_rate_bucket to take effect.
|
||||||
|
- name: speed
|
||||||
|
type: slider
|
||||||
|
label: Speed (OpenAI-style)
|
||||||
|
section: prosody
|
||||||
|
min: 0.25
|
||||||
|
max: 4.0
|
||||||
|
step: 0.05
|
||||||
|
optional: true
|
||||||
|
description: >
|
||||||
|
OpenAI-style rate multiplier. Mapped to speaking_rate when no
|
||||||
|
explicit speaking_rate is given; auto-enables speaking-rate
|
||||||
|
conditioning. Omit to leave pacing native.
|
||||||
|
- name: speaking_rate
|
||||||
|
type: slider
|
||||||
|
label: Speaking rate (native)
|
||||||
|
section: prosody
|
||||||
|
min: 0.25
|
||||||
|
max: 4.0
|
||||||
|
step: 0.05
|
||||||
|
optional: true
|
||||||
|
description: >
|
||||||
|
Native speaking-rate multiplier. Overrides speed if both are sent.
|
||||||
|
Omit to leave pacing native.
|
||||||
|
- name: speaking_rate_bucket
|
||||||
|
type: slider
|
||||||
|
label: Speaking-rate bucket
|
||||||
|
section: prosody
|
||||||
|
min: 0
|
||||||
|
max: 7
|
||||||
|
step: 1
|
||||||
|
optional: true
|
||||||
|
description: >
|
||||||
|
Words/sec bucket index 0..7 (0 = 0-8 wps … 7 = 40+ wps). Coarser than
|
||||||
|
speaking_rate. Omit to leave pacing native.
|
||||||
|
- name: quality_enabled
|
||||||
|
type: bool
|
||||||
|
label: Enable quality-target conditioning
|
||||||
|
section: quality
|
||||||
|
required: false
|
||||||
|
default: true
|
||||||
|
description: >
|
||||||
|
Advanced — turn quality-target conditioning on (on by default in
|
||||||
|
Zonos). Gates quality_values.
|
||||||
|
- name: quality_values
|
||||||
|
type: json
|
||||||
|
label: Quality metric targets (advanced)
|
||||||
|
section: quality
|
||||||
|
optional: true
|
||||||
|
description: >
|
||||||
|
Advanced — raw metric targets Zonos buckets internally, e.g.
|
||||||
|
{lufs: -23, trailing_silence_s: 0.1}. Keys: lufs, estimated_snr,
|
||||||
|
max_pause, estimated_bandlimit_hz, leading_silence_s,
|
||||||
|
trailing_silence_s. Omit for Zonos's defaults.
|
||||||
|
- name: temperature
|
||||||
|
type: slider
|
||||||
|
section: sampling
|
||||||
|
min: 0.0
|
||||||
|
max: 2.0
|
||||||
|
step: 0.05
|
||||||
|
default: 1.15
|
||||||
|
description: Sampling temperature. Higher = more varied. Zonos default 1.15.
|
||||||
|
- name: top_p
|
||||||
|
type: slider
|
||||||
|
label: Top-p
|
||||||
|
section: sampling
|
||||||
|
min: 0.0
|
||||||
|
max: 1.0
|
||||||
|
step: 0.05
|
||||||
|
default: 0.0
|
||||||
|
description: Nucleus sampling cutoff. 0.0 = off (Zonos default).
|
||||||
|
- name: min_p
|
||||||
|
type: slider
|
||||||
|
label: Min-p
|
||||||
|
section: sampling
|
||||||
|
min: 0.0
|
||||||
|
max: 1.0
|
||||||
|
step: 0.01
|
||||||
|
default: 0.18
|
||||||
|
description: Min-p sampling floor. Zonos default 0.18.
|
||||||
|
- name: topk
|
||||||
|
type: number
|
||||||
|
label: Top-k
|
||||||
|
section: sampling
|
||||||
|
required: false
|
||||||
|
default: 106
|
||||||
|
description: Top-k sampling cutoff. Zonos default 106.
|
||||||
|
- name: seed
|
||||||
|
type: number
|
||||||
|
section: sampling
|
||||||
|
optional: true
|
||||||
|
description: >
|
||||||
|
RNG seed for reproducible sampling. Omit for a random seed. Pins the
|
||||||
|
sampler only; emotion/quality conditioning still varies subtly.
|
||||||
|
- name: max_tokens
|
||||||
|
type: number
|
||||||
|
label: Max audio tokens
|
||||||
|
section: sampling
|
||||||
|
required: false
|
||||||
|
max: 6144
|
||||||
|
description: >
|
||||||
|
Cap on generated audio tokens (upper bound; Zonos stops at
|
||||||
|
end-of-speech). Omit to let Zonos decide.
|
||||||
|
- name: response_format
|
||||||
|
type: select
|
||||||
|
label: Response format
|
||||||
|
section: output
|
||||||
|
options: [pcm, wav]
|
||||||
|
default: pcm
|
||||||
|
description: >
|
||||||
|
pcm = raw s16le stream (lowest latency, for API consumers); wav adds
|
||||||
|
a header. The stream-audition UI forces wav for the browser <audio>.
|
||||||
|
response:
|
||||||
|
type: audio
|
||||||
|
mime_from_field: response_format
|
||||||
|
reproducibility:
|
||||||
|
seedable: true
|
||||||
|
deterministic: false
|
||||||
|
seed_field: seed
|
||||||
|
notes: >
|
||||||
|
Temperature-sampled; seed pins the sampler but emotion/quality
|
||||||
|
conditioning still varies subtly run-to-run.
|
||||||
|
estimated_latency:
|
||||||
|
cold_start_s: 3
|
||||||
|
warm_per_unit: "streaming; first audio in a couple seconds warm, then near-realtime on the 3090"
|
||||||
|
license: Apache-2.0
|
||||||
|
notes: |
|
||||||
|
OpenAI-compatible streaming gateway (local/zonos-gateway:0.1.0) fronting a
|
||||||
|
stock Zonos engine on the 3090 (irv-ml1 device 0). The LiteLLM `ext-tts`
|
||||||
|
alias routes here. Fields mirror the gateway's /v1/dials schema (24 params;
|
||||||
|
the CATALOG-CONTRACT blessed source for defaults/ranges) and /v1/voices.
|
||||||
|
Deliberately omits repetition_window / repetition_penalty / codebooks — the
|
||||||
|
wrapper rejects them and they are the "70s of silence" footgun. Presets seed
|
||||||
|
the dials before explicit overrides win. New service (experimental) — flip to
|
||||||
|
ready after the first verified generation + browser audition through 8890.
|
||||||
|
|
||||||
# Reproducibility audit — answers per service: (a) seedable, (b) model
|
# Reproducibility audit — answers per service: (a) seedable, (b) model
|
||||||
# deterministic without seed, (c) image tag mutable (security/reproducibility risk).
|
# deterministic without seed, (c) image tag mutable (security/reproducibility risk).
|
||||||
reproducibility_audit:
|
reproducibility_audit:
|
||||||
@@ -2412,3 +2695,8 @@ reproducibility_audit:
|
|||||||
model_deterministic: true
|
model_deterministic: true
|
||||||
image_tag_mutable: false
|
image_tag_mutable: false
|
||||||
notes: "Adapter echoes the seed used (reproducibility.seed_field=seed). Byte-stable same-GPU; bf16 may drift cross-GPU. local/zonos-api:v1 built FROM local/zonos (pin ZONOS_SHA for true repro)."
|
notes: "Adapter echoes the seed used (reproducibility.seed_field=seed). Byte-stable same-GPU; bf16 may drift cross-GPU. local/zonos-api:v1 built FROM local/zonos (pin ZONOS_SHA for true repro)."
|
||||||
|
- service: zonos-gateway
|
||||||
|
seedable: true
|
||||||
|
model_deterministic: false
|
||||||
|
image_tag_mutable: true
|
||||||
|
notes: "Seed pins the sampler (reproducibility.seed_field=seed) but emotion/quality conditioning still varies subtly run-to-run — not byte-exact. Streaming (s16le PCM / WAV). Distinct from the `zonos` adapter: this is the OpenAI-compatible gateway on :8890 behind the LiteLLM `ext-tts` alias. image local/zonos-gateway:0.1.0 is tag-pinned + mutable — pin a digest for true repro."
|
||||||
|
|||||||
Reference in New Issue
Block a user