catalog: add dia + csm service entries (status: down) + repro audits
dia (:8200) — OpenAI-compat /v1/audio/speech, seedable (not byte-exact), Apache-2.0 weights. Clean catalog fit; flip to ready after first exercise. csm (:8201) — OpenAI-compat, but NO seed + temperature-sampled = non-reproducible (contract's fix-before-adding case), catalogued by operator direction with a reproducibility caveat + gated-license warning; belongs at experimental once running. Fields read from each wrapper's API docs (2026-05-31), to confirm against live OpenAPI/Pydantic at deploy. No catalog_version bump (add-service = no bump). NOTE: services.schema.json is stale (pre-existing — 13 errors; live catalog uses lifecycle, schema predates it); regen via dump_schema.py.
This commit is contained in:
@@ -1234,6 +1234,180 @@ services:
|
|||||||
/history. Until then, expose ComfyUI as an external link in the UI.
|
/history. Until then, expose ComfyUI as an external link in the UI.
|
||||||
User state at /worktank/comfyui/basedir/.
|
User state at /worktank/comfyui/basedir/.
|
||||||
|
|
||||||
|
- id: dia
|
||||||
|
name: Dia / Dia2 TTS
|
||||||
|
description: >
|
||||||
|
Nari Labs' dialogue TTS — multi-speaker turn-taking in one pass with
|
||||||
|
[S1]/[S2] speaker tags and nonverbals (laughs)/(coughs)/(sighs).
|
||||||
|
Dia2 family (1B streaming, 2B high-quality) selectable on the host.
|
||||||
|
category: tts
|
||||||
|
version: 1
|
||||||
|
status: down
|
||||||
|
host: irv-ml1
|
||||||
|
lifecycle:
|
||||||
|
stack: dia
|
||||||
|
vram_gb: 7
|
||||||
|
gpu_device_id: 0
|
||||||
|
endpoint: http://10.100.79.3:8200/v1/audio/speech
|
||||||
|
method: POST
|
||||||
|
content_type: application/json
|
||||||
|
model:
|
||||||
|
id: nari-labs/Dia-1.6B
|
||||||
|
revision: null
|
||||||
|
image: local/dia:v1
|
||||||
|
fields:
|
||||||
|
- name: input
|
||||||
|
type: textarea
|
||||||
|
label: Text ([S1]/[S2] dialogue + nonverbals)
|
||||||
|
required: true
|
||||||
|
max_length: 5000
|
||||||
|
description: >
|
||||||
|
[S1]/[S2] tags mark speaker turns; nonverbals like (laughs),
|
||||||
|
(coughs), (sighs), (clears throat) go inline. For a clone voice,
|
||||||
|
prepend the reference transcript manually.
|
||||||
|
- name: model
|
||||||
|
type: select
|
||||||
|
options: [dia-1.6b]
|
||||||
|
default: dia-1.6b
|
||||||
|
description: >
|
||||||
|
Ignored by /v1/audio/speech (serves whatever checkpoint is loaded).
|
||||||
|
Switch Dia 1.6B / Dia2-1B / Dia2-2B via the wrapper config or Web UI.
|
||||||
|
- name: voice
|
||||||
|
type: select
|
||||||
|
label: Voice / mode
|
||||||
|
default: S1
|
||||||
|
description: >
|
||||||
|
S1 | S2 | dialogue, or a predefined/clone reference filename under
|
||||||
|
/worktank/dia/reference_audio (e.g. "my_ref.wav").
|
||||||
|
- name: response_format
|
||||||
|
type: select
|
||||||
|
options: [opus, wav]
|
||||||
|
default: opus
|
||||||
|
- name: speed
|
||||||
|
type: slider
|
||||||
|
min: 0.5
|
||||||
|
max: 2.0
|
||||||
|
step: 0.05
|
||||||
|
default: 1.0
|
||||||
|
description: Post-generation playback speed multiplier.
|
||||||
|
- name: seed
|
||||||
|
type: number
|
||||||
|
required: false
|
||||||
|
default: -1
|
||||||
|
description: -1 = random; any integer for repeatable (not byte-exact) output.
|
||||||
|
response:
|
||||||
|
type: audio
|
||||||
|
mime_from_field: response_format
|
||||||
|
reproducibility:
|
||||||
|
seedable: true
|
||||||
|
deterministic: false
|
||||||
|
notes: >
|
||||||
|
Seed gives consistent voice/prosody but NOT byte-exact output —
|
||||||
|
upstream notes float-arithmetic variance across hardware/versions.
|
||||||
|
A richer custom /tts endpoint exposes cfg_scale/temperature/top_p/
|
||||||
|
cfg_filter_top_k for finer control.
|
||||||
|
estimated_latency:
|
||||||
|
cold_start_s: 10
|
||||||
|
warm_per_unit: "dialogue one-pass; ~realtime on the 3090"
|
||||||
|
license: "Apache-2.0 (Dia weights); MIT (devnen wrapper)"
|
||||||
|
notes: |
|
||||||
|
OpenAI-compat /v1/audio/speech plus a richer custom /tts (cfg_scale 3.0,
|
||||||
|
temperature 1.3, top_p 0.95, cfg_filter_top_k 35, max_tokens,
|
||||||
|
split_text/chunk_size). Fields read from devnen documentation.md
|
||||||
|
(2026-05-31); confirm against live OpenAPI/Pydantic at deploy. Deployed
|
||||||
|
to irv-ml1 but PARKED (asset-engine orchestrates start) — flip to ready
|
||||||
|
after the first successful generation. Pin DIA_SHA before build.
|
||||||
|
|
||||||
|
- id: csm
|
||||||
|
name: Sesame CSM (conversational)
|
||||||
|
description: >
|
||||||
|
Sesame's Conversational Speech Model (CSM-1B) — context-aware speech
|
||||||
|
(Llama backbone + Mimi codec). Usable as plain TTS, but its edge is
|
||||||
|
cross-turn prosody for voice agents, not narration.
|
||||||
|
category: tts
|
||||||
|
version: 1
|
||||||
|
status: down
|
||||||
|
host: irv-ml1
|
||||||
|
lifecycle:
|
||||||
|
stack: csm
|
||||||
|
vram_gb: 8
|
||||||
|
gpu_device_id: 0
|
||||||
|
endpoint: http://10.100.79.3:8201/v1/audio/speech
|
||||||
|
method: POST
|
||||||
|
content_type: application/json
|
||||||
|
model:
|
||||||
|
id: sesame/csm-1b
|
||||||
|
revision: null
|
||||||
|
image: local/csm:v1
|
||||||
|
fields:
|
||||||
|
- name: model
|
||||||
|
type: select
|
||||||
|
options: [csm-1b]
|
||||||
|
default: csm-1b
|
||||||
|
- name: input
|
||||||
|
type: textarea
|
||||||
|
label: Text
|
||||||
|
required: true
|
||||||
|
max_length: 5000
|
||||||
|
- name: voice
|
||||||
|
type: select
|
||||||
|
label: Voice
|
||||||
|
default: alloy
|
||||||
|
description: >
|
||||||
|
alloy | echo | fable | onyx | nova | shimmer, or a cloned voice ID.
|
||||||
|
- name: response_format
|
||||||
|
type: select
|
||||||
|
options: [mp3, opus, aac, flac, wav]
|
||||||
|
default: mp3
|
||||||
|
- name: speed
|
||||||
|
type: slider
|
||||||
|
min: 0.5
|
||||||
|
max: 2.0
|
||||||
|
step: 0.05
|
||||||
|
default: 1.0
|
||||||
|
- name: temperature
|
||||||
|
type: slider
|
||||||
|
min: 0.0
|
||||||
|
max: 1.0
|
||||||
|
step: 0.05
|
||||||
|
default: 0.8
|
||||||
|
description: Sampling temperature — higher = more variation.
|
||||||
|
- name: topk
|
||||||
|
type: number
|
||||||
|
required: false
|
||||||
|
description: Top-k sampling cutoff (1–100).
|
||||||
|
- name: max_audio_length_ms
|
||||||
|
type: number
|
||||||
|
required: false
|
||||||
|
default: 90000
|
||||||
|
description: Max generated audio length in milliseconds.
|
||||||
|
response:
|
||||||
|
type: audio
|
||||||
|
mime_from_field: response_format
|
||||||
|
reproducibility:
|
||||||
|
seedable: false
|
||||||
|
deterministic: false
|
||||||
|
notes: >
|
||||||
|
⚠️ No seed param and temperature-sampled → output varies run-to-run.
|
||||||
|
Per this contract that is normally a fix-before-adding bug; catalogued
|
||||||
|
by operator direction. asset-engine "regenerate/fork" will NOT
|
||||||
|
reproduce a prior CSM take. Clean fix = upstream seed support.
|
||||||
|
estimated_latency:
|
||||||
|
cold_start_s: 12
|
||||||
|
warm_per_unit: "~realtime on the 3090"
|
||||||
|
license: "Sesame CSM license (gated, non-OSI); MIT (wrapper)"
|
||||||
|
license_warning: |
|
||||||
|
sesame/csm-1b is GATED under Sesame's own (non-OSI) license — accept
|
||||||
|
terms on HF and review before any commercial/redistribution use. Host
|
||||||
|
needs CSM_HF_TOKEN set before first start.
|
||||||
|
notes: |
|
||||||
|
Conversational speech model — context-aware prosody for voice agents,
|
||||||
|
not a narration reader. Fields read from the phildougherty wrapper
|
||||||
|
README/API (2026-05-31); confirm against live OpenAPI at deploy.
|
||||||
|
Deployed to irv-ml1 but PARKED; gated-model token required before
|
||||||
|
first start. Because it is non-deterministic, when started it belongs
|
||||||
|
at status: experimental (not ready). Pin CSM_SHA before build.
|
||||||
|
|
||||||
# Reproducibility audit — answers per service: (a) seedable, (b) model
|
# Reproducibility audit — answers per service: (a) seedable, (b) model
|
||||||
# deterministic without seed, (c) image tag mutable (security/reproducibility risk).
|
# deterministic without seed, (c) image tag mutable (security/reproducibility risk).
|
||||||
reproducibility_audit:
|
reproducibility_audit:
|
||||||
@@ -1304,3 +1478,13 @@ reproducibility_audit:
|
|||||||
model_deterministic: true
|
model_deterministic: true
|
||||||
image_tag_mutable: false
|
image_tag_mutable: false
|
||||||
notes: "Reproducibility requires persisting full workflow JSON + seed."
|
notes: "Reproducibility requires persisting full workflow JSON + seed."
|
||||||
|
- service: dia
|
||||||
|
seedable: true
|
||||||
|
model_deterministic: false
|
||||||
|
image_tag_mutable: false
|
||||||
|
notes: "Seed gives consistent voice, not byte-exact (float variance). Build pin via DIA_SHA."
|
||||||
|
- service: csm
|
||||||
|
seedable: false
|
||||||
|
model_deterministic: false
|
||||||
|
image_tag_mutable: false
|
||||||
|
notes: "⚠️ No seed + temperature-sampled → non-reproducible. Catalogued experimental by operator direction; upstream seed support is the fix. Gated Sesame license."
|
||||||
|
|||||||
Reference in New Issue
Block a user