From 8eac39b6ce822e4e3a8ae2e5cbd8b02770715a05 Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Sun, 31 May 2026 11:18:46 -0700 Subject: [PATCH] catalog: add dia + csm service entries (status: down) + repro audits MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit dia (:8200) — OpenAI-compat /v1/audio/speech, seedable (not byte-exact), Apache-2.0 weights. Clean catalog fit; flip to ready after first exercise. csm (:8201) — OpenAI-compat, but NO seed + temperature-sampled = non-reproducible (contract's fix-before-adding case), catalogued by operator direction with a reproducibility caveat + gated-license warning; belongs at experimental once running. Fields read from each wrapper's API docs (2026-05-31), to confirm against live OpenAPI/Pydantic at deploy. No catalog_version bump (add-service = no bump). NOTE: services.schema.json is stale (pre-existing — 13 errors; live catalog uses lifecycle, schema predates it); regen via dump_schema.py. --- docs/asset-engine/services.yaml | 184 ++++++++++++++++++++++++++++++++ 1 file changed, 184 insertions(+) diff --git a/docs/asset-engine/services.yaml b/docs/asset-engine/services.yaml index e194624..6c7bb8b 100644 --- a/docs/asset-engine/services.yaml +++ b/docs/asset-engine/services.yaml @@ -1234,6 +1234,180 @@ services: /history. Until then, expose ComfyUI as an external link in the UI. User state at /worktank/comfyui/basedir/. + - id: dia + name: Dia / Dia2 TTS + description: > + Nari Labs' dialogue TTS — multi-speaker turn-taking in one pass with + [S1]/[S2] speaker tags and nonverbals (laughs)/(coughs)/(sighs). + Dia2 family (1B streaming, 2B high-quality) selectable on the host. + category: tts + version: 1 + status: down + host: irv-ml1 + lifecycle: + stack: dia + vram_gb: 7 + gpu_device_id: 0 + endpoint: http://10.100.79.3:8200/v1/audio/speech + method: POST + content_type: application/json + model: + id: nari-labs/Dia-1.6B + revision: null + image: local/dia:v1 + fields: + - name: input + type: textarea + label: Text ([S1]/[S2] dialogue + nonverbals) + required: true + max_length: 5000 + description: > + [S1]/[S2] tags mark speaker turns; nonverbals like (laughs), + (coughs), (sighs), (clears throat) go inline. For a clone voice, + prepend the reference transcript manually. + - name: model + type: select + options: [dia-1.6b] + default: dia-1.6b + description: > + Ignored by /v1/audio/speech (serves whatever checkpoint is loaded). + Switch Dia 1.6B / Dia2-1B / Dia2-2B via the wrapper config or Web UI. + - name: voice + type: select + label: Voice / mode + default: S1 + description: > + S1 | S2 | dialogue, or a predefined/clone reference filename under + /worktank/dia/reference_audio (e.g. "my_ref.wav"). + - name: response_format + type: select + options: [opus, wav] + default: opus + - name: speed + type: slider + min: 0.5 + max: 2.0 + step: 0.05 + default: 1.0 + description: Post-generation playback speed multiplier. + - name: seed + type: number + required: false + default: -1 + description: -1 = random; any integer for repeatable (not byte-exact) output. + response: + type: audio + mime_from_field: response_format + reproducibility: + seedable: true + deterministic: false + notes: > + Seed gives consistent voice/prosody but NOT byte-exact output — + upstream notes float-arithmetic variance across hardware/versions. + A richer custom /tts endpoint exposes cfg_scale/temperature/top_p/ + cfg_filter_top_k for finer control. + estimated_latency: + cold_start_s: 10 + warm_per_unit: "dialogue one-pass; ~realtime on the 3090" + license: "Apache-2.0 (Dia weights); MIT (devnen wrapper)" + notes: | + OpenAI-compat /v1/audio/speech plus a richer custom /tts (cfg_scale 3.0, + temperature 1.3, top_p 0.95, cfg_filter_top_k 35, max_tokens, + split_text/chunk_size). Fields read from devnen documentation.md + (2026-05-31); confirm against live OpenAPI/Pydantic at deploy. Deployed + to irv-ml1 but PARKED (asset-engine orchestrates start) — flip to ready + after the first successful generation. Pin DIA_SHA before build. + + - id: csm + name: Sesame CSM (conversational) + description: > + Sesame's Conversational Speech Model (CSM-1B) — context-aware speech + (Llama backbone + Mimi codec). Usable as plain TTS, but its edge is + cross-turn prosody for voice agents, not narration. + category: tts + version: 1 + status: down + host: irv-ml1 + lifecycle: + stack: csm + vram_gb: 8 + gpu_device_id: 0 + endpoint: http://10.100.79.3:8201/v1/audio/speech + method: POST + content_type: application/json + model: + id: sesame/csm-1b + revision: null + image: local/csm:v1 + fields: + - name: model + type: select + options: [csm-1b] + default: csm-1b + - name: input + type: textarea + label: Text + required: true + max_length: 5000 + - name: voice + type: select + label: Voice + default: alloy + description: > + alloy | echo | fable | onyx | nova | shimmer, or a cloned voice ID. + - name: response_format + type: select + options: [mp3, opus, aac, flac, wav] + default: mp3 + - name: speed + type: slider + min: 0.5 + max: 2.0 + step: 0.05 + default: 1.0 + - name: temperature + type: slider + min: 0.0 + max: 1.0 + step: 0.05 + default: 0.8 + description: Sampling temperature — higher = more variation. + - name: topk + type: number + required: false + description: Top-k sampling cutoff (1–100). + - name: max_audio_length_ms + type: number + required: false + default: 90000 + description: Max generated audio length in milliseconds. + response: + type: audio + mime_from_field: response_format + reproducibility: + seedable: false + deterministic: false + notes: > + ⚠️ No seed param and temperature-sampled → output varies run-to-run. + Per this contract that is normally a fix-before-adding bug; catalogued + by operator direction. asset-engine "regenerate/fork" will NOT + reproduce a prior CSM take. Clean fix = upstream seed support. + estimated_latency: + cold_start_s: 12 + warm_per_unit: "~realtime on the 3090" + license: "Sesame CSM license (gated, non-OSI); MIT (wrapper)" + license_warning: | + sesame/csm-1b is GATED under Sesame's own (non-OSI) license — accept + terms on HF and review before any commercial/redistribution use. Host + needs CSM_HF_TOKEN set before first start. + notes: | + Conversational speech model — context-aware prosody for voice agents, + not a narration reader. Fields read from the phildougherty wrapper + README/API (2026-05-31); confirm against live OpenAPI at deploy. + Deployed to irv-ml1 but PARKED; gated-model token required before + first start. Because it is non-deterministic, when started it belongs + at status: experimental (not ready). Pin CSM_SHA before build. + # Reproducibility audit — answers per service: (a) seedable, (b) model # deterministic without seed, (c) image tag mutable (security/reproducibility risk). reproducibility_audit: @@ -1304,3 +1478,13 @@ reproducibility_audit: model_deterministic: true image_tag_mutable: false notes: "Reproducibility requires persisting full workflow JSON + seed." + - service: dia + seedable: true + model_deterministic: false + image_tag_mutable: false + notes: "Seed gives consistent voice, not byte-exact (float variance). Build pin via DIA_SHA." + - service: csm + seedable: false + model_deterministic: false + image_tag_mutable: false + notes: "⚠️ No seed + temperature-sampled → non-reproducible. Catalogued experimental by operator direction; upstream seed support is the fix. Gated Sesame license."