catalog(fish-s2): expose full control surface (was text + references only)
Fish-S2 rendered as an essentially blank form — only text + a references JSON blob — despite being the fleet's richest-control engine. Expose the real fish-speech ServeTTSRequest levers: temperature/top_p/repetition_penalty, latency (normal|balanced), seed, format (wav/mp3/opus), + advanced cloning (references/reference_id) and chunking (max_new_tokens/chunk_length/normalize/ use_memory_cache). Defaults from upstream schema.py, verified live 2026-06-01 (no /openapi.json; Kui server). Sections basic/sampling/advanced. Also corrected: seedable false->true (/tts has seed); cold_start 8s->240s (measured compile warmup); vram 4->10GB (~9GB BF16 weights); dropped the misleading '~150ms TTFB' for honest ~realtime throughput. version 1->2.
This commit is contained in:
+120
-16
@@ -616,13 +616,14 @@ services:
|
|||||||
name: Fish Audio S2-Pro
|
name: Fish Audio S2-Pro
|
||||||
description: >
|
description: >
|
||||||
Fishaudio's richest-paralinguistic English TTS (15,000+ inline tags).
|
Fishaudio's richest-paralinguistic English TTS (15,000+ inline tags).
|
||||||
Trained 10M+ hours, dual-AR, ~150ms streaming TTFB. Released March 2026.
|
Trained 10M+ hours, dual-AR. Released March 2026. Heavy: ~240s compile
|
||||||
|
warmup on cold start, ~realtime throughput once warm.
|
||||||
category: tts
|
category: tts
|
||||||
version: 1
|
version: 2
|
||||||
host: irv-ml1
|
host: irv-ml1
|
||||||
lifecycle:
|
lifecycle:
|
||||||
stack: fish-s2
|
stack: fish-s2
|
||||||
vram_gb: 4
|
vram_gb: 10
|
||||||
gpu_device_id: 1
|
gpu_device_id: 1
|
||||||
endpoint: http://10.100.79.3:8195/v1/tts
|
endpoint: http://10.100.79.3:8195/v1/tts
|
||||||
method: POST
|
method: POST
|
||||||
@@ -631,10 +632,19 @@ services:
|
|||||||
id: fishaudio/s2-pro
|
id: fishaudio/s2-pro
|
||||||
revision: null
|
revision: null
|
||||||
image: local/fish-s2:v1
|
image: local/fish-s2:v1
|
||||||
|
section_groups:
|
||||||
|
- id: basic
|
||||||
|
label: Text & output
|
||||||
|
- id: sampling
|
||||||
|
label: Sampling & latency
|
||||||
|
hint: fish-speech ServeTTSRequest defaults. Lower temperature/top_p for steadier reads; latency=balanced for lower latency.
|
||||||
|
- id: advanced
|
||||||
|
label: Advanced & voice cloning
|
||||||
fields:
|
fields:
|
||||||
- name: text
|
- name: text
|
||||||
type: textarea
|
type: textarea
|
||||||
label: Text (with optional [tags])
|
label: Text (with optional [tags])
|
||||||
|
section: basic
|
||||||
required: true
|
required: true
|
||||||
max_length: 5000
|
max_length: 5000
|
||||||
description: >
|
description: >
|
||||||
@@ -644,29 +654,123 @@ services:
|
|||||||
without emotional context — pair with an emotion tag. Example:
|
without emotional context — pair with an emotion tag. Example:
|
||||||
`[excited] Welcome back! [whisper] I have a secret. [laugh]
|
`[excited] Welcome back! [whisper] I have a secret. [laugh]
|
||||||
Just kidding.`
|
Just kidding.`
|
||||||
|
- name: format
|
||||||
|
type: select
|
||||||
|
section: basic
|
||||||
|
options: [wav, mp3, opus]
|
||||||
|
default: wav
|
||||||
|
description: >
|
||||||
|
Output container (upstream also accepts raw `pcm`, omitted here).
|
||||||
|
wav + mp3 verified live 2026-06-01.
|
||||||
|
- name: temperature
|
||||||
|
type: slider
|
||||||
|
section: sampling
|
||||||
|
min: 0.1
|
||||||
|
max: 1.0
|
||||||
|
step: 0.05
|
||||||
|
default: 0.8
|
||||||
|
description: Sampling temperature; lower = steadier. fish-speech default 0.8.
|
||||||
|
- name: top_p
|
||||||
|
type: slider
|
||||||
|
section: sampling
|
||||||
|
min: 0.1
|
||||||
|
max: 1.0
|
||||||
|
step: 0.05
|
||||||
|
default: 0.8
|
||||||
|
description: Nucleus sampling cutoff. fish-speech default 0.8.
|
||||||
|
- name: repetition_penalty
|
||||||
|
type: slider
|
||||||
|
section: sampling
|
||||||
|
min: 0.9
|
||||||
|
max: 2.0
|
||||||
|
step: 0.05
|
||||||
|
default: 1.1
|
||||||
|
description: >
|
||||||
|
Penalizes repeats; raise if you hear stutters/loops. fish-speech
|
||||||
|
default 1.1.
|
||||||
|
- name: latency
|
||||||
|
type: select
|
||||||
|
section: sampling
|
||||||
|
options: [normal, balanced]
|
||||||
|
default: normal
|
||||||
|
description: >
|
||||||
|
`balanced` trades a little quality for lower latency — useful given
|
||||||
|
S2-Pro's ~realtime throughput. fish-speech default `normal`.
|
||||||
|
- name: seed
|
||||||
|
type: number
|
||||||
|
section: sampling
|
||||||
|
required: false
|
||||||
|
description: Omit for a random seed; a fixed integer repeats a take.
|
||||||
- name: references
|
- name: references
|
||||||
type: json
|
type: json
|
||||||
label: Voice References
|
label: Voice References (clone)
|
||||||
|
section: advanced
|
||||||
optional: true
|
optional: true
|
||||||
description: >
|
description: >
|
||||||
Array of {audio: "/app/references/<file>.wav", text: "transcript"}.
|
Array of {audio: "/app/references/<file>.wav", text: "transcript"} for
|
||||||
Files under /worktank/fish-s2/references/ on host.
|
zero-shot cloning. Files under /worktank/fish-s2/references/ on host.
|
||||||
|
- name: reference_id
|
||||||
|
type: text
|
||||||
|
label: Stored reference id
|
||||||
|
section: advanced
|
||||||
|
optional: true
|
||||||
|
description: >
|
||||||
|
Alternative to inline `references`: id of a server-stored reference
|
||||||
|
set (a folder under the server's references/). Leave blank to use
|
||||||
|
`references` or the default voice.
|
||||||
|
- name: max_new_tokens
|
||||||
|
type: number
|
||||||
|
section: advanced
|
||||||
|
required: false
|
||||||
|
default: 1024
|
||||||
|
description: Max audio tokens per chunk. fish-speech default 1024.
|
||||||
|
- name: chunk_length
|
||||||
|
type: slider
|
||||||
|
section: advanced
|
||||||
|
min: 100
|
||||||
|
max: 1000
|
||||||
|
step: 10
|
||||||
|
default: 200
|
||||||
|
description: Target chunk length in chars when splitting. fish-speech default 200.
|
||||||
|
- name: normalize
|
||||||
|
type: bool
|
||||||
|
section: advanced
|
||||||
|
default: true
|
||||||
|
description: Text normalization (numbers, punctuation). fish-speech default on.
|
||||||
|
- name: use_memory_cache
|
||||||
|
type: select
|
||||||
|
section: advanced
|
||||||
|
options: ["off", "on"]
|
||||||
|
default: "off"
|
||||||
|
description: >
|
||||||
|
`on` caches the reference encoding to speed up repeated calls with the
|
||||||
|
same voice. fish-speech default off.
|
||||||
response:
|
response:
|
||||||
type: audio
|
type: audio
|
||||||
mime: audio/wav
|
mime_from_field: format
|
||||||
reproducibility:
|
reproducibility:
|
||||||
seedable: false
|
seedable: true
|
||||||
deterministic: true
|
deterministic: false
|
||||||
|
seed_field: seed
|
||||||
|
notes: >
|
||||||
|
/v1/tts exposes `seed` (omit = random); a fixed seed repeats a take.
|
||||||
|
Temperature/top_p sampled → not guaranteed byte-exact. (Prior catalog
|
||||||
|
claimed no seed support — corrected against the upstream ServeTTSRequest
|
||||||
|
schema + a live generation 2026-06-01.)
|
||||||
estimated_latency:
|
estimated_latency:
|
||||||
cold_start_s: 8
|
cold_start_s: 240
|
||||||
warm_per_unit: "~150ms TTFB"
|
warm_per_unit: "~realtime throughput on the A6000 (heavy). latency=balanced trades quality for lower latency."
|
||||||
license: "Research-only (Fishaudio terms; non-commercial)"
|
license: "Research-only (Fishaudio terms; non-commercial)"
|
||||||
license_warning: |
|
license_warning: |
|
||||||
Research/internal only. Not clear for commercial use. Same flag as
|
Research/internal only. Not clear for commercial use. Same flag as
|
||||||
Voxtral and SAO — UI must surface when output destined for products.
|
Voxtral and SAO — UI must surface when output destined for products.
|
||||||
notes: |
|
notes: |
|
||||||
NOT OpenAI-compatible: only /v1/tts. No /v1/audio/voices, no /v1/models.
|
NOT OpenAI-compatible: only /v1/tts (no /v1/audio/voices, /v1/models, or
|
||||||
Voice discovery is manual (host filesystem).
|
/openapi.json). Voice discovery is manual (host filesystem). Full control
|
||||||
|
surface (temperature/top_p/repetition_penalty/latency/seed/format + the
|
||||||
|
cloning fields) sourced from fish-speech's ServeTTSRequest schema and
|
||||||
|
verified live 2026-06-01; prior catalog exposed only text + references.
|
||||||
|
HEAVY: ~240s torch.compile warmup measured on cold start; budget the wait.
|
||||||
|
|
||||||
- id: kyutai-tts
|
- id: kyutai-tts
|
||||||
name: Kyutai TTS 1.6B EN/FR
|
name: Kyutai TTS 1.6B EN/FR
|
||||||
@@ -1860,10 +1964,10 @@ reproducibility_audit:
|
|||||||
image_tag_mutable: false
|
image_tag_mutable: false
|
||||||
notes: "instruct field broken for English; XML tags only."
|
notes: "instruct field broken for English; XML tags only."
|
||||||
- service: fish-s2
|
- service: fish-s2
|
||||||
seedable: false
|
seedable: true
|
||||||
model_deterministic: true
|
model_deterministic: false
|
||||||
image_tag_mutable: false
|
image_tag_mutable: false
|
||||||
notes: "Research-only license — non-commercial."
|
notes: "/v1/tts exposes seed (omit=random); temperature/top_p sampled, not byte-exact. Research-only license — non-commercial."
|
||||||
- service: kyutai-tts
|
- service: kyutai-tts
|
||||||
seedable: false
|
seedable: false
|
||||||
model_deterministic: true
|
model_deterministic: true
|
||||||
|
|||||||
Reference in New Issue
Block a user