catalog: ace-step v5 — defaults audit against upstream Gradio UI

asset_engine consumer audited the entire ace-step entry's defaults
and slider ranges against acestep/ui/components.py (althing thread
01KRCN0SHP9YJGQD58EE95DC5P). The catalog had been authored from
documentation rather than from source; ten defaults were wrong and
several slider ranges were either too narrow or impractically wide.

Defaults changed (catalog -> upstream-authoritative):
  infer_step                 20    -> 60
  guidance_scale             7.5   -> 15.0
  cfg_type                   cfg   -> apg
  omega_scale                0.5   -> 10.0
  guidance_interval          0.0   -> 0.5
  guidance_interval_decay    1.0   -> 0.0
  min_guidance_scale         1.0   -> 3.0
  use_erg_tag                false -> true
  use_erg_diffusion          false -> true
  actual_seeds               [42]  -> []   (random per call)

Slider ranges adopted from upstream where reasonable; bounded
locally where upstream's range is so wide it's unusable as a UI
slider:
  guidance_scale         [1.0, 15.0]   -> [0.0, 30.0]   (upstream)
  guidance_scale_text    [0.0, 15.0]   -> [0.0, 10.0]   (upstream)
  guidance_scale_lyric   [0.0, 15.0]   -> [0.0, 10.0]   (upstream)
  lora_weight            [0.0, 2.0]    -> [-3.0, 3.0]   (upstream)
  audio_duration         [5.0, 600.0]  -> [5.0, 240.0]  (upstream max)
  omega_scale            [0.0, 1.0]    -> [-10.0, 30.0] (UI bound; upstream is [-100, 100])
  min_guidance_scale     [0.0, 10.0]   -> [0.0, 20.0]   (UI bound; upstream is [0, 200])

Verified empty-string actual_seeds path against the live pipeline
source: pipeline_ace_step.py:set_seeds() falls through to
torch.randint when manual_seeds is "" (string, no comma, not all
digits). Smoked end-to-end: HTTP 200 in 11s, real WAV bytes back.

Reproducibility gap honestly documented in the entry's
reproducibility.notes and the actual_seeds field description: with
the new default `actual_seeds: []`, the wrapper rolls a random seed
inside the pipeline but doesn't capture or surface the chosen seed
back through the response. Default-defaulted assets cannot be
regenerated bit-exact; users requiring reproducibility must set
actual_seeds explicitly. Wrapper enhancement to surface the chosen
seed via X-Actual-Seeds header + a CatalogResponse.header_accessories
schema field is the planned fix.

ace-step bumped version 4 -> 5. catalog_version stays at 1 (no
schema changes).

Also added a "source-of-truth precedence" subsection to
CATALOG-CONTRACT.md's service-authoring notes, codifying the
read-order (Pydantic model > handler/pipeline code > Gradio UI >
README). Three ace-step bugs in three rounds (missing field, wrong
enums, stranded bytes, wrong defaults — really four) all share the
same root cause: catalog authored from doc surfaces that lie by
omission.
This commit is contained in:
vh
2026-05-11 16:17:25 -07:00
parent f8ecc6c047
commit f020049769
2 changed files with 89 additions and 33 deletions
+17 -4
View File
@@ -144,10 +144,23 @@ the consumer. This list is what we audit when planning a
When **adding** a service to the catalog: When **adding** a service to the catalog:
- The Pydantic model in code is the ground truth for `fields:`. Read - **Source-of-truth precedence.** OpenAPI is convenient but often lies
the actual server code, not the README; READMEs rot. The by omission (bare `string` types where the dispatch chain enforces
reproducibility audit at the bottom of `services.yaml` MUST get a a specific enum, missing required-only-in-the-handler fields, etc).
matching entry. The actual ground truth is, in order:
1. The Pydantic input model — for field shape, types, required-ness.
2. The handler / pipeline code — for enum dispatches and runtime
validation.
3. The Gradio UI / client code (when present) — for blessed
defaults and slider ranges, since the model authors picked
these for human-facing UX.
4. The README — least reliable; rots fastest.
Every catalog change should cite which of these was read. "OpenAPI
said X" is not enough on its own — past audits have caught three
bugs in a single service (ace-step) that all looked correct in
OpenAPI.
- The reproducibility audit at the bottom of `services.yaml` MUST get
a matching entry for any new service.
- If a service can't produce reproducible output (no seed AND - If a service can't produce reproducible output (no seed AND
non-deterministic model), that's a **bug in the service** — fix it non-deterministic model), that's a **bug in the service** — fix it
there before adding to the catalog. The asset-engine relies on there before adding to the catalog. The asset-engine relies on
+72 -29
View File
@@ -752,7 +752,7 @@ services:
Apache-2.0 hybrid diffusion+LLM music generation. Multi-minute lyric-aware Apache-2.0 hybrid diffusion+LLM music generation. Multi-minute lyric-aware
songs with vocals + instrumentation. songs with vocals + instrumentation.
category: music category: music
version: 4 version: 5
host: irv-ml1 host: irv-ml1
endpoint: http://10.100.79.3:8210/generate endpoint: http://10.100.79.3:8210/generate
method: POST method: POST
@@ -800,20 +800,29 @@ services:
- name: audio_duration - name: audio_duration
type: slider type: slider
min: 5.0 min: 5.0
max: 600.0 max: 240.0
default: 30.0 default: 30.0
label: Duration (seconds) label: Duration (seconds)
section: basic section: basic
description: >
Upstream caps at 240s (the model's training horizon). Lower bound
5s is our choice — upstream uses -1 as a "random duration" sentinel
which is hostile UX for a slider. Default 30s also kept (upstream
uses -1; explicit 30 is the better first-time-user experience).
- name: infer_step - name: infer_step
type: number type: number
default: 20 default: 60
label: Inference Steps label: Inference Steps
section: generation section: generation
description: >
Upstream Gradio default is 60 (matches benchmark numbers in the
README). Lower values (20-30) are useful for "preview" passes;
higher (80-100) marginal returns.
- name: guidance_scale - name: guidance_scale
type: slider type: slider
min: 1.0 min: 0.0
max: 15.0 max: 30.0
default: 7.5 default: 15.0
section: generation section: generation
- name: scheduler_type - name: scheduler_type
type: select type: select
@@ -827,44 +836,67 @@ services:
- name: cfg_type - name: cfg_type
type: select type: select
options: [apg, cfg, cfg_star] options: [apg, cfg, cfg_star]
default: cfg default: apg
section: generation section: generation
description: > description: >
Classifier-free guidance variant. `cfg` is the standard formulation; Classifier-free guidance variant. Upstream Gradio default is `apg`
`apg` (adaptive projected guidance) and `cfg_star` are advanced (adaptive projected guidance); `cfg` is the standard SD-style
alternatives — see the upstream pipeline source for trade-offs. formulation; `cfg_star` is an advanced alternative. See the
upstream pipeline source for trade-offs.
- name: omega_scale - name: omega_scale
type: slider type: slider
min: 0.0 min: -10.0
max: 1.0 max: 30.0
default: 0.5 default: 10.0
section: generation section: generation
description: >
Upstream technically allows [-100, 100] but values that wide are
unusable as a slider. UI-bounded to [-10, 30] which covers the
typical zone with headroom. Hit the API directly for extremes.
- name: actual_seeds - name: actual_seeds
type: json type: json
label: Seeds label: Seeds (empty = random)
default: [42] default: []
section: generation section: generation
description: >
Empty list = wrapper sends empty string to pipeline = pipeline
picks a random seed per batch element. Explicit seeds (e.g. [42]
or [42, 137, 9999]) for reproducibility.
REPRODUCIBILITY GAP (queued for follow-up): the pipeline returns
the chosen seed in its result dict, but our wrapper currently
throws it away. Assets generated with the default `[]` cannot
currently be regenerated. Workaround: set actual_seeds explicitly
when reproducibility matters. Wrapper enhancement to surface
random-resolved seeds via X-Actual-Seeds header + a catalog
schema field for header→accessory capture is the planned fix.
- name: guidance_interval - name: guidance_interval
type: slider type: slider
min: 0.0 min: 0.0
max: 1.0 max: 1.0
default: 0.0 default: 0.5
section: conditioning section: conditioning
- name: guidance_interval_decay - name: guidance_interval_decay
type: slider type: slider
min: 0.0 min: 0.0
max: 1.0 max: 1.0
default: 1.0 default: 0.0
section: conditioning section: conditioning
description: >
Upstream Gradio default is 0.0 (no decay). Catalog v4 had this at
1.0 (full decay) — wrong; produced under-conditioned outputs.
- name: min_guidance_scale - name: min_guidance_scale
type: slider type: slider
min: 0.0 min: 0.0
max: 10.0 max: 20.0
default: 1.0 default: 3.0
section: conditioning section: conditioning
description: >
Upstream technically allows up to 200; UI-bounded to 20 (covers
the typical zone). Hit the API directly for extremes.
- name: use_erg_tag - name: use_erg_tag
type: bool type: bool
default: false default: true
section: conditioning section: conditioning
- name: use_erg_lyric - name: use_erg_lyric
type: bool type: bool
@@ -872,7 +904,7 @@ services:
section: conditioning section: conditioning
- name: use_erg_diffusion - name: use_erg_diffusion
type: bool type: bool
default: false default: true
section: conditioning section: conditioning
- name: oss_steps - name: oss_steps
type: json type: json
@@ -881,13 +913,13 @@ services:
- name: guidance_scale_text - name: guidance_scale_text
type: slider type: slider
min: 0.0 min: 0.0
max: 15.0 max: 10.0
default: 0.0 default: 0.0
section: conditioning section: conditioning
- name: guidance_scale_lyric - name: guidance_scale_lyric
type: slider type: slider
min: 0.0 min: 0.0
max: 15.0 max: 10.0
default: 0.0 default: 0.0
section: conditioning section: conditioning
- name: audio2audio_enable - name: audio2audio_enable
@@ -912,10 +944,13 @@ services:
section: lora section: lora
- name: lora_weight - name: lora_weight
type: slider type: slider
min: 0.0 min: -3.0
max: 2.0 max: 3.0
default: 1.0 default: 1.0
section: lora section: lora
description: >
Negative weights are legitimate (apply the LoRA in inverse).
Upstream Gradio range adopted verbatim.
- name: audio_format - name: audio_format
type: select type: select
options: [wav, mp3, flac] options: [wav, mp3, flac]
@@ -961,10 +996,18 @@ services:
notes: > notes: >
actual_seeds parameter exposed; identical seeds + params = identical audio. actual_seeds parameter exposed; identical seeds + params = identical audio.
Local infer-api.py patches upstream's broken 24-arg pipeline signature Local infer-api.py patches upstream's broken 24-arg pipeline signature
(was 18 in upstream — caused crashes with audio_duration in `format` slot) (v2: was 18 in upstream — caused crashes with audio_duration in `format`
AND inline-streams the generated audio bytes (was returning a JSON slot) AND inline-streams the generated audio bytes (v4: was returning a
path reference to a file inside the container, which was unreachable JSON path reference to a file inside the container, which was
from outside). unreachable from outside).
REPRODUCIBILITY GAP (v5): default `actual_seeds: []` triggers random
seed selection inside the pipeline. The chosen seed IS available in
the pipeline's return dict (`actual_seeds` key) but our wrapper
doesn't capture or surface it — so default-defaulted assets cannot be
regenerated bit-exact. Set actual_seeds explicitly when reproducibility
is required. Wrapper enhancement to surface chosen seeds via response
header + a catalog schema for header→accessory capture is queued.
estimated_latency: estimated_latency:
cold_start_s: 30 cold_start_s: 30
warm_per_unit: "~10–60s depending on audio_duration + infer_step" warm_per_unit: "~10–60s depending on audio_duration + infer_step"