Smoke testing in the asset_engine consumer surfaced an
UnboundLocalError 500 from ace-step (althing thread
01KRCJF7NGMXYE9F62Q1A6KFD4 msg 3). Root cause: this catalog had
invented enum values for scheduler_type and cfg_type that don't
exist in the upstream pipeline.
Read pipeline_ace_step.py inside the running container:
scheduler_type dispatch:
if == "euler": scheduler = FlowMatchEulerDiscreteScheduler(...)
elif== "heun": scheduler = FlowMatchHeunDiscreteScheduler(...)
elif== "pingpong": scheduler = FlowMatchPingPongScheduler(...)
# no else -> "linear" / "squared" / "sqrt" leave scheduler unbound
cfg_type dispatch:
accepts: apg | cfg | cfg_star
Catalog had:
scheduler_type: [linear, squared, sqrt] / default linear <- all invalid
cfg_type: [none, cfg, cfg_rw] / default cfg <- only cfg works
Fixed:
scheduler_type: [euler, heun, pingpong] / default euler
cfg_type: [apg, cfg, cfg_star] / default cfg
Bumped ace-step version 2 -> 3. Existing assets generated under v2
with scheduler_type=linear cannot reproduce (the value is now invalid);
v2 assets with the accidentally-valid cfg_type=cfg + a corrected
scheduler can be regenerated under v3 by mapping linear -> euler.
catalog_version stays at 1 (no schema change).
Verified end-to-end against live ace-step on irv-ml1:
POST /generate { scheduler_type: euler, cfg_type: cfg, ... }
-> 200, output_path returned, ~8s wall time
Lesson: OpenAPI introspection isn't enough for accurate catalog
authoring. Upstream OpenAPI returns bare `string` for both fields.
Reading the actual dispatch code is the only way to capture the
allowed values. Will sweep the other 11 service entries against
their implementations before P2 (scale to all services) lands.
asset_engine consumer (althing thread 01KRCJF7NGMXYE9F62Q1A6KFD4)
needed structure for ace-step's 27-field form. Two additive Pydantic
changes — backward-compatible, no catalog_version bump per the
policy table:
- CatalogField.section: str | None = None
- CatalogService.section_groups: list[CatalogSectionGroup] = []
- new CatalogSectionGroup model: {id, label, hint?}
Validator: every Field.section value must reference a declared
section_groups[].id within the same service; section_groups[].id
values are unique. CATALOG-CONTRACT.md updated with both the new
service-fields row and a versioning-policy row covering
"add optional Field/Service keys -> no bump."
ace-step entry rewritten to use the new schema:
- bumped version 1 -> 2
- declared 6 section groups (basic / generation / conditioning /
a2a / lora / output) with hints
- tagged every field with a section
- added previously-missing checkpoint_path (required: true,
default: "/app/checkpoints" — the container's mount path).
Wrapper-side cleanup (default in infer-api.py) queued as
follow-up.
- changed lyrics from optional: true -> required: true with
default "" to match upstream's `lyrics: str` shape (empty
string satisfies it).
JSON Schema regenerated.
Pydantic-model side of this change lives in asset_engine at
src/asset_engine/catalog.py — committed there separately.
Consumer-side renderer for the JSON-envelope + timestamps shape
shipped (althing thread 01KRCF4W66X3, msg 5). Smoke + regression
clean. Per the contract on the entry's notes block, flipping to
ready now that the renderer is in place.
asset_engine consumer needed to render kokoro-captioned, whose wire
shape is a JSON envelope carrying base64-encoded audio plus a
structured timestamps array. Modeling it as response.type=json
would force either a per-service-id renderer (forbidden by
brief §1.7) or extending the closed response-type vocabulary
(forbidden by brief §2.2 without a coordinated bump).
Resolution (per althing thread 01KRCF4W66X3): keep response.type
closed at the existing six values and decompose at the response
*field* level instead — the same flexibility seam already used by
mime / mime_from_field / output_field. Adds three optional keys:
- audio_field: JSON key holding base64-encoded audio bytes
- audio_format_field: JSON key holding the decoded audio MIME
- timestamps_field: JSON key holding a structured timestamps array
(independent of type, declared by any service emitting time-
aligned markers)
Validators in CatalogResponse enforce sane combinations:
- audio_field requires response.type=audio
- audio_field forbids mime_from_field
- audio_format_field requires audio_field
This is additive and backward-compatible — no catalog_version bump,
existing services parse unchanged. CATALOG-CONTRACT.md updated with
the new rows in the response-field table and a versioning-policy
row codifying that adding optional keys to response: doesn't bump.
kokoro-captioned re-shaped to use the new schema:
response:
type: audio
audio_field: audio
audio_format_field: audio_format
timestamps_field: timestamps
And marked status: experimental until the asset_engine consumer's
audio-with-timestamps renderer ships.
JSON Schema regenerated to reflect the new Pydantic shape.
Pydantic-model side of this change lives in the asset_engine repo
at src/asset_engine/catalog.py — committed there separately.
Per a request from the asset_engine consumer (althing thread
01KRCF4W66X3N24B01FF2Y7V3D), and verified against the live kokoro
OpenAPI + exercised endpoints:
* kokoro: version 1 → 2; adds three fields surfaced by the upstream
schema but not previously declared:
- speed (slider 0.25–4.0, default 1.0)
- volume_multiplier (slider 0.5–2.0, default 1.0; UI-bounded
since upstream is unbounded — noted in description)
- lang_code (text, optional override of the voice-name-derived
language hint)
* kokoro-captioned: new service entry wrapping
/dev/captioned_speech. Same model + image as kokoro proper but
separate catalog entry because the response shape is structured
JSON (audio inline as base64 + word-level timestamps), not raw
audio bytes. Verified shape captured in reproducibility.notes
so future consumers don't have to re-discover it. response.type
= json (consumer renders custom: player + subtitle overlay).
* reproducibility_audit: row added for kokoro-captioned.
Deferred (separate from this commit):
- kokoro-blend-voice. /v1/audio/voices/combine returns 403 on the
default config (allow_local_voice_saving=False); even with the
flag flipped it writes to a temp dir, not /worktank/kokoro/user_voices.
The persistent blend mechanism in this fleet is
playbooks/blend-kokoro-voice.yaml. Ad-hoc blending already works
through /v1/audio/speech via the inline syntax voice="a(w)+b(w)";
consumer can surface that as a UI affordance without any
catalog change.
catalog_version stays at 1 (no field-type vocabulary changes).
JSON Schema regeneration produced byte-identical output.
Adds the supporting infra around the service catalog now that it
has external consumers (the asset_engine UI being the first; CLIs,
monitoring, other services may follow):
- CATALOG-CONTRACT.md: the consumer-facing contract. Defines
versioning policy (catalog_version vs per-service version),
closed field-type and response-type vocabularies, recommended
vendor+drift-check sync workflow, known-consumers list, service
authoring notes.
- services.schema.json: JSON Schema (draft 2020-12) for the
catalog. Generated from the Pydantic model in
~/development/asset_engine/src/asset_engine/catalog.py via
`uv run scripts/dump_schema.py --publish`. Lets non-Python
consumers validate against the same shape.
- services.yaml: adds catalog_version: 1 at the root and reframes
the file's header to call out its first-class-contract status.
Quotes a vibevoice label that contained an unescaped colon
(caught by the asset_engine's strict YAML parser on first sync).
services.yaml: form-generator contract for the forthcoming
asset-generation UI. 13 inference services on irv-ml1 (TTS, ASR,
SFX, music) catalogued with field schemas extracted from Pydantic
models, response types, reproducibility audit, and license
warnings. ComfyUI flagged catalog-deferred (workflow-DAG API
doesn't fit a form-based UI without a per-asset-type wrapper).
design-brief.md: the prompt to give a frontend-design agent before
any pixels. Locks in the data-model decisions whose later cost is
asymmetric (asset-as-first-class entity, content-addressed output
storage, reproducibility hard requirement, job table, auth as a
no-op DI seam, API surface ≠ UI surface, schema versioning,
tags/collections plumbed in v1 with no UI). Defines a closed
field-type vocabulary (8 types) and response-renderer vocabulary
(6 types) — agent isn't allowed to extend them. Pre-decides the
required UI surfaces; leaves IA, library-nav pattern, long-job
UX, and big-form ergonomics open for the agent to opine on.