catalog-contract: add response-decomposition fields (audio_field, timestamps_field, audio_format_field)
asset_engine consumer needed to render kokoro-captioned, whose wire
shape is a JSON envelope carrying base64-encoded audio plus a
structured timestamps array. Modeling it as response.type=json
would force either a per-service-id renderer (forbidden by
brief §1.7) or extending the closed response-type vocabulary
(forbidden by brief §2.2 without a coordinated bump).
Resolution (per althing thread 01KRCF4W66X3): keep response.type
closed at the existing six values and decompose at the response
*field* level instead — the same flexibility seam already used by
mime / mime_from_field / output_field. Adds three optional keys:
- audio_field: JSON key holding base64-encoded audio bytes
- audio_format_field: JSON key holding the decoded audio MIME
- timestamps_field: JSON key holding a structured timestamps array
(independent of type, declared by any service emitting time-
aligned markers)
Validators in CatalogResponse enforce sane combinations:
- audio_field requires response.type=audio
- audio_field forbids mime_from_field
- audio_format_field requires audio_field
This is additive and backward-compatible — no catalog_version bump,
existing services parse unchanged. CATALOG-CONTRACT.md updated with
the new rows in the response-field table and a versioning-policy
row codifying that adding optional keys to response: doesn't bump.
kokoro-captioned re-shaped to use the new schema:
response:
type: audio
audio_field: audio
audio_format_field: audio_format
timestamps_field: timestamps
And marked status: experimental until the asset_engine consumer's
audio-with-timestamps renderer ships.
JSON Schema regenerated to reflect the new Pydantic shape.
Pydantic-model side of this change lives in the asset_engine repo
at src/asset_engine/catalog.py — committed there separately.
This commit is contained in:
@@ -46,10 +46,13 @@ reproducibility_audit: [Audit] # one entry per service
|
|||||||
| `model.revision` | string \| null | no | SHA when known |
|
| `model.revision` | string \| null | no | SHA when known |
|
||||||
| `model.image` | string | yes | container image ref this is hosted from |
|
| `model.image` | string | yes | container image ref this is hosted from |
|
||||||
| `fields` | list[Field] | no | request parameters; empty for catalog-deferred |
|
| `fields` | list[Field] | no | request parameters; empty for catalog-deferred |
|
||||||
| `response.type` | enum: audio, image, video, text, json, file | yes | response renderer hint |
|
| `response.type` | enum: audio, image, video, text, json, file | yes | renderer dispatch (closed vocabulary) |
|
||||||
| `response.mime` | string | no | static response MIME |
|
| `response.mime` | string | no | static response MIME (raw-bytes wire) |
|
||||||
| `response.mime_from_field` | string | no | name of a field whose value determines the MIME |
|
| `response.mime_from_field` | string | no | name of a field whose value determines the MIME (raw-bytes wire) |
|
||||||
| `response.output_field` | string | no | for JSON responses, the key holding the asset |
|
| `response.output_field` | string | no | for text/json responses, JSON key holding the asset |
|
||||||
|
| `response.audio_field` | string | no | for `type: audio` with JSON-envelope wire, JSON key holding base64-encoded audio bytes |
|
||||||
|
| `response.audio_format_field` | string | no | for `type: audio` with JSON-envelope wire, JSON key holding the decoded audio MIME (e.g. "audio/wav") |
|
||||||
|
| `response.timestamps_field` | string | no | JSON key holding a structured timestamps array — independent of type, declared by any service that emits time-aligned markers alongside its primary output |
|
||||||
| `reproducibility.seedable` | bool | yes | does the endpoint accept a seed? |
|
| `reproducibility.seedable` | bool | yes | does the endpoint accept a seed? |
|
||||||
| `reproducibility.deterministic` | bool | yes | same params → same bytes? |
|
| `reproducibility.deterministic` | bool | yes | same params → same bytes? |
|
||||||
| `reproducibility.notes` | string | no | gotchas |
|
| `reproducibility.notes` | string | no | gotchas |
|
||||||
@@ -91,6 +94,7 @@ Same change-management as field types.
|
|||||||
|----------------------------------------------|-----------------------|
|
|----------------------------------------------|-----------------------|
|
||||||
| Add a new service | nothing |
|
| Add a new service | nothing |
|
||||||
| Add a non-required field to an existing service | service `version:` |
|
| Add a non-required field to an existing service | service `version:` |
|
||||||
|
| Add an optional key to the `response:` schema (e.g. audio_field, timestamps_field) | nothing — additive, backward-compatible |
|
||||||
| Change a field's type, range, or default | service `version:` |
|
| Change a field's type, range, or default | service `version:` |
|
||||||
| Remove a service | service `version:` (sentinel: removed=true), then drop in next catalog_version bump |
|
| Remove a service | service `version:` (sentinel: removed=true), then drop in next catalog_version bump |
|
||||||
| Add a new entry to the field-type vocabulary | `catalog_version:` |
|
| Add a new entry to the field-type vocabulary | `catalog_version:` |
|
||||||
|
|||||||
@@ -332,6 +332,7 @@
|
|||||||
},
|
},
|
||||||
"CatalogResponse": {
|
"CatalogResponse": {
|
||||||
"additionalProperties": false,
|
"additionalProperties": false,
|
||||||
|
"description": "How to interpret the inference response.\n\nThe `type` is the *renderer dispatch* \u2014 what kind of asset the user\nultimately sees (audio player, text block, etc). The other fields\ndescribe how to *extract* that asset from the wire shape:\n\n - Raw-bytes wire (e.g. /v1/audio/speech returns raw audio):\n type: audio\n mime: audio/wav (or mime_from_field for dynamic)\n\n - JSON-envelope wire with base64-encoded asset (e.g. captioned\n speech returns {audio: <b64>, audio_format: ..., timestamps: ...}):\n type: audio\n audio_field: audio # JSON key with the base64 bytes\n audio_format_field: audio_format # JSON key with the MIME\n timestamps_field: timestamps # optional structured aside\n\n - JSON-envelope wire with text payload (e.g. ASR transcript):\n type: text\n output_field: text # JSON key with the text body\n timestamps_field: segments # optional, for ASR-with-timestamps\n\nRules:\n - When audio_field is set, the renderer expects a JSON envelope and\n will base64-decode that key. The wire MIME is application/json\n regardless of `mime`; `mime_from_field` is incompatible.\n - timestamps_field is independent of type \u2014 any service that\n emits structured timestamps alongside its primary output can\n declare it; the renderer extends accordingly.",
|
||||||
"properties": {
|
"properties": {
|
||||||
"type": {
|
"type": {
|
||||||
"enum": [
|
"enum": [
|
||||||
@@ -380,6 +381,42 @@
|
|||||||
],
|
],
|
||||||
"default": null,
|
"default": null,
|
||||||
"title": "Output Field"
|
"title": "Output Field"
|
||||||
|
},
|
||||||
|
"audio_field": {
|
||||||
|
"anyOf": [
|
||||||
|
{
|
||||||
|
"type": "string"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "null"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"default": null,
|
||||||
|
"title": "Audio Field"
|
||||||
|
},
|
||||||
|
"audio_format_field": {
|
||||||
|
"anyOf": [
|
||||||
|
{
|
||||||
|
"type": "string"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "null"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"default": null,
|
||||||
|
"title": "Audio Format Field"
|
||||||
|
},
|
||||||
|
"timestamps_field": {
|
||||||
|
"anyOf": [
|
||||||
|
{
|
||||||
|
"type": "string"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "null"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"default": null,
|
||||||
|
"title": "Timestamps Field"
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
"required": [
|
"required": [
|
||||||
|
|||||||
@@ -111,11 +111,12 @@ services:
|
|||||||
name: Kokoro Captioned Speech
|
name: Kokoro Captioned Speech
|
||||||
description: >
|
description: >
|
||||||
Kokoro TTS with word-level timestamps returned alongside the audio.
|
Kokoro TTS with word-level timestamps returned alongside the audio.
|
||||||
For subtitle generation and video sync. Same model as `kokoro`; this
|
For subtitle generation and video sync. Same model as `kokoro`;
|
||||||
is a separate catalog entry because the response shape is structured
|
separate catalog entry because the wire shape is a JSON envelope
|
||||||
JSON (audio inline + timestamps), not raw audio bytes.
|
carrying base64-encoded audio plus a structured timestamps array.
|
||||||
category: tts
|
category: tts
|
||||||
version: 1
|
version: 1
|
||||||
|
status: experimental
|
||||||
host: irv-ml1
|
host: irv-ml1
|
||||||
endpoint: http://10.100.79.3:8193/dev/captioned_speech
|
endpoint: http://10.100.79.3:8193/dev/captioned_speech
|
||||||
method: POST
|
method: POST
|
||||||
@@ -153,32 +154,43 @@ services:
|
|||||||
label: Language code
|
label: Language code
|
||||||
required: false
|
required: false
|
||||||
response:
|
response:
|
||||||
type: json
|
# Stays in the closed type vocabulary: from a renderer-dispatch
|
||||||
mime: application/json
|
# standpoint this IS audio. The audio_field/audio_format_field/
|
||||||
|
# timestamps_field decomposition tells consumers how to extract
|
||||||
|
# those parts from the JSON envelope wire shape — added to the
|
||||||
|
# catalog schema in 2026-05 specifically to support response
|
||||||
|
# shapes like this one without extending the type vocab.
|
||||||
|
type: audio
|
||||||
|
audio_field: audio
|
||||||
|
audio_format_field: audio_format
|
||||||
|
timestamps_field: timestamps
|
||||||
reproducibility:
|
reproducibility:
|
||||||
seedable: false
|
seedable: false
|
||||||
deterministic: true
|
deterministic: true
|
||||||
notes: >
|
notes: >
|
||||||
Same determinism story as kokoro proper. Response shape (verified
|
Same determinism story as kokoro proper. Verified wire shape
|
||||||
2026-05-11 against live API):
|
(2026-05-11 against live API):
|
||||||
{
|
{
|
||||||
"audio": "<base64-encoded bytes in response_format>",
|
"audio": "<base64-encoded bytes in response_format>",
|
||||||
"audio_format": "audio/wav" (or matching response_format),
|
"audio_format": "audio/wav" (or matching response_format),
|
||||||
"timestamps": [{"word": str, "start_time": float, "end_time": float}, ...]
|
"timestamps": [{"word": str, "start_time": float, "end_time": float}, ...]
|
||||||
}
|
}
|
||||||
Consumer must base64-decode `audio` to play; `timestamps` drives
|
Consumer base64-decodes `audio` to play; `timestamps` drives
|
||||||
subtitle/karaoke UI. Audio_format string in the payload is
|
subtitle/karaoke UI. The response decomposition fields above
|
||||||
authoritative for the decoded bytes.
|
encode this so the renderer doesn't need per-service-id branches.
|
||||||
estimated_latency:
|
estimated_latency:
|
||||||
cold_start_s: 2
|
cold_start_s: 2
|
||||||
warm_per_unit: "~same as kokoro proper, plus minor overhead for timestamp emission"
|
warm_per_unit: "~same as kokoro proper, plus minor overhead for timestamp emission"
|
||||||
license: Apache-2.0
|
license: Apache-2.0
|
||||||
notes: |
|
notes: |
|
||||||
`return_timestamps` and `stream` upstream params are deliberately
|
`return_timestamps` and `stream` upstream params deliberately
|
||||||
omitted from the catalog: timestamps must be on for this endpoint
|
omitted from the catalog: timestamps must be on for this endpoint
|
||||||
to be meaningful, and streaming + JSON-with-base64 don't compose.
|
to be meaningful, and streaming + JSON-with-base64 don't compose.
|
||||||
`download_format` / `return_download_link` skipped — same reasoning
|
`download_format` / `return_download_link` skipped — same as kokoro
|
||||||
as kokoro proper.
|
proper.
|
||||||
|
|
||||||
|
status: experimental until the consumer's audio-with-timestamps
|
||||||
|
renderer ships. Once present, flip to status: ready.
|
||||||
|
|
||||||
- id: chatterbox
|
- id: chatterbox
|
||||||
name: Chatterbox Turbo TTS
|
name: Chatterbox Turbo TTS
|
||||||
|
|||||||
Reference in New Issue
Block a user