fix: Web Audio streaming playback — fixes Safari NotSupportedError

Operator confirmed the "TTS blocked" was NotSupportedError on Safari — WebKit refuses a
streaming 0xFFFFFFFF-length WAV via <audio src> (can't compute duration/seek), exactly
as infra-ops warned. Replaced the <audio src> playback with a Web Audio path that works
in all engines:

- speakOnDone: fetch the chunked /api/tts stream, skip the WAV header to the data chunk,
  decode int16 LE PCM -> Float32, and schedule the samples GAPLESSLY into an AudioContext
  as they arrive (BufferSource per chunk, playAt += buf.duration). Progressive, TTFA
  ~0.5s. Decoding the raw PCM ourselves sidesteps every WAV-container quirk.
- unlock: an AudioContext starts suspended; Safari + Chrome need resume() from a user
  gesture. _unlockTtsAudio() now resumes the ctx on the first interaction anywhere +
  toggle-on + submit, so it's running before the ~15s-delayed speak-on-done.
- cancelTts: aborts the fetch + stops all scheduled BufferSource nodes.

Validated in Chromium (Playwright, strict autoplay): 43 nodes scheduled, 5.1s of PCM
decoded, ctx "running" 6.5s post-gesture, zero errors. Headless WebKit can't launch here
(missing system libs — an infra-ops install), so the operator's live Safari is the final
check; the code is standard Web Audio Safari has supported for years.

Contract FN client:speakOnDone updated (Web Audio; the Safari NotSupportedError reason).
This commit is contained in:
vh
2026-08-02 07:21:48 -07:00
parent 677b03327d
commit 9041f1f402
2 changed files with 90 additions and 60 deletions
@@ -201,19 +201,26 @@ pin_kb_context(question: str, agent_id: str | None, *, client) -> list[dict] #
searches in-voice natively. searches in-voice natively.
``` ```
### FN client: speakOnDone (index.html — STREAMING, DEC-2) ### FN client: speakOnDone (index.html — Web Audio STREAMING, DEC-2)
``` ```
on SSE `done`: on SSE `done`:
if !ttsEnabled(): return # INV-TTS-2 if !ttsEnabled(): return # INV-TTS-2
cancelTts() # INV-TTS-3: audio.pause()+removeAttribute(src)+load() cancelTts() # INV-TTS-3: abort fetch + stop scheduled nodes
audio.src = "/api/tts?text=&agent_id=&p=&a=" # GET streaming URL (text sliced to the 2000 cap) fetch("/api/tts?text=&agent_id=&p=&a=") -> reader # chunked stream (text sliced to the 2000 cap)
audio.play() -> ▶ voiced ; .catch -> "playback blocked" # progressive; failure non-fatal (INV-TTS-4) loop: read chunk -> skip WAV header up to the data chunk -> int16 LE PCM -> Float32 -> AudioBuffer ->
BufferSource.start(playAt) scheduled GAPLESSLY -> playAt += buf.duration # progressive, TTFA ~0.5s
first scheduled node -> "▶ voiced"; any failure -> ticker + skip (INV-TTS-4)
AUTOPLAY UNLOCK (the load-bearing fix for "no audio"): speak fires play() in an async callback seconds after WHY Web Audio, not <audio src>: Safari/WebKit REFUSES a streaming 0xFFFFFFFF-length WAV via <audio src>
the keypress, past the browser's transient-activation window, so a bare play() is blocked. _unlockTtsAudio() (NotSupportedError — it can't compute duration/seek), which was the operator's live failure. Decoding the raw
plays a tiny silent WAV inside a REAL gesture (toggle-on + each prompt submit), which grants the <audio> int16 PCM ourselves and scheduling it into an AudioContext sidesteps every WAV-container quirk and works in all
element a PERSISTENT "may-play" flag so the later streaming play() isn't refused. Validated in Chromium with engines. Validated in Chromium: 43 nodes scheduled, 5.1s decoded, no error.
--autoplay-policy=document-user-activation-required: play succeeds 6.5s after the gesture, currentTime advances.
AUTOPLAY UNLOCK: an AudioContext starts "suspended"; Safari + Chrome require resume() to originate from a user
gesture (then it stays running). _unlockTtsAudio() resumes it on the FIRST interaction anywhere (document
pointerdown/keydown) + toggle-on + each submit, so it's running before the ~15s-delayed speak-on-done. Validated:
ctx is "running" 6.5s after the gesture (past the transient-activation window). Page served no-store so a stale
cache can't hide these updates.
``` ```
## Slice plan ## Slice plan
+74 -51
View File
@@ -1994,74 +1994,97 @@ async function cancelTurn() {
}); });
})(); })();
// ---- auto-TTS: voiced, affect-modulated STREAMING playback on turn `done` ----- // ---- auto-TTS: voiced, affect-modulated STREAMING playback via Web Audio ------------
// The <audio> element streams the chunked GET /api/tts response and plays as it arrives // Fetch the chunked GET /api/tts stream, decode its int16 PCM, and schedule the samples
// (TTFA ~0.5s), never waiting for the whole clip. Opt-in (INV-TTS-2), one stream at a // GAPLESSLY into an AudioContext as they arrive (TTFA ~0.5s). Web Audio, NOT <audio src>,
// time (INV-TTS-3: a new turn / toggle-off aborts the prior <audio> load, dropping the // because Safari/WebKit REFUSES a streaming 0xFFFFFFFF-length WAV via <audio src>
// server stream), non-blocking (INV-TTS-4: any failure logs to the ticker, never the turn). // (NotSupportedError) — decoding the raw PCM ourselves sidesteps every WAV-container quirk
// and works in all engines. Opt-in (INV-TTS-2), one stream at a time (INV-TTS-3: a new
// turn aborts the fetch + stops scheduled nodes), non-blocking (INV-TTS-4).
function ttsEnabled() { function ttsEnabled() {
try { return localStorage.getItem("ratatoskr-tts") === "1"; } catch (_) { return false; } try { return localStorage.getItem("ratatoskr-tts") === "1"; } catch (_) { return false; }
} }
// A short silent WAV (PCM/mono/44.1k/16-bit) used only to UNLOCK the <audio> element. let _ttsCtx = null, _ttsAbort = null, _ttsNodes = [];
// speak-on-done fires play() in an async callback seconds after the user's keypress, by function _ttsAudioCtx() {
// which point the browser's autoplay policy has revoked the activation and blocks it. if (!_ttsCtx) {
// Playing this once inside a real gesture (toggle-on, each prompt submit) grants the const AC = window.AudioContext || window.webkitAudioContext;
// element a persistent "may play" flag so the later real playback isn't blocked. if (AC) _ttsCtx = new AC();
const _TTS_SILENT = (() => { }
const N = 256, data = 2 * N, buf = new Uint8Array(44 + data), dv = new DataView(buf.buffer); return _ttsCtx;
buf.set([0x52, 0x49, 0x46, 0x46], 0); dv.setUint32(4, 36 + data, true); }
buf.set([0x57, 0x41, 0x56, 0x45], 8); buf.set([0x66, 0x6d, 0x74, 0x20], 12); // Unlock: resume the AudioContext inside a user gesture (Safari + Chrome both require the
dv.setUint32(16, 16, true); dv.setUint16(20, 1, true); dv.setUint16(22, 1, true); // resume to originate from an interaction; once running it stays running). Fired on the
dv.setUint32(24, 44100, true); dv.setUint32(28, 88200, true); // FIRST interaction anywhere, so it's ready before the delayed speak-on-done.
dv.setUint16(32, 2, true); dv.setUint16(34, 16, true);
buf.set([0x64, 0x61, 0x74, 0x61], 36); dv.setUint32(40, data, true);
let s = ""; for (const x of buf) s += String.fromCharCode(x);
return "data:audio/wav;base64," + btoa(s);
})();
let _ttsUnlocked = false;
function _unlockTtsAudio() { function _unlockTtsAudio() {
if (_ttsUnlocked) return; // once per page is enough const ctx = _ttsAudioCtx();
const a = $("tts-audio"); if (!a) return; if (ctx && ctx.state === "suspended") ctx.resume().catch(() => {});
try {
a.src = _TTS_SILENT;
const p = a.play();
if (p && p.then) p.then(() => {
_ttsUnlocked = true;
try { a.pause(); a.removeAttribute("src"); a.load(); } catch (_) {}
}).catch(() => {});
} catch (_) {}
} }
// Unlock on the FIRST user interaction ANYWHERE on the page — not just toggle/submit —
// so the element has autoplay permission before the delayed speak-on-done play(), no
// matter how the operator first touches the page. _unlockTtsAudio self-guards, so these
// are no-ops once granted.
document.addEventListener("pointerdown", _unlockTtsAudio, true); document.addEventListener("pointerdown", _unlockTtsAudio, true);
document.addEventListener("keydown", _unlockTtsAudio, true); document.addEventListener("keydown", _unlockTtsAudio, true);
function cancelTts() { function cancelTts() {
// Stop + drop the current <audio> src; load() aborts the in-flight GET stream, which if (_ttsAbort) { try { _ttsAbort.abort(); } catch (_) {} _ttsAbort = null; }
// drops the server proxy (releasing its synth lock). No object URLs to revoke — the for (const n of _ttsNodes) { try { n.stop(); } catch (_) {} try { n.disconnect(); } catch (_) {} }
// src is a streaming /api/tts URL, not a blob. _ttsNodes = [];
const a = $("tts-audio");
if (a) { try { a.pause(); a.removeAttribute("src"); a.load(); } catch (_) {} }
} }
function speakOnDone(text, agentId, pad) { function _findDataChunk(u8) { // offset of the "data" chunk id in a WAV header, or -1
for (let i = 0; i + 4 <= u8.length; i++)
if (u8[i] === 0x64 && u8[i + 1] === 0x61 && u8[i + 2] === 0x74 && u8[i + 3] === 0x61) return i;
return -1;
}
function _u8concat(a, b) { // always returns a FRESH array (byteOffset 0) so Int16Array aligns
const out = new Uint8Array(a.length + b.length); out.set(a, 0); out.set(b, a.length); return out;
}
async function speakOnDone(text, agentId, pad) {
const clip = (text || "").trim(); const clip = (text || "").trim();
if (!clip) return; if (!clip) return;
cancelTts(); // INV-TTS-3: stop any prior stream cancelTts(); // INV-TTS-3: stop any prior stream
const a = $("tts-audio"); if (!a) return; const ctx = _ttsAudioCtx();
if (!ctx) { tickerAdd("err", "tts", "no audio ctx"); return; }
if (ctx.state === "suspended") { try { await ctx.resume(); } catch (_) {} }
const params = new URLSearchParams({ text: clip.slice(0, 2000) }); // matches the server cap const params = new URLSearchParams({ text: clip.slice(0, 2000) }); // matches the server cap
if (agentId) params.set("agent_id", agentId); if (agentId) params.set("agent_id", agentId);
if (pad && typeof pad.pleasure === "number" && typeof pad.arousal === "number") { if (pad && typeof pad.pleasure === "number" && typeof pad.arousal === "number") {
params.set("p", pad.pleasure); params.set("a", pad.arousal); // affect dials (DEC-7) params.set("p", pad.pleasure); params.set("a", pad.arousal); // affect dials (DEC-7)
} }
// Progressive playback: point <audio> at the chunked stream; it plays as bytes arrive. const ctrl = new AbortController(); _ttsAbort = ctrl;
a.src = "/api/tts?" + params.toString(); let resp;
// play() may be refused by the autoplay policy in this async callback; the prompt/toggle try { resp = await fetch("/api/tts?" + params.toString(), { signal: ctrl.signal }); }
// gesture unlock (_unlockTtsAudio) satisfies it, and the catch keeps a refusal non-fatal. catch (_) { return; } // aborted / network → silent skip (INV-TTS-4)
a.play().then(() => tickerAdd("ok", "tts", "▶ voiced")) if (!resp.ok || !resp.body) { tickerAdd("err", "tts", "unavailable " + resp.status); return; }
.catch((e) => tickerAdd("err", "tts", "blocked: " + ((e && e.name) || "?"))); const reader = resp.body.getReader();
// NotAllowedError = autoplay (unlock didn't take); NotSupportedError = the browser const SR = 44100;
// refused the streaming WAV; AbortError = a new turn superseded it. let playAt = ctx.currentTime + 0.06, started = false, headerDone = false;
let acc = new Uint8Array(0), carry = new Uint8Array(0);
try {
while (true) {
const { done, value } = await reader.read();
if (done || ctrl !== _ttsAbort) break; // finished, or superseded by a new turn
let bytes = value;
if (!headerDone) { // skip the WAV header (up to + incl the data id/size)
acc = _u8concat(acc, bytes);
const di = _findDataChunk(acc);
if (di < 0 || di + 8 > acc.length) continue; // header spans chunks — keep accumulating
bytes = acc.subarray(di + 8); headerDone = true; acc = null;
}
const u8 = _u8concat(carry, bytes); // prepend the odd-byte carry; fresh + aligned
const even = u8.length - (u8.length & 1);
carry = u8.slice(even); // stash a trailing odd byte for next chunk
if (even < 2) continue;
const pcm = new Int16Array(u8.buffer, 0, even / 2); // int16 LE (all real browsers are LE)
const f32 = new Float32Array(pcm.length);
for (let i = 0; i < pcm.length; i++) f32[i] = pcm[i] / 32768;
const ab = ctx.createBuffer(1, f32.length, SR);
ab.getChannelData(0).set(f32);
const node = ctx.createBufferSource();
node.buffer = ab; node.connect(ctx.destination);
if (playAt < ctx.currentTime) playAt = ctx.currentTime; // underrun guard
node.start(playAt); playAt += ab.duration;
_ttsNodes.push(node);
node.onended = () => { const i = _ttsNodes.indexOf(node); if (i >= 0) _ttsNodes.splice(i, 1); };
if (!started) { started = true; tickerAdd("ok", "tts", "▶ voiced"); }
}
} catch (_) { /* aborted / stream error → whatever's already scheduled finishes */ }
} }
// ---- 🔊 toggle (mirrors theme / cot-toggle; default OFF, persisted) ---- // ---- 🔊 toggle (mirrors theme / cot-toggle; default OFF, persisted) ----
(function () { (function () {