A/B of the live STT seat (fv-ml1 GPU 0, sherpa-onnx int8 v3) against nvidia/parakeet-unified-en-0.6b, measured on GPU 3 with the seat's own image, k2-fsa's published unified int8 export, fp32/fp16 exports made with k2-fsa's recipe, v2 int8, and NeMo 3.0.0 (fp32, bf16 autocast, bf16 weights). - Seat int8 graph runs on one CPU thread (cpu/wall 1.00, GPU 2-9%). - unified-en under NeMo: -121/-234/-530 ms vs the seat at 1-3/3-8/8-20 s (paired, n=120/bin; floor <=6 ms; +50 ms positive control reads +52-54). - unified-en WER lower in every runtime: -0.7 pp clean, -1.5 pp other, -3.2 to -4.4 pp AMI (paired CIs exclude 0). - Seat defects found: hard 400 s input ceiling (HTTP 500), truncation after a quiet 1.5 s pause, and severe long-window dropouts (int8 v3 only). - B-bf16w needs +0.8 to +1.5 GB over the seat's 1,690 MiB on GPU 0. Raw requests, hypotheses, manifests and the full harness under services/parakeet-ab-2026-09-30/. No deploy; live seat untouched apart from 240 light test requests.
29 lines
1.4 KiB
Python
29 lines
1.4 KiB
Python
"""Positive-control variants of the same 40 utterances: the 1.5 s gap filled with white noise at -60 and
|
|
-50 dBFS RMS instead of digital zeros (a room-tone pause rather than a gated one), and a 0.75 s zero gap.
|
|
Reference unchanged. Writes data/pc-n60.jsonl, data/pc-n50.jsonl, data/pc-z075.jsonl."""
|
|
import json
|
|
import os
|
|
|
|
import numpy as np
|
|
import soundfile as sf
|
|
|
|
D = "/tank/spikes/parakeet-ab/data"
|
|
pc = [json.loads(l) for l in open(f"{D}/pc.jsonl")]
|
|
clean = {json.loads(l)["id"]: json.loads(l) for l in open(f"{D}/ls-clean.jsonl")}
|
|
rng = np.random.default_rng(20260930)
|
|
for tag, mode, level, span in (("pc-n60", "noise", -60, 1.5), ("pc-n50", "noise", -50, 1.5), ("pc-z075", "zero", None, 0.75)):
|
|
os.makedirs(f"{D}/{tag}", exist_ok=True)
|
|
rows = []
|
|
for r in pc:
|
|
a, sr = sf.read(clean[r["id"]]["wav"].replace("/ab/", "/tank/spikes/parakeet-ab/", 1), dtype="float32")
|
|
s0 = int(0.40 * len(a)); s1 = s0 + int(span * sr)
|
|
b = a.copy()
|
|
b[s0:s1] = (rng.standard_normal(s1 - s0).astype(np.float32) * 10 ** (level / 20)) if mode == "noise" else 0.0
|
|
p = f"{D}/{tag}/{r['id']}.wav"
|
|
sf.write(p, b, sr, subtype="PCM_16")
|
|
rows.append(dict(r, wav=p.replace("/tank/spikes/parakeet-ab/", "/ab/", 1), silence=[round(s0 / sr, 3), round(s1 / sr, 3)], fill=f"{mode} {level}"))
|
|
with open(f"{D}/{tag}.jsonl", "w") as f:
|
|
for r in rows:
|
|
f.write(json.dumps(r) + "\n")
|
|
print(tag, len(rows))
|