Files
esh-pfi-infrastructure/services/parakeet-ab-2026-09-30/code/pilot.py
T
vh a6c1d3c454 docs(parakeet): seat A/B vs parakeet-unified-en-0.6b - latency is the int8-on-CPU runtime; unified wins WER
A/B of the live STT seat (fv-ml1 GPU 0, sherpa-onnx int8 v3) against
nvidia/parakeet-unified-en-0.6b, measured on GPU 3 with the seat's own image,
k2-fsa's published unified int8 export, fp32/fp16 exports made with k2-fsa's
recipe, v2 int8, and NeMo 3.0.0 (fp32, bf16 autocast, bf16 weights).

- Seat int8 graph runs on one CPU thread (cpu/wall 1.00, GPU 2-9%).
- unified-en under NeMo: -121/-234/-530 ms vs the seat at 1-3/3-8/8-20 s
  (paired, n=120/bin; floor <=6 ms; +50 ms positive control reads +52-54).
- unified-en WER lower in every runtime: -0.7 pp clean, -1.5 pp other,
  -3.2 to -4.4 pp AMI (paired CIs exclude 0).
- Seat defects found: hard 400 s input ceiling (HTTP 500), truncation after
  a quiet 1.5 s pause, and severe long-window dropouts (int8 v3 only).
- B-bf16w needs +0.8 to +1.5 GB over the seat's 1,690 MiB on GPU 0.

Raw requests, hypotheses, manifests and the full harness under
services/parakeet-ab-2026-09-30/. No deploy; live seat untouched apart
from 240 light test requests.
2026-09-30 18:51:44 -07:00

15 lines
754 B
Python

import json, sys, time
sys.path.insert(0, "/tank/spikes/parakeet-ab/code")
from bench import post
url = sys.argv[1]
rows = {json.loads(l)["id"]: json.loads(l) for l in open("/tank/spikes/parakeet-ab/data/lat.jsonl")}
from bench import hostpath; w = lambda i: open(hostpath(rows[i]["wav"]), "rb").read()
def show(tag, i):
r = post(url, w(i)); print(f"{tag:10s} {i:12s} dur {rows[i]['dur']:6.2f}s e2e {r['e2e_ms']:8.1f} server {r['server_ms']}", flush=True)
for k in range(5): show("repeat", "b3_8_10")
for k in range(5): show("fresh", f"b3_8_{k:02d}")
for k in range(5): show("fresh1-3", f"b1_3_{k:02d}")
for k in range(3): show("repeat", "b1_3_00")
for k in range(3): show("fresh20", f"b20_60_{k:02d}")
for k in range(3): show("repeat", "b20_60_00")