docs(parakeet): seat A/B vs parakeet-unified-en-0.6b - latency is the int8-on-CPU runtime; unified wins WER

A/B of the live STT seat (fv-ml1 GPU 0, sherpa-onnx int8 v3) against
nvidia/parakeet-unified-en-0.6b, measured on GPU 3 with the seat's own image,
k2-fsa's published unified int8 export, fp32/fp16 exports made with k2-fsa's
recipe, v2 int8, and NeMo 3.0.0 (fp32, bf16 autocast, bf16 weights).

- Seat int8 graph runs on one CPU thread (cpu/wall 1.00, GPU 2-9%).
- unified-en under NeMo: -121/-234/-530 ms vs the seat at 1-3/3-8/8-20 s
  (paired, n=120/bin; floor <=6 ms; +50 ms positive control reads +52-54).
- unified-en WER lower in every runtime: -0.7 pp clean, -1.5 pp other,
  -3.2 to -4.4 pp AMI (paired CIs exclude 0).
- Seat defects found: hard 400 s input ceiling (HTTP 500), truncation after
  a quiet 1.5 s pause, and severe long-window dropouts (int8 v3 only).
- B-bf16w needs +0.8 to +1.5 GB over the seat's 1,690 MiB on GPU 0.

Raw requests, hypotheses, manifests and the full harness under
services/parakeet-ab-2026-09-30/. No deploy; live seat untouched apart
from 240 light test requests.
This commit is contained in:
vh
2026-09-30 18:51:44 -07:00
parent 11174ffea1
commit a6c1d3c454
64 changed files with 23022 additions and 0 deletions
+9
View File
@@ -0,0 +1,9 @@
#!/usr/bin/env bash
# Light pass on the LIVE seat (GPU 0): talk-shaped 1-20 s clips only, 1 s apart, abort on the first non-200.
cd /tank/spikes/parakeet-ab
st() { echo "$1 $(date +%T) restarts/started/health=$(docker inspect -f "{{.RestartCount}} {{.State.StartedAt}} {{.State.Health.Status}}" parakeet) seat=$(nvidia-smi --query-compute-apps=pid,used_memory --format=csv,noheader | grep 1594431) gpu0free=$(nvidia-smi -i 0 --query-gpu=memory.free --format=csv,noheader)"; }
st BEFORE
python3 code/bench.py lat http://127.0.0.1:8300/v1/audio/transcriptions A-live data/lat.jsonl out/raw/lat-live.jsonl --rounds 2 --round-offset 0 --jitter-ms 490 --pause 1.0 --bins b1_3,b3_8,b8_20 --stop-on-error
echo "rc=$?"
st AFTER
echo LIVE PASS DONE