Files
esh-pfi-infrastructure/services/parakeet-ab-2026-09-30/code/gw_analysis.py
T
vh a6c1d3c454 docs(parakeet): seat A/B vs parakeet-unified-en-0.6b - latency is the int8-on-CPU runtime; unified wins WER
A/B of the live STT seat (fv-ml1 GPU 0, sherpa-onnx int8 v3) against
nvidia/parakeet-unified-en-0.6b, measured on GPU 3 with the seat's own image,
k2-fsa's published unified int8 export, fp32/fp16 exports made with k2-fsa's
recipe, v2 int8, and NeMo 3.0.0 (fp32, bf16 autocast, bf16 weights).

- Seat int8 graph runs on one CPU thread (cpu/wall 1.00, GPU 2-9%).
- unified-en under NeMo: -121/-234/-530 ms vs the seat at 1-3/3-8/8-20 s
  (paired, n=120/bin; floor <=6 ms; +50 ms positive control reads +52-54).
- unified-en WER lower in every runtime: -0.7 pp clean, -1.5 pp other,
  -3.2 to -4.4 pp AMI (paired CIs exclude 0).
- Seat defects found: hard 400 s input ceiling (HTTP 500), truncation after
  a quiet 1.5 s pause, and severe long-window dropouts (int8 v3 only).
- B-bf16w needs +0.8 to +1.5 GB over the seat's 1,690 MiB on GPU 0.

Raw requests, hypotheses, manifests and the full harness under
services/parakeet-ab-2026-09-30/. No deploy; live seat untouched apart
from 240 light test requests.
2026-09-30 18:51:44 -07:00

43 lines
2.1 KiB
Python

"""Live seat (A), gateway and cross-site passes, paired against the GPU-3 arms on identical trimmed input.
A-live: fv-ml1 loopback to :8300, rounds 0-1. A-gateway: nh3-dev -> LiteLLM ext-stt (ana-docker) -> :8300,
round 0. A-nh3direct: nh3-dev -> :8300, round 0. 1-20 s bins only (the live seat never got 20-60 s clips).
usage: gw_analysis.py RAW_DIR"""
import json
import sys
import numpy as np
RAW = sys.argv[1]
rows = [json.loads(l) for f in ("lat-live.jsonl", "lat-gw.jsonl", "lat.jsonl") for l in open(f"{RAW}/{f}")]
rows = [r for r in rows if r.get("mode") == "lat" and r["status"] == 200]
by = {}
for r in rows:
by.setdefault(r["arm"], {})[(r["round"], r["id"])] = r
rng = np.random.default_rng(7)
def ci(d):
d = np.asarray(d)
b = [np.median(d[rng.integers(0, len(d), len(d))]) for _ in range(2000)]
return np.percentile(b, [2.5, 97.5])
BINS = ("b1_3", "b3_8", "b8_20")
for arm in ("A-live", "A-gateway", "A-nh3direct"):
for b in BINS:
x = [r["e2e_ms"] for r in by[arm].values() if r["bin"] == b]
print(f"{arm:12s} {b:6s} n={len(x):3d} e2e p50 {np.median(x):7.1f} p90 {np.percentile(x, 90):7.1f} max {max(x):7.1f}")
print("paired: median diff [95% CI]")
for a, ref in (("A-live", "ab-a1"), ("A-gateway", "A-live"), ("A-nh3direct", "A-live"), ("A-gateway", "A-nh3direct"),
("A-gateway", "ab-b32"), ("A-live", "ab-b32")):
for b in BINS:
ks = [k for k in by[a] if k in by[ref] and by[a][k]["bin"] == b]
assert all(by[a][k].get("trim_ms") == by[ref][k].get("trim_ms") for k in ks)
d = [by[a][k]["e2e_ms"] - by[ref][k]["e2e_ms"] for k in ks]
lo, hi = ci(d)
print(f"{a:12s} - {ref:12s} {b:6s} pairs={len(ks):3d} {np.median(d):+8.1f} [{lo:+7.1f},{hi:+7.1f}]")
ks = [k for k in by["A-live"] if k in by["ab-a1"]]
print("A-live vs ab-a1 identical text:", sum(by["A-live"][k]["text"] == by["ab-a1"][k]["text"] for k in ks), "/", len(ks))
ks = [k for k in by["A-gateway"] if k in by["A-live"]]
print("A-gateway vs A-live identical text:", sum(by["A-gateway"][k]["text"] == by["A-live"][k]["text"] for k in ks), "/", len(ks))