A/B of the live STT seat (fv-ml1 GPU 0, sherpa-onnx int8 v3) against nvidia/parakeet-unified-en-0.6b, measured on GPU 3 with the seat's own image, k2-fsa's published unified int8 export, fp32/fp16 exports made with k2-fsa's recipe, v2 int8, and NeMo 3.0.0 (fp32, bf16 autocast, bf16 weights). - Seat int8 graph runs on one CPU thread (cpu/wall 1.00, GPU 2-9%). - unified-en under NeMo: -121/-234/-530 ms vs the seat at 1-3/3-8/8-20 s (paired, n=120/bin; floor <=6 ms; +50 ms positive control reads +52-54). - unified-en WER lower in every runtime: -0.7 pp clean, -1.5 pp other, -3.2 to -4.4 pp AMI (paired CIs exclude 0). - Seat defects found: hard 400 s input ceiling (HTTP 500), truncation after a quiet 1.5 s pause, and severe long-window dropouts (int8 v3 only). - B-bf16w needs +0.8 to +1.5 GB over the seat's 1,690 MiB on GPU 0. Raw requests, hypotheses, manifests and the full harness under services/parakeet-ab-2026-09-30/. No deploy; live seat untouched apart from 240 light test requests.
16 lines
928 B
Python
16 lines
928 B
Python
"""CPU seconds the server process burns per request vs the request's wall time (host /proc; Linux CLK_TCK=100)."""
|
|
import json, os, sys, time, subprocess
|
|
sys.path.insert(0, "/tank/spikes/parakeet-ab/code")
|
|
from bench import post, hostpath
|
|
url, pid = sys.argv[1], int(sys.argv[2])
|
|
ids = sys.argv[3].split(",")
|
|
rows = {json.loads(l)["id"]: json.loads(l) for l in open("/tank/spikes/parakeet-ab/data/lat.jsonl")}
|
|
def cpu(p):
|
|
f = open(f"/proc/{p}/stat").read().rsplit(")", 1)[1].split()
|
|
return (int(f[11]) + int(f[12])) / os.sysconf("SC_CLK_TCK")
|
|
def threads(p): return len(os.listdir(f"/proc/{p}/task"))
|
|
for i in ids:
|
|
wav = open(hostpath(rows[i]["wav"]), "rb").read()
|
|
c0 = cpu(pid); r = post(url, wav); c1 = cpu(pid)
|
|
print(f"{i:10s} dur {rows[i]['dur']:6.2f}s server {r['server_ms']:8.1f} ms process CPU {1000*(c1-c0):8.1f} ms cpu/wall {(c1-c0)*1000/r['server_ms']:.2f} threads {threads(pid)}", flush=True)
|