A/B of the live STT seat (fv-ml1 GPU 0, sherpa-onnx int8 v3) against nvidia/parakeet-unified-en-0.6b, measured on GPU 3 with the seat's own image, k2-fsa's published unified int8 export, fp32/fp16 exports made with k2-fsa's recipe, v2 int8, and NeMo 3.0.0 (fp32, bf16 autocast, bf16 weights). - Seat int8 graph runs on one CPU thread (cpu/wall 1.00, GPU 2-9%). - unified-en under NeMo: -121/-234/-530 ms vs the seat at 1-3/3-8/8-20 s (paired, n=120/bin; floor <=6 ms; +50 ms positive control reads +52-54). - unified-en WER lower in every runtime: -0.7 pp clean, -1.5 pp other, -3.2 to -4.4 pp AMI (paired CIs exclude 0). - Seat defects found: hard 400 s input ceiling (HTTP 500), truncation after a quiet 1.5 s pause, and severe long-window dropouts (int8 v3 only). - B-bf16w needs +0.8 to +1.5 GB over the seat's 1,690 MiB on GPU 0. Raw requests, hypotheses, manifests and the full harness under services/parakeet-ab-2026-09-30/. No deploy; live seat untouched apart from 240 light test requests.
35 lines
1.5 KiB
Python
35 lines
1.5 KiB
Python
"""Memory a FRESH seat-config instance (A') needs per input length: ascending lengths, one request each,
|
|
GPU memory of its PID read 2 s after each response. ORT's arena never shrinks, so the reading after a
|
|
request is the high-water mark so far. Prefixes of the public SCOTUS audio (16 kHz mono).
|
|
usage: len_sweep.py ARM PORT SECONDS..."""
|
|
import io
|
|
import json
|
|
import subprocess
|
|
import sys
|
|
import time
|
|
|
|
import soundfile as sf
|
|
|
|
sys.path.insert(0, "/tank/spikes/parakeet-ab/code")
|
|
from bench import post # noqa: E402
|
|
|
|
arm, port = sys.argv[1], sys.argv[2]
|
|
lens = [float(x) for x in sys.argv[3:]]
|
|
pids = set(subprocess.run(["docker", "top", arm, "-eo", "pid"], capture_output=True, text=True).stdout.split()[1:])
|
|
a, sr = sf.read("/tank/spikes/scriberr-slicer/public/scotus.wav", dtype="float32")
|
|
|
|
|
|
def mem():
|
|
out = subprocess.run(["nvidia-smi", "--query-compute-apps=pid,used_memory", "--format=csv,noheader,nounits"],
|
|
capture_output=True, text=True).stdout
|
|
return sum(int(l.split(",")[1]) for l in out.splitlines() if l.split(",")[0].strip() in pids)
|
|
|
|
|
|
print(json.dumps(dict(arm=arm, seconds=0, status=None, mib=mem())), flush=True)
|
|
for s in lens:
|
|
bio = io.BytesIO()
|
|
sf.write(bio, a[: int(s * sr)], sr, format="WAV", subtype="PCM_16")
|
|
r = post(f"http://127.0.0.1:{port}/v1/audio/transcriptions", bio.getvalue(), timeout=1800)
|
|
time.sleep(2)
|
|
print(json.dumps(dict(arm=arm, seconds=s, status=r["status"], e2e_ms=r["e2e_ms"], mib=mem())), flush=True)
|