Files
esh-pfi-infrastructure/services/parakeet-ab-2026-09-30/code/memtrace.py
T
vh a6c1d3c454 docs(parakeet): seat A/B vs parakeet-unified-en-0.6b - latency is the int8-on-CPU runtime; unified wins WER
A/B of the live STT seat (fv-ml1 GPU 0, sherpa-onnx int8 v3) against
nvidia/parakeet-unified-en-0.6b, measured on GPU 3 with the seat's own image,
k2-fsa's published unified int8 export, fp32/fp16 exports made with k2-fsa's
recipe, v2 int8, and NeMo 3.0.0 (fp32, bf16 autocast, bf16 weights).

- Seat int8 graph runs on one CPU thread (cpu/wall 1.00, GPU 2-9%).
- unified-en under NeMo: -121/-234/-530 ms vs the seat at 1-3/3-8/8-20 s
  (paired, n=120/bin; floor <=6 ms; +50 ms positive control reads +52-54).
- unified-en WER lower in every runtime: -0.7 pp clean, -1.5 pp other,
  -3.2 to -4.4 pp AMI (paired CIs exclude 0).
- Seat defects found: hard 400 s input ceiling (HTTP 500), truncation after
  a quiet 1.5 s pause, and severe long-window dropouts (int8 v3 only).
- B-bf16w needs +0.8 to +1.5 GB over the seat's 1,690 MiB on GPU 0.

Raw requests, hypotheses, manifests and the full harness under
services/parakeet-ab-2026-09-30/. No deploy; live seat untouched apart
from 240 light test requests.
2026-09-30 18:51:44 -07:00

27 lines
1.3 KiB
Python

"""Trace one arm's GPU memory against the requests it had served: when did the arena grow, and after what?
usage: memtrace.py ARM (reads out/raw/pids.json, mem.csv, first/lat/conc jsonl)"""
import datetime, json, sys
AB = "/tank/spikes/parakeet-ab"
arm = sys.argv[1]
pids = set(map(str, json.load(open(f"{AB}/out/raw/pids.json"))[arm]))
series = []
for l in open(f"{AB}/out/raw/mem.csv"):
p = [x.strip() for x in l.split(",")]
if len(p) == 4 and p[2] in pids:
series.append((datetime.datetime.strptime(p[0], "%Y/%m/%d %H:%M:%S.%f").timestamp(), int(p[3])))
reqs = []
for f in ("first", "lat", "conc"):
for l in open(f"{AB}/out/raw/{f}.jsonl"):
r = json.loads(l)
if r.get("arm") == arm and "e2e_ms" in r:
reqs.append((r["t_wall"], round(r["dur"] - r.get("trim_ms", 0) / 1000, 2), r.get("mode", "first")))
reqs.sort()
prev, i, longest, last = None, 0, 0.0, None
for t, m in series:
while i < len(reqs) and reqs[i][0] <= t:
longest = max(longest, reqs[i][1]); last = reqs[i]; i += 1
if m != prev:
ts = datetime.datetime.fromtimestamp(t).strftime("%H:%M:%S")
print(f"{ts} {m:6d} MiB longest served so far {longest:5.1f}s last request {last[1] if last else '-'}s ({last[2] if last else '-'})")
prev = m