A/B of the live STT seat (fv-ml1 GPU 0, sherpa-onnx int8 v3) against nvidia/parakeet-unified-en-0.6b, measured on GPU 3 with the seat's own image, k2-fsa's published unified int8 export, fp32/fp16 exports made with k2-fsa's recipe, v2 int8, and NeMo 3.0.0 (fp32, bf16 autocast, bf16 weights). - Seat int8 graph runs on one CPU thread (cpu/wall 1.00, GPU 2-9%). - unified-en under NeMo: -121/-234/-530 ms vs the seat at 1-3/3-8/8-20 s (paired, n=120/bin; floor <=6 ms; +50 ms positive control reads +52-54). - unified-en WER lower in every runtime: -0.7 pp clean, -1.5 pp other, -3.2 to -4.4 pp AMI (paired CIs exclude 0). - Seat defects found: hard 400 s input ceiling (HTTP 500), truncation after a quiet 1.5 s pause, and severe long-window dropouts (int8 v3 only). - B-bf16w needs +0.8 to +1.5 GB over the seat's 1,690 MiB on GPU 0. Raw requests, hypotheses, manifests and the full harness under services/parakeet-ab-2026-09-30/. No deploy; live seat untouched apart from 240 light test requests.
17 lines
714 B
Bash
Executable File
17 lines
714 B
Bash
Executable File
#!/usr/bin/env bash
|
|
# Long-form for ONE arm, with a GPU 3 headroom guard for Scriberr (an on-demand tenant of GPU 3, ~5.5 GB
|
|
# per job): a watcher kills the arm container if free memory on GPU 3 drops under 8 GiB.
|
|
# usage: long_guarded.sh ARM PORT
|
|
cd /tank/spikes/parakeet-ab
|
|
arm=$1 port=$2
|
|
( while docker inspect "$arm" >/dev/null 2>&1; do
|
|
f=$(nvidia-smi -i 3 --query-gpu=memory.free --format=csv,noheader,nounits)
|
|
if [ "$f" -lt 8192 ]; then echo "GUARD: GPU3 free ${f} MiB < 8192, stopping $arm" ; docker rm -f "$arm" >/dev/null; break; fi
|
|
sleep 0.5
|
|
done ) &
|
|
w=$!
|
|
echo "phase start $(date +%s.%N) $arm"
|
|
python3 code/run_long.py "$arm:$port"
|
|
echo "phase end $(date +%s.%N) $arm"
|
|
kill $w 2>/dev/null
|