Files
esh-pfi-infrastructure/services/parakeet-ab-2026-09-30/code/mk_env.sh
T
vh a6c1d3c454 docs(parakeet): seat A/B vs parakeet-unified-en-0.6b - latency is the int8-on-CPU runtime; unified wins WER
A/B of the live STT seat (fv-ml1 GPU 0, sherpa-onnx int8 v3) against
nvidia/parakeet-unified-en-0.6b, measured on GPU 3 with the seat's own image,
k2-fsa's published unified int8 export, fp32/fp16 exports made with k2-fsa's
recipe, v2 int8, and NeMo 3.0.0 (fp32, bf16 autocast, bf16 weights).

- Seat int8 graph runs on one CPU thread (cpu/wall 1.00, GPU 2-9%).
- unified-en under NeMo: -121/-234/-530 ms vs the seat at 1-3/3-8/8-20 s
  (paired, n=120/bin; floor <=6 ms; +50 ms positive control reads +52-54).
- unified-en WER lower in every runtime: -0.7 pp clean, -1.5 pp other,
  -3.2 to -4.4 pp AMI (paired CIs exclude 0).
- Seat defects found: hard 400 s input ceiling (HTTP 500), truncation after
  a quiet 1.5 s pause, and severe long-window dropouts (int8 v3 only).
- B-bf16w needs +0.8 to +1.5 GB over the seat's 1,690 MiB on GPU 0.

Raw requests, hypotheses, manifests and the full harness under
services/parakeet-ab-2026-09-30/. No deploy; live seat untouched apart
from 240 light test requests.
2026-09-30 18:51:44 -07:00

12 lines
928 B
Bash
Executable File

#!/usr/bin/env bash
# Serving env for parakeet-unified-en-0.6b under NeMo torch: same pins as the investigation's nemo300
# (torch 2.8 cu128, nemo_toolkit[asr]==3.0.0) plus the seat's HTTP stack. Built inside the scriberr image
# (its /usr/bin/python3 is 3.13). Own uv cache; the investigation's env and cache are not touched.
set -e
export UV_CACHE_DIR=/ab/envs/.uvcache UV_LINK_MODE=copy
cd /ab/envs
uv venv -q --python /usr/bin/python3 nemo300-serve
VIRTUAL_ENV=/ab/envs/nemo300-serve uv pip install -q --index-url https://download.pytorch.org/whl/cu128 --extra-index-url https://pypi.org/simple \
"torch==2.8.*" "torchaudio==2.8.*" "nemo_toolkit[asr]==3.0.0" fastapi "uvicorn[standard]" python-multipart soundfile httpx 2>&1 | tail -3
/ab/envs/nemo300-serve/bin/python -c "import nemo, torch, fastapi; print('nemo', nemo.__version__, 'torch', torch.__version__, 'fastapi', fastapi.__version__, torch.cuda.is_available())"