A/B of the live STT seat (fv-ml1 GPU 0, sherpa-onnx int8 v3) against nvidia/parakeet-unified-en-0.6b, measured on GPU 3 with the seat's own image, k2-fsa's published unified int8 export, fp32/fp16 exports made with k2-fsa's recipe, v2 int8, and NeMo 3.0.0 (fp32, bf16 autocast, bf16 weights). - Seat int8 graph runs on one CPU thread (cpu/wall 1.00, GPU 2-9%). - unified-en under NeMo: -121/-234/-530 ms vs the seat at 1-3/3-8/8-20 s (paired, n=120/bin; floor <=6 ms; +50 ms positive control reads +52-54). - unified-en WER lower in every runtime: -0.7 pp clean, -1.5 pp other, -3.2 to -4.4 pp AMI (paired CIs exclude 0). - Seat defects found: hard 400 s input ceiling (HTTP 500), truncation after a quiet 1.5 s pause, and severe long-window dropouts (int8 v3 only). - B-bf16w needs +0.8 to +1.5 GB over the seat's 1,690 MiB on GPU 0. Raw requests, hypotheses, manifests and the full harness under services/parakeet-ab-2026-09-30/. No deploy; live seat untouched apart from 240 light test requests.
24 lines
1.3 KiB
Python
24 lines
1.3 KiB
Python
"""fp16 copy of an fp32 sherpa-onnx transducer export: onnxconverter-common float16, keep_io_types=True
|
|
(inputs/outputs stay fp32, so sherpa-onnx feeds it exactly as before); metadata carried over.
|
|
Shape inference runs first BY PATH (infer_shapes_path handles the >2 GB fp32 encoder), so the converter
|
|
sees every intermediate type; without it a scalar Mul in pre_encode is left fp32 and the graph won't load.
|
|
usage: convert_fp16.py SRC_DIR DST_DIR"""
|
|
import os, shutil, sys, tempfile
|
|
import onnx
|
|
from onnx.shape_inference import infer_shapes_path
|
|
from onnxconverter_common import float16
|
|
src, dst = sys.argv[1:3]
|
|
os.makedirs(dst, exist_ok=True)
|
|
for m in ("encoder", "decoder", "joiner"):
|
|
inferred = f"{src}/{m}.inferred.onnx"
|
|
infer_shapes_path(f"{src}/{m}.onnx", inferred)
|
|
model = onnx.load(inferred)
|
|
# the conv subsampling front (pre_encode, ~0.1 % of the FLOPs) stays fp32: the converter mis-types its
|
|
# length-mask Cast/Mul otherwise
|
|
keep32 = [n.name for n in model.graph.node if n.name.startswith("/pre_encode/")]
|
|
m16 = float16.convert_float_to_float16(model, keep_io_types=True, disable_shape_infer=True, node_block_list=keep32)
|
|
onnx.save(m16, f"{dst}/{m}.fp16.onnx")
|
|
os.remove(inferred)
|
|
shutil.copy(f"{src}/tokens.txt", f"{dst}/tokens.txt")
|
|
print("ok", dst)
|