Files
esh-pfi-infrastructure/docs/pfi/parakeet-seat-ab-2026-09-30.md
T

41 KiB
Raw Blame History

Parakeet seat A/B: the live seat vs parakeet-unified-en-0.6b (2026-09-30)

Prime's ask, via the coordinator, 2026-09-30: "a/b the one on fv-ml1's general seat against the unified new one in jun for speed and accuracy for english. that seat demands low latency and any win, even 50ms, is load-bearing." A/B only. Nothing was deployed. The live parakeet container was not restarted or reconfigured: it got 240 light requests (§ 4.3). Every candidate ran as a transient container on fv-ml1 GPU 3, and all of them are gone. Raw data and code: services/parakeet-ab-2026-09-30/.

1. Headline

  • Yes on both counts, with one big qualifier: the latency win comes from the runtime, not the model.
  • Latency. parakeet-unified-en-0.6b under NeMo torch is 121, 234 and 530 ms faster than the seat at 1–3, 3–8 and 8–20 s (median of 120 paired requests per bin; 95 % CIs ±4 to ±35 ms). The A-vs-A floor is ≤ 6 ms (§ 4.2), and the +50 ms positive control reads +52 to +54 ms at every length, so a 50 ms difference is resolvable at every length. That arm answers a 3–8 s turn in 27 ms against the seat's 260 ms, or 308 ms on the live seat itself.
  • Accuracy. In every runtime tried, unified-en has a lower English WER than the seat on every set: about 0.7 pp lower on LibriSpeech test-clean, 1.5 pp on test-other and 3–4 pp on AMI meetings. Every paired CI excludes zero.
  • The same model in the seat's runtime is NOT faster. k2-fsa publish unified-en as an int8 sherpa-onnx export, and it runs in the seat's own image unchanged (arm C), but it is 8–71 ms slower than the seat. The reason is the runtime: the seat's int8 graph runs on ONE CPU thread. CPU time equals wall time on every request and GPU utilisation is 2–9 %. The proof is the seat's own v3 weights, exported to fp32 with k2-fsa's recipe, in the seat's own image: 4–12× faster, and more accurate on test-other and AMI.
  • Best arm measured: unified-en, NeMo, bf16 weights (B-bf16w). Its p50 is 23 / 27 / 33 / 62 ms against the seat's 144 / 260 / 565 / 2,027 ms on the same card. WER is identical to fp32 (399/400 utterances with the same edits). It needs 2.5–3.2 GB, which is +0.8 to +1.5 GB over today's 1,690 MiB, so it does not fit GPU 0 in place (§ 6).
  • Long form. On 6-minute files sent whole, which is inside what the seat can take, the seat's config lost 320–1,676 words of clean speech on the audiobook, depending only on where the cuts fell. unified-en lost none (§ 5.4).
  • Two defects in the live seat surfaced on the way, independent of any switch:
    • It cannot take a file longer than 400 s (6 min 40 s). The ONNX export freezes the relative-position table at 5,000 frames, so anything longer returns HTTP 500 (§ 5.4).
    • On a 1.5 s stretch of digital silence inside an utterance it stops transcribing, and the rest of the utterance is lost. That is the v3 int8 export only: the same weights in fp32 do not do it, nor do unified or v2 (§ 5.3).

2. What was compared

Everything except arm A ran as a transient container on fv-ml1 GPU 3, bound to 127.0.0.1.

  • Every sherpa-onnx arm runs the seat's own image (local/parakeet:sherpa-onnx-v4, sherpa-onnx 1.12.39+cuda12.cudnn9, ORT CUDA EP) with the seat's env (PROVIDER=cuda, NUM_THREADS=1).
  • They also run the seat's app.py, plus a timing header (services/parakeet-ab-2026-09-30/code/app_timed.py). Its only differences from stacks/parakeet/app.py are the header, an injected-delay knob (for the positive control) and a model file-suffix knob.
  • The live container runs exactly stacks/parakeet/app.py (sha256 6a0bd695…, checked inside the container).
  • The NeMo arms use the same HTTP shape: the model loads at import, one warm-up decode runs on 1 s of silence, and async handlers call a blocking decode, so requests are serialised exactly as in the seat.
arm model weights / export runtime where
A parakeet-tdt-0.6b-v3 (TDT, 25 languages) k2-fsa int8 (the live /tank/parakeet/models) the live seat GPU 0, :8300
A′ #1, #2 same the same files, read-only seat image GPU 3
A′ +50 ms same same seat image, 50 ms sleep added per request GPU 3 (positive control)
A′ plain same same seat image, its own app.py untouched GPU 3 (null control)
A′ 16 threads same same seat image, NUM_THREADS=16 (a config-only change) GPU 3
A-fp32 the same v3 weights fp32 ONNX, our export with k2-fsa's recipe seat image GPU 3
C parakeet-unified-en-0.6b (RNN-T, English) k2-fsa int8, published seat image GPU 3
C-fp32 same fp32 ONNX, our export with k2-fsa's recipe seat image GPU 3
C-fp16 same C-fp32 converted to fp16 (pre_encode kept fp32) seat image GPU 3
D parakeet-tdt-0.6b-v2 (TDT, English) k2-fsa int8, published seat image GPU 3
B-fp32 parakeet-unified-en-0.6b the .nemo, HF fe53cd88 NeMo 3.0.0 + torch 2.8 cu128 GPU 3
B-bf16 same same same, torch.autocast bf16 (fp32 weights) GPU 3
B-bf16w same same, encoder/decoder/joint weights cast to bf16, loaded via CPU same GPU 3

Arm C needed no export of ours: k2-fsa publish it.

  • Asset. sherpa-onnx-nemo-parakeet-unified-en-0.6b-int8-non-streaming.tar.bz2 is on the sherpa-onnx asr-models release (2026-04-27; sha256 99f63605…e150, matched).
  • Provenance. It comes from scripts/nemo/parakeet-unified-en-0.6b/export_onnx.py (k2-fsa PR #3556, sherpa-onnx 1.13.0).
  • Shape. It is a plain EncDecRNNTBPEModel with full-context attention (att_context_size [-1,-1]; "contains only the non-streaming part").
  • It loads and runs in the seat's pinned sherpa-onnx 1.12.39 unchanged.
  • Out of scope. k2-fsa also publish 240, 560 and 1,120 ms streaming variants.
  • No Hugging Face mirror. An authenticated call on csukuangfj/…unified… returns 404.

The fp32 and fp16 arms were added because of what the pilot found (§ 4.1).

  • A-fp32 and C-fp32 are k2-fsa's export script with only the final quantize_dynamic removed: the same encoder/decoder/joint.export() and the same metadata. Their tokens.txt is byte-identical to the seat's and to k2-fsa's.
  • C-fp16 is C-fp32 put through onnxconverter-common 1.16.0 with fp32 I/O. The conv subsampling front had to stay fp32 or the graph will not load.
  • v3 in fp16 is broken (every token <unk>), so it has no arm.

Not measured: k2-fsa's v2 fp16 tarball, TensorRT, torch.compile, batching, and the unified streaming exports.

3. What the seat actually serves

  • Callers.
    • talk (latency-critical). tts-stack's voice loop on nh3-dev: push-to-talk and hands-free VAD turns. The browser decimates 48 kHz to 16 kHz mono WAV with a 10 MiB body cap, and it calls through LiteLLM ext-stt on the all-agents-local key (python-httpx, a new client per call).
    • tts-stack batch tooling (the rest). venus_score.py calls :8300 directly; it sent the 308 requests on 2026-09-28 1327–1332 PDT, none of which are in the gateway log. The accent and tag probes sent 226 requests of 48 kHz WAV between 09-15 and 09-22.
    • whisper-1 has had zero calls.
  • Volume. 1,112 transcriptions from 2026-09-15 to 09-30 (the container's whole life), all of them 200. talk makes a handful to a few dozen turns a day. Concurrency is effectively 1.
  • Sizes. Neither log records audio length, so the length mix is inferred.
    • The gateway's 58 ext-stt calls since 09-23 took p50 0.31 s, p90 0.47 s, max 0.71 s end to end. Its log is capped at 10,000 rows, so it only reaches back to 09-23.
    • Read against the measured gateway path (§ 4.3), that is a median talk turn of about 3–5 s and a longest of about 10 s.
    • That puts most turns in the 1–8 s bins, with a tail into 8–20 s. The 20–60 s bin is batch tooling.
  • Content. English: Prime's voice turns and English WER tooling. No caller reads a language field, and the seat never returned one.

4. Latency

4.1 Why the seat is slow (the pilot finding that reshaped the arms)

  • The seat's int8 graph runs on the CPU. With NUM_THREADS=1, the process's CPU time equals the request's wall time on every clip (ratio 0.97–1.05, from host /proc), while GPU 3's SM utilisation sits at 2–9 %.
  • The cause is ORT's int8 path. k2-fsa's export uses quantize_dynamic (MatMulInteger / DynamicQuantizeLinear), which ORT's CUDA EP leaves to the CPU.
  • More threads help, but poorly. At NUM_THREADS=16 the seat is about 40 % faster and burns 12–29 cores spinning.
  • The same weights exported in fp32 run on the GPU. In the same image they are 4–12× faster.
  • Consequence. The model and the runtime are separate variables. Arm C, the new model in the old runtime, cannot show the new model's speed.

4.2 Harness

  • Client. Stdlib Python on the fv-ml1 host (loopback), one new HTTP connection per request (as talk's per-call httpx client does), multipart file=turn.wav + model=ext-stt, every body built before the clock starts.
  • What is timed. End-to-end (e2e) runs from connect to the last byte. Server-side is the x-ab-decode-ms header: audio decode, features, model and text inside the handler. HTTP adds 1.3–5.8 ms on top of server-side. The live seat has no header, so A is e2e only.
  • Clips. 20 per bin from LibriSpeech test-clean, at evenly spaced duration quantiles:
    • 1–3 s: 1.8–3.0 s;
    • 3–8 s: 3.1–7.8 s;
    • 8–20 s: 8.1–18.9 s;
    • 20–60 s: consecutive utterances of one chapter joined with 0.25 s gaps (22–59 s).
  • Every request has a never-seen length. ORT tunes per input shape and NeMo's CUDA-graph decoder captures per maximum length, so repeating a clip flatters both. In the pilot, the first call at a new shape cost the seat +30–100 ms and fp32 ONNX +20–45 ms.
    • Each request therefore trims a seeded 0–490 ms (10 ms steps) off the tail.
    • The trim is seeded by round and clip, never by arm, so in a given round every arm received byte-identical input, and every delta is a median of paired per-request differences.
  • Warm, single stream. 6 rounds × 80 clips, so 120 requests per bin per arm. Within a round the arms run in a seeded shuffled order, so drift (thermal, a Scriberr job) lands on all of them.
  • Concurrency 4. Per arm and bin, 40 requests from 4 workers back to back.
  • Cold. Each arm was started fresh, one at a time, and timed from docker StartedAt to the first healthy /healthz. Then one never-seen clip went in per bin, in ascending length.
  • Statistics. p50/p90/p99 over requests, with 95 % bootstrap CIs (2,000 resamples). p99 over 120 requests is roughly the second-largest sample: it is reported, and it is weak.
  • Blocks. Block 1 ran 12 arms. Block 2 ran the memory-lean NeMo variants with B-fp32, B-bf16 and both fp32 ONNX arms as anchors.

Controls (block 1), as the paired median difference against A′ #1, in ms [95 % CI]:

1–3 s 3–8 s 8–20 s 20–60 s
floor: A′ #2 (an identical second instance) +1.0 [+0.7, +2.1] +1.5 [+0.9, +2.9] +1.8 [+0.3, +3.2] +2.2 [−4.1, +7.6]
null: A′ plain (the image's own app, no header) +4.4 [+3.9, +5.1] +5.8 [+4.9, +6.8] +6.6 [+4.7, +10.4] +16.8 [+13.2, +21.0]
positive: A′ +50 ms injected +52.6 [+52.1, +53.1] +53.2 [+52.5, +53.8] +52.3 [+50.8, +53.2] +53.9 [+50.6, +56.1]
GPU-arm floor: B-fp32 loaded via CPU vs B-fp32 (block 2) +0.0 +0.1 −0.1 +1.2
  • The null control registered a small effect: +4 to +17 ms, about 1–3 %.
    • With one instance per condition I cannot tell how much of it is the code path (the plain app's dict return goes through FastAPI's validation) and how much is per-instance placement of a CPU-bound single thread on a multi-socket host.
    • So it is treated as part of the floor. The A-vs-A floor is ≤ 6 ms for 1–20 s and ≤ 21 ms for 20–60 s on the CPU-bound arms, and about 1 ms on the GPU arms.
  • Sensitivity. The positive control reads +50 ms (plus about 3 ms that the floor explains) at every length, with CI half-widths ≤ 3 ms. A 50 ms difference is therefore resolvable at every length, and so is anything above about 10 ms (1–20 s) or about 25 ms (20–60 s).

4.3 Results: warm, single stream

End-to-end milliseconds, p50 / p90 / p99, 120 requests per bin, fv-ml1 loopback, GPU 3 unless noted. Block 1 except where marked.

arm 1–3 s 3–8 s 8–20 s 20–60 s
A, the live seat (GPU 0, n = 40 per bin) 187 / 221 / 232 max 308 / 413 / 454 max 626 / 862 / 956 max not sent
A, via LiteLLM ext-stt from nh3-dev (n = 20) 257 / 297 / 309 max 388 / 494 / 511 max 711 / 968 / 1,095 max not sent
A′, the seat's config 144 / 170 / 184 260 / 370 / 384 565 / 814 / 896 2,027 / 2,634 / 2,768
A′ 16 threads 85 / 98 / 104 143 / 185 / 197 318 / 438 / 507 1,171 / 1,546 / 1,610
A-fp32 (v3 fp32 ONNX, seat image) 38 / 41 / 44 47 / 53 / 57 67 / 80 / 87 167 / 219 / 231
C (unified int8, seat image) 156 / 180 / 197 271 / 389 / 414 591 / 845 / 942 2,101 / 2,795 / 2,943
C-fp32 (unified fp32 ONNX, seat image) 41 / 43 / 67 50 / 58 / 63 76 / 91 / 98 186 / 248 / 260
C-fp16 89 / 91 / 116 99 / 107 / 111 124 / 141 / 148 244 / 307 / 318
D (v2 int8, seat image) 150 / 174 / 188 264 / 379 / 390 565 / 806 / 901 2,004 / 2,718 / 2,767
B-fp32 (unified, NeMo) 23 / 25 / 29 27 / 29 / 30 34 / 40 / 62 77 / 98 / 258
B-bf16 (autocast) 29 / 31 / 55 32 / 36 / 39 39 / 43 / 46 67 / 83 / 95
B-bf16w (bf16 weights; block 2) 23 / 26 / 51 27 / 30 / 43 33 / 38 / 207 62 / 77 / 269

Paired median difference against A′, in ms [95 % CI]. Negative means faster than the seat:

arm 1–3 s 3–8 s 8–20 s 20–60 s
A′ 16 threads −61 [−63, −58] −119 [−129, −112] −249 [−265, −233] −844 [−910, −747]
A-fp32 −107 [−114, −104] −213 [−235, −205] −495 [−528, −468] −1,859 [−1,979, −1,616]
C (unified int8) +8 [+7, +8] +12 [+11, +14] +22 [+19, +24] +71 [+62, +80]
C-fp32 −104 [−109, −101] −209 [−227, −200] −488 [−517, −461] −1,839 [−1,942, −1,599]
C-fp16 −55 [−61, −51] −160 [−181, −153] −439 [−472, −410] −1,780 [−1,880, −1,527]
D (v2 int8) +3 [+2, +3] +4 [+3, +4] +3 [+2, +5] +4 [+0, +10]
B-fp32 −121 [−128, −117] −234 [−254, −226] −530 [−565, −499] −1,947 [−2,058, −1,699]
B-bf16 −115 [−122, −112] −229 [−249, −218] −527 [−561, −491] −1,959 [−2,064, −1,706]
B-bf16w vs B-fp32 (block 2, paired) +0.5 [+0.4, +0.6] +0.9 [+0.4, +1.2] −0.3 [−0.8, +0.3] −13.7 [−15.2, −12.2]
A (live) vs A′ +37 [+34, +38] +40 [+38, +41] +42 [+36, +47] —
A via gateway vs A (live) +56 [+50, +76] +76 [+59, +79] +86 [+79, +97] —
  • C is slower than the seat, not faster. The difference sits just above the floor at 1–20 s and is clear at 20–60 s, and it is consistent with RNN-T decoding every frame where TDT skips frames. D is the same speed as the seat (+3 to +4 ms, inside the floor).
  • The live seat is 37–42 ms slower than its byte-identical copy on GPU 3. They gave identical text on 120/120 inputs. GPU 0 is the card the vLLM seats share, and its allocator is pinned at 1,690 MiB with 101 MiB free. Which of those costs the 40 ms is not measured.
  • The gateway path adds 56–86 ms over loopback. About 24–29 ms of it is nh3-dev → fv-ml1 transit (measured direct to :8300 from nh3-dev). The LiteLLM hop itself is 36–61 ms and grows with the body. An earlier note put this hop "below harness resolution (±30 ms)"; paired, it resolves.
  • So talk's real path today is p50 257 / 388 / 711 ms for 1–3 / 3–8 / 8–20 s turns. With B on the same path that would be about 80 / 100 / 120 ms. That is a derived figure: B's loopback e2e plus the measured gateway-path cost. It excludes any GPU 0 placement penalty; the seat itself pays +37–42 ms for its GPU 0 placement.
  • Live-seat safety. The live seat got only 1–20 s clips (120 on loopback, 60 via the gateway, 60 direct), one per second, with an abort on the first non-200. All 240 returned 200. Its restart count stayed at 0, it stayed healthy, and its GPU memory went from 1,690 to 1,692 MiB.

4.4 Concurrency 4, cold start and first call

Concurrency 4: e2e p50 in ms, then throughput in × realtime. Every server serialises requests (the seat's async handler blocks the event loop), so e2e is about 4× the decode time.

arm 1–3 s 3–8 s 8–20 s 20–60 s ×realtime
A′ (seat config) 569 1,062 2,304 8,313 15–21×
A′ 16 threads 335 572 1,299 4,647 26–37×
A-fp32 138 181 265 662 63–254×
C 636 1,137 2,415 8,763 14–20×
C-fp32 148 199 295 786 59–218×
D 610 1,104 2,350 8,644 15–21×
B-fp32 85 98 129 308 101–561×
B-bf16w (block 2) 84 97 124 236 101–709×

Cold start: from container start to healthy.

  • The seat image takes 45–48 s for every ONNX arm. That is the seat's own warm-up decode compiling CUDA kernels; it measured 45.1 s on the live seat's boot.
  • NeMo takes 12–13 s.

First call after start (one never-seen clip per bin, in ascending length) compared with that arm's warm p50:

  • The seat: 223 / 357 / 649 / 2,904 ms, against 144 / 260 / 565 / 2,027 ms warm.
  • B-fp32: 60 / 44 / 38 / 410 ms, against 23 / 27 / 34 / 77 ms warm.
  • B's 410 ms is NeMo's CUDA-graph decoder capturing for a new maximum length. It is paid once per new maximum, which is also the 260–300 ms p99 in B's 20–60 s bin.
  • The fix is a warm-up decode at the longest length the seat should serve. It costs nothing, though I have not measured it.

5. Accuracy (English, against ground truth)

Sets. All are public, with licences read at source:

  • LibriSpeech test-clean and test-other: 400 utterances each, a seeded random sample (openslr/librispeech_asr@71cacbfb, CC-BY-4.0). 49.9 and 45.0 minutes; 8,144 and 7,370 words.
  • AMI meetings, IHM headset mics: 400 test utterances of ≥ 1 s, a seeded random sample covering all 16 test meetings (edinburghcstr/ami@46f28f25, CC-BY-4.0). 23.8 minutes, 4,121 words. This is the conversational set.
  • Long form: the investigation's two public files with their ground truth (§ 5.4).

Scoring. Whisper's English normaliser (transformers 4.53.3, the investigation's version, no spelling map), applied to the whole utterance on both sides, then exact Levenshtein. Long-form uses the investigation's own gtscore.py, copied verbatim (sha256 5012a552…).

5.1 WER

WER % seat (v3 int8) v3 fp32 ONNX C: unified int8 C-fp32 C-fp16 D: v2 int8 B-fp32 B-bf16 B-bf16w
LS test-clean 2.70 2.43 1.94 1.89 1.89 2.11 1.98 1.98 1.97
LS test-other 4.56 3.31 2.93 3.03 3.01 3.39 3.11 3.11 3.09
AMI IHM 12.69 10.48 9.51 9.37 9.32 10.02 8.37 8.28 8.30
AMI empty outputs (of 400) 26 18 17 15 15 15 10 9 9

Paired bootstrap of WER minus the seat's WER, in pp [95 % CI]:

test-clean test-other AMI
C (unified int8, seat runtime) −0.76 [−1.11, −0.41] −1.63 [−2.15, −1.14] −3.18 [−5.13, −1.48]
C-fp32 −0.81 [−1.18, −0.45] −1.53 [−2.06, −1.04] −3.32 [−5.25, −1.60]
B-fp32 / B-bf16w −0.72 / −0.74 [≈ −1.1, −0.35] −1.45 / −1.47 [≈ −2.0, −0.95] −4.32 / −4.39 [≈ −6.5, −2.4]
D (v2 int8) −0.59 [−0.96, −0.26] −1.17 [−1.68, −0.68] −2.67 [−4.42, −1.05]
A-fp32 (the same v3 weights, fp32) −0.27 [−0.60, +0.05] −1.25 [−1.70, −0.85] −2.21 [−3.68, −0.89]
  • unified-en beats the seat on English in every runtime, on every set.
  • The three unified runtimes tie on read speech. C, C-fp32 and B are within ±0.18 pp of each other on LibriSpeech, and every CI includes zero.
  • On AMI, NeMo is about 1 pp better than the sherpa runtimes: C − B = +1.14 [+0.28, +2.16]; C-fp32 − B = +1.00 [+0.24, +1.92]. That is the same weights through a different front end and decoder.
  • The seat's int8 quantisation of v3 itself costs 1.2 pp on test-other and 2.2 pp on AMI.
  • bf16 weights cost nothing. B-bf16w gives the same edit counts as B-fp32 on 399/400 utterances per set.

5.2 Controls and determinism

  • Scorer self-test (positive and null). One deleted, one substituted and one inserted word each register exactly once, and casing, punctuation or a filler register zero. PASS.
  • Determinism.
    • A′ #1 = A′ #2 = A′ at 16 threads: 400/400 identical text.
    • The live seat = A′ on 120/120 inputs, so A′'s accuracy is the seat's.
    • The gateway path = the live seat on 60/60.
    • NeMo's direct decode path = NeMo's transcribe() on 400/400, so the thin wrapper loses nothing.
  • Positive control (the pipeline).
    • The same 40 test-clean utterances (≥ 6 s) had 1.5 s of digital silence placed at 40 % of their length, with the reference unchanged.
    • Every arm registered extra deletions on 40 of 40 utterances: from 1–2 to 167–306 words.
  • Null control. All 400 test-clean utterances at −0.5 dB. The WER change sits inside a CI that includes zero for every arm (the seat: +0.04 [−0.08, +0.16] pp).
  • Sensitivity. With 400 utterances, a paired WER difference of about 0.35 pp (test-clean), 0.5 pp (test-other) or 1.5 pp (AMI) is resolvable. Smaller ones are not.

5.3 The seat loses the rest of an utterance after a quiet pause (int8 v3 only)

The positive control should cost each arm the four or so words under the 1.5 s silence.

  • Every arm except the seat lost 167–185 words across the 40 utterances. The seat lost 306.
  • On 6 of the 40 its output simply stops at the silence, and the rest of the utterance (about 21 words each, on average) is gone. Example, with the silence at 8.1–9.6 s of 20.3 s:
    • v3 int8 (the seat): "…than you could help running if you heard the wheel."
    • v3 fp32 (the same weights, the same runtime): "…than you could help running if you heard little on the other end of the house. The voice would go to your heart…"
  • The same 40 utterances with other fillers. The count is utterances that lost ≥ 8 words more than without the gap:
gap fill seat (v3 int8) v3 fp32 C (unified int8) B-bf16w D (v2 int8)
1.5 s digital silence 6/40 (306 deletions) 1/40 (185) 1/40 (181) 1/40 (177) 0/40 (167)
1.5 s white noise at −60 dBFS 6/40 (337) 2/40 (196) 1/40 (180) 1/40 (176) 0/40 (165)
1.5 s white noise at −50 dBFS 0/40 (170) 4/40 (227) 1/40 (178) 1/40 (179) 0/40 (171)
0.75 s digital silence 0/40 (88) 0/40 (78) 0/40 (86) 0/40 (84) 0/40 (73)

Reading.

  • When it happens. The seat truncates after a 1.5 s pause that is silent or nearly so (≤ −60 dBFS). A louder room tone (−50 dBFS) or a shorter gap (0.75 s) does not trigger it.
  • v3 has a gap sensitivity of its own, and int8 moves where it bites: fp32 shows 4/40 at −50 dBFS.
  • unified and v2 are steady across all four fills.
  • Why it could matter for talk. Browser capture with noise suppression can produce near-digital silence in a pause, so a 1.5 s mid-sentence pause in a push-to-talk turn could cost the rest of the turn. That is not measured on real talk audio, which is private and was not used.
  • Basis. Measured: 40 utterances × 4 fills, deterministic arms.

5.4 Long form, and the seat's 400 s ceiling

The ceiling.

  • Cause. k2-fsa's export bakes the encoder's relative-position table at pos_emb_max_len 5000, and our fp32 exports inherit it.
  • Effect. An input over 5,000 encoder frames (400 s) fails in layer 0's attention: /layers.0/self_attn/Add_2: right operand cannot broadcast … {1,8,T,T} vs {1,8,T,9999}. The seat then returns HTTP 500 after 2–11 s.
  • Measured on a fresh seat-config instance: 395 s → 200, 410 s → 500. The 24- and 30-minute files → 500 (4 of 4). The fp32 unified export returned 500 on the 24-minute file too.
  • Every sherpa arm carries the same table. NeMo extends its table on the fly, so B has no ceiling (below).

The seat's GPU memory climbs with length. Fresh instance, ascending lengths, ORT arena high-water mark:

input length 15 s 30 s 45–60 s 90–150 s 180–300 s 360–395 s 410 s
seat-config arena (MiB) 924 1,180 1,692 2,716 3,742 7,838 HTTP 500
e2e (s) 0.8 1.6 2.3–2.9 4.7–7.4 9.5–15.5 19.4–20.5 —
  • Read the table as indicative. It is one ascending sequence on one fresh instance. ORT's arena steps depend on allocation history, not only on length: in block 1 the same config reached 1,700 MiB after an 11.6 s request.
  • On GPU 0 the live seat has 1,791 MiB, exactly the 45–60 s level. Its 1,690 MiB is that level.
  • Anything that needs the next step (90 s here) needs memory GPU 0 does not have. Whether ORT's allocation back-off squeezes such a request through anyway was not tested; it is not safe to test on production.
  • So the live seat's practical long-file limit lies between about 1 minute and 400 s.

Long-form accuracy inside the envelope.

  • Windows. Both public files cut into pieces the seat can take: ≤ 375 s each, cut at pauses, with ground truth sliced from the investigation's timed GT.
  • Method. Every window goes in whole, in one request, in two placements: first cut at 360 s (9 windows), then at 180 s (11 windows). Scoring uses the investigation's gtscore and its crosstalk rule (diarized overlap ≥ 10 % = crosstalk).

Clean-speech words dropped, then WER, for placement 1 / placement 2:

Wilde: dropped Wilde: WER % SCOTUS: dropped SCOTUS: WER %
seat (v3 int8) 1,676 / 320 48.7 / 11.5 196 / 74 9.4 / 7.0
v3 fp32 57 / 16 4.0 / 2.7 0 / 11 4.4 / 3.9
C (unified int8) 0 / 0 2.5 / 2.4 13 / 0 5.0 / 5.1
C-fp32 (placement 1 only) 0 2.4 13 5.0
B-fp32 (placement 1 only) 0 2.3 13 5.0
B-bf16w 0 / 0 2.3 / 2.1 13 / 0 5.0 / 4.8
D (v2 int8) 133 / 1,009 6.1 / 29.1 0 / 12 4.7 / 5.0

For scale (the investigation, Scriberr's 120 s slicer, 8 placements): v3 lost 140 [48–240] (Wilde) and 66 (SCOTUS); unified lost 31 [0–92] and 22.

  • The seat's long-context collapse is severe at 6-minute windows, and it moves with the cuts. It lost 1,676 words or 320 on the audiobook, depending only on where the cuts fell. The same v3 weights in fp32 lose 16–57. int8 makes the known v3 flaw far worse.
  • unified-en lost no clean speech on Wilde in any of 4 window runs (2 placements × int8 and NeMo). On SCOTUS it lost one 13-word stretch in one of the two placements.
  • Two placements show a large gap, not a confidence interval. The investigation used 8.
  • v2 collapses on the audiobook too (133 and 1,009 words), as the investigation found.
  • On SCOTUS the best WER is v3 fp32 (3.9–4.4 %). unified is 4.8–5.1 % and the seat 7.0–9.4 %.

One request for the whole file, beyond 400 s. Only NeMo can do it.

  • Full attention needs more than 37 GB for 24 minutes. The headroom guard stopped it.
  • NeMo's long-audio mode works. With local attention ±128, B-bf16w transcribed:
    • the 24-minute Wilde in 1.6–2.0 s, WER 3.77–3.85 %, one clean dropout of 57 words;
    • the 30-minute SCOTUS in 2.3–2.6 s, WER 4.99–5.02 %, one clean dropout of 13 words.
  • No insertion runs. The text was not byte-identical across the two runs (local attention plus bf16), but the dropouts were the same in both.

6. Memory and where each candidate could live

Per-process GPU memory in MiB, from nvidia-smi every 200 ms, filtered to each arm's own PIDs.

arm at rest after warm-up serving peak (utterance block) highest seen (incl. load) extra over the seat's 1,690 MiB (rest / serving / highest)
A, the live seat 1,690 (1,692 after this test) — — 0
A′, the seat's config on a roomy card 922–932, then 1,690–1,700 after the first 11.6 s request 4,764–5,798 † 5,798 † —
C (unified int8) / D (v2 int8) 916 → 1,684 after the first 11.6 s request 4,758 † 4,758 † ≈ 0 (the same footprint as the seat)
A-fp32 / C-fp32 3,922–3,932 / 3,874 6,994–7,004 † / 6,946 † 7,004 / 6,946 +2,240 / +5,310 / +5,310
C-fp16 2,776 4,824 † 4,824 +1,090 / +3,130 / +3,130
B-fp32 (.nemo restored straight to GPU) 3,258–3,836 4,800 5,620 (load transient, about 1 s) +2,150 / +3,110 / +3,930
B-fp32, restored via CPU 3,836 4,800 4,800 +2,150 / +3,110 / +3,110
B-bf16 (autocast) 4,802 5,126 5,620 +3,110 / +3,440 / +3,930
B-bf16w (bf16 weights, via CPU) 2,492 2,794 3,194 (load) +800 / +1,100 / +1,500

† ORT's arena grows in 1,024 MiB steps under sustained never-seen-length traffic, whatever the request length.

  • What triggers a step. Steps came after 2.8 s requests as often as after 55 s ones. The seat's config held about 1,700 MiB through 59 s requests until one arrived. Every ORT arm stepped this way, int8 and fp32 alike.
  • The seat's real working set is about 1.7 GB. On a roomy card that grows to 4.8–5.8 GB over about 600 requests. On GPU 0 the live seat stays at 1,690 MiB because there is nothing to grow into, and it has served 1,112 requests without one failure.
  • ORT "serving peaks" are arena policy, not need. An ORT replacement on GPU 0 would behave the same way, but its floor is its at-rest figure.
  • NeMo is different. Its plateau, 4,800 MiB for fp32 and 2,794 MiB for bf16w, is torch's cache holding the longest request's activations, and that includes 59 s clips. A seat capped at talk's ≤ 12 s turns would plateau lower; I did not measure how much lower.
  • The bf16w load transient (3,194) should be avoidable by casting to bf16 on the CPU before the move. That is reasoned, not measured.

Fit verdict for GPU 0. The seat's 1,690 MiB plus 101 MiB free gives 1,791 MiB available in place.

  • C and D fit in place but are not faster, so there is nothing to gain.
  • None of the fast arms fits in place. The cheapest is B-bf16w at about +1.1 GB (serving), or +1.5 GB (its load transient as measured).
    • By the coordinator's figure of about 0.95 GB per 0.01 of gen-small's gpu_memory_utilization, that is about 0.012–0.016 of gen-small; I did not measure the KV trade.
    • The fp32 routes need +2.2 GB at rest and more as the arena steps. They would also want an arena cap.
  • Other homes, as observed around 1745 PDT and not checked for reservations:
    • GPU 1 showed 6,625 MiB free (intern-decision 8.8 GB plus vLLM seats), enough for B-bf16w's 3.2 GB peak. Whether that headroom is spoken for, I do not know.
    • GPU 3 is the card kept empty for a full-card seat. Scriberr uses it on demand and needs about 5.5 GB per job.

7. Output features callers might rely on

seat (v3 int8) unified (any runtime) v2 int8
Response body {"text": …} only the same (the wrappers are identical) the same
Casing yes (94.5 % of outputs carry upper case) yes (96–97 %) yes
Commas yes (73 %) yes (62–65 %) yes (74 %)
Sentence-final .?! 85 % of test-clean outputs end with one (AMI 82 %) 43–46 %: it drops the final period about half the time (AMI 54–56 %) 80.5 % (AMI 85 %)
Timestamps none returned. sherpa-onnx computes token times, and the wrapper discards them. none (NeMo can return word and segment times; the wrapper does not ask) none
Language field none (not returned; response_format and language ignored) none none
Languages 25 European English only English only
Longest input 400 s, then HTTP 500 (§ 5.4) ONNX: the same 400 s. NeMo: no ceiling (§ 5.4) 400 s
Empty outputs on AMI 26/400 9–17/400 15/400

English-only is a real change, but no current caller depends on another language. talk, the WER tooling and the accent probes are English. A non-English clip sent to unified-en would come back as English-ish text, not as an error.

8. Recommendation

claim strength basis reversibility
The seat's latency is its runtime (int8 on one CPU thread), not its model insist measured: CPU/wall 1.00, GPU 2–9 %; the same v3 weights in fp32 are 4–12× faster in the same image n/a (a finding)
Move the seat off the int8-on-CPU graph strongly recommend measured: every GPU-native arm saves 100–530 ms per 1–20 s turn (CIs ±4 to ±35 ms; floor ≤ 6 ms), and the int8 graph also costs WER and drops the rest of an utterance after a gap (§ 5.3) reversible (image or alias swap)
The destination: unified-en under NeMo with bf16 weights (B-bf16w), over v3 as fp32 ONNX recommend measured: fastest arm (−121 / −234 / −530 ms vs the seat; 16–42 ms faster than fp32 ONNX), the best AMI WER, no 400 s ceiling, the smallest fast footprint. Against it: English only, NVIDIA Open Model License, about 1.1–1.5 GB of GPU 0 to find, a heavier image reversible (old image and alias kept)
If English-only or the licence is a blocker: v3 exported fp32 in the seat's own image lean measured: −107 / −213 / −495 ms, 1.2–2.2 pp better than today; but +2.2 GB, an arena that steps, and still the 400 s ceiling; the export is ours (k2-fsa do not publish v3 fp32) reversible
Do NOT adopt C (k2-fsa's unified int8 in the seat runtime) recommend against measured: 8–71 ms SLOWER than the seat; its accuracy gain is real but available faster elsewhere reversible
Do NOT adopt D (v2 int8) as the English alternative recommend against measured: the same speed as the seat (+3 ms, inside the floor); better WER than v3 int8 but worse than unified reversible
Do NOT use the fp16 ONNX conversion lean against measured: 50–60 ms slower than fp32 on short turns (per-shape tuning); v3 fp16 is broken outright reversible
Stopgap with no model change: NUM_THREADS=16 on the live seat lean measured: −61 / −119 / −249 ms and identical text, but it burns 12–29 host cores per request and needs a seat recreate (~45 s warm-up) reversible (one env var)

9. What a switch to B-bf16w would take

  • Runtime. NeMo 3.0.0 + torch 2.8 (cu128) + FastAPI/uvicorn.
    • The wrapper exists: services/parakeet-ab-2026-09-30/code/serve_nemo.py, run with DTYPE=bf16w LOAD_CPU=1. It keeps the seat's endpoints and {"text": …} body, and it returns the same text as NeMo's transcribe().
    • NeMo's card says 2.7.3, but released 2.7.3 lacks this encoder's att_chunk_context_size, so it needs 3.0.0. The .nemo also lacks a validation_ds config, which needs the two-line shim the wrapper carries (both as found by the investigation).
  • Image. None exists yet. This A/B ran the wrapper from a uv env mounted into the scriberr:local-blackwell image.
    • A seat image would be a CUDA 12.8 runtime base plus that env: torch and the NVIDIA wheels come to several GB, against the current seat's 5.09 GB image.
    • Also needed before shipping: a warm-up at the longest served length (the CUDA-graph capture), a bf16 cast before the move to GPU (the load transient), and, for files over a few minutes, a switch to local attention (LOCAL_ATT=128,128).
    • With local attention a 30-minute file took 2.6 s in one request, and memory grows linearly instead of quadratically (§ 5.4).
  • Weights. nvidia/parakeet-unified-en-0.6b @ fe53cd885760c96b6a5f51a0bfd362cb4584a98b, already pinned in /tank/aimodels/huggingface (sha256 ec23ed91…, mount read-only).
  • Where it lives. GPU 0 needs about +1.1 GB serving, or +1.5 GB with the load transient as measured, from a vLLM seat's KV (gen-small is the coordinator's suggested donor). Alternatively it could live on GPU 1, if that card's 6.6 GB of free headroom is really free (§ 6).
  • Licence. The NVIDIA Open Model License Agreement (the card says "ready for commercial/non-commercial use"; read at the card, the agreement's own text not reviewed here). That replaces CC-BY-4.0 for this seat. Heads-up only; the call is Prime's.
  • Consumers. None need a code change: same port, same body, and the LiteLLM alias is unchanged. The visible differences:
    • fewer sentence-final periods;
    • English only;
    • inputs over 400 s now work instead of returning 500.

10. Provenance

artifact id / revision sha256 licence
live seat weights (v3 int8) /tank/parakeet/models (k2-fsa sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8) encoder acfc2b44…, decoder 179e50c4…, joiner 3164c13f…, tokens d5854467… CC-BY-4.0
unified int8 (arm C) GitHub k2-fsa/sherpa-onnx release asr-models, …unified-en-0.6b-int8-non-streaming.tar.bz2 tarball 99f63605…e150 (= GitHub digest); encoder 6716910b… NVIDIA Open Model License
v2 int8 (arm D) the same release, …parakeet-tdt-0.6b-v2-int8.tar.bz2 tarball 157c157b…61e1ad (= GitHub digest); encoder a32b12d1… CC-BY-4.0
unified .nemo (arm B, and C-fp32/fp16 source) HF nvidia/parakeet-unified-en-0.6b @ fe53cd88… ec23ed91… (= the investigation's pin) NVIDIA Open Model License
v3 .nemo (A-fp32 source) Scriberr's env copy (HF 541d1f99, per the investigation) 3cbdc858… CC-BY-4.0
our fp32 exports k2-fsa recipe @ sherpa-onnx 040afe36 minus quantisation unified encoder.weights d1559afb…; v3 encoder.weights 9a22d372… as their source
LibriSpeech test parquets HF openslr/librispeech_asr @ 71cacbfb… clean 7113aa4c…, other 38e0c86a… (= LFS oids) CC-BY-4.0
AMI IHM test parquets (4) HF edinburghcstr/ami @ 46f28f25… d95920dc…, 07a2b1c4…, 83cde21e…, 3312bb79… (= LFS oids) CC-BY-4.0
  • Every HF id was verified with an authenticated API call before pulling. A known-phantom repo returned 404 as a check of the check.
  • GitHub had no token anywhere (none on nh3-dev, on fv-ml1 or in the vault).
    • GitHub answers 404, not 401, for a missing public asset, and the release API returned each asset's size and sha256 digest.
    • Every download matched that digest. That is the positive proof the authenticated call stands in for.
  • Full hashes: services/parakeet-ab-2026-09-30/results/model-sha256.txt and provenance-fetch.json.

11. Host changes, cleanup and reproduction

  • fv-ml1.
    • Spike dir /tank/spikes/parakeet-ab/. It holds the downloads, the extracted and exported models (about 12 GB), the test sets, two uv envs and the raw outputs.
    • About 40 transient ab-* containers, all --rm, on GPU 3 only, bound to 127.0.0.1 and labelled ab=parakeet-2026-09-30. All are removed. GPU 3 holds no process of ours.
    • The memory sampler (nvidia-smi -lms 200) is stopped.
    • The live parakeet container got 180 requests (§ 4.3). It saw no restart and no config change.
    • Untouched: Scriberr, intern-decision, the vLLM seats and LiteLLM's config.
  • ops-log. create-spike-dir and transient-containers records (2026-09-30 1611 and 1638 PDT), plus a cleanup record.
  • The spike dir is kept for re-runs and holds no private data. Whether to delete it, about 25 GB with envs, caches and test audio, is Prime's call.
  • Reproduce. code/ in this directory is the full harness:
    • mk_env.sh, prep_data.py, export_onnx.py, convert_fp16.py, make_windows.py, make_pc2.py;
    • arm.sh (start an arm), run_lat.py (the latency blocks), run_acc.py, longwin.py;
    • analyze_lat.py, gw_analysis.py, acc_summary.py, score.py, longscore.py, memtrace.py.