A/B of the live STT seat (fv-ml1 GPU 0, sherpa-onnx int8 v3) against nvidia/parakeet-unified-en-0.6b, measured on GPU 3 with the seat's own image, k2-fsa's published unified int8 export, fp32/fp16 exports made with k2-fsa's recipe, v2 int8, and NeMo 3.0.0 (fp32, bf16 autocast, bf16 weights). - Seat int8 graph runs on one CPU thread (cpu/wall 1.00, GPU 2-9%). - unified-en under NeMo: -121/-234/-530 ms vs the seat at 1-3/3-8/8-20 s (paired, n=120/bin; floor <=6 ms; +50 ms positive control reads +52-54). - unified-en WER lower in every runtime: -0.7 pp clean, -1.5 pp other, -3.2 to -4.4 pp AMI (paired CIs exclude 0). - Seat defects found: hard 400 s input ceiling (HTTP 500), truncation after a quiet 1.5 s pause, and severe long-window dropouts (int8 v3 only). - B-bf16w needs +0.8 to +1.5 GB over the seat's 1,690 MiB on GPU 0. Raw requests, hypotheses, manifests and the full harness under services/parakeet-ab-2026-09-30/. No deploy; live seat untouched apart from 240 light test requests.
587 lines
41 KiB
Markdown
587 lines
41 KiB
Markdown
# Parakeet seat A/B: the live seat vs parakeet-unified-en-0.6b (2026-09-30)
|
||
|
||
Prime's ask, via the coordinator, 2026-09-30: *"a/b the one on fv-ml1's general seat against the unified
|
||
new one in jun for speed and accuracy for english. that seat demands low latency and any win, even 50ms,
|
||
is load-bearing."* **A/B only.** Nothing was deployed. The live `parakeet` container was not restarted
|
||
or reconfigured: it got 240 light requests (§ 4.3). Every candidate ran as a transient container on
|
||
fv-ml1 GPU 3, and all of them are gone. Raw data and code: `services/parakeet-ab-2026-09-30/`.
|
||
|
||
## 1. Headline
|
||
|
||
- **Yes on both counts, with one big qualifier: the latency win comes from the runtime, not the
|
||
model.**
|
||
- **Latency.** `parakeet-unified-en-0.6b` under NeMo torch is **121, 234 and 530 ms faster** than the
|
||
seat at 1–3, 3–8 and 8–20 s (median of 120 paired requests per bin; 95 % CIs ±4 to ±35 ms).
|
||
The A-vs-A floor is ≤ 6 ms (§ 4.2), and the +50 ms positive control reads +52 to +54 ms at every
|
||
length, so **a 50 ms difference is resolvable at every length**. That arm answers a 3–8 s turn in
|
||
**27 ms** against the seat's 260 ms, or 308 ms on the live seat itself.
|
||
- **Accuracy.** In **every** runtime tried, unified-en has a lower English WER than the seat on every
|
||
set: about **0.7 pp** lower on LibriSpeech test-clean, **1.5 pp** on test-other and **3–4 pp** on
|
||
AMI meetings. Every paired CI excludes zero.
|
||
- **The same model in the seat's runtime is NOT faster.** k2-fsa publish unified-en as an int8 sherpa-onnx
|
||
export, and it runs in the seat's own image unchanged (arm C), but it is **8–71 ms slower** than the
|
||
seat. The reason is the runtime: **the seat's int8 graph runs on ONE CPU thread.** CPU time equals
|
||
wall time on every request and GPU utilisation is 2–9 %. The proof is the seat's own v3 weights,
|
||
exported to fp32 with k2-fsa's recipe, in the seat's own image: **4–12× faster**, and more accurate
|
||
on test-other and AMI.
|
||
- **Best arm measured: unified-en, NeMo, bf16 weights (`B-bf16w`).** Its p50 is 23 / 27 / 33 / 62 ms
|
||
against the seat's 144 / 260 / 565 / 2,027 ms on the same card. WER is identical to fp32 (399/400
|
||
utterances with the same edits). It needs **2.5–3.2 GB, which is +0.8 to +1.5 GB over today's
|
||
1,690 MiB**, so it does not fit GPU 0 in place (§ 6).
|
||
- **Long form.** On 6-minute files sent whole, which is inside what the seat can take, the seat's
|
||
config lost **320–1,676 words of clean speech on the audiobook**, depending only on where the cuts
|
||
fell. unified-en lost **none** (§ 5.4).
|
||
- **Two defects in the live seat surfaced on the way, independent of any switch:**
|
||
- **It cannot take a file longer than 400 s (6 min 40 s).** The ONNX export freezes the
|
||
relative-position table at 5,000 frames, so anything longer returns **HTTP 500** (§ 5.4).
|
||
- **On a 1.5 s stretch of digital silence inside an utterance it stops transcribing**, and the rest of
|
||
the utterance is lost. That is the v3 **int8** export only: the same weights in fp32 do not do it,
|
||
nor do unified or v2 (§ 5.3).
|
||
|
||
## 2. What was compared
|
||
|
||
Everything except arm A ran as a transient container on **fv-ml1 GPU 3**, bound to 127.0.0.1.
|
||
|
||
- **Every sherpa-onnx arm runs the seat's own image** (`local/parakeet:sherpa-onnx-v4`, sherpa-onnx
|
||
1.12.39+cuda12.cudnn9, ORT CUDA EP) with the seat's env (`PROVIDER=cuda`, `NUM_THREADS=1`).
|
||
- **They also run the seat's `app.py`, plus a timing header**
|
||
(`services/parakeet-ab-2026-09-30/code/app_timed.py`). Its only differences from
|
||
`stacks/parakeet/app.py` are the header, an injected-delay knob (for the positive control) and a
|
||
model file-suffix knob.
|
||
- **The live container runs exactly `stacks/parakeet/app.py`** (sha256 `6a0bd695…`, checked inside
|
||
the container).
|
||
- **The NeMo arms use the same HTTP shape**: the model loads at import, one warm-up decode runs on 1 s
|
||
of silence, and `async` handlers call a blocking decode, so requests are serialised exactly as in
|
||
the seat.
|
||
|
||
| arm | model | weights / export | runtime | where |
|
||
|---|---|---|---|---|
|
||
| **A** | parakeet-tdt-0.6b-v3 (TDT, 25 languages) | k2-fsa int8 (the live `/tank/parakeet/models`) | **the live seat** | GPU 0, :8300 |
|
||
| **A′** #1, #2 | same | the same files, read-only | seat image | GPU 3 |
|
||
| A′ +50 ms | same | same | seat image, 50 ms sleep added per request | GPU 3 (positive control) |
|
||
| A′ plain | same | same | seat image, its own `app.py` untouched | GPU 3 (null control) |
|
||
| A′ 16 threads | same | same | seat image, `NUM_THREADS=16` (a config-only change) | GPU 3 |
|
||
| **A-fp32** | the same v3 weights | fp32 ONNX, our export with k2-fsa's recipe | seat image | GPU 3 |
|
||
| **C** | parakeet-unified-en-0.6b (RNN-T, English) | **k2-fsa int8, published** | seat image | GPU 3 |
|
||
| **C-fp32** | same | fp32 ONNX, our export with k2-fsa's recipe | seat image | GPU 3 |
|
||
| C-fp16 | same | C-fp32 converted to fp16 (`pre_encode` kept fp32) | seat image | GPU 3 |
|
||
| **D** | parakeet-tdt-0.6b-v2 (TDT, English) | k2-fsa int8, published | seat image | GPU 3 |
|
||
| **B-fp32** | parakeet-unified-en-0.6b | the `.nemo`, HF `fe53cd88` | NeMo 3.0.0 + torch 2.8 cu128 | GPU 3 |
|
||
| B-bf16 | same | same | same, `torch.autocast` bf16 (fp32 weights) | GPU 3 |
|
||
| **B-bf16w** | same | same, encoder/decoder/joint weights cast to bf16, loaded via CPU | same | GPU 3 |
|
||
|
||
**Arm C needed no export of ours: k2-fsa publish it.**
|
||
|
||
- **Asset.** `sherpa-onnx-nemo-parakeet-unified-en-0.6b-int8-non-streaming.tar.bz2` is on the
|
||
sherpa-onnx `asr-models` release (2026-04-27; sha256 `99f63605…e150`, matched).
|
||
- **Provenance.** It comes from `scripts/nemo/parakeet-unified-en-0.6b/export_onnx.py` (k2-fsa
|
||
PR #3556, sherpa-onnx 1.13.0).
|
||
- **Shape.** It is a plain `EncDecRNNTBPEModel` with full-context attention
|
||
(`att_context_size [-1,-1]`; "contains only the non-streaming part").
|
||
- **It loads and runs in the seat's pinned sherpa-onnx 1.12.39 unchanged.**
|
||
- **Out of scope.** k2-fsa also publish 240, 560 and 1,120 ms streaming variants.
|
||
- **No Hugging Face mirror.** An authenticated call on `csukuangfj/…unified…` returns 404.
|
||
|
||
**The fp32 and fp16 arms were added because of what the pilot found** (§ 4.1).
|
||
|
||
- **`A-fp32` and `C-fp32`** are k2-fsa's export script with only the final `quantize_dynamic` removed:
|
||
the same `encoder/decoder/joint.export()` and the same metadata. Their `tokens.txt` is
|
||
byte-identical to the seat's and to k2-fsa's.
|
||
- **`C-fp16`** is C-fp32 put through onnxconverter-common 1.16.0 with fp32 I/O. The conv subsampling
|
||
front had to stay fp32 or the graph will not load.
|
||
- **v3 in fp16 is broken** (every token `<unk>`), so it has no arm.
|
||
|
||
**Not measured:** k2-fsa's v2 fp16 tarball, TensorRT, `torch.compile`, batching, and the unified
|
||
streaming exports.
|
||
|
||
## 3. What the seat actually serves
|
||
|
||
- **Callers.**
|
||
- **`talk` (latency-critical).** tts-stack's voice loop on nh3-dev: push-to-talk and hands-free
|
||
VAD turns. The browser decimates 48 kHz to **16 kHz mono WAV** with a 10 MiB body cap, and it
|
||
calls through LiteLLM `ext-stt` on the `all-agents-local` key (python-httpx, a new client per
|
||
call).
|
||
- **tts-stack batch tooling (the rest).** `venus_score.py` calls `:8300` directly; it sent the 308
|
||
requests on 2026-09-28 20:27–20:32 UTC, none of which are in the gateway log. The accent and tag
|
||
probes sent 226 requests of 48 kHz WAV between 09-15 and 09-22.
|
||
- **`whisper-1` has had zero calls.**
|
||
- **Volume.** 1,112 transcriptions from 2026-09-15 to 09-30 (the container's whole life), all of
|
||
them 200. `talk` makes a handful to a few dozen turns a day. Concurrency is effectively 1.
|
||
- **Sizes.** Neither log records audio length, so the length mix is inferred.
|
||
- The gateway's 58 `ext-stt` calls since 09-23 took **p50 0.31 s, p90 0.47 s, max 0.71 s** end to
|
||
end. Its log is capped at 10,000 rows, so it only reaches back to 09-23.
|
||
- Read against the measured gateway path (§ 4.3), that is a median `talk` turn of about **3–5 s**
|
||
and a longest of about **10 s**.
|
||
- That puts most turns in the **1–8 s** bins, with a tail into 8–20 s. The 20–60 s bin is batch
|
||
tooling.
|
||
- **Content.** English: Prime's voice turns and English WER tooling. No caller reads a language
|
||
field, and the seat never returned one.
|
||
|
||
## 4. Latency
|
||
|
||
### 4.1 Why the seat is slow (the pilot finding that reshaped the arms)
|
||
|
||
- **The seat's int8 graph runs on the CPU.** With `NUM_THREADS=1`, the process's CPU time equals the
|
||
request's wall time on every clip (ratio 0.97–1.05, from host `/proc`), while GPU 3's SM
|
||
utilisation sits at **2–9 %**.
|
||
- **The cause is ORT's int8 path.** k2-fsa's export uses `quantize_dynamic` (`MatMulInteger` /
|
||
`DynamicQuantizeLinear`), which ORT's CUDA EP leaves to the CPU.
|
||
- **More threads help, but poorly.** At `NUM_THREADS=16` the seat is about 40 % faster and burns
|
||
12–29 cores spinning.
|
||
- **The same weights exported in fp32 run on the GPU.** In the same image they are 4–12× faster.
|
||
- **Consequence.** The model and the runtime are separate variables. Arm C, the new model in the old
|
||
runtime, cannot show the new model's speed.
|
||
|
||
### 4.2 Harness
|
||
|
||
- **Client.** Stdlib Python on the fv-ml1 host (loopback), **one new HTTP connection per request**
|
||
(as `talk`'s per-call httpx client does), multipart `file=turn.wav` + `model=ext-stt`, every body
|
||
built before the clock starts.
|
||
- **What is timed.** **End-to-end (e2e)** runs from connect to the last byte. **Server-side** is the
|
||
`x-ab-decode-ms` header: audio decode, features, model and text inside the handler. HTTP adds
|
||
1.3–5.8 ms on top of server-side. The live seat has no header, so A is e2e only.
|
||
- **Clips.** 20 per bin from LibriSpeech test-clean, at evenly spaced duration quantiles:
|
||
- 1–3 s: 1.8–3.0 s;
|
||
- 3–8 s: 3.1–7.8 s;
|
||
- 8–20 s: 8.1–18.9 s;
|
||
- 20–60 s: consecutive utterances of one chapter joined with 0.25 s gaps (22–59 s).
|
||
- **Every request has a never-seen length.** ORT tunes per input shape and NeMo's CUDA-graph decoder
|
||
captures per maximum length, so repeating a clip flatters both. In the pilot, the first call at a
|
||
new shape cost the seat +30–100 ms and fp32 ONNX +20–45 ms.
|
||
- Each request therefore trims a seeded 0–490 ms (10 ms steps) off the tail.
|
||
- The trim is seeded by round and clip, never by arm, so **in a given round every arm received
|
||
byte-identical input**, and every delta is a median of paired per-request differences.
|
||
- **Warm, single stream.** 6 rounds × 80 clips, so **120 requests per bin per arm**. Within a round
|
||
the arms run in a seeded shuffled order, so drift (thermal, a Scriberr job) lands on all of them.
|
||
- **Concurrency 4.** Per arm and bin, 40 requests from 4 workers back to back.
|
||
- **Cold.** Each arm was started fresh, **one at a time**, and timed from docker `StartedAt` to the
|
||
first healthy `/healthz`. Then one never-seen clip went in per bin, in ascending length.
|
||
- **Statistics.** p50/p90/p99 over requests, with 95 % bootstrap CIs (2,000 resamples). p99 over 120
|
||
requests is roughly the second-largest sample: it is reported, and it is weak.
|
||
- **Blocks.** Block 1 ran 12 arms. Block 2 ran the memory-lean NeMo variants with B-fp32, B-bf16 and
|
||
both fp32 ONNX arms as anchors.
|
||
|
||
**Controls (block 1), as the paired median difference against A′ #1, in ms [95 % CI]:**
|
||
|
||
| | 1–3 s | 3–8 s | 8–20 s | 20–60 s |
|
||
|---|---|---|---|---|
|
||
| floor: A′ #2 (an identical second instance) | +1.0 [+0.7, +2.1] | +1.5 [+0.9, +2.9] | +1.8 [+0.3, +3.2] | +2.2 [−4.1, +7.6] |
|
||
| null: A′ plain (the image's own app, no header) | +4.4 [+3.9, +5.1] | +5.8 [+4.9, +6.8] | +6.6 [+4.7, +10.4] | +16.8 [+13.2, +21.0] |
|
||
| **positive: A′ +50 ms injected** | **+52.6** [+52.1, +53.1] | **+53.2** [+52.5, +53.8] | **+52.3** [+50.8, +53.2] | **+53.9** [+50.6, +56.1] |
|
||
| GPU-arm floor: B-fp32 loaded via CPU vs B-fp32 (block 2) | +0.0 | +0.1 | −0.1 | +1.2 |
|
||
|
||
- **The null control registered a small effect: +4 to +17 ms, about 1–3 %.**
|
||
- With one instance per condition I cannot tell how much of it is the code path (the plain app's
|
||
`dict` return goes through FastAPI's validation) and how much is per-instance placement of a
|
||
CPU-bound single thread on a multi-socket host.
|
||
- So it is treated as part of the floor. **The A-vs-A floor is ≤ 6 ms for 1–20 s and ≤ 21 ms for
|
||
20–60 s on the CPU-bound arms, and about 1 ms on the GPU arms.**
|
||
- **Sensitivity.** The positive control reads +50 ms (plus about 3 ms that the floor explains) at
|
||
every length, with CI half-widths ≤ 3 ms. **A 50 ms difference is therefore resolvable at every
|
||
length**, and so is anything above about 10 ms (1–20 s) or about 25 ms (20–60 s).
|
||
|
||
### 4.3 Results: warm, single stream
|
||
|
||
End-to-end milliseconds, **p50 / p90 / p99**, 120 requests per bin, fv-ml1 loopback, GPU 3 unless
|
||
noted. Block 1 except where marked.
|
||
|
||
| arm | 1–3 s | 3–8 s | 8–20 s | 20–60 s |
|
||
|---|---|---|---|---|
|
||
| **A, the live seat (GPU 0, n = 40 per bin)** | **187** / 221 / 232 max | **308** / 413 / 454 max | **626** / 862 / 956 max | not sent |
|
||
| **A, via LiteLLM `ext-stt` from nh3-dev (n = 20)** | **257** / 297 / 309 max | **388** / 494 / 511 max | **711** / 968 / 1,095 max | not sent |
|
||
| **A′, the seat's config** | **144** / 170 / 184 | **260** / 370 / 384 | **565** / 814 / 896 | **2,027** / 2,634 / 2,768 |
|
||
| A′ 16 threads | 85 / 98 / 104 | 143 / 185 / 197 | 318 / 438 / 507 | 1,171 / 1,546 / 1,610 |
|
||
| A-fp32 (v3 fp32 ONNX, seat image) | 38 / 41 / 44 | 47 / 53 / 57 | 67 / 80 / 87 | 167 / 219 / 231 |
|
||
| C (unified int8, seat image) | 156 / 180 / 197 | 271 / 389 / 414 | 591 / 845 / 942 | 2,101 / 2,795 / 2,943 |
|
||
| C-fp32 (unified fp32 ONNX, seat image) | 41 / 43 / 67 | 50 / 58 / 63 | 76 / 91 / 98 | 186 / 248 / 260 |
|
||
| C-fp16 | 89 / 91 / 116 | 99 / 107 / 111 | 124 / 141 / 148 | 244 / 307 / 318 |
|
||
| D (v2 int8, seat image) | 150 / 174 / 188 | 264 / 379 / 390 | 565 / 806 / 901 | 2,004 / 2,718 / 2,767 |
|
||
| **B-fp32 (unified, NeMo)** | **23** / 25 / 29 | **27** / 29 / 30 | **34** / 40 / 62 | **77** / 98 / 258 |
|
||
| B-bf16 (autocast) | 29 / 31 / 55 | 32 / 36 / 39 | 39 / 43 / 46 | 67 / 83 / 95 |
|
||
| **B-bf16w (bf16 weights; block 2)** | **23** / 26 / 51 | **27** / 30 / 43 | **33** / 38 / 207 | **62** / 77 / 269 |
|
||
|
||
**Paired median difference against A′, in ms [95 % CI]. Negative means faster than the seat:**
|
||
|
||
| arm | 1–3 s | 3–8 s | 8–20 s | 20–60 s |
|
||
|---|---|---|---|---|
|
||
| A′ 16 threads | −61 [−63, −58] | −119 [−129, −112] | −249 [−265, −233] | −844 [−910, −747] |
|
||
| A-fp32 | −107 [−114, −104] | −213 [−235, −205] | −495 [−528, −468] | −1,859 [−1,979, −1,616] |
|
||
| C (unified int8) | **+8 [+7, +8]** | **+12 [+11, +14]** | **+22 [+19, +24]** | **+71 [+62, +80]** |
|
||
| C-fp32 | −104 [−109, −101] | −209 [−227, −200] | −488 [−517, −461] | −1,839 [−1,942, −1,599] |
|
||
| C-fp16 | −55 [−61, −51] | −160 [−181, −153] | −439 [−472, −410] | −1,780 [−1,880, −1,527] |
|
||
| D (v2 int8) | +3 [+2, +3] | +4 [+3, +4] | +3 [+2, +5] | +4 [+0, +10] |
|
||
| **B-fp32** | **−121 [−128, −117]** | **−234 [−254, −226]** | **−530 [−565, −499]** | **−1,947 [−2,058, −1,699]** |
|
||
| B-bf16 | −115 [−122, −112] | −229 [−249, −218] | −527 [−561, −491] | −1,959 [−2,064, −1,706] |
|
||
| B-bf16w vs B-fp32 (block 2, paired) | +0.5 [+0.4, +0.6] | +0.9 [+0.4, +1.2] | −0.3 [−0.8, +0.3] | −13.7 [−15.2, −12.2] |
|
||
| **A (live) vs A′** | **+37 [+34, +38]** | **+40 [+38, +41]** | **+42 [+36, +47]** | — |
|
||
| A via gateway vs A (live) | +56 [+50, +76] | +76 [+59, +79] | +86 [+79, +97] | — |
|
||
|
||
- **C is slower than the seat, not faster.** The difference sits just above the floor at 1–20 s and is
|
||
clear at 20–60 s, and it is consistent with RNN-T decoding every frame where TDT skips frames.
|
||
**D is the same speed as the seat** (+3 to +4 ms, inside the floor).
|
||
- **The live seat is 37–42 ms slower than its byte-identical copy on GPU 3.** They gave identical text
|
||
on 120/120 inputs. GPU 0 is the card the vLLM seats share, and its allocator is pinned at 1,690 MiB
|
||
with 101 MiB free. Which of those costs the 40 ms is not measured.
|
||
- **The gateway path adds 56–86 ms over loopback.** About 24–29 ms of it is nh3-dev → fv-ml1 transit
|
||
(measured direct to `:8300` from nh3-dev). **The LiteLLM hop itself is 36–61 ms** and grows with
|
||
the body. An earlier note put this hop "below harness resolution (±30 ms)"; paired, it resolves.
|
||
- **So `talk`'s real path today is p50 257 / 388 / 711 ms** for 1–3 / 3–8 / 8–20 s turns. With B on
|
||
the same path that would be about **80 / 100 / 120 ms**. That is a derived figure: B's loopback e2e
|
||
plus the measured gateway-path cost. It excludes any GPU 0 placement penalty; the seat itself pays
|
||
+37–42 ms for its GPU 0 placement.
|
||
- **Live-seat safety.** The live seat got only 1–20 s clips (120 on loopback, 60 via the gateway, 60
|
||
direct), one per second, with an abort on the first non-200. All 240 returned 200. Its restart
|
||
count stayed at 0, it stayed healthy, and its GPU memory went from 1,690 to 1,692 MiB.
|
||
|
||
### 4.4 Concurrency 4, cold start and first call
|
||
|
||
**Concurrency 4**: e2e p50 in ms, then throughput in × realtime. Every server serialises requests
|
||
(the seat's `async` handler blocks the event loop), so e2e is about 4× the decode time.
|
||
|
||
| arm | 1–3 s | 3–8 s | 8–20 s | 20–60 s | ×realtime |
|
||
|---|---|---|---|---|---|
|
||
| A′ (seat config) | 569 | 1,062 | 2,304 | 8,313 | 15–21× |
|
||
| A′ 16 threads | 335 | 572 | 1,299 | 4,647 | 26–37× |
|
||
| A-fp32 | 138 | 181 | 265 | 662 | 63–254× |
|
||
| C | 636 | 1,137 | 2,415 | 8,763 | 14–20× |
|
||
| C-fp32 | 148 | 199 | 295 | 786 | 59–218× |
|
||
| D | 610 | 1,104 | 2,350 | 8,644 | 15–21× |
|
||
| **B-fp32** | **85** | **98** | **129** | **308** | 101–561× |
|
||
| **B-bf16w (block 2)** | **84** | **97** | **124** | **236** | 101–709× |
|
||
|
||
**Cold start**: from container start to healthy.
|
||
|
||
- **The seat image takes 45–48 s for every ONNX arm.** That is the seat's own warm-up decode
|
||
compiling CUDA kernels; it measured 45.1 s on the live seat's boot.
|
||
- **NeMo takes 12–13 s.**
|
||
|
||
**First call after start** (one never-seen clip per bin, in ascending length) compared with that arm's
|
||
warm p50:
|
||
|
||
- **The seat:** 223 / 357 / 649 / 2,904 ms, against 144 / 260 / 565 / 2,027 ms warm.
|
||
- **B-fp32:** 60 / 44 / 38 / **410** ms, against 23 / 27 / 34 / 77 ms warm.
|
||
- **B's 410 ms** is NeMo's CUDA-graph decoder capturing for a new maximum length. It is paid once per
|
||
new maximum, which is also the 260–300 ms p99 in B's 20–60 s bin.
|
||
- **The fix is a warm-up decode at the longest length** the seat should serve. It costs nothing,
|
||
though I have not measured it.
|
||
|
||
## 5. Accuracy (English, against ground truth)
|
||
|
||
**Sets.** All are public, with licences read at source:
|
||
|
||
- **LibriSpeech test-clean and test-other:** 400 utterances each, a seeded random sample
|
||
(`openslr/librispeech_asr@71cacbfb`, CC-BY-4.0). 49.9 and 45.0 minutes; 8,144 and 7,370 words.
|
||
- **AMI meetings, IHM headset mics:** 400 test utterances of ≥ 1 s, a seeded random sample covering
|
||
all 16 test meetings (`edinburghcstr/ami@46f28f25`, CC-BY-4.0). 23.8 minutes, 4,121 words. This is
|
||
the conversational set.
|
||
- **Long form:** the investigation's two public files with their ground truth (§ 5.4).
|
||
|
||
**Scoring.** Whisper's English normaliser (transformers **4.53.3**, the investigation's version, no
|
||
spelling map), applied to the whole utterance on both sides, then exact Levenshtein. Long-form uses
|
||
the investigation's own `gtscore.py`, copied verbatim (sha256 `5012a552…`).
|
||
|
||
### 5.1 WER
|
||
|
||
| WER % | **seat (v3 int8)** | v3 fp32 ONNX | **C: unified int8** | C-fp32 | C-fp16 | **D: v2 int8** | **B-fp32** | B-bf16 | **B-bf16w** |
|
||
|---|---|---|---|---|---|---|---|---|---|
|
||
| LS test-clean | **2.70** | 2.43 | 1.94 | **1.89** | 1.89 | 2.11 | 1.98 | 1.98 | 1.97 |
|
||
| LS test-other | **4.56** | 3.31 | **2.93** | 3.03 | 3.01 | 3.39 | 3.11 | 3.11 | 3.09 |
|
||
| AMI IHM | **12.69** | 10.48 | 9.51 | 9.37 | 9.32 | 10.02 | 8.37 | **8.28** | 8.30 |
|
||
| AMI empty outputs (of 400) | 26 | 18 | 17 | 15 | 15 | 15 | 10 | 9 | 9 |
|
||
|
||
**Paired bootstrap of WER minus the seat's WER, in pp [95 % CI]:**
|
||
|
||
| | test-clean | test-other | AMI |
|
||
|---|---|---|---|
|
||
| C (unified int8, seat runtime) | −0.76 [−1.11, −0.41] | −1.63 [−2.15, −1.14] | −3.18 [−5.13, −1.48] |
|
||
| C-fp32 | −0.81 [−1.18, −0.45] | −1.53 [−2.06, −1.04] | −3.32 [−5.25, −1.60] |
|
||
| **B-fp32 / B-bf16w** | **−0.72 / −0.74** [≈ −1.1, −0.35] | **−1.45 / −1.47** [≈ −2.0, −0.95] | **−4.32 / −4.39** [≈ −6.5, −2.4] |
|
||
| D (v2 int8) | −0.59 [−0.96, −0.26] | −1.17 [−1.68, −0.68] | −2.67 [−4.42, −1.05] |
|
||
| A-fp32 (the same v3 weights, fp32) | −0.27 [−0.60, +0.05] | **−1.25 [−1.70, −0.85]** | **−2.21 [−3.68, −0.89]** |
|
||
|
||
- **unified-en beats the seat on English in every runtime, on every set.**
|
||
- **The three unified runtimes tie on read speech.** C, C-fp32 and B are within ±0.18 pp of each
|
||
other on LibriSpeech, and every CI includes zero.
|
||
- **On AMI, NeMo is about 1 pp better than the sherpa runtimes:** C − B = +1.14 [+0.28, +2.16];
|
||
C-fp32 − B = +1.00 [+0.24, +1.92]. That is the same weights through a different front end and
|
||
decoder.
|
||
- **The seat's int8 quantisation of v3 itself costs 1.2 pp on test-other and 2.2 pp on AMI.**
|
||
- **bf16 weights cost nothing.** B-bf16w gives the same edit counts as B-fp32 on 399/400 utterances
|
||
per set.
|
||
|
||
### 5.2 Controls and determinism
|
||
|
||
- **Scorer self-test (positive and null).** One deleted, one substituted and one inserted word each
|
||
register exactly once, and casing, punctuation or a filler register zero. PASS.
|
||
- **Determinism.**
|
||
- A′ #1 = A′ #2 = A′ at 16 threads: 400/400 identical text.
|
||
- **The live seat = A′ on 120/120 inputs**, so A′'s accuracy is the seat's.
|
||
- The gateway path = the live seat on 60/60.
|
||
- NeMo's direct decode path = NeMo's `transcribe()` on 400/400, so the thin wrapper loses nothing.
|
||
- **Positive control (the pipeline).**
|
||
- The same 40 test-clean utterances (≥ 6 s) had 1.5 s of digital silence placed at 40 % of their
|
||
length, with the reference unchanged.
|
||
- **Every arm registered extra deletions on 40 of 40 utterances:** from 1–2 to 167–306 words.
|
||
- **Null control.** All 400 test-clean utterances at −0.5 dB. The WER change sits inside a CI that
|
||
includes zero for every arm (the seat: +0.04 [−0.08, +0.16] pp).
|
||
- **Sensitivity.** With 400 utterances, a paired WER difference of about 0.35 pp (test-clean),
|
||
0.5 pp (test-other) or 1.5 pp (AMI) is resolvable. Smaller ones are not.
|
||
|
||
### 5.3 The seat loses the rest of an utterance after a quiet pause (int8 v3 only)
|
||
|
||
The positive control should cost each arm the four or so words under the 1.5 s silence.
|
||
|
||
- **Every arm except the seat lost 167–185 words** across the 40 utterances. **The seat lost 306.**
|
||
- **On 6 of the 40 its output simply stops at the silence**, and the rest of the utterance (about 21
|
||
words each, on average) is gone. Example, with the silence at 8.1–9.6 s of 20.3 s:
|
||
- **v3 int8 (the seat):** "…than you could help running if you heard the wheel."
|
||
- **v3 fp32 (the same weights, the same runtime):** "…than you could help running if you heard
|
||
little on the other end of the house. The voice would go to your heart…"
|
||
- **The same 40 utterances with other fillers.** The count is utterances that lost ≥ 8 words more
|
||
than without the gap:
|
||
|
||
| gap fill | seat (v3 int8) | v3 fp32 | C (unified int8) | B-bf16w | D (v2 int8) |
|
||
|---|---|---|---|---|---|
|
||
| 1.5 s digital silence | **6/40** (306 deletions) | 1/40 (185) | 1/40 (181) | 1/40 (177) | 0/40 (167) |
|
||
| 1.5 s white noise at −60 dBFS | **6/40** (337) | 2/40 (196) | 1/40 (180) | 1/40 (176) | 0/40 (165) |
|
||
| 1.5 s white noise at −50 dBFS | 0/40 (170) | 4/40 (227) | 1/40 (178) | 1/40 (179) | 0/40 (171) |
|
||
| 0.75 s digital silence | 0/40 (88) | 0/40 (78) | 0/40 (86) | 0/40 (84) | 0/40 (73) |
|
||
|
||
**Reading.**
|
||
|
||
- **When it happens.** The seat truncates after a 1.5 s pause that is silent or nearly so
|
||
(≤ −60 dBFS). A louder room tone (−50 dBFS) or a shorter gap (0.75 s) does not trigger it.
|
||
- **v3 has a gap sensitivity of its own,** and int8 moves where it bites: fp32 shows 4/40 at
|
||
−50 dBFS.
|
||
- **unified and v2 are steady** across all four fills.
|
||
- **Why it could matter for `talk`.** Browser capture with noise suppression can produce near-digital
|
||
silence in a pause, so a 1.5 s mid-sentence pause in a push-to-talk turn could cost the rest of the
|
||
turn. That is not measured on real `talk` audio, which is private and was not used.
|
||
- **Basis.** Measured: 40 utterances × 4 fills, deterministic arms.
|
||
|
||
### 5.4 Long form, and the seat's 400 s ceiling
|
||
|
||
**The ceiling.**
|
||
|
||
- **Cause.** k2-fsa's export bakes the encoder's relative-position table at `pos_emb_max_len 5000`,
|
||
and our fp32 exports inherit it.
|
||
- **Effect.** An input over 5,000 encoder frames (400 s) fails in layer 0's attention:
|
||
`/layers.0/self_attn/Add_2: right operand cannot broadcast … {1,8,T,T} vs {1,8,T,9999}`. The seat
|
||
then returns **HTTP 500** after 2–11 s.
|
||
- **Measured on a fresh seat-config instance:** 395 s → 200, 410 s → 500. The 24- and 30-minute files
|
||
→ 500 (4 of 4). The fp32 unified export returned 500 on the 24-minute file too.
|
||
- **Every sherpa arm carries the same table.** NeMo extends its table on the fly, so B has no ceiling
|
||
(below).
|
||
|
||
**The seat's GPU memory climbs with length.** Fresh instance, ascending lengths, ORT arena high-water
|
||
mark:
|
||
|
||
| input length | 15 s | 30 s | 45–60 s | 90–150 s | 180–300 s | 360–395 s | 410 s |
|
||
|---|---|---|---|---|---|---|---|
|
||
| seat-config arena (MiB) | 924 | 1,180 | 1,692 | 2,716 | 3,742 | 7,838 | HTTP 500 |
|
||
| e2e (s) | 0.8 | 1.6 | 2.3–2.9 | 4.7–7.4 | 9.5–15.5 | 19.4–20.5 | — |
|
||
|
||
- **Read the table as indicative.** It is one ascending sequence on one fresh instance. ORT's arena
|
||
steps depend on allocation history, not only on length: in block 1 the same config reached
|
||
1,700 MiB after an 11.6 s request.
|
||
- **On GPU 0 the live seat has 1,791 MiB, exactly the 45–60 s level.** Its 1,690 MiB is that level.
|
||
- **Anything that needs the next step (90 s here) needs memory GPU 0 does not have.** Whether ORT's
|
||
allocation back-off squeezes such a request through anyway was not tested; it is not safe to test on
|
||
production.
|
||
- **So the live seat's practical long-file limit lies between about 1 minute and 400 s.**
|
||
|
||
**Long-form accuracy inside the envelope.**
|
||
|
||
- **Windows.** Both public files cut into pieces the seat can take: ≤ 375 s each, cut at pauses, with
|
||
ground truth sliced from the investigation's timed GT.
|
||
- **Method.** Every window goes in whole, in one request, in **two placements**: first cut at 360 s
|
||
(9 windows), then at 180 s (11 windows). Scoring uses the investigation's `gtscore` and its crosstalk
|
||
rule (diarized overlap ≥ 10 % = crosstalk).
|
||
|
||
**Clean-speech words dropped, then WER, for placement 1 / placement 2:**
|
||
|
||
| | Wilde: dropped | Wilde: WER % | SCOTUS: dropped | SCOTUS: WER % |
|
||
|---|---|---|---|---|
|
||
| **seat (v3 int8)** | **1,676 / 320** | **48.7 / 11.5** | **196 / 74** | **9.4 / 7.0** |
|
||
| v3 fp32 | 57 / 16 | 4.0 / 2.7 | 0 / 11 | 4.4 / 3.9 |
|
||
| C (unified int8) | **0 / 0** | 2.5 / 2.4 | 13 / 0 | 5.0 / 5.1 |
|
||
| C-fp32 (placement 1 only) | 0 | 2.4 | 13 | 5.0 |
|
||
| B-fp32 (placement 1 only) | 0 | 2.3 | 13 | 5.0 |
|
||
| **B-bf16w** | **0 / 0** | **2.3 / 2.1** | 13 / 0 | 5.0 / 4.8 |
|
||
| D (v2 int8) | 133 / **1,009** | 6.1 / 29.1 | 0 / 12 | 4.7 / 5.0 |
|
||
|
||
*For scale (the investigation, Scriberr's 120 s slicer, 8 placements):* v3 lost 140 [48–240] (Wilde)
|
||
and 66 (SCOTUS); unified lost 31 [0–92] and 22.
|
||
|
||
- **The seat's long-context collapse is severe at 6-minute windows, and it moves with the cuts.** It
|
||
lost 1,676 words or 320 on the audiobook, depending only on where the cuts fell. The same v3 weights
|
||
in fp32 lose 16–57. **int8 makes the known v3 flaw far worse.**
|
||
- **unified-en lost no clean speech on Wilde in any of 4 window runs** (2 placements × int8 and NeMo).
|
||
On SCOTUS it lost one 13-word stretch in one of the two placements.
|
||
- **Two placements show a large gap, not a confidence interval.** The investigation used 8.
|
||
- **v2 collapses on the audiobook too** (133 and 1,009 words), as the investigation found.
|
||
- **On SCOTUS the best WER is v3 fp32 (3.9–4.4 %).** unified is 4.8–5.1 % and the seat 7.0–9.4 %.
|
||
|
||
**One request for the whole file, beyond 400 s.** Only NeMo can do it.
|
||
|
||
- **Full attention needs more than 37 GB for 24 minutes.** The headroom guard stopped it.
|
||
- **NeMo's long-audio mode works.** With local attention ±128, B-bf16w transcribed:
|
||
- the 24-minute Wilde in **1.6–2.0 s**, WER 3.77–3.85 %, one clean dropout of 57 words;
|
||
- the 30-minute SCOTUS in **2.3–2.6 s**, WER 4.99–5.02 %, one clean dropout of 13 words.
|
||
- **No insertion runs.** The text was not byte-identical across the two runs (local attention plus
|
||
bf16), but the dropouts were the same in both.
|
||
|
||
## 6. Memory and where each candidate could live
|
||
|
||
**Per-process GPU memory in MiB, from `nvidia-smi` every 200 ms, filtered to each arm's own PIDs.**
|
||
|
||
| arm | at rest after warm-up | serving peak (utterance block) | highest seen (incl. load) | **extra over the seat's 1,690 MiB** (rest / serving / highest) |
|
||
|---|---|---|---|---|
|
||
| A, the live seat | 1,690 (1,692 after this test) | — | — | 0 |
|
||
| A′, the seat's config on a roomy card | 922–932, then **1,690–1,700 after the first 11.6 s request** | 4,764–5,798 † | 5,798 † | — |
|
||
| C (unified int8) / D (v2 int8) | 916 → 1,684 after the first 11.6 s request | 4,758 † | 4,758 † | **≈ 0** (the same footprint as the seat) |
|
||
| A-fp32 / C-fp32 | 3,922–3,932 / 3,874 | 6,994–7,004 † / 6,946 † | 7,004 / 6,946 | +2,240 / +5,310 / +5,310 |
|
||
| C-fp16 | 2,776 | 4,824 † | 4,824 | +1,090 / +3,130 / +3,130 |
|
||
| B-fp32 (`.nemo` restored straight to GPU) | 3,258–3,836 | 4,800 | **5,620** (load transient, about 1 s) | +2,150 / +3,110 / +3,930 |
|
||
| B-fp32, restored via CPU | 3,836 | 4,800 | 4,800 | +2,150 / +3,110 / +3,110 |
|
||
| B-bf16 (autocast) | 4,802 | 5,126 | 5,620 | +3,110 / +3,440 / +3,930 |
|
||
| **B-bf16w (bf16 weights, via CPU)** | **2,492** | **2,794** | **3,194** (load) | **+800 / +1,100 / +1,500** |
|
||
|
||
**† ORT's arena grows in 1,024 MiB steps under sustained never-seen-length traffic, whatever the request
|
||
length.**
|
||
|
||
- **What triggers a step.** Steps came after 2.8 s requests as often as after 55 s ones. The seat's
|
||
config held about 1,700 MiB through 59 s requests until one arrived. Every ORT arm stepped this way,
|
||
int8 and fp32 alike.
|
||
- **The seat's real working set is about 1.7 GB.** On a roomy card that grows to 4.8–5.8 GB over about
|
||
600 requests. On GPU 0 the live seat stays at 1,690 MiB because there is nothing to grow into, and
|
||
it has served 1,112 requests without one failure.
|
||
- **ORT "serving peaks" are arena policy, not need.** An ORT replacement on GPU 0 would behave the same
|
||
way, but its floor is its at-rest figure.
|
||
- **NeMo is different.** Its plateau, 4,800 MiB for fp32 and 2,794 MiB for bf16w, is torch's cache
|
||
holding the longest request's activations, and that includes 59 s clips. A seat capped at `talk`'s
|
||
≤ 12 s turns would plateau lower; I did not measure how much lower.
|
||
- **The bf16w load transient (3,194) should be avoidable** by casting to bf16 on the CPU before the
|
||
move. That is reasoned, not measured.
|
||
|
||
**Fit verdict for GPU 0.** The seat's 1,690 MiB plus 101 MiB free gives **1,791 MiB available in place**.
|
||
|
||
- **C and D fit in place but are not faster**, so there is nothing to gain.
|
||
- **None of the fast arms fits in place.** The cheapest is **B-bf16w at about +1.1 GB** (serving), or
|
||
**+1.5 GB** (its load transient as measured).
|
||
- By the coordinator's figure of about 0.95 GB per 0.01 of gen-small's `gpu_memory_utilization`,
|
||
that is **about 0.012–0.016 of gen-small**; I did not measure the KV trade.
|
||
- The fp32 routes need **+2.2 GB at rest and more as the arena steps**. They would also want an
|
||
arena cap.
|
||
- **Other homes, as observed at 17:4x and not checked for reservations:**
|
||
- **GPU 1 showed 6,625 MiB free** (intern-decision 8.8 GB plus vLLM seats), enough for B-bf16w's
|
||
3.2 GB peak. Whether that headroom is spoken for, I do not know.
|
||
- **GPU 3** is the card kept empty for a full-card seat. Scriberr uses it on demand and needs about
|
||
5.5 GB per job.
|
||
|
||
## 7. Output features callers might rely on
|
||
|
||
| | seat (v3 int8) | unified (any runtime) | v2 int8 |
|
||
|---|---|---|---|
|
||
| Response body | `{"text": …}` only | the same (the wrappers are identical) | the same |
|
||
| Casing | yes (94.5 % of outputs carry upper case) | yes (96–97 %) | yes |
|
||
| Commas | yes (73 %) | yes (62–65 %) | yes (74 %) |
|
||
| **Sentence-final `.?!`** | **85 % of test-clean outputs end with one** (AMI 82 %) | **43–46 %: it drops the final period about half the time** (AMI 54–56 %) | 80.5 % (AMI 85 %) |
|
||
| Timestamps | **none returned.** sherpa-onnx computes token times, and the wrapper discards them. | none (NeMo can return word and segment times; the wrapper does not ask) | none |
|
||
| Language field | none (not returned; `response_format` and `language` ignored) | none | none |
|
||
| Languages | 25 European | **English only** | English only |
|
||
| Longest input | **400 s, then HTTP 500** (§ 5.4) | ONNX: the same 400 s. NeMo: no ceiling (§ 5.4) | 400 s |
|
||
| Empty outputs on AMI | 26/400 | 9–17/400 | 15/400 |
|
||
|
||
**English-only is a real change, but no current caller depends on another language.** `talk`, the
|
||
WER tooling and the accent probes are English. A non-English clip sent to unified-en would come back
|
||
as English-ish text, not as an error.
|
||
|
||
## 8. Recommendation
|
||
|
||
| claim | strength | basis | reversibility |
|
||
|---|---|---|---|
|
||
| The seat's latency is its runtime (int8 on one CPU thread), not its model | **insist** | measured: CPU/wall 1.00, GPU 2–9 %; the same v3 weights in fp32 are 4–12× faster in the same image | n/a (a finding) |
|
||
| Move the seat off the int8-on-CPU graph | **strongly recommend** | measured: every GPU-native arm saves 100–530 ms per 1–20 s turn (CIs ±4 to ±35 ms; floor ≤ 6 ms), and the int8 graph also costs WER and drops the rest of an utterance after a gap (§ 5.3) | reversible (image or alias swap) |
|
||
| The destination: unified-en under NeMo with bf16 weights (B-bf16w), over v3 as fp32 ONNX | **recommend** | measured: fastest arm (−121 / −234 / −530 ms vs the seat; 16–42 ms faster than fp32 ONNX), the best AMI WER, no 400 s ceiling, the smallest fast footprint. Against it: English only, NVIDIA Open Model License, about 1.1–1.5 GB of GPU 0 to find, a heavier image | reversible (old image and alias kept) |
|
||
| If English-only or the licence is a blocker: v3 exported fp32 in the seat's own image | **lean** | measured: −107 / −213 / −495 ms, 1.2–2.2 pp better than today; but +2.2 GB, an arena that steps, and still the 400 s ceiling; the export is ours (k2-fsa do not publish v3 fp32) | reversible |
|
||
| Do NOT adopt C (k2-fsa's unified int8 in the seat runtime) | **recommend against** | measured: 8–71 ms SLOWER than the seat; its accuracy gain is real but available faster elsewhere | reversible |
|
||
| Do NOT adopt D (v2 int8) as the English alternative | **recommend against** | measured: the same speed as the seat (+3 ms, inside the floor); better WER than v3 int8 but worse than unified | reversible |
|
||
| Do NOT use the fp16 ONNX conversion | **lean against** | measured: 50–60 ms slower than fp32 on short turns (per-shape tuning); v3 fp16 is broken outright | reversible |
|
||
| Stopgap with no model change: `NUM_THREADS=16` on the live seat | **lean** | measured: −61 / −119 / −249 ms and identical text, but it burns 12–29 host cores per request and needs a seat recreate (~45 s warm-up) | reversible (one env var) |
|
||
|
||
## 9. What a switch to B-bf16w would take
|
||
|
||
- **Runtime.** NeMo 3.0.0 + torch 2.8 (cu128) + FastAPI/uvicorn.
|
||
- The wrapper exists: `services/parakeet-ab-2026-09-30/code/serve_nemo.py`, run with
|
||
`DTYPE=bf16w LOAD_CPU=1`. It keeps the seat's endpoints and `{"text": …}` body, and it returns
|
||
the same text as NeMo's `transcribe()`.
|
||
- NeMo's card says 2.7.3, but released 2.7.3 lacks this encoder's `att_chunk_context_size`, so it
|
||
needs **3.0.0**. The `.nemo` also lacks a `validation_ds` config, which needs the two-line shim
|
||
the wrapper carries (both as found by the investigation).
|
||
- **Image.** None exists yet. This A/B ran the wrapper from a uv env mounted into the
|
||
`scriberr:local-blackwell` image.
|
||
- A seat image would be a CUDA 12.8 runtime base plus that env: torch and the NVIDIA wheels come to
|
||
several GB, against the current seat's 5.09 GB image.
|
||
- Also needed before shipping: a warm-up at the longest served length (the CUDA-graph capture), a
|
||
bf16 cast before the move to GPU (the load transient), and, for files over a few minutes, a
|
||
switch to local attention (`LOCAL_ATT=128,128`).
|
||
- With local attention a 30-minute file took 2.6 s in one request, and memory grows linearly
|
||
instead of quadratically (§ 5.4).
|
||
- **Weights.** `nvidia/parakeet-unified-en-0.6b` @ `fe53cd885760c96b6a5f51a0bfd362cb4584a98b`, already
|
||
pinned in `/tank/aimodels/huggingface` (sha256 `ec23ed91…`, mount read-only).
|
||
- **Where it lives.** GPU 0 needs about +1.1 GB serving, or +1.5 GB with the load transient as
|
||
measured, from a vLLM seat's KV (gen-small is the coordinator's suggested donor). Alternatively it
|
||
could live on GPU 1, if that card's 6.6 GB of free headroom is really free (§ 6).
|
||
- **Licence.** The NVIDIA Open Model License Agreement (the card says "ready for
|
||
commercial/non-commercial use"; read at the card, the agreement's own text not reviewed here). That
|
||
replaces CC-BY-4.0 for this seat. Heads-up only; the call is Prime's.
|
||
- **Consumers.** None need a code change: same port, same body, and the LiteLLM alias is unchanged.
|
||
The visible differences:
|
||
- fewer sentence-final periods;
|
||
- English only;
|
||
- inputs over 400 s now work instead of returning 500.
|
||
|
||
## 10. Provenance
|
||
|
||
| artifact | id / revision | sha256 | licence |
|
||
|---|---|---|---|
|
||
| live seat weights (v3 int8) | `/tank/parakeet/models` (k2-fsa `sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8`) | encoder `acfc2b44…`, decoder `179e50c4…`, joiner `3164c13f…`, tokens `d5854467…` | CC-BY-4.0 |
|
||
| unified int8 (arm C) | GitHub `k2-fsa/sherpa-onnx` release `asr-models`, `…unified-en-0.6b-int8-non-streaming.tar.bz2` | tarball `99f63605…e150` (= GitHub digest); encoder `6716910b…` | NVIDIA Open Model License |
|
||
| v2 int8 (arm D) | the same release, `…parakeet-tdt-0.6b-v2-int8.tar.bz2` | tarball `157c157b…61e1ad` (= GitHub digest); encoder `a32b12d1…` | CC-BY-4.0 |
|
||
| unified `.nemo` (arm B, and C-fp32/fp16 source) | HF `nvidia/parakeet-unified-en-0.6b` @ `fe53cd88…` | `ec23ed91…` (= the investigation's pin) | NVIDIA Open Model License |
|
||
| v3 `.nemo` (A-fp32 source) | Scriberr's env copy (HF `541d1f99`, per the investigation) | `3cbdc858…` | CC-BY-4.0 |
|
||
| our fp32 exports | k2-fsa recipe @ sherpa-onnx `040afe36` minus quantisation | unified encoder.weights `d1559afb…`; v3 encoder.weights `9a22d372…` | as their source |
|
||
| LibriSpeech test parquets | HF `openslr/librispeech_asr` @ `71cacbfb…` | clean `7113aa4c…`, other `38e0c86a…` (= LFS oids) | CC-BY-4.0 |
|
||
| AMI IHM test parquets (4) | HF `edinburghcstr/ami` @ `46f28f25…` | `d95920dc…`, `07a2b1c4…`, `83cde21e…`, `3312bb79…` (= LFS oids) | CC-BY-4.0 |
|
||
|
||
- **Every HF id was verified with an authenticated API call before pulling.** A known-phantom repo
|
||
returned 404 as a check of the check.
|
||
- **GitHub had no token anywhere** (none on nh3-dev, on fv-ml1 or in the vault).
|
||
- GitHub answers 404, not 401, for a missing public asset, and the release API returned each asset's
|
||
size and sha256 `digest`.
|
||
- Every download matched that digest. That is the positive proof the authenticated call stands in
|
||
for.
|
||
- **Full hashes:** `services/parakeet-ab-2026-09-30/results/model-sha256.txt` and
|
||
`provenance-fetch.json`.
|
||
|
||
## 11. Host changes, cleanup and reproduction
|
||
|
||
- **fv-ml1.**
|
||
- **Spike dir `/tank/spikes/parakeet-ab/`.** It holds the downloads, the extracted and exported
|
||
models (about 12 GB), the test sets, two uv envs and the raw outputs.
|
||
- **About 40 transient `ab-*` containers**, all `--rm`, on **GPU 3 only**, bound to 127.0.0.1 and
|
||
labelled `ab=parakeet-2026-09-30`. **All are removed. GPU 3 holds no process of ours.**
|
||
- **The memory sampler (`nvidia-smi -lms 200`) is stopped.**
|
||
- **The live `parakeet` container got 180 requests** (§ 4.3). It saw no restart and no config
|
||
change.
|
||
- **Untouched:** Scriberr, intern-decision, the vLLM seats and LiteLLM's config.
|
||
- **ops-log.** `create-spike-dir` and `transient-containers` records (2026-09-30 23:11 and 23:38 UTC),
|
||
plus a cleanup record.
|
||
- **The spike dir is kept** for re-runs and holds no private data. Whether to delete it, about 25 GB
|
||
with envs, caches and test audio, is Prime's call.
|
||
- **Reproduce.** `code/` in this directory is the full harness:
|
||
- `mk_env.sh`, `prep_data.py`, `export_onnx.py`, `convert_fp16.py`, `make_windows.py`, `make_pc2.py`;
|
||
- `arm.sh` (start an arm), `run_lat.py` (the latency blocks), `run_acc.py`, `longwin.py`;
|
||
- `analyze_lat.py`, `gw_analysis.py`, `acc_summary.py`, `score.py`, `longscore.py`, `memtrace.py`.
|