Deployed 2026-09-30 1211 PT by pointing SCRIBERR_IMAGE at the patched tag (.env backed up as .env.bak-20260930-pre-slicer1; rollback is the unpatched scriberr:local-blackwell). PrepareEnvironment rewrote the env's parakeet_transcribe_buffered.py from the embed (sha256 matches the patched source). One live run on GPU 1 beside intern-decision peaked at 5,496 MiB. Memory records the open Parakeet mid-chunk dropout finding and the held upstream PR.
13 KiB
Scriberr Parakeet slicer bench (2026-09-30)
Prime's ruling, 2026-09-30: "build the slicer". This document records how the
pause-aware slicer patch (stacks/scriberr/patches/0001-parakeet-pause-aware-slicer.patch)
was measured and why the shipped variant was chosen. Metrics only: two of
the recordings are Prime's and private, so no transcript text appears here or
anywhere in git. Those recordings, and every transcript made from them, stay on
fv-ml1 in /tank/spikes/scriberr-slicer/private/ (mode 700).
Recordings (4 files, 118 minutes)
All four were converted to 16 kHz mono WAV with the image's ffmpeg, which is what Scriberr feeds Parakeet.
| id | length | what | provenance / licence |
|---|---|---|---|
| p1 | 35.3 min | Prime's upload, conversational | private |
| p2 | 22.3 min | Prime's upload, conversational | private |
| scotus | 30.0 min (first half hour) | U.S. Supreme Court oral argument, Loper Bright Enterprises v. Raimondo, No. 22-451, argued 2024-01-17; spontaneous multi-speaker speech with interruptions | https://www.supremecourt.gov/media/audio/mp3files/22-451.mp3 (sha256 7e704b1f…5dc2); U.S. Government work, public domain (17 U.S.C. §105) |
| wilde | 24.4 min | LibriVox The Trial of Oscar Wilde (dramatic reading), section 1; several readers, courtroom dialogue | https://archive.org/download/trialofoscarwilde_1601_librivox/trialofoscarwilde_01_anon_64kb.mp3 (sha256 492f4c56…3420); public domain (LibriVox, PD Mark 1.0) |
Harness
- Image and env:
scriberr:local-blackwell(the live image), the live/tank/scriberr/whisperx-envmounted read-only,PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, invoked exactly as Scriberr does (uv run --native-tls --project /app/whisperx-env/parakeet python …). It runs as uid 1002 withUSERset, not as appuser 10001; that changes file ownership only. - GPU: transient containers on GPU 3 only (verified: the container sees
one card, UUID
GPU-186dacf4…). GPU 3 was at 2 MiB before and after. - Reference: upstream's standard script on the whole file in one pass with
local attention (
--context-left 255 --context-right 255), which has no cuts. It needs >16 GB, which is why it cannot run on GPU 1. - Variants run through the real
transcribe_buffered(); the only harness change is that the model is loaded once per process instead of once per call. The CLI memory runs below reproduce the harness output byte for byte. - Placements: every variant at three maximum lengths, 120, 110 and 100 s. Parakeet is deterministic, so a repeat run cannot supply variance; moving the cuts can. Each (file, variant) therefore has n = 3 different cut placements.
Metric
Each variant's words are aligned against the reference, after lowercasing and stripping punctuation (difflib opcodes, then exact Levenshtein inside each non-matching block). Each error gets a time: the hypothesis word's start for a substitution or insertion, the reference word's start for a deletion.
- near-cut / elsewhere word errors (S/I/D): near-cut means within ±3 s of any of that variant's cuts (for overlap variants, the cut is the stitch point at the middle of the overlap, and both chunk edges lie within ±3 s of it).
- dropped / duplicated at cuts: deletions near a cut, and insertions near a cut that repeat an adjacent word.
- error events: errors clustered with gaps ≤ 1 s; rate per minute elsewhere.
- damaged cuts (the decision metric): cuts with at least one error within ±3 s, against damaged phantoms, the same test at points midway between the variant's own cuts (same count, same spacing, as far from any cut as the audio gets). The phantom rate is the background any slicer's cuts are judged against.
Why the decision metric is not raw word counts
The negative control caught it. Word errors elsewhere varied by up to ±50 % between slicers of the same length (p1: 131 to 215). Most of that comes from a few unstable stretches, the same stretches for every variant (p1: 976–992 s, 1291–1298 s, 928–934 s), where the reference and any slicing disagree by 20–60 words depending on context, at arbitrary distances from any cut. That is heavy-tailed noise, not cut damage. Counting events, and asking per cut whether anything near it went wrong, is robust to it; the word counts are still reported.
Variants
| name | overlap | pause search | stitch |
|---|---|---|---|
| fixed | 0 | off | none: upstream's slicer (patched code with both off reproduces it byte for byte) |
| pause | 0 | 25 s | none needed |
| overlap | 4 s | off | v1: each word kept by the chunk whose cut-to-cut range holds its start time (the brief's rule) |
| both | 4 s | 25 s | v1 |
| overlap2 | 4 s | off | v2: hand over at the nearest word both chunks transcribed alike, within 0.5 s (shipped as v3, identical output) |
| both2 | 4 s | 25 s | v2 |
| both8 | 8 s | 25 s | v2 |
v3 is v2 after code review, and it is what ships. Anchors are paired by time first (same text and within 0.5 s, so a longer text match elsewhere in the overlap cannot crowd out the true anchor), punctuation alone never anchors, and the anchor's copy comes from whichever chunk keeps word starts in time order. Re-run on all 12 overlap runs (4 files × 3 placements), v3's output is byte-identical to v2's (and "both3" to "both2"): the reviewer's cases did not occur in this audio, so every v2 number below holds for the shipped code.
v2 exists because v1 failed its own goal. At 26 of 108 "both" stitches the same word sat on both sides of the cut, the two copies' start times 0.00–0.09 s apart (one encoder frame): a word that follows a pause is timestamped anywhere in the pause, and a pause is exactly where a pause-aware cut lands. Start time at the midpoint is the worst possible stitch rule for pause cuts.
Results
Pooled (4 files × 3 placements)
| variant | damaged cuts | damaged phantoms | excess | by length 120 / 110 / 100 | near-cut words | dropped | duplicated | near-cut events | events elsewhere /min |
|---|---|---|---|---|---|---|---|---|---|
| fixed (upstream) | 93/179 = 52 % | 34/179 = 19 % | +0.33 | 0.48 / 0.52 / 0.55 | 223 | 56 | 17 | 102 | 2.08 |
| pause | 53/201 = 26 % | 36/201 = 18 % | +0.085 | 0.33 / 0.20 / 0.27 | 113 | 52 | 0 | 59 | 2.23 |
| overlap (v1) | 52/184 = 28 % | 33/184 = 18 % | +0.10 | 0.30 / 0.28 / 0.27 | 117 | 25 | 18 | 60 | 2.00 |
| both (v1) | 75/207 = 36 % | 38/207 = 18 % | +0.18 | 0.39 / 0.35 / 0.36 | 126 | 21 | 40 | 87 | 2.35 |
| overlap2 | 41/184 = 22 % | 33/184 = 18 % | +0.04 | 0.21 / 0.22 / 0.24 | 101 | 24 | 2 | 47 | 2.00 |
| both2 | 52/207 = 25 % | 38/207 = 18 % | +0.07 | 0.32 / 0.20 / 0.24 | 89 | 22 | 2 | 54 | 2.35 |
| both8 | 48/222 = 22 % | 46/222 = 21 % | +0.01 | 0.23 / 0.23 / 0.19 | 120 | 42 | 0 | 54 | 2.18 |
Per file (summed over the 3 placements)
| file | variant | damaged cuts | phantom | near S/I/D | near words | else words/min | drop | dup | events near | events else/min |
|---|---|---|---|---|---|---|---|---|---|---|
| p1 | fixed | 27/57 | 11/57 | 21/27/9 | 57 | 4.16 | 9 | 10 | 29 | 2.15 |
| p1 | pause | 19/64 | 8/64 | 20/3/25 | 48 | 5.62 | 25 | 0 | 20 | 2.39 |
| p1 | overlap | 20/59 | 16/59 | 16/26/4 | 46 | 4.66 | 4 | 8 | 23 | 2.07 |
| p1 | both | 29/67 | 13/67 | 23/25/4 | 52 | 5.70 | 4 | 17 | 35 | 2.56 |
| p1 | overlap2 | 15/59 | 16/59 | 16/19/4 | 39 | 4.66 | 4 | 0 | 17 | 2.07 |
| p1 | both2 | 21/67 | 13/67 | 23/8/5 | 36 | 5.70 | 5 | 0 | 21 | 2.56 |
| p1 | both8 | 18/70 | 15/70 | 16/2/8 | 26 | 6.22 | 8 | 0 | 18 | 2.40 |
| p2 | fixed | 14/36 | 4/36 | 10/7/14 | 31 | 4.45 | 14 | 2 | 16 | 1.39 |
| p2 | pause | 10/40 | 6/40 | 11/1/1 | 13 | 2.71 | 1 | 0 | 10 | 1.56 |
| p2 | overlap | 6/36 | 6/36 | 4/2/0 | 6 | 3.56 | 0 | 2 | 6 | 1.47 |
| p2 | both | 16/41 | 6/41 | 13/10/0 | 23 | 3.11 | 0 | 8 | 18 | 1.64 |
| p2 | overlap2 | 4/36 | 6/36 | 4/0/0 | 4 | 3.56 | 0 | 0 | 4 | 1.47 |
| p2 | both2 | 10/41 | 6/41 | 13/2/0 | 15 | 3.11 | 0 | 0 | 11 | 1.64 |
| p2 | both8 | 9/43 | 7/43 | 9/1/1 | 11 | 3.17 | 1 | 0 | 9 | 1.58 |
| scotus | fixed | 27/47 | 15/47 | 14/49/21 | 84 | 9.47 | 21 | 4 | 30 | 3.34 |
| scotus | pause | 18/54 | 17/54 | 13/9/10 | 32 | 8.63 | 10 | 0 | 22 | 3.36 |
| scotus | overlap | 20/49 | 6/49 | 16/23/11 | 50 | 9.43 | 11 | 5 | 24 | 3.01 |
| scotus | both | 19/55 | 14/55 | 9/13/5 | 27 | 9.46 | 5 | 8 | 21 | 3.39 |
| scotus | overlap2 | 18/49 | 6/49 | 15/20/10 | 45 | 9.43 | 10 | 2 | 21 | 3.01 |
| scotus | both2 | 15/55 | 14/55 | 9/7/5 | 21 | 9.46 | 5 | 2 | 15 | 3.39 |
| scotus | both8 | 14/61 | 19/61 | 14/31/10 | 55 | 8.19 | 10 | 0 | 19 | 3.30 |
| wilde | fixed | 25/39 | 4/39 | 9/30/12 | 51 | 6.91 | 12 | 1 | 27 | 1.03 |
| wilde | pause | 6/43 | 5/43 | 4/0/16 | 20 | 6.95 | 16 | 0 | 7 | 1.23 |
| wilde | overlap | 6/40 | 5/40 | 2/3/10 | 15 | 6.74 | 10 | 3 | 7 | 1.13 |
| wilde | both | 11/44 | 5/44 | 4/8/12 | 24 | 9.43 | 12 | 7 | 13 | 1.43 |
| wilde | overlap2 | 4/40 | 5/40 | 3/0/10 | 13 | 6.74 | 10 | 0 | 5 | 1.13 |
| wilde | both2 | 6/44 | 5/44 | 4/1/12 | 17 | 9.43 | 12 | 0 | 7 | 1.43 |
| wilde | both8 | 7/48 | 5/48 | 5/0/23 | 28 | 6.80 | 23 | 0 | 8 | 1.04 |
Where the remaining near-cut errors sit
Signed distance from the cut, pooled over files: upstream's errors pile up within ±0.5 s (chopped words; at 100 s, 24 insertions in [−0.5, 0) alone). Pause-only leaves deletions 1–3 s before its cuts, where the left chunk ends (8/9, 3/4, 7/8 in [−3, −1) at the three lengths). overlap2's residue is spread evenly over ±3 s, which is what background looks like. both2 and both8 are clean at the handover but keep some of pause-only's pre-cut deletions.
Controls
- A-vs-A: the same variant twice gives byte-identical words, segments and text (fixed-120, both-120, both2-120 on all 4 files, in one process), and the shipped script run three times as separate CLI processes matches the harness output byte for byte (p1 and scotus). Decoding is deterministic, so run-to-run variance is zero; the variance that matters comes from cut placement.
- Harness validity: patched code with overlap and pause search off equals upstream's unmodified script on all 4 files; v2 without overlap equals v1.
- Positive control: upstream's fixed cutter damages 52 % of its cuts against a 19 % background (+0.33, about 7 standard errors), per file 47–64 % against 11–25 %. The instrument sees cut damage.
- Negative control: the phantom (background) rate is 18–21 % for every variant, and error events away from cuts run at 2.0–2.35 per minute for all of them. Word counts away from cuts do not agree across slicers (see "Parakeet drops stretches" below), which is why they are not the decision metric.
- Sensitivity floor: at ~180–220 cuts per variant the 2-standard-error band on a difference of damaged-cut rates is ±0.08 pooled and ±0.17 for one file. No slicer can be measured below the background (~18–21 %). Word-count differences under ~60 words are unresolvable, because a single dropped stretch (next section) is 10–60 words.
Verdict
Every variant that overlaps with the v2 handover, and pause-only, beats
upstream by 0.26–0.30 in damaged-cut rate, far outside the ±0.08 floor. Among
them the differences are inside the floor. overlap2 ships: tied best on
damaged cuts, fewest near-cut error events (47 vs 102), near-zero duplicates,
nothing concentrated at the handover, and the simplest mechanism. Pause search
measured neutral once the stitch was fixed, so it stays in the patch as an
opt-in --pause-search, off by default.
Memory (GPU 3, production invocation, nvidia-smi every 0.2 s)
| run | n | per-process peak |
|---|---|---|
| upstream script, 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB (reproduces the 2026-09-30 budget figure) |
| shipped script (v2), 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB |
| shipped script (v3, final), 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB; all 3 CLI outputs byte-identical to the harness |
| shipped script, 120 s, scotus (public) | 3 | 5,496 / 5,496 / 5,496 MiB |
| positive control: shipped script, 300 s, scotus | 1 | 7,056 MiB (deterministic; 300 s was n=3 in the earlier table) |
live, GPU 1, deployed container, docker exec as appuser, 120 s, p1 |
1 | 5,496 MiB (2026-09-30 1212 PT, beside intern-decision; 49 s for 35 min of audio; seam check OK) |
Zero spread; the peak is set by the 120 s maximum, and the overlap sits inside it. The scotus file reads the same peak as p1, so the rebuild script uses it as its public memory fixture. A spike shorter than the 0.2 s sample period could be missed.
Parakeet drops stretches of speech (separate finding, not the slicer)
Every chunked variant, upstream's included, sometimes skips a run of ≥10 consecutive words in the middle of a chunk, and which runs it skips changes chaotically with the cut placement. Wilde at 120 s loses nothing, at 110 s 171 words, and overlap2 at 110 s loses one 60-second stretch. Across 4 files × 3 placements:
| variant | runs of ≥10 words lost | words lost |
|---|---|---|
| fixed (upstream) | 15 | 599 |
| pause | 17 | 511 |
| overlap2 | 12 | 656 |
| both2 | 14 | 719 |
| both8 | 12 | 560 |
The whole-file local-attention reference does the same: runs where every slicer has words the reference lacks (13–16 per 12 comparisons, 340–400 words). Today's production setting (upstream, 120 s) lost 85 words in 2 runs on p2. The slicer neither causes nor cures it; it needs its own investigation (decoder settings, chunk length, or model), which was out of scope here.