Files
esh-pfi-infrastructure/docs/pfi/scriberr-slicer-bench-2026-09-30.md
T
vh ee3db68db1 feat(scriberr): overlap-and-stitch Parakeet slicer patch, rebuild script, bench
Carry patches/0001 on our Scriberr build (upstream a353078): adjacent
buffered chunks overlap by 4 s inside --chunk-len and hand over at a word
both chunks transcribed alike, instead of cutting at fixed marks with no
overlap. Pause-aware cutting is included as an opt-in (--pause-search);
it measured neutral once the stitch was right. The Go<->Python CLI and
JSON seam is unchanged.

Bench (4 recordings, 118 min, 3 cut placements each, against a no-cut
whole-file reference; metrics only, private audio stays on fv-ml1):
cuts with an error within +-3 s fall from 52% (93/179) to 22% (41/184)
against a 19% background; floor +-0.08. Positive control: upstream's
cutter +0.33 over background. A-vs-A byte-identical in-process and
across CLI processes. Peak GPU memory unchanged at 5,496 MiB (n=3).
Also found: Parakeet skips runs of >=10 words mid-chunk with any
slicer, upstream's included; not addressed here.

scripts/scriberr-rebuild clones a pinned upstream sha into a new
/opt/docker/src dir, git-apply-checks the patches, builds a distinct
tag, and checks embed, unit tests, the JSON seam (scriberr-seam-check.py)
and the memory budget on idle GPU 3. Deploy stays manual. The upstream
PR is prepared under patches/upstream-pr/ and not opened.
2026-09-30 12:10:32 -07:00

13 KiB
Raw Blame History

Scriberr Parakeet slicer bench (2026-09-30)

Prime's ruling, 2026-09-30: "build the slicer". This document records how the pause-aware slicer patch (stacks/scriberr/patches/0001-parakeet-pause-aware-slicer.patch) was measured and why the shipped variant was chosen. Metrics only: two of the recordings are Prime's and private, so no transcript text appears here or anywhere in git. Those recordings, and every transcript made from them, stay on fv-ml1 in /tank/spikes/scriberr-slicer/private/ (mode 700).

Recordings (4 files, 118 minutes)

All four were converted to 16 kHz mono WAV with the image's ffmpeg, which is what Scriberr feeds Parakeet.

id length what provenance / licence
p1 35.3 min Prime's upload, conversational private
p2 22.3 min Prime's upload, conversational private
scotus 30.0 min (first half hour) U.S. Supreme Court oral argument, Loper Bright Enterprises v. Raimondo, No. 22-451, argued 2024-01-17; spontaneous multi-speaker speech with interruptions https://www.supremecourt.gov/media/audio/mp3files/22-451.mp3 (sha256 7e704b1f…5dc2); U.S. Government work, public domain (17 U.S.C. §105)
wilde 24.4 min LibriVox The Trial of Oscar Wilde (dramatic reading), section 1; several readers, courtroom dialogue https://archive.org/download/trialofoscarwilde_1601_librivox/trialofoscarwilde_01_anon_64kb.mp3 (sha256 492f4c56…3420); public domain (LibriVox, PD Mark 1.0)

Harness

  • Image and env: scriberr:local-blackwell (the live image), the live /tank/scriberr/whisperx-env mounted read-only, PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, invoked exactly as Scriberr does (uv run --native-tls --project /app/whisperx-env/parakeet python …). It runs as uid 1002 with USER set, not as appuser 10001; that changes file ownership only.
  • GPU: transient containers on GPU 3 only (verified: the container sees one card, UUID GPU-186dacf4…). GPU 3 was at 2 MiB before and after.
  • Reference: upstream's standard script on the whole file in one pass with local attention (--context-left 255 --context-right 255), which has no cuts. It needs >16 GB, which is why it cannot run on GPU 1.
  • Variants run through the real transcribe_buffered(); the only harness change is that the model is loaded once per process instead of once per call. The CLI memory runs below reproduce the harness output byte for byte.
  • Placements: every variant at three maximum lengths, 120, 110 and 100 s. Parakeet is deterministic, so a repeat run cannot supply variance; moving the cuts can. Each (file, variant) therefore has n = 3 different cut placements.

Metric

Each variant's words are aligned against the reference, after lowercasing and stripping punctuation (difflib opcodes, then exact Levenshtein inside each non-matching block). Each error gets a time: the hypothesis word's start for a substitution or insertion, the reference word's start for a deletion.

  • near-cut / elsewhere word errors (S/I/D): near-cut means within ±3 s of any of that variant's cuts (for overlap variants, the cut is the stitch point at the middle of the overlap, and both chunk edges lie within ±3 s of it).
  • dropped / duplicated at cuts: deletions near a cut, and insertions near a cut that repeat an adjacent word.
  • error events: errors clustered with gaps ≤ 1 s; rate per minute elsewhere.
  • damaged cuts (the decision metric): cuts with at least one error within ±3 s, against damaged phantoms, the same test at points midway between the variant's own cuts (same count, same spacing, as far from any cut as the audio gets). The phantom rate is the background any slicer's cuts are judged against.

Why the decision metric is not raw word counts

The negative control caught it. Word errors elsewhere varied by up to ±50 % between slicers of the same length (p1: 131 to 215). Most of that comes from a few unstable stretches, the same stretches for every variant (p1: 976–992 s, 1291–1298 s, 928–934 s), where the reference and any slicing disagree by 20–60 words depending on context, at arbitrary distances from any cut. That is heavy-tailed noise, not cut damage. Counting events, and asking per cut whether anything near it went wrong, is robust to it; the word counts are still reported.

Variants

name overlap pause search stitch
fixed 0 off none: upstream's slicer (patched code with both off reproduces it byte for byte)
pause 0 25 s none needed
overlap 4 s off v1: each word kept by the chunk whose cut-to-cut range holds its start time (the brief's rule)
both 4 s 25 s v1
overlap2 4 s off v2: hand over at the nearest word both chunks transcribed alike, within 0.5 s (shipped as v3, identical output)
both2 4 s 25 s v2
both8 8 s 25 s v2

v3 is v2 after code review, and it is what ships. Anchors are paired by time first (same text and within 0.5 s, so a longer text match elsewhere in the overlap cannot crowd out the true anchor), punctuation alone never anchors, and the anchor's copy comes from whichever chunk keeps word starts in time order. Re-run on all 12 overlap runs (4 files × 3 placements), v3's output is byte-identical to v2's (and "both3" to "both2"): the reviewer's cases did not occur in this audio, so every v2 number below holds for the shipped code.

v2 exists because v1 failed its own goal. At 26 of 108 "both" stitches the same word sat on both sides of the cut, the two copies' start times 0.00–0.09 s apart (one encoder frame): a word that follows a pause is timestamped anywhere in the pause, and a pause is exactly where a pause-aware cut lands. Start time at the midpoint is the worst possible stitch rule for pause cuts.

Results

Pooled (4 files × 3 placements)

variant damaged cuts damaged phantoms excess by length 120 / 110 / 100 near-cut words dropped duplicated near-cut events events elsewhere /min
fixed (upstream) 93/179 = 52 % 34/179 = 19 % +0.33 0.48 / 0.52 / 0.55 223 56 17 102 2.08
pause 53/201 = 26 % 36/201 = 18 % +0.085 0.33 / 0.20 / 0.27 113 52 0 59 2.23
overlap (v1) 52/184 = 28 % 33/184 = 18 % +0.10 0.30 / 0.28 / 0.27 117 25 18 60 2.00
both (v1) 75/207 = 36 % 38/207 = 18 % +0.18 0.39 / 0.35 / 0.36 126 21 40 87 2.35
overlap2 41/184 = 22 % 33/184 = 18 % +0.04 0.21 / 0.22 / 0.24 101 24 2 47 2.00
both2 52/207 = 25 % 38/207 = 18 % +0.07 0.32 / 0.20 / 0.24 89 22 2 54 2.35
both8 48/222 = 22 % 46/222 = 21 % +0.01 0.23 / 0.23 / 0.19 120 42 0 54 2.18

Per file (summed over the 3 placements)

file variant damaged cuts phantom near S/I/D near words else words/min drop dup events near events else/min
p1 fixed 27/57 11/57 21/27/9 57 4.16 9 10 29 2.15
p1 pause 19/64 8/64 20/3/25 48 5.62 25 0 20 2.39
p1 overlap 20/59 16/59 16/26/4 46 4.66 4 8 23 2.07
p1 both 29/67 13/67 23/25/4 52 5.70 4 17 35 2.56
p1 overlap2 15/59 16/59 16/19/4 39 4.66 4 0 17 2.07
p1 both2 21/67 13/67 23/8/5 36 5.70 5 0 21 2.56
p1 both8 18/70 15/70 16/2/8 26 6.22 8 0 18 2.40
p2 fixed 14/36 4/36 10/7/14 31 4.45 14 2 16 1.39
p2 pause 10/40 6/40 11/1/1 13 2.71 1 0 10 1.56
p2 overlap 6/36 6/36 4/2/0 6 3.56 0 2 6 1.47
p2 both 16/41 6/41 13/10/0 23 3.11 0 8 18 1.64
p2 overlap2 4/36 6/36 4/0/0 4 3.56 0 0 4 1.47
p2 both2 10/41 6/41 13/2/0 15 3.11 0 0 11 1.64
p2 both8 9/43 7/43 9/1/1 11 3.17 1 0 9 1.58
scotus fixed 27/47 15/47 14/49/21 84 9.47 21 4 30 3.34
scotus pause 18/54 17/54 13/9/10 32 8.63 10 0 22 3.36
scotus overlap 20/49 6/49 16/23/11 50 9.43 11 5 24 3.01
scotus both 19/55 14/55 9/13/5 27 9.46 5 8 21 3.39
scotus overlap2 18/49 6/49 15/20/10 45 9.43 10 2 21 3.01
scotus both2 15/55 14/55 9/7/5 21 9.46 5 2 15 3.39
scotus both8 14/61 19/61 14/31/10 55 8.19 10 0 19 3.30
wilde fixed 25/39 4/39 9/30/12 51 6.91 12 1 27 1.03
wilde pause 6/43 5/43 4/0/16 20 6.95 16 0 7 1.23
wilde overlap 6/40 5/40 2/3/10 15 6.74 10 3 7 1.13
wilde both 11/44 5/44 4/8/12 24 9.43 12 7 13 1.43
wilde overlap2 4/40 5/40 3/0/10 13 6.74 10 0 5 1.13
wilde both2 6/44 5/44 4/1/12 17 9.43 12 0 7 1.43
wilde both8 7/48 5/48 5/0/23 28 6.80 23 0 8 1.04

Where the remaining near-cut errors sit

Signed distance from the cut, pooled over files: upstream's errors pile up within ±0.5 s (chopped words; at 100 s, 24 insertions in [−0.5, 0) alone). Pause-only leaves deletions 1–3 s before its cuts, where the left chunk ends (8/9, 3/4, 7/8 in [−3, −1) at the three lengths). overlap2's residue is spread evenly over ±3 s, which is what background looks like. both2 and both8 are clean at the handover but keep some of pause-only's pre-cut deletions.

Controls

  • A-vs-A: the same variant twice gives byte-identical words, segments and text (fixed-120, both-120, both2-120 on all 4 files, in one process), and the shipped script run three times as separate CLI processes matches the harness output byte for byte (p1 and scotus). Decoding is deterministic, so run-to-run variance is zero; the variance that matters comes from cut placement.
  • Harness validity: patched code with overlap and pause search off equals upstream's unmodified script on all 4 files; v2 without overlap equals v1.
  • Positive control: upstream's fixed cutter damages 52 % of its cuts against a 19 % background (+0.33, about 7 standard errors), per file 47–64 % against 11–25 %. The instrument sees cut damage.
  • Negative control: the phantom (background) rate is 18–21 % for every variant, and error events away from cuts run at 2.0–2.35 per minute for all of them. Word counts away from cuts do not agree across slicers (see "Parakeet drops stretches" below), which is why they are not the decision metric.
  • Sensitivity floor: at ~180–220 cuts per variant the 2-standard-error band on a difference of damaged-cut rates is ±0.08 pooled and ±0.17 for one file. No slicer can be measured below the background (~18–21 %). Word-count differences under ~60 words are unresolvable, because a single dropped stretch (next section) is 10–60 words.

Verdict

Every variant that overlaps with the v2 handover, and pause-only, beats upstream by 0.26–0.30 in damaged-cut rate, far outside the ±0.08 floor. Among them the differences are inside the floor. overlap2 ships: tied best on damaged cuts, fewest near-cut error events (47 vs 102), near-zero duplicates, nothing concentrated at the handover, and the simplest mechanism. Pause search measured neutral once the stitch was fixed, so it stays in the patch as an opt-in --pause-search, off by default.

Memory (GPU 3, production invocation, nvidia-smi every 0.2 s)

run n per-process peak
upstream script, 120 s, p1 3 5,496 / 5,496 / 5,496 MiB (reproduces the 2026-09-30 budget figure)
shipped script (v2), 120 s, p1 3 5,496 / 5,496 / 5,496 MiB
shipped script (v3, final), 120 s, p1 3 5,496 / 5,496 / 5,496 MiB; all 3 CLI outputs byte-identical to the harness
shipped script, 120 s, scotus (public) 3 5,496 / 5,496 / 5,496 MiB
positive control: shipped script, 300 s, scotus 1 7,056 MiB (deterministic; 300 s was n=3 in the earlier table)

Zero spread; the peak is set by the 120 s maximum, and the overlap sits inside it. The scotus file reads the same peak as p1, so the rebuild script uses it as its public memory fixture. A spike shorter than the 0.2 s sample period could be missed.

Parakeet drops stretches of speech (separate finding, not the slicer)

Every chunked variant, upstream's included, sometimes skips a run of ≥10 consecutive words in the middle of a chunk, and which runs it skips changes chaotically with the cut placement. Wilde at 120 s loses nothing, at 110 s 171 words, and overlap2 at 110 s loses one 60-second stretch. Across 4 files × 3 placements:

variant runs of ≥10 words lost words lost
fixed (upstream) 15 599
pause 17 511
overlap2 12 656
both2 14 719
both8 12 560

The whole-file local-attention reference does the same: runs where every slicer has words the reference lacks (13–16 per 12 comparisons, 340–400 words). Today's production setting (upstream, 120 s) lost 85 words in 2 runs on p2. The slicer neither causes nor cures it; it needs its own investigation (decoder settings, chunk length, or model), which was out of scope here.