Files
esh-pfi-infrastructure/docs/pfi/scriberr-slicer-bench-2026-09-30.md
T
vh ee3db68db1 feat(scriberr): overlap-and-stitch Parakeet slicer patch, rebuild script, bench
Carry patches/0001 on our Scriberr build (upstream a353078): adjacent
buffered chunks overlap by 4 s inside --chunk-len and hand over at a word
both chunks transcribed alike, instead of cutting at fixed marks with no
overlap. Pause-aware cutting is included as an opt-in (--pause-search);
it measured neutral once the stitch was right. The Go<->Python CLI and
JSON seam is unchanged.

Bench (4 recordings, 118 min, 3 cut placements each, against a no-cut
whole-file reference; metrics only, private audio stays on fv-ml1):
cuts with an error within +-3 s fall from 52% (93/179) to 22% (41/184)
against a 19% background; floor +-0.08. Positive control: upstream's
cutter +0.33 over background. A-vs-A byte-identical in-process and
across CLI processes. Peak GPU memory unchanged at 5,496 MiB (n=3).
Also found: Parakeet skips runs of >=10 words mid-chunk with any
slicer, upstream's included; not addressed here.

scripts/scriberr-rebuild clones a pinned upstream sha into a new
/opt/docker/src dir, git-apply-checks the patches, builds a distinct
tag, and checks embed, unit tests, the JSON seam (scriberr-seam-check.py)
and the memory budget on idle GPU 3. Deploy stays manual. The upstream
PR is prepared under patches/upstream-pr/ and not opened.
2026-09-30 12:10:32 -07:00

221 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Scriberr Parakeet slicer bench (2026-09-30)
Prime's ruling, 2026-09-30: "build the slicer". This document records how the
pause-aware slicer patch (`stacks/scriberr/patches/0001-parakeet-pause-aware-slicer.patch`)
was measured and why the shipped variant was chosen. **Metrics only**: two of
the recordings are Prime's and private, so no transcript text appears here or
anywhere in git. Those recordings, and every transcript made from them, stay on
fv-ml1 in `/tank/spikes/scriberr-slicer/private/` (mode 700).
## Recordings (4 files, 118 minutes)
All four were converted to 16 kHz mono WAV with the image's ffmpeg, which is
what Scriberr feeds Parakeet.
| id | length | what | provenance / licence |
|---|---|---|---|
| p1 | 35.3 min | Prime's upload, conversational | private |
| p2 | 22.3 min | Prime's upload, conversational | private |
| scotus | 30.0 min (first half hour) | U.S. Supreme Court oral argument, *Loper Bright Enterprises v. Raimondo*, No. 22-451, argued 2024-01-17; spontaneous multi-speaker speech with interruptions | `https://www.supremecourt.gov/media/audio/mp3files/22-451.mp3` (sha256 `7e704b1f…5dc2`); U.S. Government work, public domain (17 U.S.C. §105) |
| wilde | 24.4 min | LibriVox *The Trial of Oscar Wilde* (dramatic reading), section 1; several readers, courtroom dialogue | `https://archive.org/download/trialofoscarwilde_1601_librivox/trialofoscarwilde_01_anon_64kb.mp3` (sha256 `492f4c56…3420`); public domain (LibriVox, PD Mark 1.0) |
## Harness
- **Image and env:** `scriberr:local-blackwell` (the live image), the live
`/tank/scriberr/whisperx-env` mounted **read-only**, `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`,
invoked exactly as Scriberr does (`uv run --native-tls --project /app/whisperx-env/parakeet python …`).
It runs as uid 1002 with `USER` set, not as appuser 10001; that changes file
ownership only.
- **GPU:** transient containers on **GPU 3 only** (verified: the container sees
one card, UUID `GPU-186dacf4…`). GPU 3 was at 2 MiB before and after.
- **Reference:** upstream's standard script on the whole file in one pass with
local attention (`--context-left 255 --context-right 255`), which has no cuts.
It needs >16 GB, which is why it cannot run on GPU 1.
- **Variants** run through the real `transcribe_buffered()`; the only harness
change is that the model is loaded once per process instead of once per call.
The CLI memory runs below reproduce the harness output byte for byte.
- **Placements:** every variant at three maximum lengths, 120, 110 and 100 s.
Parakeet is deterministic, so a repeat run cannot supply variance; moving the
cuts can. Each (file, variant) therefore has n = 3 different cut placements.
## Metric
Each variant's words are aligned against the reference, after lowercasing and
stripping punctuation (difflib opcodes, then exact Levenshtein inside each
non-matching block). Each error gets a time: the hypothesis word's start for a
substitution or insertion, the reference word's start for a deletion.
- **near-cut / elsewhere word errors (S/I/D):** near-cut means within ±3 s of
any of that variant's cuts (for overlap variants, the cut is the stitch point
at the middle of the overlap, and both chunk edges lie within ±3 s of it).
- **dropped / duplicated at cuts:** deletions near a cut, and insertions near a
cut that repeat an adjacent word.
- **error events:** errors clustered with gaps ≤ 1 s; rate per minute
elsewhere.
- **damaged cuts (the decision metric):** cuts with at least one error within
±3 s, against **damaged phantoms**, the same test at points midway between
the variant's own cuts (same count, same spacing, as far from any cut as the
audio gets). The phantom rate is the background any slicer's cuts are judged
against.
### Why the decision metric is not raw word counts
The negative control caught it. Word errors elsewhere varied by up to ±50 %
between slicers of the same length (p1: 131 to 215). Most of that comes from a
few **unstable stretches**, the same stretches for every variant (p1: 976–992 s,
1291–1298 s, 928–934 s), where the reference and any slicing disagree by 20–60
words depending on context, at arbitrary distances from any cut. That is heavy-tailed
noise, not cut damage. Counting events, and asking per cut whether anything near
it went wrong, is robust to it; the word counts are still reported.
## Variants
| name | overlap | pause search | stitch |
|---|---|---|---|
| fixed | 0 | off | none: upstream's slicer (patched code with both off reproduces it byte for byte) |
| pause | 0 | 25 s | none needed |
| overlap | 4 s | off | v1: each word kept by the chunk whose cut-to-cut range holds its start time (the brief's rule) |
| both | 4 s | 25 s | v1 |
| **overlap2** | 4 s | off | **v2: hand over at the nearest word both chunks transcribed alike, within 0.5 s** (shipped as v3, identical output) |
| both2 | 4 s | 25 s | v2 |
| both8 | 8 s | 25 s | v2 |
**v3** is v2 after code review, and it is what ships. Anchors are paired by
time first (same text *and* within 0.5 s, so a longer text match elsewhere in
the overlap cannot crowd out the true anchor), punctuation alone never anchors,
and the anchor's copy comes from whichever chunk keeps word starts in time order.
Re-run on all 12 overlap runs (4 files × 3 placements), **v3's output is
byte-identical to v2's** (and "both3" to "both2"): the reviewer's cases did not
occur in this audio, so every v2 number below holds for the shipped code.
v2 exists because v1 failed its own goal. At 26 of 108 "both" stitches the same
word sat on both sides of the cut, the two copies' start times 0.00–0.09 s apart
(one encoder frame): a word that follows a pause is timestamped anywhere in the
pause, and a pause is exactly where a pause-aware cut lands. Start time at the
midpoint is the worst possible stitch rule for pause cuts.
## Results
### Pooled (4 files × 3 placements)
| variant | damaged cuts | damaged phantoms | excess | by length 120 / 110 / 100 | near-cut words | dropped | duplicated | near-cut events | events elsewhere /min |
|---|---|---|---|---|---|---|---|---|---|
| fixed (upstream) | 93/179 = 52 % | 34/179 = 19 % | +0.33 | 0.48 / 0.52 / 0.55 | 223 | 56 | 17 | 102 | 2.08 |
| pause | 53/201 = 26 % | 36/201 = 18 % | +0.085 | 0.33 / 0.20 / 0.27 | 113 | 52 | 0 | 59 | 2.23 |
| overlap (v1) | 52/184 = 28 % | 33/184 = 18 % | +0.10 | 0.30 / 0.28 / 0.27 | 117 | 25 | 18 | 60 | 2.00 |
| both (v1) | 75/207 = 36 % | 38/207 = 18 % | +0.18 | 0.39 / 0.35 / 0.36 | 126 | 21 | 40 | 87 | 2.35 |
| **overlap2** | **41/184 = 22 %** | 33/184 = 18 % | **+0.04** | 0.21 / 0.22 / 0.24 | 101 | 24 | 2 | **47** | 2.00 |
| both2 | 52/207 = 25 % | 38/207 = 18 % | +0.07 | 0.32 / 0.20 / 0.24 | 89 | 22 | 2 | 54 | 2.35 |
| both8 | 48/222 = 22 % | 46/222 = 21 % | +0.01 | 0.23 / 0.23 / 0.19 | 120 | 42 | 0 | 54 | 2.18 |
### Per file (summed over the 3 placements)
| file | variant | damaged cuts | phantom | near S/I/D | near words | else words/min | drop | dup | events near | events else/min |
|---|---|---|---|---|---|---|---|---|---|---|
| p1 | fixed | 27/57 | 11/57 | 21/27/9 | 57 | 4.16 | 9 | 10 | 29 | 2.15 |
| p1 | pause | 19/64 | 8/64 | 20/3/25 | 48 | 5.62 | 25 | 0 | 20 | 2.39 |
| p1 | overlap | 20/59 | 16/59 | 16/26/4 | 46 | 4.66 | 4 | 8 | 23 | 2.07 |
| p1 | both | 29/67 | 13/67 | 23/25/4 | 52 | 5.70 | 4 | 17 | 35 | 2.56 |
| p1 | overlap2 | 15/59 | 16/59 | 16/19/4 | 39 | 4.66 | 4 | 0 | 17 | 2.07 |
| p1 | both2 | 21/67 | 13/67 | 23/8/5 | 36 | 5.70 | 5 | 0 | 21 | 2.56 |
| p1 | both8 | 18/70 | 15/70 | 16/2/8 | 26 | 6.22 | 8 | 0 | 18 | 2.40 |
| p2 | fixed | 14/36 | 4/36 | 10/7/14 | 31 | 4.45 | 14 | 2 | 16 | 1.39 |
| p2 | pause | 10/40 | 6/40 | 11/1/1 | 13 | 2.71 | 1 | 0 | 10 | 1.56 |
| p2 | overlap | 6/36 | 6/36 | 4/2/0 | 6 | 3.56 | 0 | 2 | 6 | 1.47 |
| p2 | both | 16/41 | 6/41 | 13/10/0 | 23 | 3.11 | 0 | 8 | 18 | 1.64 |
| p2 | overlap2 | 4/36 | 6/36 | 4/0/0 | 4 | 3.56 | 0 | 0 | 4 | 1.47 |
| p2 | both2 | 10/41 | 6/41 | 13/2/0 | 15 | 3.11 | 0 | 0 | 11 | 1.64 |
| p2 | both8 | 9/43 | 7/43 | 9/1/1 | 11 | 3.17 | 1 | 0 | 9 | 1.58 |
| scotus | fixed | 27/47 | 15/47 | 14/49/21 | 84 | 9.47 | 21 | 4 | 30 | 3.34 |
| scotus | pause | 18/54 | 17/54 | 13/9/10 | 32 | 8.63 | 10 | 0 | 22 | 3.36 |
| scotus | overlap | 20/49 | 6/49 | 16/23/11 | 50 | 9.43 | 11 | 5 | 24 | 3.01 |
| scotus | both | 19/55 | 14/55 | 9/13/5 | 27 | 9.46 | 5 | 8 | 21 | 3.39 |
| scotus | overlap2 | 18/49 | 6/49 | 15/20/10 | 45 | 9.43 | 10 | 2 | 21 | 3.01 |
| scotus | both2 | 15/55 | 14/55 | 9/7/5 | 21 | 9.46 | 5 | 2 | 15 | 3.39 |
| scotus | both8 | 14/61 | 19/61 | 14/31/10 | 55 | 8.19 | 10 | 0 | 19 | 3.30 |
| wilde | fixed | 25/39 | 4/39 | 9/30/12 | 51 | 6.91 | 12 | 1 | 27 | 1.03 |
| wilde | pause | 6/43 | 5/43 | 4/0/16 | 20 | 6.95 | 16 | 0 | 7 | 1.23 |
| wilde | overlap | 6/40 | 5/40 | 2/3/10 | 15 | 6.74 | 10 | 3 | 7 | 1.13 |
| wilde | both | 11/44 | 5/44 | 4/8/12 | 24 | 9.43 | 12 | 7 | 13 | 1.43 |
| wilde | overlap2 | 4/40 | 5/40 | 3/0/10 | 13 | 6.74 | 10 | 0 | 5 | 1.13 |
| wilde | both2 | 6/44 | 5/44 | 4/1/12 | 17 | 9.43 | 12 | 0 | 7 | 1.43 |
| wilde | both8 | 7/48 | 5/48 | 5/0/23 | 28 | 6.80 | 23 | 0 | 8 | 1.04 |
### Where the remaining near-cut errors sit
Signed distance from the cut, pooled over files: upstream's errors pile up
within ±0.5 s (chopped words; at 100 s, 24 insertions in [−0.5, 0) alone).
Pause-only leaves deletions 1–3 s **before** its cuts, where the left chunk ends
(8/9, 3/4, 7/8 in [−3, −1) at the three lengths). overlap2's residue is spread
evenly over ±3 s, which is what background looks like. both2 and both8 are clean
at the handover but keep some of pause-only's pre-cut deletions.
### Controls
- **A-vs-A:** the same variant twice gives byte-identical words, segments and
text (fixed-120, both-120, both2-120 on all 4 files, in one process), and the
shipped script run three times as separate CLI processes matches the harness
output byte for byte (p1 and scotus). Decoding is deterministic, so run-to-run
variance is zero; the variance that matters comes from cut placement.
- **Harness validity:** patched code with overlap and pause search off equals
upstream's unmodified script on all 4 files; v2 without overlap equals v1.
- **Positive control:** upstream's fixed cutter damages 52 % of its cuts
against a 19 % background (+0.33, about 7 standard errors), per file 47–64 %
against 11–25 %. The instrument sees cut damage.
- **Negative control:** the phantom (background) rate is 18–21 % for every
variant, and error events away from cuts run at 2.0–2.35 per minute for all
of them. **Word counts away from cuts do not agree across slicers** (see
"Parakeet drops stretches" below), which is why they are not the decision metric.
- **Sensitivity floor:** at ~180–220 cuts per variant the 2-standard-error band
on a difference of damaged-cut rates is **±0.08** pooled and **±0.17** for
one file. No slicer can be measured below the background (~18–21 %). Word-count
differences under ~60 words are unresolvable, because a single dropped stretch
(next section) is 10–60 words.
### Verdict
Every variant that overlaps with the v2 handover, and pause-only, beats
upstream by 0.26–0.30 in damaged-cut rate, far outside the ±0.08 floor. Among
them the differences are **inside** the floor. overlap2 ships: tied best on
damaged cuts, fewest near-cut error events (47 vs 102), near-zero duplicates,
nothing concentrated at the handover, and the simplest mechanism. Pause search
measured neutral once the stitch was fixed, so it stays in the patch as an
opt-in `--pause-search`, off by default.
## Memory (GPU 3, production invocation, nvidia-smi every 0.2 s)
| run | n | per-process peak |
|---|---|---|
| upstream script, 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB (reproduces the 2026-09-30 budget figure) |
| shipped script (v2), 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB |
| shipped script (v3, final), 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB; all 3 CLI outputs byte-identical to the harness |
| shipped script, 120 s, scotus (public) | 3 | 5,496 / 5,496 / 5,496 MiB |
| positive control: shipped script, **300 s**, scotus | 1 | 7,056 MiB (deterministic; 300 s was n=3 in the earlier table) |
Zero spread; the peak is set by the 120 s maximum, and the overlap sits inside
it. The scotus file reads the same peak as p1, so the rebuild script uses it
as its public memory fixture. A spike shorter than the 0.2 s sample period
could be missed.
## Parakeet drops stretches of speech (separate finding, not the slicer)
Every chunked variant, **upstream's included**, sometimes skips a run of ≥10
consecutive words in the middle of a chunk, and which runs it skips changes
chaotically with the cut placement. Wilde at 120 s loses nothing, at 110 s
171 words, and overlap2 at 110 s loses one 60-second stretch. Across 4 files × 3
placements:
| variant | runs of ≥10 words lost | words lost |
|---|---|---|
| fixed (upstream) | 15 | 599 |
| pause | 17 | 511 |
| overlap2 | 12 | 656 |
| both2 | 14 | 719 |
| both8 | 12 | 560 |
The whole-file local-attention reference does the same: runs where every
slicer has words the reference lacks (13–16 per 12 comparisons, 340–400 words).
Today's production setting (upstream, 120 s) lost 85 words in 2 runs on p2. The
slicer neither causes nor cures it; it needs its own investigation (decoder
settings, chunk length, or model), which was out of scope here.