Deployed 2026-09-30 1211 PT by pointing SCRIBERR_IMAGE at the patched tag (.env backed up as .env.bak-20260930-pre-slicer1; rollback is the unpatched scriberr:local-blackwell). PrepareEnvironment rewrote the env's parakeet_transcribe_buffered.py from the embed (sha256 matches the patched source). One live run on GPU 1 beside intern-decision peaked at 5,496 MiB. Memory records the open Parakeet mid-chunk dropout finding and the held upstream PR.
222 lines
13 KiB
Markdown
222 lines
13 KiB
Markdown
# Scriberr Parakeet slicer bench (2026-09-30)
|
||
|
||
Prime's ruling, 2026-09-30: "build the slicer". This document records how the
|
||
pause-aware slicer patch (`stacks/scriberr/patches/0001-parakeet-pause-aware-slicer.patch`)
|
||
was measured and why the shipped variant was chosen. **Metrics only**: two of
|
||
the recordings are Prime's and private, so no transcript text appears here or
|
||
anywhere in git. Those recordings, and every transcript made from them, stay on
|
||
fv-ml1 in `/tank/spikes/scriberr-slicer/private/` (mode 700).
|
||
|
||
## Recordings (4 files, 118 minutes)
|
||
|
||
All four were converted to 16 kHz mono WAV with the image's ffmpeg, which is
|
||
what Scriberr feeds Parakeet.
|
||
|
||
| id | length | what | provenance / licence |
|
||
|---|---|---|---|
|
||
| p1 | 35.3 min | Prime's upload, conversational | private |
|
||
| p2 | 22.3 min | Prime's upload, conversational | private |
|
||
| scotus | 30.0 min (first half hour) | U.S. Supreme Court oral argument, *Loper Bright Enterprises v. Raimondo*, No. 22-451, argued 2024-01-17; spontaneous multi-speaker speech with interruptions | `https://www.supremecourt.gov/media/audio/mp3files/22-451.mp3` (sha256 `7e704b1f…5dc2`); U.S. Government work, public domain (17 U.S.C. §105) |
|
||
| wilde | 24.4 min | LibriVox *The Trial of Oscar Wilde* (dramatic reading), section 1; several readers, courtroom dialogue | `https://archive.org/download/trialofoscarwilde_1601_librivox/trialofoscarwilde_01_anon_64kb.mp3` (sha256 `492f4c56…3420`); public domain (LibriVox, PD Mark 1.0) |
|
||
|
||
## Harness
|
||
|
||
- **Image and env:** `scriberr:local-blackwell` (the live image), the live
|
||
`/tank/scriberr/whisperx-env` mounted **read-only**, `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`,
|
||
invoked exactly as Scriberr does (`uv run --native-tls --project /app/whisperx-env/parakeet python …`).
|
||
It runs as uid 1002 with `USER` set, not as appuser 10001; that changes file
|
||
ownership only.
|
||
- **GPU:** transient containers on **GPU 3 only** (verified: the container sees
|
||
one card, UUID `GPU-186dacf4…`). GPU 3 was at 2 MiB before and after.
|
||
- **Reference:** upstream's standard script on the whole file in one pass with
|
||
local attention (`--context-left 255 --context-right 255`), which has no cuts.
|
||
It needs >16 GB, which is why it cannot run on GPU 1.
|
||
- **Variants** run through the real `transcribe_buffered()`; the only harness
|
||
change is that the model is loaded once per process instead of once per call.
|
||
The CLI memory runs below reproduce the harness output byte for byte.
|
||
- **Placements:** every variant at three maximum lengths, 120, 110 and 100 s.
|
||
Parakeet is deterministic, so a repeat run cannot supply variance; moving the
|
||
cuts can. Each (file, variant) therefore has n = 3 different cut placements.
|
||
|
||
## Metric
|
||
|
||
Each variant's words are aligned against the reference, after lowercasing and
|
||
stripping punctuation (difflib opcodes, then exact Levenshtein inside each
|
||
non-matching block). Each error gets a time: the hypothesis word's start for a
|
||
substitution or insertion, the reference word's start for a deletion.
|
||
|
||
- **near-cut / elsewhere word errors (S/I/D):** near-cut means within ±3 s of
|
||
any of that variant's cuts (for overlap variants, the cut is the stitch point
|
||
at the middle of the overlap, and both chunk edges lie within ±3 s of it).
|
||
- **dropped / duplicated at cuts:** deletions near a cut, and insertions near a
|
||
cut that repeat an adjacent word.
|
||
- **error events:** errors clustered with gaps ≤ 1 s; rate per minute
|
||
elsewhere.
|
||
- **damaged cuts (the decision metric):** cuts with at least one error within
|
||
±3 s, against **damaged phantoms**, the same test at points midway between
|
||
the variant's own cuts (same count, same spacing, as far from any cut as the
|
||
audio gets). The phantom rate is the background any slicer's cuts are judged
|
||
against.
|
||
|
||
### Why the decision metric is not raw word counts
|
||
|
||
The negative control caught it. Word errors elsewhere varied by up to ±50 %
|
||
between slicers of the same length (p1: 131 to 215). Most of that comes from a
|
||
few **unstable stretches**, the same stretches for every variant (p1: 976–992 s,
|
||
1291–1298 s, 928–934 s), where the reference and any slicing disagree by 20–60
|
||
words depending on context, at arbitrary distances from any cut. That is heavy-tailed
|
||
noise, not cut damage. Counting events, and asking per cut whether anything near
|
||
it went wrong, is robust to it; the word counts are still reported.
|
||
|
||
## Variants
|
||
|
||
| name | overlap | pause search | stitch |
|
||
|---|---|---|---|
|
||
| fixed | 0 | off | none: upstream's slicer (patched code with both off reproduces it byte for byte) |
|
||
| pause | 0 | 25 s | none needed |
|
||
| overlap | 4 s | off | v1: each word kept by the chunk whose cut-to-cut range holds its start time (the brief's rule) |
|
||
| both | 4 s | 25 s | v1 |
|
||
| **overlap2** | 4 s | off | **v2: hand over at the nearest word both chunks transcribed alike, within 0.5 s** (shipped as v3, identical output) |
|
||
| both2 | 4 s | 25 s | v2 |
|
||
| both8 | 8 s | 25 s | v2 |
|
||
|
||
**v3** is v2 after code review, and it is what ships. Anchors are paired by
|
||
time first (same text *and* within 0.5 s, so a longer text match elsewhere in
|
||
the overlap cannot crowd out the true anchor), punctuation alone never anchors,
|
||
and the anchor's copy comes from whichever chunk keeps word starts in time order.
|
||
Re-run on all 12 overlap runs (4 files × 3 placements), **v3's output is
|
||
byte-identical to v2's** (and "both3" to "both2"): the reviewer's cases did not
|
||
occur in this audio, so every v2 number below holds for the shipped code.
|
||
|
||
v2 exists because v1 failed its own goal. At 26 of 108 "both" stitches the same
|
||
word sat on both sides of the cut, the two copies' start times 0.00–0.09 s apart
|
||
(one encoder frame): a word that follows a pause is timestamped anywhere in the
|
||
pause, and a pause is exactly where a pause-aware cut lands. Start time at the
|
||
midpoint is the worst possible stitch rule for pause cuts.
|
||
|
||
## Results
|
||
|
||
### Pooled (4 files × 3 placements)
|
||
|
||
| variant | damaged cuts | damaged phantoms | excess | by length 120 / 110 / 100 | near-cut words | dropped | duplicated | near-cut events | events elsewhere /min |
|
||
|---|---|---|---|---|---|---|---|---|---|
|
||
| fixed (upstream) | 93/179 = 52 % | 34/179 = 19 % | +0.33 | 0.48 / 0.52 / 0.55 | 223 | 56 | 17 | 102 | 2.08 |
|
||
| pause | 53/201 = 26 % | 36/201 = 18 % | +0.085 | 0.33 / 0.20 / 0.27 | 113 | 52 | 0 | 59 | 2.23 |
|
||
| overlap (v1) | 52/184 = 28 % | 33/184 = 18 % | +0.10 | 0.30 / 0.28 / 0.27 | 117 | 25 | 18 | 60 | 2.00 |
|
||
| both (v1) | 75/207 = 36 % | 38/207 = 18 % | +0.18 | 0.39 / 0.35 / 0.36 | 126 | 21 | 40 | 87 | 2.35 |
|
||
| **overlap2** | **41/184 = 22 %** | 33/184 = 18 % | **+0.04** | 0.21 / 0.22 / 0.24 | 101 | 24 | 2 | **47** | 2.00 |
|
||
| both2 | 52/207 = 25 % | 38/207 = 18 % | +0.07 | 0.32 / 0.20 / 0.24 | 89 | 22 | 2 | 54 | 2.35 |
|
||
| both8 | 48/222 = 22 % | 46/222 = 21 % | +0.01 | 0.23 / 0.23 / 0.19 | 120 | 42 | 0 | 54 | 2.18 |
|
||
|
||
### Per file (summed over the 3 placements)
|
||
|
||
| file | variant | damaged cuts | phantom | near S/I/D | near words | else words/min | drop | dup | events near | events else/min |
|
||
|---|---|---|---|---|---|---|---|---|---|---|
|
||
| p1 | fixed | 27/57 | 11/57 | 21/27/9 | 57 | 4.16 | 9 | 10 | 29 | 2.15 |
|
||
| p1 | pause | 19/64 | 8/64 | 20/3/25 | 48 | 5.62 | 25 | 0 | 20 | 2.39 |
|
||
| p1 | overlap | 20/59 | 16/59 | 16/26/4 | 46 | 4.66 | 4 | 8 | 23 | 2.07 |
|
||
| p1 | both | 29/67 | 13/67 | 23/25/4 | 52 | 5.70 | 4 | 17 | 35 | 2.56 |
|
||
| p1 | overlap2 | 15/59 | 16/59 | 16/19/4 | 39 | 4.66 | 4 | 0 | 17 | 2.07 |
|
||
| p1 | both2 | 21/67 | 13/67 | 23/8/5 | 36 | 5.70 | 5 | 0 | 21 | 2.56 |
|
||
| p1 | both8 | 18/70 | 15/70 | 16/2/8 | 26 | 6.22 | 8 | 0 | 18 | 2.40 |
|
||
| p2 | fixed | 14/36 | 4/36 | 10/7/14 | 31 | 4.45 | 14 | 2 | 16 | 1.39 |
|
||
| p2 | pause | 10/40 | 6/40 | 11/1/1 | 13 | 2.71 | 1 | 0 | 10 | 1.56 |
|
||
| p2 | overlap | 6/36 | 6/36 | 4/2/0 | 6 | 3.56 | 0 | 2 | 6 | 1.47 |
|
||
| p2 | both | 16/41 | 6/41 | 13/10/0 | 23 | 3.11 | 0 | 8 | 18 | 1.64 |
|
||
| p2 | overlap2 | 4/36 | 6/36 | 4/0/0 | 4 | 3.56 | 0 | 0 | 4 | 1.47 |
|
||
| p2 | both2 | 10/41 | 6/41 | 13/2/0 | 15 | 3.11 | 0 | 0 | 11 | 1.64 |
|
||
| p2 | both8 | 9/43 | 7/43 | 9/1/1 | 11 | 3.17 | 1 | 0 | 9 | 1.58 |
|
||
| scotus | fixed | 27/47 | 15/47 | 14/49/21 | 84 | 9.47 | 21 | 4 | 30 | 3.34 |
|
||
| scotus | pause | 18/54 | 17/54 | 13/9/10 | 32 | 8.63 | 10 | 0 | 22 | 3.36 |
|
||
| scotus | overlap | 20/49 | 6/49 | 16/23/11 | 50 | 9.43 | 11 | 5 | 24 | 3.01 |
|
||
| scotus | both | 19/55 | 14/55 | 9/13/5 | 27 | 9.46 | 5 | 8 | 21 | 3.39 |
|
||
| scotus | overlap2 | 18/49 | 6/49 | 15/20/10 | 45 | 9.43 | 10 | 2 | 21 | 3.01 |
|
||
| scotus | both2 | 15/55 | 14/55 | 9/7/5 | 21 | 9.46 | 5 | 2 | 15 | 3.39 |
|
||
| scotus | both8 | 14/61 | 19/61 | 14/31/10 | 55 | 8.19 | 10 | 0 | 19 | 3.30 |
|
||
| wilde | fixed | 25/39 | 4/39 | 9/30/12 | 51 | 6.91 | 12 | 1 | 27 | 1.03 |
|
||
| wilde | pause | 6/43 | 5/43 | 4/0/16 | 20 | 6.95 | 16 | 0 | 7 | 1.23 |
|
||
| wilde | overlap | 6/40 | 5/40 | 2/3/10 | 15 | 6.74 | 10 | 3 | 7 | 1.13 |
|
||
| wilde | both | 11/44 | 5/44 | 4/8/12 | 24 | 9.43 | 12 | 7 | 13 | 1.43 |
|
||
| wilde | overlap2 | 4/40 | 5/40 | 3/0/10 | 13 | 6.74 | 10 | 0 | 5 | 1.13 |
|
||
| wilde | both2 | 6/44 | 5/44 | 4/1/12 | 17 | 9.43 | 12 | 0 | 7 | 1.43 |
|
||
| wilde | both8 | 7/48 | 5/48 | 5/0/23 | 28 | 6.80 | 23 | 0 | 8 | 1.04 |
|
||
|
||
### Where the remaining near-cut errors sit
|
||
|
||
Signed distance from the cut, pooled over files: upstream's errors pile up
|
||
within ±0.5 s (chopped words; at 100 s, 24 insertions in [−0.5, 0) alone).
|
||
Pause-only leaves deletions 1–3 s **before** its cuts, where the left chunk ends
|
||
(8/9, 3/4, 7/8 in [−3, −1) at the three lengths). overlap2's residue is spread
|
||
evenly over ±3 s, which is what background looks like. both2 and both8 are clean
|
||
at the handover but keep some of pause-only's pre-cut deletions.
|
||
|
||
### Controls
|
||
|
||
- **A-vs-A:** the same variant twice gives byte-identical words, segments and
|
||
text (fixed-120, both-120, both2-120 on all 4 files, in one process), and the
|
||
shipped script run three times as separate CLI processes matches the harness
|
||
output byte for byte (p1 and scotus). Decoding is deterministic, so run-to-run
|
||
variance is zero; the variance that matters comes from cut placement.
|
||
- **Harness validity:** patched code with overlap and pause search off equals
|
||
upstream's unmodified script on all 4 files; v2 without overlap equals v1.
|
||
- **Positive control:** upstream's fixed cutter damages 52 % of its cuts
|
||
against a 19 % background (+0.33, about 7 standard errors), per file 47–64 %
|
||
against 11–25 %. The instrument sees cut damage.
|
||
- **Negative control:** the phantom (background) rate is 18–21 % for every
|
||
variant, and error events away from cuts run at 2.0–2.35 per minute for all
|
||
of them. **Word counts away from cuts do not agree across slicers** (see
|
||
"Parakeet drops stretches" below), which is why they are not the decision metric.
|
||
- **Sensitivity floor:** at ~180–220 cuts per variant the 2-standard-error band
|
||
on a difference of damaged-cut rates is **±0.08** pooled and **±0.17** for
|
||
one file. No slicer can be measured below the background (~18–21 %). Word-count
|
||
differences under ~60 words are unresolvable, because a single dropped stretch
|
||
(next section) is 10–60 words.
|
||
|
||
### Verdict
|
||
|
||
Every variant that overlaps with the v2 handover, and pause-only, beats
|
||
upstream by 0.26–0.30 in damaged-cut rate, far outside the ±0.08 floor. Among
|
||
them the differences are **inside** the floor. overlap2 ships: tied best on
|
||
damaged cuts, fewest near-cut error events (47 vs 102), near-zero duplicates,
|
||
nothing concentrated at the handover, and the simplest mechanism. Pause search
|
||
measured neutral once the stitch was fixed, so it stays in the patch as an
|
||
opt-in `--pause-search`, off by default.
|
||
|
||
## Memory (GPU 3, production invocation, nvidia-smi every 0.2 s)
|
||
|
||
| run | n | per-process peak |
|
||
|---|---|---|
|
||
| upstream script, 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB (reproduces the 2026-09-30 budget figure) |
|
||
| shipped script (v2), 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB |
|
||
| shipped script (v3, final), 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB; all 3 CLI outputs byte-identical to the harness |
|
||
| shipped script, 120 s, scotus (public) | 3 | 5,496 / 5,496 / 5,496 MiB |
|
||
| positive control: shipped script, **300 s**, scotus | 1 | 7,056 MiB (deterministic; 300 s was n=3 in the earlier table) |
|
||
| **live, GPU 1**, deployed container, `docker exec` as appuser, 120 s, p1 | 1 | **5,496 MiB** (2026-09-30 1212 PT, beside intern-decision; 49 s for 35 min of audio; seam check OK) |
|
||
|
||
Zero spread; the peak is set by the 120 s maximum, and the overlap sits inside
|
||
it. The scotus file reads the same peak as p1, so the rebuild script uses it
|
||
as its public memory fixture. A spike shorter than the 0.2 s sample period
|
||
could be missed.
|
||
|
||
## Parakeet drops stretches of speech (separate finding, not the slicer)
|
||
|
||
Every chunked variant, **upstream's included**, sometimes skips a run of ≥10
|
||
consecutive words in the middle of a chunk, and which runs it skips changes
|
||
chaotically with the cut placement. Wilde at 120 s loses nothing, at 110 s
|
||
171 words, and overlap2 at 110 s loses one 60-second stretch. Across 4 files × 3
|
||
placements:
|
||
|
||
| variant | runs of ≥10 words lost | words lost |
|
||
|---|---|---|
|
||
| fixed (upstream) | 15 | 599 |
|
||
| pause | 17 | 511 |
|
||
| overlap2 | 12 | 656 |
|
||
| both2 | 14 | 719 |
|
||
| both8 | 12 | 560 |
|
||
|
||
The whole-file local-attention reference does the same: runs where every
|
||
slicer has words the reference lacks (13–16 per 12 comparisons, 340–400 words).
|
||
Today's production setting (upstream, 120 s) lost 85 words in 2 runs on p2. The
|
||
slicer neither causes nor cures it; it needs its own investigation (decoder
|
||
settings, chunk length, or model), which was out of scope here.
|