feat(scriberr): overlap-and-stitch Parakeet slicer patch, rebuild script, bench

Carry patches/0001 on our Scriberr build (upstream a353078): adjacent
buffered chunks overlap by 4 s inside --chunk-len and hand over at a word
both chunks transcribed alike, instead of cutting at fixed marks with no
overlap. Pause-aware cutting is included as an opt-in (--pause-search);
it measured neutral once the stitch was right. The Go<->Python CLI and
JSON seam is unchanged.

Bench (4 recordings, 118 min, 3 cut placements each, against a no-cut
whole-file reference; metrics only, private audio stays on fv-ml1):
cuts with an error within +-3 s fall from 52% (93/179) to 22% (41/184)
against a 19% background; floor +-0.08. Positive control: upstream's
cutter +0.33 over background. A-vs-A byte-identical in-process and
across CLI processes. Peak GPU memory unchanged at 5,496 MiB (n=3).
Also found: Parakeet skips runs of >=10 words mid-chunk with any
slicer, upstream's included; not addressed here.

scripts/scriberr-rebuild clones a pinned upstream sha into a new
/opt/docker/src dir, git-apply-checks the patches, builds a distinct
tag, and checks embed, unit tests, the JSON seam (scriberr-seam-check.py)
and the memory budget on idle GPU 3. Deploy stays manual. The upstream
PR is prepared under patches/upstream-pr/ and not opened.
This commit is contained in:
vh
2026-09-30 12:10:32 -07:00
parent ab62644315
commit ee3db68db1
8 changed files with 1556 additions and 7 deletions
@@ -0,0 +1,220 @@
# Scriberr Parakeet slicer bench (2026-09-30)
Prime's ruling, 2026-09-30: "build the slicer". This document records how the
pause-aware slicer patch (`stacks/scriberr/patches/0001-parakeet-pause-aware-slicer.patch`)
was measured and why the shipped variant was chosen. **Metrics only**: two of
the recordings are Prime's and private, so no transcript text appears here or
anywhere in git. Those recordings, and every transcript made from them, stay on
fv-ml1 in `/tank/spikes/scriberr-slicer/private/` (mode 700).
## Recordings (4 files, 118 minutes)
All four were converted to 16 kHz mono WAV with the image's ffmpeg, which is
what Scriberr feeds Parakeet.
| id | length | what | provenance / licence |
|---|---|---|---|
| p1 | 35.3 min | Prime's upload, conversational | private |
| p2 | 22.3 min | Prime's upload, conversational | private |
| scotus | 30.0 min (first half hour) | U.S. Supreme Court oral argument, *Loper Bright Enterprises v. Raimondo*, No. 22-451, argued 2024-01-17; spontaneous multi-speaker speech with interruptions | `https://www.supremecourt.gov/media/audio/mp3files/22-451.mp3` (sha256 `7e704b1f…5dc2`); U.S. Government work, public domain (17 U.S.C. §105) |
| wilde | 24.4 min | LibriVox *The Trial of Oscar Wilde* (dramatic reading), section 1; several readers, courtroom dialogue | `https://archive.org/download/trialofoscarwilde_1601_librivox/trialofoscarwilde_01_anon_64kb.mp3` (sha256 `492f4c56…3420`); public domain (LibriVox, PD Mark 1.0) |
## Harness
- **Image and env:** `scriberr:local-blackwell` (the live image), the live
`/tank/scriberr/whisperx-env` mounted **read-only**, `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`,
invoked exactly as Scriberr does (`uv run --native-tls --project /app/whisperx-env/parakeet python …`).
It runs as uid 1002 with `USER` set, not as appuser 10001; that changes file
ownership only.
- **GPU:** transient containers on **GPU 3 only** (verified: the container sees
one card, UUID `GPU-186dacf4…`). GPU 3 was at 2 MiB before and after.
- **Reference:** upstream's standard script on the whole file in one pass with
local attention (`--context-left 255 --context-right 255`), which has no cuts.
It needs >16 GB, which is why it cannot run on GPU 1.
- **Variants** run through the real `transcribe_buffered()`; the only harness
change is that the model is loaded once per process instead of once per call.
The CLI memory runs below reproduce the harness output byte for byte.
- **Placements:** every variant at three maximum lengths, 120, 110 and 100 s.
Parakeet is deterministic, so a repeat run cannot supply variance; moving the
cuts can. Each (file, variant) therefore has n = 3 different cut placements.
## Metric
Each variant's words are aligned against the reference, after lowercasing and
stripping punctuation (difflib opcodes, then exact Levenshtein inside each
non-matching block). Each error gets a time: the hypothesis word's start for a
substitution or insertion, the reference word's start for a deletion.
- **near-cut / elsewhere word errors (S/I/D):** near-cut means within ±3 s of
any of that variant's cuts (for overlap variants, the cut is the stitch point
at the middle of the overlap, and both chunk edges lie within ±3 s of it).
- **dropped / duplicated at cuts:** deletions near a cut, and insertions near a
cut that repeat an adjacent word.
- **error events:** errors clustered with gaps ≤ 1 s; rate per minute
elsewhere.
- **damaged cuts (the decision metric):** cuts with at least one error within
±3 s, against **damaged phantoms**, the same test at points midway between
the variant's own cuts (same count, same spacing, as far from any cut as the
audio gets). The phantom rate is the background any slicer's cuts are judged
against.
### Why the decision metric is not raw word counts
The negative control caught it. Word errors elsewhere varied by up to ±50 %
between slicers of the same length (p1: 131 to 215). Most of that comes from a
few **unstable stretches**, the same stretches for every variant (p1: 976–992 s,
1291–1298 s, 928–934 s), where the reference and any slicing disagree by 20–60
words depending on context, at arbitrary distances from any cut. That is heavy-tailed
noise, not cut damage. Counting events, and asking per cut whether anything near
it went wrong, is robust to it; the word counts are still reported.
## Variants
| name | overlap | pause search | stitch |
|---|---|---|---|
| fixed | 0 | off | none: upstream's slicer (patched code with both off reproduces it byte for byte) |
| pause | 0 | 25 s | none needed |
| overlap | 4 s | off | v1: each word kept by the chunk whose cut-to-cut range holds its start time (the brief's rule) |
| both | 4 s | 25 s | v1 |
| **overlap2** | 4 s | off | **v2: hand over at the nearest word both chunks transcribed alike, within 0.5 s** (shipped as v3, identical output) |
| both2 | 4 s | 25 s | v2 |
| both8 | 8 s | 25 s | v2 |
**v3** is v2 after code review, and it is what ships. Anchors are paired by
time first (same text *and* within 0.5 s, so a longer text match elsewhere in
the overlap cannot crowd out the true anchor), punctuation alone never anchors,
and the anchor's copy comes from whichever chunk keeps word starts in time order.
Re-run on all 12 overlap runs (4 files × 3 placements), **v3's output is
byte-identical to v2's** (and "both3" to "both2"): the reviewer's cases did not
occur in this audio, so every v2 number below holds for the shipped code.
v2 exists because v1 failed its own goal. At 26 of 108 "both" stitches the same
word sat on both sides of the cut, the two copies' start times 0.00–0.09 s apart
(one encoder frame): a word that follows a pause is timestamped anywhere in the
pause, and a pause is exactly where a pause-aware cut lands. Start time at the
midpoint is the worst possible stitch rule for pause cuts.
## Results
### Pooled (4 files × 3 placements)
| variant | damaged cuts | damaged phantoms | excess | by length 120 / 110 / 100 | near-cut words | dropped | duplicated | near-cut events | events elsewhere /min |
|---|---|---|---|---|---|---|---|---|---|
| fixed (upstream) | 93/179 = 52 % | 34/179 = 19 % | +0.33 | 0.48 / 0.52 / 0.55 | 223 | 56 | 17 | 102 | 2.08 |
| pause | 53/201 = 26 % | 36/201 = 18 % | +0.085 | 0.33 / 0.20 / 0.27 | 113 | 52 | 0 | 59 | 2.23 |
| overlap (v1) | 52/184 = 28 % | 33/184 = 18 % | +0.10 | 0.30 / 0.28 / 0.27 | 117 | 25 | 18 | 60 | 2.00 |
| both (v1) | 75/207 = 36 % | 38/207 = 18 % | +0.18 | 0.39 / 0.35 / 0.36 | 126 | 21 | 40 | 87 | 2.35 |
| **overlap2** | **41/184 = 22 %** | 33/184 = 18 % | **+0.04** | 0.21 / 0.22 / 0.24 | 101 | 24 | 2 | **47** | 2.00 |
| both2 | 52/207 = 25 % | 38/207 = 18 % | +0.07 | 0.32 / 0.20 / 0.24 | 89 | 22 | 2 | 54 | 2.35 |
| both8 | 48/222 = 22 % | 46/222 = 21 % | +0.01 | 0.23 / 0.23 / 0.19 | 120 | 42 | 0 | 54 | 2.18 |
### Per file (summed over the 3 placements)
| file | variant | damaged cuts | phantom | near S/I/D | near words | else words/min | drop | dup | events near | events else/min |
|---|---|---|---|---|---|---|---|---|---|---|
| p1 | fixed | 27/57 | 11/57 | 21/27/9 | 57 | 4.16 | 9 | 10 | 29 | 2.15 |
| p1 | pause | 19/64 | 8/64 | 20/3/25 | 48 | 5.62 | 25 | 0 | 20 | 2.39 |
| p1 | overlap | 20/59 | 16/59 | 16/26/4 | 46 | 4.66 | 4 | 8 | 23 | 2.07 |
| p1 | both | 29/67 | 13/67 | 23/25/4 | 52 | 5.70 | 4 | 17 | 35 | 2.56 |
| p1 | overlap2 | 15/59 | 16/59 | 16/19/4 | 39 | 4.66 | 4 | 0 | 17 | 2.07 |
| p1 | both2 | 21/67 | 13/67 | 23/8/5 | 36 | 5.70 | 5 | 0 | 21 | 2.56 |
| p1 | both8 | 18/70 | 15/70 | 16/2/8 | 26 | 6.22 | 8 | 0 | 18 | 2.40 |
| p2 | fixed | 14/36 | 4/36 | 10/7/14 | 31 | 4.45 | 14 | 2 | 16 | 1.39 |
| p2 | pause | 10/40 | 6/40 | 11/1/1 | 13 | 2.71 | 1 | 0 | 10 | 1.56 |
| p2 | overlap | 6/36 | 6/36 | 4/2/0 | 6 | 3.56 | 0 | 2 | 6 | 1.47 |
| p2 | both | 16/41 | 6/41 | 13/10/0 | 23 | 3.11 | 0 | 8 | 18 | 1.64 |
| p2 | overlap2 | 4/36 | 6/36 | 4/0/0 | 4 | 3.56 | 0 | 0 | 4 | 1.47 |
| p2 | both2 | 10/41 | 6/41 | 13/2/0 | 15 | 3.11 | 0 | 0 | 11 | 1.64 |
| p2 | both8 | 9/43 | 7/43 | 9/1/1 | 11 | 3.17 | 1 | 0 | 9 | 1.58 |
| scotus | fixed | 27/47 | 15/47 | 14/49/21 | 84 | 9.47 | 21 | 4 | 30 | 3.34 |
| scotus | pause | 18/54 | 17/54 | 13/9/10 | 32 | 8.63 | 10 | 0 | 22 | 3.36 |
| scotus | overlap | 20/49 | 6/49 | 16/23/11 | 50 | 9.43 | 11 | 5 | 24 | 3.01 |
| scotus | both | 19/55 | 14/55 | 9/13/5 | 27 | 9.46 | 5 | 8 | 21 | 3.39 |
| scotus | overlap2 | 18/49 | 6/49 | 15/20/10 | 45 | 9.43 | 10 | 2 | 21 | 3.01 |
| scotus | both2 | 15/55 | 14/55 | 9/7/5 | 21 | 9.46 | 5 | 2 | 15 | 3.39 |
| scotus | both8 | 14/61 | 19/61 | 14/31/10 | 55 | 8.19 | 10 | 0 | 19 | 3.30 |
| wilde | fixed | 25/39 | 4/39 | 9/30/12 | 51 | 6.91 | 12 | 1 | 27 | 1.03 |
| wilde | pause | 6/43 | 5/43 | 4/0/16 | 20 | 6.95 | 16 | 0 | 7 | 1.23 |
| wilde | overlap | 6/40 | 5/40 | 2/3/10 | 15 | 6.74 | 10 | 3 | 7 | 1.13 |
| wilde | both | 11/44 | 5/44 | 4/8/12 | 24 | 9.43 | 12 | 7 | 13 | 1.43 |
| wilde | overlap2 | 4/40 | 5/40 | 3/0/10 | 13 | 6.74 | 10 | 0 | 5 | 1.13 |
| wilde | both2 | 6/44 | 5/44 | 4/1/12 | 17 | 9.43 | 12 | 0 | 7 | 1.43 |
| wilde | both8 | 7/48 | 5/48 | 5/0/23 | 28 | 6.80 | 23 | 0 | 8 | 1.04 |
### Where the remaining near-cut errors sit
Signed distance from the cut, pooled over files: upstream's errors pile up
within ±0.5 s (chopped words; at 100 s, 24 insertions in [−0.5, 0) alone).
Pause-only leaves deletions 1–3 s **before** its cuts, where the left chunk ends
(8/9, 3/4, 7/8 in [−3, −1) at the three lengths). overlap2's residue is spread
evenly over ±3 s, which is what background looks like. both2 and both8 are clean
at the handover but keep some of pause-only's pre-cut deletions.
### Controls
- **A-vs-A:** the same variant twice gives byte-identical words, segments and
text (fixed-120, both-120, both2-120 on all 4 files, in one process), and the
shipped script run three times as separate CLI processes matches the harness
output byte for byte (p1 and scotus). Decoding is deterministic, so run-to-run
variance is zero; the variance that matters comes from cut placement.
- **Harness validity:** patched code with overlap and pause search off equals
upstream's unmodified script on all 4 files; v2 without overlap equals v1.
- **Positive control:** upstream's fixed cutter damages 52 % of its cuts
against a 19 % background (+0.33, about 7 standard errors), per file 47–64 %
against 11–25 %. The instrument sees cut damage.
- **Negative control:** the phantom (background) rate is 18–21 % for every
variant, and error events away from cuts run at 2.0–2.35 per minute for all
of them. **Word counts away from cuts do not agree across slicers** (see
"Parakeet drops stretches" below), which is why they are not the decision metric.
- **Sensitivity floor:** at ~180–220 cuts per variant the 2-standard-error band
on a difference of damaged-cut rates is **±0.08** pooled and **±0.17** for
one file. No slicer can be measured below the background (~18–21 %). Word-count
differences under ~60 words are unresolvable, because a single dropped stretch
(next section) is 10–60 words.
### Verdict
Every variant that overlaps with the v2 handover, and pause-only, beats
upstream by 0.26–0.30 in damaged-cut rate, far outside the ±0.08 floor. Among
them the differences are **inside** the floor. overlap2 ships: tied best on
damaged cuts, fewest near-cut error events (47 vs 102), near-zero duplicates,
nothing concentrated at the handover, and the simplest mechanism. Pause search
measured neutral once the stitch was fixed, so it stays in the patch as an
opt-in `--pause-search`, off by default.
## Memory (GPU 3, production invocation, nvidia-smi every 0.2 s)
| run | n | per-process peak |
|---|---|---|
| upstream script, 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB (reproduces the 2026-09-30 budget figure) |
| shipped script (v2), 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB |
| shipped script (v3, final), 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB; all 3 CLI outputs byte-identical to the harness |
| shipped script, 120 s, scotus (public) | 3 | 5,496 / 5,496 / 5,496 MiB |
| positive control: shipped script, **300 s**, scotus | 1 | 7,056 MiB (deterministic; 300 s was n=3 in the earlier table) |
Zero spread; the peak is set by the 120 s maximum, and the overlap sits inside
it. The scotus file reads the same peak as p1, so the rebuild script uses it
as its public memory fixture. A spike shorter than the 0.2 s sample period
could be missed.
## Parakeet drops stretches of speech (separate finding, not the slicer)
Every chunked variant, **upstream's included**, sometimes skips a run of ≥10
consecutive words in the middle of a chunk, and which runs it skips changes
chaotically with the cut placement. Wilde at 120 s loses nothing, at 110 s
171 words, and overlap2 at 110 s loses one 60-second stretch. Across 4 files × 3
placements:
| variant | runs of ≥10 words lost | words lost |
|---|---|---|
| fixed (upstream) | 15 | 599 |
| pause | 17 | 511 |
| overlap2 | 12 | 656 |
| both2 | 14 | 719 |
| both8 | 12 | 560 |
The whole-file local-attention reference does the same: runs where every
slicer has words the reference lacks (13–16 per 12 comparisons, 340–400 words).
Today's production setting (upstream, 120 s) lost 85 words in 2 runs on p2. The
slicer neither causes nor cures it; it needs its own investigation (decoder
settings, chunk length, or model), which was out of scope here.
+263
View File
@@ -0,0 +1,263 @@
#!/usr/bin/env bash
# scriberr-rebuild — rebuild Scriberr's Blackwell image at a PINNED upstream sha
# with our patches applied, then prove the result before anyone deploys it.
#
# Runs on nh3-dev and drives fv-ml1 over ssh. Deploy is a SEPARATE manual step
# (stacks/scriberr/patches/README.md § Deploy); this script never touches the
# live container, its .env, or GPU 1.
#
# Why this exists: upstream publishes no sm_120 image, so we build from source,
# and we carry a patch to the Parakeet slicer (stacks/scriberr/patches/). Upstream
# moves slowly, so an upgrade should be one command plus a verdict.
#
# Stages (each prints PASS/FAIL; the first FAIL stops the run):
# clone clean shallow clone of upstream at the pinned sha, in a NEW dir
# /opt/docker/src/scriberr-<sha7>-<suffix> (reused only if it already
# holds that sha with every patch applied)
# patch `git apply --check` then `git apply`, patch by patch; a conflict
# stops the run loudly and names the patch
# build docker build -f Dockerfile.cuda.12.9 -t scriberr:local-blackwell-<sha7>-<suffix>
# — a DISTINCT tag, so the running image is never overwritten
# embed the patched script's exact bytes are inside the new Go binary
# (Scriberr rewrites the env's copy from this embed on every start)
# unit the slicer's pure-function tests, under the live env's numpy/librosa
# seam patched script, production invocation, short fixture at
# --chunk-len 10, JSON validated against the Go struct
# memory same on a long recording at --chunk-len 120 on an IDLE GPU,
# nvidia-smi sampled every 0.2 s; per-process peak <= the budget
#
# Usage:
# scripts/scriberr-rebuild [--sha SHA40] [--suffix NAME] [--gpu N]
# [--budget MIB] [--memory-audio PATH] [--reuse-image]
#
# Defaults: --sha PINNED_SHA below, --suffix slicer1, --gpu 3, --budget 5496,
# --memory-audio the public 30-min SCOTUS fixture. The GPU must be idle
# (< 100 MiB used), which in practice means GPU 3; GPUs 0-2 run live seats.
set -euo pipefail
PINNED_SHA=a353078fd96b8aca4002681813524b7397c90df1 # upstream HEAD 2026-09-20
UPSTREAM=https://github.com/rishikanthc/Scriberr.git
HOST=${SCRIBERR_REBUILD_HOST:-infra-ops@10.251.50.54}
ENV_DIR=/tank/scriberr/whisperx-env # live env, always mounted READ-ONLY
TOOLS=/opt/docker/src/scriberr-rebuild # fixtures + seam checker on fv-ml1
SCRIPT_REL=internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
TEST_REL=internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
SEAM_AUDIO_REL=tests/data/AMI-Corpus-IB4002.Mix-Headset-clip.wav
SHA=$PINNED_SHA SUFFIX=slicer1 GPU=3 BUDGET=5496 REUSE_IMAGE=0
MEM_AUDIO=$TOOLS/fixtures/scotus-22-451-first30m.wav
while [ $# -gt 0 ]; do
case $1 in
--sha) SHA=$2; shift 2 ;;
--suffix) SUFFIX=$2; shift 2 ;;
--gpu) GPU=$2; shift 2 ;;
--budget) BUDGET=$2; shift 2 ;;
--memory-audio) MEM_AUDIO=$2; shift 2 ;;
--reuse-image) REUSE_IMAGE=1; shift ;;
-h|--help) sed -n '2,/^set -euo/p' "$0" | sed '$d; s/^# \{0,1\}//'; exit 0 ;;
*) echo "unknown argument: $1 (see --help)" >&2; exit 2 ;;
esac
done
[[ $SHA =~ ^[0-9a-f]{40}$ ]] || { echo "--sha must be a full 40-char sha (GitHub fetches by full sha)" >&2; exit 2; }
[[ $SUFFIX =~ ^[a-z0-9][a-z0-9.-]*$ ]] || { echo "--suffix must be [a-z0-9.-]" >&2; exit 2; }
[[ $GPU =~ ^[0-9]+$ && $BUDGET =~ ^[0-9]+$ ]] || { echo "--gpu and --budget must be integers" >&2; exit 2; }
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
PATCH_DIR=$REPO_ROOT/stacks/scriberr/patches
mapfile -t PATCHES < <(find "$PATCH_DIR" -maxdepth 1 -name '*.patch' | sort)
[ ${#PATCHES[@]} -gt 0 ] || { echo "no patches in $PATCH_DIR" >&2; exit 2; }
# The only paths a reused build dir may differ from upstream in.
mapfile -t PATCHED_PATHS < <(sed -n 's#^+++ b/##p' "${PATCHES[@]}" | sort -u)
SHA7=${SHA:0:7}
TAG=scriberr:local-blackwell-$SHA7-$SUFFIX
BUILD_DIR=/opt/docker/src/scriberr-$SHA7-$SUFFIX
PATCH_SUM=$(cat "${PATCHES[@]}" | sha256sum | cut -c1-16)
CNAME=scriberr-rebuild-$SHA7-$SUFFIX # every GPU container, so cleanup can find it
SAMPLES=/tmp/$CNAME.mem.csv
BUILD_LOG=/tmp/$CNAME.build.log
MIN_FREE_GB=20 # under Docker's root dir (zroot); an image adds ~0.1-6 GB
# One multiplexed ssh connection serves the ~15 remote calls of a run.
CM=(-o BatchMode=yes -o ControlMaster=auto -o "ControlPath=$HOME/.ssh/cm-%C" -o ControlPersist=120)
SSH=(ssh "${CM[@]}" "$HOST")
SCP=(scp -q "${CM[@]}")
RESULTS=()
pass() { RESULTS+=("PASS $1 $2"); echo "== PASS $1: $2"; }
fail() {
RESULTS+=("FAIL $1 $2"); echo "== FAIL $1: $2" >&2
summary; exit 1
}
summary() {
echo; echo "scriberr-rebuild upstream=$SHA7 patches=$PATCH_SUM tag=$TAG"
printf ' %s\n' "${RESULTS[@]}"
}
record() { # best effort: the audit trail must not become a failure point
"$REPO_ROOT/scripts/ops-log" record --host fv-ml1 --action "$1" --target "$2" \
--detail "$3" >/dev/null 2>&1 || echo "(ops-log record failed; continuing)" >&2
}
remote() { "${SSH[@]}" bash -s -- "$@"; }
# Whatever happens (a FAIL, Ctrl-C, a dropped connection), never leave the
# nvidia-smi sampler or a transient GPU container behind on fv-ml1.
cleanup() {
"${SSH[@]}" "[ -s $SAMPLES.pid ] && kill \$(cat $SAMPLES.pid) 2>/dev/null; : >$SAMPLES.pid; \
docker rm -f $CNAME >/dev/null 2>&1; true" 2>/dev/null || true
}
trap cleanup EXIT
trap 'exit 130' INT TERM
# An unguarded remote call that fails must still say so and print the table.
set -E
trap 'echo "== ABORT: unexpected failure at line $LINENO (see output above)" >&2; summary' ERR
# Every container run: the new image, the live env READ-ONLY, the build tree
# read-only, and nothing else writable but the container's own /tmp.
DOCKER_RUN="docker run --rm --user 1002:1003 -e HOME=/tmp -e USER=infra-ops -e LOGNAME=infra-ops \
-e PYTHONDONTWRITEBYTECODE=1 -e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True -e UV_LINK_MODE=copy \
-v $ENV_DIR:/app/whisperx-env:ro -v $BUILD_DIR:/src:ro -v $TOOLS:/tools:ro --entrypoint bash"
UVRUN="uv run --native-tls --project /app/whisperx-env/parakeet"
# ── clone ──────────────────────────────────────────────────────────────────
if out=$(remote "$BUILD_DIR" "$UPSTREAM" "$SHA" "$TOOLS" "${PATCHED_PATHS[@]}" 2>&1 <<'EOF'
set -euo pipefail
dir=$1 upstream=$2 sha=$3 tools=$4
shift 4
sudo -n install -d -o infra-ops -g infra-ops "$tools" "$tools/fixtures" | cat
if [ -e "$dir" ]; then
[ "$(git -C "$dir" rev-parse HEAD 2>/dev/null)" = "$sha" ] \
|| { echo "$dir exists but is not a checkout of $sha; remove it by hand (sudo -n rm -rf $dir) or pick --suffix"; exit 1; }
extra=$({ git -C "$dir" diff --name-only HEAD; git -C "$dir" ls-files --others --exclude-standard; } \
| sort -u | grep -vxF -f <(printf '%s\n' "$@") || true)
[ -z "$extra" ] || { echo "$dir has changes outside the patches ($extra); remove it by hand or pick --suffix"; exit 1; }
echo "reusing $dir"
exit 0
fi
sudo -n install -d -o infra-ops -g infra-ops "$dir" | cat
cd "$dir"
git init -q
git remote add origin "$upstream"
git fetch -q --depth 1 origin "$sha" \
|| { echo "fetch of $sha failed; $dir is left empty, remove it by hand (sudo -n rm -rf $dir)"; exit 1; }
git checkout -q --detach FETCH_HEAD
[ "$(git rev-parse HEAD)" = "$sha" ] || { echo "checked out $(git rev-parse HEAD), wanted $sha"; exit 1; }
echo "cloned $sha into $dir"
EOF
); then pass clone "$out"; else fail clone "$out"; fi
if [[ $out == reusing* ]]; then REUSED=1; else REUSED=0; record create "$BUILD_DIR" "scriberr-rebuild: clean clone of upstream $SHA7"; fi
# ── patch ──────────────────────────────────────────────────────────────────
for p in "${PATCHES[@]}"; do
name=$(basename "$p")
# Already applied (a reused dir)? `apply --reverse --check` succeeds only then.
if "${SSH[@]}" "cd $BUILD_DIR && git apply --reverse --check -" <"$p" >/dev/null 2>&1; then
pass patch "$name already applied"
continue
fi
if ! out=$("${SSH[@]}" "cd $BUILD_DIR && git apply --check -" <"$p" 2>&1); then
if [ "$REUSED" = 1 ]; then
fail patch "$name does not apply to the REUSED $BUILD_DIR, which holds an older state of the patches; pick a new --suffix or remove the dir by hand (sudo -n rm -rf $BUILD_DIR):
$out"
fi
fail patch "$name DOES NOT APPLY to upstream $SHA7 — rebase the patch before upgrading:
$out"
fi
"${SSH[@]}" "cd $BUILD_DIR && git apply -" <"$p" || fail patch "$name: git apply failed after a clean --check"
record patch "$BUILD_DIR" "scriberr-rebuild: git apply $name"
pass patch "$name applied"
done
# ── build ──────────────────────────────────────────────────────────────────
if "${SSH[@]}" "docker image inspect $TAG >/dev/null 2>&1"; then
[ "$REUSE_IMAGE" = 1 ] || fail build "$TAG already exists; pass --reuse-image to test it, or pick a new --suffix"
pass build "reusing existing $TAG"
else
free=$("${SSH[@]}" "df -BG --output=avail \$(docker info -f '{{.DockerRootDir}}') | tail -1") \
|| fail build "could not read free space on fv-ml1"
free=${free//[!0-9]/}
[ "${free:-0}" -ge "$MIN_FREE_GB" ] \
|| fail build "only ${free:-?} GB free under Docker's root dir, need $MIN_FREE_GB; remove superseded scriberr tags first (patches/README.md)"
echo "building $TAG (several minutes; log on fv-ml1 at $BUILD_LOG)"
if "${SSH[@]}" "docker build -f $BUILD_DIR/Dockerfile.cuda.12.9 -t $TAG \
--label org.phasefinal.scriberr.upstream=$SHA --label org.phasefinal.scriberr.patches=$PATCH_SUM \
$BUILD_DIR >$BUILD_LOG 2>&1"; then
record build "$TAG" "scriberr-rebuild: upstream $SHA7 + patches $PATCH_SUM; old images kept"
pass build "$TAG ($free GB was free)"
else
"${SSH[@]}" "tail -25 $BUILD_LOG" >&2 || true
fail build "docker build failed (log tail above)"
fi
fi
# ── embed ──────────────────────────────────────────────────────────────────
if out=$(remote "$TAG" "$BUILD_DIR" "$SCRIPT_REL" 2>&1 <<'EOF'
docker run --rm -v "$2":/src:ro --entrypoint python3 "$1" -c "
import sys
script = open('/src/$3', 'rb').read()
sys.exit(0 if script in open('/app/scriberr', 'rb').read() else 1)"
EOF
); then
pass embed "patched $(basename "$SCRIPT_REL") is byte-identical inside /app/scriberr"
else
fail embed "the binary does not embed the patched script ${out:+($out)}"
fi
# ── unit ───────────────────────────────────────────────────────────────────
if out=$("${SSH[@]}" "$DOCKER_RUN $TAG -c 'cd /tmp && $UVRUN --with pytest \
python -m pytest -q -p no:cacheprovider /src/$TEST_REL 2>&1 | tail -3'" 2>&1) \
&& grep -q ' passed' <<<"$out" && ! grep -Eq 'failed|error' <<<"$out"; then
pass unit "$(tail -1 <<<"$out")"
else
fail unit "$out"
fi
# The checker travels with this script, so the host copy is refreshed each run.
"${SCP[@]}" "$REPO_ROOT/scripts/scriberr-seam-check.py" "$HOST:$TOOLS/seam-check.py" \
|| fail seam "could not copy the seam checker to fv-ml1:$TOOLS"
record update "$TOOLS/seam-check.py" "scriberr-rebuild: refreshed the seam checker"
# One GPU run of the patched script under the production invocation, output
# kept inside the container, validated there. Prints the seam checker's line.
gpu_run() { # $1 audio path on host, $2 --chunk-len, $3 --min-chunks
local audio_dir; audio_dir=$(dirname "$1")
"${SSH[@]}" "$DOCKER_RUN --name $CNAME --gpus '\"device=$GPU\"' -e NVIDIA_VISIBLE_DEVICES=$GPU \
-v $audio_dir:/audio:ro $TAG -c 'cd /tmp && $UVRUN python /src/$SCRIPT_REL /audio/$(basename "$1") \
--output /tmp/out.json --chunk-len $2 >/tmp/run.log 2>&1 || { tail -5 /tmp/run.log; exit 1; }; \
python3 /tools/seam-check.py /tmp/out.json --min-chunks $3'"
}
gpu_idle() {
local used
used=$("${SSH[@]}" "nvidia-smi -i $GPU --query-gpu=memory.used --format=csv,noheader,nounits") \
|| fail "$1" "could not read GPU $GPU memory on fv-ml1"
used=${used//[!0-9]/}
[ -n "$used" ] && [ "$used" -lt 100 ] \
|| fail "$1" "GPU $GPU is not idle (${used:-?} MiB used); refusing to share a live card"
}
# ── seam ───────────────────────────────────────────────────────────────────
gpu_idle seam
if out=$(gpu_run "$BUILD_DIR/$SEAM_AUDIO_REL" 10 2 2>&1); then pass seam "$(tail -1 <<<"$out")"
else fail seam "$out"; fi
# ── memory ─────────────────────────────────────────────────────────────────
"${SSH[@]}" "test -s $MEM_AUDIO" || fail memory "memory audio $MEM_AUDIO not found on fv-ml1"
gpu_idle memory
"${SSH[@]}" "nohup nvidia-smi -i $GPU --query-compute-apps=pid,used_memory \
--format=csv,noheader,nounits -lms 200 </dev/null >$SAMPLES 2>/dev/null & echo \$! >$SAMPLES.pid" \
|| fail memory "could not start the nvidia-smi sampler"
record run "$CNAME" "transient $TAG on GPU $GPU, --chunk-len 120, env ro; removed on exit"
rc=0; out=$(gpu_run "$MEM_AUDIO" 120 2 2>&1) || rc=$? # `||`, not set +e: keeps the ERR trap quiet
"${SSH[@]}" "kill \$(cat $SAMPLES.pid) 2>/dev/null; : >$SAMPLES.pid" || true
read -r pids peak < <("${SSH[@]}" \
"awk -F', *' 'NF==2 {if (!(\$1 in p)) {p[\$1]=1; n++}; if (\$2+0>m) m=\$2+0} END {print n+0, m+0}' $SAMPLES") \
|| fail memory "no GPU samples could be read back from $SAMPLES"
[ "$rc" = 0 ] || fail memory "memory run failed: $out"
[ "$pids" = 1 ] || fail memory "saw $pids processes on GPU $GPU during the run; the peak is not attributable"
if [ "$peak" -le "$BUDGET" ]; then
pass memory "peak $peak MiB <= budget $BUDGET MiB (0.2 s samples, GPU $GPU); $(tail -1 <<<"$out")"
else
fail memory "peak $peak MiB > budget $BUDGET MiB — do NOT deploy beside intern-decision"
fi
summary
echo
echo "VERDICT: PASS. Deploy is manual: stacks/scriberr/patches/README.md § Deploy (SCRIBERR_IMAGE=$TAG)."
+86
View File
@@ -0,0 +1,86 @@
#!/usr/bin/env python3
"""Validate a parakeet_transcribe_buffered.py result against the Go seam.
Scriberr's parakeet_adapter.go (parseResult) unmarshals this JSON into a struct
with typed fields; a float where Go expects an int, or a missing key, fails the
job. This checks the shape Go reads plus the stitching invariants the slicer
patch promises. Stdlib only, so it runs under any python3. Prints counts, never
transcript text.
usage: scriberr-seam-check.py RESULT.json [--min-chunks N]
"""
import argparse
import json
import sys
NUMBER = (int, float)
def fail(msg):
print(f"SEAM FAIL: {msg}")
sys.exit(1)
def check_items(items, text_key, name):
for i, item in enumerate(items):
if not isinstance(item, dict):
fail(f"{name}[{i}] is not an object")
if not isinstance(item.get(text_key), str):
fail(f"{name}[{i}].{text_key} is not a string")
for key in ("start_offset", "end_offset"):
if type(item.get(key)) is not int:
fail(f"{name}[{i}].{key} is not an integer (Go field is int)")
for key in ("start", "end"):
if not isinstance(item.get(key), NUMBER) or isinstance(item.get(key), bool):
fail(f"{name}[{i}].{key} is not a number")
if item["start"] > item["end"]:
fail(f"{name}[{i}] starts after it ends")
def main():
parser = argparse.ArgumentParser(description="Validate a buffered Parakeet result for Go.")
parser.add_argument("result", help="result JSON written by parakeet_transcribe_buffered.py")
parser.add_argument("--min-chunks", type=int, default=1,
help="fail unless the run used at least this many chunks")
args = parser.parse_args()
min_chunks = args.min_chunks
try:
data = json.load(open(args.result, encoding="utf-8"))
except (OSError, ValueError) as e:
fail(f"cannot read {args.result}: {e}")
if not isinstance(data, dict):
fail("the result is not a JSON object")
required = {"transcription": str, "language": str, "word_timestamps": list,
"segment_timestamps": list, "audio_file": str, "model": str}
for key, kind in required.items():
if not isinstance(data.get(key), kind):
fail(f"'{key}' missing or not {kind.__name__}")
if data.get("buffered") is not True:
fail("'buffered' is not true")
if not isinstance(data.get("chunk_duration_secs"), NUMBER):
fail("'chunk_duration_secs' is not a number")
if type(data.get("num_chunks")) is not int or data["num_chunks"] < min_chunks:
fail(f"'num_chunks' is not an integer >= {min_chunks}")
words, segments = data["word_timestamps"], data["segment_timestamps"]
if not words or not data["transcription"].strip():
fail("empty transcript")
check_items(words, "word", "word_timestamps")
check_items(segments, "segment", "segment_timestamps")
starts = [w["start"] for w in words]
if starts != sorted(starts):
fail("word start times go backwards (a stitch repeated or reordered words)")
joined = " ".join(w["word"] for w in words)
if data["transcription"] != joined:
fail("'transcription' is not the stitched words joined by spaces")
if " ".join(s["segment"] for s in segments) != joined:
fail("segments do not cover the stitched words exactly once, in order")
print(f"SEAM OK: {len(words)} words, {len(segments)} segments, "
f"{data['num_chunks']} chunks, cuts at {len(data.get('cut_times', []))} points")
if __name__ == "__main__":
main()
+5
View File
@@ -4,6 +4,11 @@
# ── Image ──────────────────────────────────────────────────────────────── # ── Image ────────────────────────────────────────────────────────────────
# Built locally from Dockerfile.cuda.12.9 — see the compose header for why # Built locally from Dockerfile.cuda.12.9 — see the compose header for why
# the published scriberr-cuda image is NOT usable on these Blackwell cards. # the published scriberr-cuda image is NOT usable on these Blackwell cards.
# Patched builds come from scripts/scriberr-rebuild and are tagged
# scriberr:local-blackwell-<upstream sha7>-<suffix> (e.g. -a353078-slicer1)
# so every build keeps its own tag. Deploying = pointing this at a new tag;
# rollback = pointing it back. See stacks/scriberr/patches/README.md.
# (Unset falls back to the original unpatched scriberr:local-blackwell.)
SCRIBERR_IMAGE=scriberr:local-blackwell SCRIBERR_IMAGE=scriberr:local-blackwell
# ── Network ────────────────────────────────────────────────────────────── # ── Network ──────────────────────────────────────────────────────────────
+21 -7
View File
@@ -27,15 +27,21 @@ image** — it will fail on these cards or quietly fall back to CPU.
### Rebuilding ### Rebuilding
We carry local patches (`patches/`, currently the pause-aware Parakeet
slicer), so a rebuild is one command from nh3-dev, pinned to an upstream sha:
```bash ```bash
ssh fv-ml1 scripts/scriberr-rebuild --sha <full upstream sha> --suffix slicer1
cd /tank/scriberr/src/Scriberr
git pull
docker build -f Dockerfile.cuda.12.9 -t scriberr:local-blackwell .
cd /opt/docker/compose/scriberr && docker compose up -d
``` ```
Source checkout lives on `/tank`, not the root pool — see storage below. It makes a clean clone in `/opt/docker/src/scriberr-<sha7>-<suffix>` on
fv-ml1, `git apply --check`s the patches (a conflict stops it), builds
`scriberr:local-blackwell-<sha7>-<suffix>` beside the old images, and checks
the embed, the unit tests, the Go↔Python JSON seam, and the GPU memory budget.
Deploying it is a separate manual step: `patches/README.md` § Deploy.
The old checkout at `/tank/scriberr/src/Scriberr` (lkraven-owned, shallow)
built the original `scriberr:local-blackwell` and is left as it was.
## Deploy ## Deploy
@@ -62,7 +68,7 @@ mounts are bind-mounted onto `/tank` (4+ TB) instead of named volumes:
|---|---|---| |---|---|---|
| `/tank/scriberr/data` | `/app/data` | SQLite DB, uploads, transcripts | | `/tank/scriberr/data` | `/app/data` | SQLite DB, uploads, transcripts |
| `/tank/scriberr/whisperx-env` | `/app/whisperx-env` | Python env + model weights | | `/tank/scriberr/whisperx-env` | `/app/whisperx-env` | Python env + model weights |
| `/tank/scriberr/src/Scriberr` | — | build checkout | | `/tank/scriberr/src/Scriberr` | — | original build checkout (patched builds: `/opt/docker/src/scriberr-<sha7>-<suffix>`) |
Both are owned by uid/gid 1000 to match `PUID`/`PGID`. Both are owned by uid/gid 1000 to match `PUID`/`PGID`.
@@ -164,3 +170,11 @@ Peak GPU memory on a 35-minute file:
- The two env knobs are read by upstream's Go code (`parakeet_adapter.go`), so - The two env knobs are read by upstream's Go code (`parakeet_adapter.go`), so
they survive image upgrades for as long as upstream keeps them. Re-measure the they survive image upgrades for as long as upstream keeps them. Re-measure the
peak after any upgrade. peak after any upgrade.
- **The slicer itself is patched** (`patches/0001-parakeet-pause-aware-slicer.patch`,
2026-09-30). Adjacent 120 s slices now overlap by 4 s and are stitched at a word
both transcribed, which cut the share of cuts with an error nearby from 52 % to
22 % against a 19 % background. The overlap sits *inside* the 120 s, so the
peak is unchanged (5,496 MiB, n=3). See `patches/README.md` and
`docs/pfi/scriberr-slicer-bench-2026-09-30.md`. That bench also found that
Parakeet sometimes skips stretches of ≥10 words mid-slice, with or without the
patch; that is still open.
@@ -0,0 +1,686 @@
From 2dafe7ce81c217609ff2c8616e43b0d72255e170 Mon Sep 17 00:00:00 2001
From: Vuong Hoang <vh@phasefinal.com>
Date: Wed, 30 Sep 2026 11:41:07 -0700
Subject: [PATCH] fix(parakeet): overlap buffered chunks and stitch at an
agreed word
parakeet_transcribe_buffered.py cut long audio at fixed --chunk-len marks
with no overlap, so a word straddling a mark was chopped, dropped or
transcribed twice. Adjacent chunks now overlap by --overlap seconds
(default 4, counted inside --chunk-len so no chunk grows), and in each
overlap the chunks hand over at the word nearest the cut that both
transcribed with the same text at nearly the same time (within 0.5 s),
keeping whichever copy leaves the words in time order. With no such word
they split at the cut. Splitting both chunks at the cut by word start time
is not enough on its own: a word after a pause can be timestamped anywhere
in the pause, so the two chunks may place it on opposite sides of the cut.
--pause-search N (opt-in) also moves each cut back to the quietest 0.3 s
in the last N seconds before the limit. Measured neutral on top of the
overlap, so it is off by default.
The CLI and JSON the Go adapter reads are unchanged; the new flags are
optional, and the JSON gains overlap_secs, pause_search_secs and
cut_times. --overlap 0 reproduces the previous output exactly. NeMo is
now imported inside transcribe_buffered() so the slicing and stitching
helpers can be unit-tested without a GPU. If a chunk ever returns text
without word timestamps, its text is kept rather than dropped.
---
.../py/nvidia/parakeet_transcribe_buffered.py | 221 ++++++++++--
.../py/nvidia/tests/test_parakeet_slicing.py | 329 ++++++++++++++++++
2 files changed, 525 insertions(+), 25 deletions(-)
create mode 100644 internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
diff --git a/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py b/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
index 29d5047..ba755c1 100644
--- a/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
+++ b/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
@@ -2,6 +2,10 @@
"""
NVIDIA Parakeet buffered inference for long audio files.
Splits audio into chunks to avoid GPU memory issues.
+
+Adjacent chunks overlap slightly, and in each overlap the chunks hand over at
+a word both transcribed alike, so a word near a cut is neither chopped, dropped
+nor repeated. Optionally (--pause-search) each cut also moves into a pause.
"""
import argparse
@@ -12,37 +16,178 @@ import librosa
import soundfile as sf
import numpy as np
from pathlib import Path
-import nemo.collections.asr as nemo_asr
+DEFAULT_OVERLAP_SECS = 4.0
+DEFAULT_PAUSE_SEARCH_SECS = 0.0 # opt-in; measured no gain on top of the overlap
+QUIET_WINDOW_SECS = 0.3
+SAME_WORD_SECS = 0.5
+
+
+def plan_slices(audio, sr, max_chunk_secs, overlap_secs=0.0, search_secs=0.0):
+ """Choose where to cut `audio` so no chunk exceeds `max_chunk_secs`.
-def split_audio_file(audio_path, chunk_duration_secs=300):
- """Split audio file into chunks of specified duration."""
+ Returns (spans, cuts): `cuts` are the sample indices where one chunk's
+ share of the audio ends and the next one's begins; `spans` are the
+ (start, end) samples actually transcribed, each cut-to-cut range widened
+ by half the overlap on both sides. The overlap counts towards the limit.
+
+ With `search_secs` > 0, each cut moves back from the limit to the middle
+ of the quietest QUIET_WINDOW_SECS window within the last `search_secs`.
+ """
+ overlap_secs = max(0.0, min(overlap_secs, max_chunk_secs / 4))
+ step = int((max_chunk_secs - overlap_secs) * sr)
+ if step < 1:
+ raise ValueError(f"chunk length must be positive, got {max_chunk_secs}s")
+ search = min(int(search_secs * sr), step // 2)
+ window = max(1, int(QUIET_WINDOW_SECS * sr))
+
+ cuts = []
+ position = 0
+ while len(audio) - position > step:
+ cut = position + step
+ if search > window:
+ cut = _quietest_point(audio, cut - search, cut, window)
+ cuts.append(cut)
+ position = cut
+
+ pad = int(overlap_secs * sr) // 2
+ edges = [0] + cuts + [len(audio)]
+ spans = [(max(0, start - pad), min(len(audio), end + pad))
+ for start, end in zip(edges, edges[1:])]
+ return spans, cuts
+
+
+def _quietest_point(audio, start, end, window):
+ """Middle of the lowest-energy `window` samples within audio[start:end].
+
+ Ties go to the latest window, which keeps chunks as long as allowed.
+ """
+ x = audio[start:end].astype(np.float64)
+ cumulative = np.concatenate(([0.0], np.cumsum(x * x)))
+ energy = cumulative[window:] - cumulative[:-window]
+ latest_min = len(energy) - 1 - int(np.argmin(energy[::-1]))
+ return start + latest_min + window // 2
+
+
+def stitch_slices(slice_results, cut_times, chunk_spans=None):
+ """Merge per-chunk (words, segments), already shifted to absolute time.
+
+ `chunk_spans` gives each chunk's (start, end) in seconds; omit it when the
+ chunks do not overlap. Each chunk contributes the words between its two
+ handovers. A handover is at the cut, unless the chunks overlap: then it
+ moves to the nearest word in the overlap that both chunks transcribed
+ alike, at the same time, and the left chunk keeps the words before it,
+ the right chunk that word and the ones after. (Splitting both chunks at
+ the cut is fragile: a word that follows a pause can be timestamped
+ anywhere in the pause, so the two chunks may put it on opposite sides of
+ the cut and keep it twice, or not at all.)
+
+ Segments are trimmed to the words their chunk keeps, and dropped if none.
+ """
+ first = [0] * len(slice_results)
+ last = [len(chunk_words) for chunk_words, _ in slice_results]
+ for k, cut in enumerate(cut_times):
+ overlap = (chunk_spans[k + 1][0], chunk_spans[k][1]) if chunk_spans else (cut, cut)
+ last[k], first[k + 1] = _handover(
+ slice_results[k][0], slice_results[k + 1][0], cut, overlap
+ )
+
+ words, segments = [], []
+ for (chunk_words, chunk_segments), lo, hi in zip(slice_results, first, last):
+ words.extend(chunk_words[lo:hi])
+ for seg, (start, stop) in zip(chunk_segments, _segment_ranges(chunk_words, chunk_segments)):
+ kept = chunk_words[max(start, lo):min(stop, hi)]
+ if start == stop: # a segment without words: keep it where its chunk does
+ if lo <= start < hi:
+ segments.append(seg)
+ elif len(kept) == stop - start:
+ segments.append(seg)
+ elif kept:
+ segments.append({
+ **seg,
+ "segment": " ".join(w["word"] for w in kept),
+ "start_offset": kept[0]["start_offset"],
+ "end_offset": kept[-1]["end_offset"],
+ "start": kept[0]["start"],
+ "end": kept[-1]["end"],
+ })
+ return words, segments
+
+
+def _handover(left, right, cut, overlap):
+ """(i, j): the left chunk keeps left[:i] and the right chunk right[j:].
+
+ Anchors are words in the overlap that both chunks transcribed with the same
+ text at nearly the same time. At the anchor nearest the cut, the right
+ chunk's copy is kept, or the left chunk's if that is what keeps the words
+ in time order. Without an anchor, both chunks split at the cut.
+ """
+ i = sum(1 for w in left if w["start"] < cut)
+ j = sum(1 for w in right if w["start"] < cut)
+ in_order = lambda a, b: a == 0 or b == len(right) or left[a - 1]["start"] <= right[b]["start"]
+ anchors = []
+ for p, lw in enumerate(left):
+ text = _normalize(lw["word"])
+ if not text or not overlap[0] <= lw["start"] < overlap[1]:
+ continue
+ partners = [q for q, rw in enumerate(right) if _normalize(rw["word"]) == text
+ and abs(rw["start"] - lw["start"]) <= SAME_WORD_SECS]
+ if partners:
+ q = min(partners, key=lambda q: abs(right[q]["start"] - lw["start"]))
+ options = [h for h in ((p, q), (p + 1, q + 1)) if in_order(*h)]
+ if options:
+ anchors.append((abs(lw["start"] + right[q]["start"] - 2 * cut), options[0]))
+ if anchors:
+ i, j = min(anchors)[1]
+ return i, j
+
+
+def _normalize(word):
+ return "".join(c for c in word.lower() if c.isalnum() or c == "'")
+
+
+def _segment_ranges(words, segments):
+ """[start, stop) word indices of each segment, matched in order by frame offsets."""
+ ranges, i = [], 0
+ for seg in segments:
+ while i < len(words) and words[i]["start_offset"] < seg["start_offset"]:
+ i += 1
+ start = i
+ while i < len(words) and words[i]["end_offset"] <= seg["end_offset"]:
+ i += 1
+ ranges.append((start, i))
+ return ranges
+
+
+def split_audio_file(audio_path, chunk_duration_secs=300, overlap_secs=0.0, search_secs=0.0):
+ """Split audio file into chunks of at most chunk_duration_secs."""
audio, sr = librosa.load(audio_path, sr=None, mono=True)
- total_duration = len(audio) / sr
- chunk_samples = int(chunk_duration_secs * sr)
+ spans, cuts = plan_slices(audio, sr, chunk_duration_secs, overlap_secs, search_secs)
chunks = []
- for start_sample in range(0, len(audio), chunk_samples):
- end_sample = min(start_sample + chunk_samples, len(audio))
+ for start_sample, end_sample in spans:
chunk_audio = audio[start_sample:end_sample]
- start_time = start_sample / sr
chunks.append({
'audio': chunk_audio,
- 'start_time': start_time,
+ 'start_time': start_sample / sr,
'duration': len(chunk_audio) / sr
})
- return chunks, sr
+ return chunks, sr, [cut / sr for cut in cuts]
def transcribe_buffered(
audio_path: str,
output_file: str = None,
chunk_duration_secs: float = 300, # 5 minutes default
+ overlap_secs: float = DEFAULT_OVERLAP_SECS,
+ pause_search_secs: float = DEFAULT_PAUSE_SEARCH_SECS,
):
"""
Transcribe long audio by splitting into chunks and merging results.
"""
+ import nemo.collections.asr as nemo_asr
+
# Determine model path
model_filename = "parakeet-tdt-0.6b-v3.nemo"
model_path = None
@@ -79,13 +224,15 @@ def transcribe_buffered(
asr_model.change_decoding_strategy(dec_cfg)
print("✓ CUDA graphs disabled successfully")
- print(f"Splitting audio into {chunk_duration_secs}s chunks...")
- chunks, sr = split_audio_file(audio_path, chunk_duration_secs)
+ print(f"Splitting audio into chunks of at most {chunk_duration_secs}s "
+ f"(overlap {overlap_secs}s, pause search {pause_search_secs}s)...")
+ chunks, sr, cut_times = split_audio_file(
+ audio_path, chunk_duration_secs, overlap_secs, pause_search_secs
+ )
print(f"Created {len(chunks)} chunks")
- all_words = []
- all_segments = []
- full_text = []
+ slice_results = []
+ chunk_texts = []
for i, chunk_info in enumerate(chunks):
print(f"Transcribing chunk {i+1}/{len(chunks)} (duration: {chunk_info['duration']:.1f}s)...")
@@ -104,25 +251,26 @@ def transcribe_buffered(
result_data = output[0]
chunk_text = result_data.text
- full_text.append(chunk_text)
+ chunk_texts.append(chunk_text)
+ chunk_words = []
+ chunk_segments = []
# Extract and adjust timestamps
if hasattr(result_data, 'timestamp') and result_data.timestamp:
- chunk_words = result_data.timestamp.get("word", [])
- chunk_segments = result_data.timestamp.get("segment", [])
-
# Adjust timestamps by chunk start time
- for word in chunk_words:
+ for word in result_data.timestamp.get("word", []):
word_copy = dict(word)
word_copy['start'] += chunk_info['start_time']
word_copy['end'] += chunk_info['start_time']
- all_words.append(word_copy)
+ chunk_words.append(word_copy)
- for segment in chunk_segments:
+ for segment in result_data.timestamp.get("segment", []):
seg_copy = dict(segment)
seg_copy['start'] += chunk_info['start_time']
seg_copy['end'] += chunk_info['start_time']
- all_segments.append(seg_copy)
+ chunk_segments.append(seg_copy)
+
+ slice_results.append((chunk_words, chunk_segments))
print(f"Chunk {i+1} complete: {len(chunk_text)} characters")
@@ -131,7 +279,15 @@ def transcribe_buffered(
if os.path.exists(chunk_path):
os.remove(chunk_path)
- final_text = " ".join(full_text)
+ chunk_spans = [(c['start_time'], c['start_time'] + c['duration']) for c in chunks]
+ all_words, all_segments = stitch_slices(slice_results, cut_times, chunk_spans)
+ if any(text.strip() and not words for (words, _), text in zip(slice_results, chunk_texts)):
+ # A chunk came back without word timestamps, so there is nothing to
+ # stitch it by; keep its text rather than lose it.
+ print("Warning: a chunk has text but no word timestamps; joining chunk texts")
+ final_text = " ".join(chunk_texts)
+ else:
+ final_text = " ".join(w["word"] for w in all_words)
print(f"Transcription complete: {len(final_text)} characters total")
output_data = {
@@ -144,6 +300,9 @@ def transcribe_buffered(
"buffered": True,
"chunk_duration_secs": chunk_duration_secs,
"num_chunks": len(chunks),
+ "overlap_secs": overlap_secs,
+ "pause_search_secs": pause_search_secs,
+ "cut_times": cut_times,
}
if output_file:
@@ -162,7 +321,17 @@ def main():
parser.add_argument("--output", "-o", help="Output file path", required=True)
parser.add_argument(
"--chunk-len", type=float, default=300,
- help="Chunk duration in seconds (default: 300 = 5 minutes)"
+ help="Maximum chunk duration in seconds, overlap included (default: 300 = 5 minutes)"
+ )
+ parser.add_argument(
+ "--overlap", type=float, default=DEFAULT_OVERLAP_SECS,
+ help=f"Seconds shared by adjacent chunks, capped at a quarter of --chunk-len "
+ f"(default: {DEFAULT_OVERLAP_SECS}; 0 disables)"
+ )
+ parser.add_argument(
+ "--pause-search", type=float, default=DEFAULT_PAUSE_SEARCH_SECS,
+ help=f"Seconds before each chunk limit searched for the quietest point to cut at, "
+ f"e.g. 25 (default: {DEFAULT_PAUSE_SEARCH_SECS}, cut at the limit)"
)
args = parser.parse_args()
@@ -175,6 +344,8 @@ def main():
audio_path=args.audio_file,
output_file=args.output,
chunk_duration_secs=args.chunk_len,
+ overlap_secs=args.overlap,
+ pause_search_secs=args.pause_search,
)
diff --git a/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py b/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
new file mode 100644
index 0000000..6a35947
--- /dev/null
+++ b/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
@@ -0,0 +1,329 @@
+"""Unit tests for the slicing and stitching helpers in parakeet_transcribe_buffered.py.
+
+These are pure functions: they need numpy, librosa and soundfile (imported by the
+script) but no GPU, no NeMo and no model.
+"""
+import sys
+from pathlib import Path
+
+import numpy as np
+import pytest
+
+sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
+from parakeet_transcribe_buffered import plan_slices, stitch_slices # noqa: E402
+
+SR = 16000
+
+
+def speech(seconds, seed=0):
+ """Stand-in for continuous speech: broadband noise at a speech-like level."""
+ rng = np.random.default_rng(seed)
+ return (0.1 * rng.standard_normal(int(seconds * SR))).astype(np.float32)
+
+
+def with_pauses(audio, pauses, level=0.001, seed=1):
+ """Replace each (start_s, end_s) span with low-level room noise (or zeros)."""
+ rng = np.random.default_rng(seed)
+ out = audio.copy()
+ for start, end in pauses:
+ a, b = int(start * SR), int(end * SR)
+ out[a:b] = level * rng.standard_normal(b - a)
+ return out
+
+
+def assert_valid_plan(spans, cuts, num_samples, max_secs):
+ assert spans[0][0] == 0 and spans[-1][1] == num_samples
+ assert len(spans) == len(cuts) + 1
+ for start, end in spans:
+ assert 0 < end - start <= max_secs * SR
+ for (a0, a1), (b0, b1), cut in zip(spans, spans[1:], cuts):
+ assert b0 <= cut <= a1, "each cut must lie inside both neighbouring slices"
+ assert a0 < cut < b1
+
+
+# -- plan_slices: cut placement -------------------------------------------------
+
+
+def test_audio_shorter_than_one_slice_is_not_cut():
+ audio = speech(60)
+ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=4, search_secs=25)
+ assert cuts == []
+ assert spans == [(0, len(audio))]
+
+
+def test_audio_exactly_one_slice_long_is_not_cut():
+ audio = speech(120)
+ spans, cuts = plan_slices(audio, SR, 120)
+ assert cuts == []
+ assert spans == [(0, len(audio))]
+
+
+def test_without_search_or_overlap_the_legacy_fixed_grid_is_reproduced():
+ audio = speech(300)
+ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=0, search_secs=0)
+ assert cuts == [120 * SR, 240 * SR]
+ assert spans == [(0, 120 * SR), (120 * SR, 240 * SR), (240 * SR, 300 * SR)]
+
+
+def test_cuts_land_in_the_pauses_before_the_limit():
+ pauses = [(100.0, 100.5), (215.0, 215.5), (330.0, 330.5)]
+ audio = with_pauses(speech(400), pauses)
+ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=0, search_secs=25)
+ assert len(cuts) == 3
+ for cut, (start, end) in zip(cuts, pauses):
+ assert start <= cut / SR <= end
+ assert_valid_plan(spans, cuts, len(audio), 120)
+
+
+def test_the_quietest_pause_wins():
+ # Two pauses inside the same search window; the later one is louder.
+ audio = with_pauses(speech(200), [(100.0, 100.6)], level=0.0)
+ audio = with_pauses(audio, [(115.0, 115.6)], level=0.01)
+ _, cuts = plan_slices(audio, SR, 120, search_secs=25)
+ assert 100.0 <= cuts[0] / SR <= 100.6
+
+
+def test_audio_with_no_pause_is_still_cut_within_the_limit():
+ audio = speech(400)
+ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=0, search_secs=25)
+ assert_valid_plan(spans, cuts, len(audio), 120)
+ edges = [0] + cuts
+ for prev, cut in zip(edges, cuts):
+ assert 95 * SR <= cut - prev <= 120 * SR, "cut must fall inside its search window"
+
+
+def test_a_pause_at_the_very_start_is_never_a_cut():
+ audio = with_pauses(speech(200), [(0.0, 5.0)], level=0.0)
+ spans, cuts = plan_slices(audio, SR, 120, search_secs=25)
+ assert len(cuts) == 1 and 95 <= cuts[0] / SR <= 120
+ assert_valid_plan(spans, cuts, len(audio), 120)
+
+
+def test_a_pause_at_the_very_end_leaves_no_empty_slice():
+ # First cut ~100.35 s, so the second search window is ~[195, 220] s and
+ # holds the start of the trailing silence (217-222 s).
+ audio = with_pauses(speech(222), [(100.0, 100.5), (217.0, 222.0)], level=0.0)
+ spans, cuts = plan_slices(audio, SR, 120, search_secs=25)
+ assert len(cuts) == 2
+ assert 217.0 <= cuts[1] / SR < 222.0
+ assert_valid_plan(spans, cuts, len(audio), 120)
+
+
+# -- plan_slices: overlap -------------------------------------------------------
+
+
+def test_overlap_is_included_in_the_slice_limit_and_centred_on_each_cut():
+ pauses = [(100.0, 100.5), (215.0, 215.5), (330.0, 330.5)]
+ audio = with_pauses(speech(400), pauses)
+ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=4, search_secs=25)
+ assert_valid_plan(spans, cuts, len(audio), 120)
+ for (_, a1), (b0, _), cut in zip(spans, spans[1:], cuts):
+ assert a1 - b0 == 4 * SR
+ assert cut - b0 == a1 - cut
+ for cut, (start, end) in zip(cuts, pauses):
+ assert start <= cut / SR <= end
+
+
+def test_overlap_without_pause_search_uses_a_fixed_grid():
+ audio = speech(300)
+ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=4, search_secs=0)
+ assert cuts == [116 * SR, 232 * SR]
+ assert spans == [(0, 118 * SR), (114 * SR, 234 * SR), (230 * SR, 300 * SR)]
+
+
+def test_short_slice_lengths_clamp_overlap_and_search():
+ # Upstream's own buffered test runs a 19 s clip with --chunk-len 10.
+ audio = speech(19)
+ spans, cuts = plan_slices(audio, SR, 10, overlap_secs=4, search_secs=25)
+ assert len(spans) >= 2
+ assert_valid_plan(spans, cuts, len(audio), 10)
+
+
+# -- stitch_slices ------------------------------------------------------------------
+
+FRAME = 0.08
+OVERLAPPING = [(0.0, 12.0), (8.0, 20.0)] # two chunks sharing 8-12 s, cut at 10 s
+
+
+def word(text, start, end, slice_start):
+ return {
+ "word": text,
+ "start_offset": round((start - slice_start) / FRAME),
+ "end_offset": round((end - slice_start) / FRAME),
+ "start": start,
+ "end": end,
+ }
+
+
+def segment(words):
+ return {
+ "segment": " ".join(w["word"] for w in words),
+ "start_offset": words[0]["start_offset"],
+ "end_offset": words[-1]["end_offset"],
+ "start": words[0]["start"],
+ "end": words[-1]["end"],
+ }
+
+
+def texts(items, key="word"):
+ return [item[key] for item in items]
+
+
+def test_no_overlap_stitching_is_plain_concatenation():
+ left = [word("one", 1.0, 1.4, 0), word("two", 5.0, 5.3, 0)]
+ right = [word("three", 10.5, 10.9, 10), word("four", 14.0, 14.4, 10)]
+ words, segments = stitch_slices(
+ [(left, [segment(left)]), (right, [segment(right)])], [10.0]
+ )
+ assert words == left + right
+ assert segments == [segment(left), segment(right)]
+
+
+def test_overlapping_slices_keep_every_word_exactly_once():
+ # Slices [0, 12] and [8, 20], cut at 10. Both transcribe the overlap and
+ # their timestamps for the same word differ by a few ms.
+ left = [
+ word("a", 1.0, 1.3, 0),
+ word("b", 5.0, 5.4, 0),
+ word("c", 9.00, 9.40, 0),
+ word("d", 10.50, 10.90, 0),
+ word("e", 11.50, 11.80, 0),
+ ]
+ right = [
+ word("c", 9.02, 9.40, 8),
+ word("d", 10.48, 10.90, 8),
+ word("e", 11.52, 11.80, 8),
+ word("f", 15.00, 15.40, 8),
+ ]
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
+ assert texts(words) == ["a", "b", "c", "d", "e", "f"]
+ assert words[2] is left[2] and words[3] is right[1]
+
+
+def test_a_word_straddling_the_cut_is_kept_once():
+ left = [word("over", 9.90, 10.30, 0), word("the", 10.40, 10.55, 0)]
+ right = [word("over", 9.92, 10.30, 8), word("the", 10.40, 10.55, 8)]
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
+ assert texts(words) == ["over", "the"]
+ # Both chunks agree on it, so the right chunk takes over from it.
+ assert words[0] is right[0] and words[1] is right[1]
+
+
+def test_a_segment_straddling_the_cut_is_trimmed_to_the_words_each_slice_owns():
+ left_words = [
+ word("Hello", 8.5, 8.9, 0),
+ word("there", 9.2, 9.6, 0),
+ word("friend.", 10.4, 10.9, 0),
+ ]
+ right_words = [
+ word("there", 9.21, 9.6, 8),
+ word("friend.", 10.41, 10.9, 8),
+ word("Bye.", 13.0, 13.4, 8),
+ ]
+ words, segments = stitch_slices(
+ [
+ (left_words, [segment(left_words)]),
+ (right_words, [segment(right_words[:2]), segment(right_words[2:])]),
+ ],
+ [10.0],
+ OVERLAPPING,
+ )
+ assert texts(words) == ["Hello", "there", "friend.", "Bye."]
+ assert texts(segments, "segment") == ["Hello there", "friend.", "Bye."]
+ assert segments[0]["end"] == 9.6 and segments[1]["start"] == 10.41
+ # A segment that needed no trimming is passed through untouched.
+ assert segments[2] == segment(right_words[2:])
+ # Every word appears in exactly one segment, in order.
+ assert " ".join(texts(segments, "segment")) == " ".join(texts(words))
+
+
+def test_a_segment_wholly_inside_the_other_slices_share_is_dropped():
+ left_words = [word("a", 2.0, 2.3, 0), word("b.", 10.6, 11.0, 0)]
+ right_words = [word("b.", 10.61, 11.0, 8), word("c", 12.0, 12.3, 8)]
+ _, segments = stitch_slices(
+ [
+ (left_words, [segment(left_words[:1]), segment(left_words[1:])]),
+ (right_words, [segment(right_words[:1]), segment(right_words[1:])]),
+ ],
+ [10.0],
+ OVERLAPPING,
+ )
+ assert texts(segments, "segment") == ["a", "b.", "c"]
+ assert segments[1]["start"] == 10.61
+
+
+def test_a_non_positive_chunk_length_is_rejected_rather_than_looping():
+ with pytest.raises(ValueError):
+ plan_slices(speech(5), SR, 0)
+
+
+def test_a_word_the_two_chunks_timestamp_either_side_of_the_cut_is_kept_once():
+ # After a pause TDT may place a word's start anywhere in the pause, so the
+ # two chunks can disagree about which side of the cut it starts on.
+ left = [word("so", 8.2, 8.5, 0), word("then", 9.98, 10.3, 0), word("we", 10.4, 10.6, 0)]
+ right = [word("so", 8.2, 8.5, 8), word("then", 10.03, 10.3, 8), word("we", 10.4, 10.6, 8)]
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
+ assert texts(words) == ["so", "then", "we"]
+ # ...and the mirror image, where splitting both at the cut would drop it.
+ left = [word("so", 8.2, 8.5, 0), word("then", 10.03, 10.3, 0), word("we", 10.4, 10.6, 0)]
+ right = [word("so", 8.2, 8.5, 8), word("then", 9.98, 10.3, 8), word("we", 10.4, 10.6, 8)]
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
+ assert texts(words) == ["so", "then", "we"]
+
+
+def test_handover_happens_at_the_agreed_word_nearest_the_cut():
+ # The chunks differ in casing/punctuation and the left one drops "really"
+ # near its end; the right chunk's version of the overlap after the cut wins.
+ left = [word("It", 8.5, 8.7, 0), word("was", 9.6, 9.9, 0), word("good,", 11.0, 11.4, 0)]
+ right = [word("it", 8.5, 8.7, 8), word("was", 9.62, 9.9, 8), word("really", 10.3, 10.7, 8),
+ word("good.", 11.0, 11.4, 8)]
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
+ assert texts(words) == ["It", "was", "really", "good."]
+ assert words[0] is left[0] and words[1] is right[1]
+
+
+def test_without_an_agreed_word_the_split_falls_back_to_the_cut():
+ left = [word("alpha", 9.0, 9.4, 0), word("beta", 10.5, 10.9, 0)]
+ right = [word("gamma", 9.1, 9.4, 8), word("delta", 10.6, 10.9, 8)]
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
+ assert texts(words) == ["alpha", "delta"]
+
+
+def test_another_occurrence_of_the_word_elsewhere_in_the_overlap_is_not_an_anchor():
+ # The chunks disagree everywhere except on "the", but the left chunk's
+ # "the" (9.0 s) and the right chunk's (11.0 s) are different words. As an
+ # anchor they would average to the cut and discard the left one.
+ left = [word("the", 9.0, 9.2, 0), word("dog", 10.5, 10.8, 0)]
+ right = [word("cat", 9.3, 9.6, 8), word("the", 11.0, 11.2, 8)]
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
+ assert [w["start"] for w in words] == [9.0, 11.0]
+
+
+def test_the_handover_never_puts_words_out_of_time_order():
+ # "y" agrees (0.45 s apart), but the right chunk's "y" (9.65) after the
+ # left chunk's "x" (9.70) would run time backwards, so the left chunk's
+ # copy is kept instead. Splitting at the cut would lose "y" altogether.
+ left = [word("x", 9.70, 9.90, 0), word("y", 10.10, 10.30, 0)]
+ right = [word("z", 9.40, 9.60, 8), word("y", 9.65, 9.90, 8), word("w", 10.6, 10.8, 8)]
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
+ assert texts(words) == ["x", "y", "w"] and words[1] is left[1]
+ starts = [w["start"] for w in words]
+ assert starts == sorted(starts)
+
+
+def test_a_co_timed_anchor_is_found_even_when_a_longer_match_lies_elsewhere():
+ # "x y" recurs later in the right chunk, a longer text match than "z", but
+ # at a different time. Only "z" is the same word in both chunks, and it
+ # straddles the cut, so splitting both at the cut would keep it twice.
+ left = [word("x", 8.2, 8.3, 0), word("y", 8.4, 8.5, 0), word("z", 9.98, 10.2, 0)]
+ right = [word("z", 10.03, 10.2, 8), word("x", 11.0, 11.1, 8), word("y", 11.2, 11.3, 8)]
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
+ assert texts(words) == ["x", "y", "z", "x", "y"]
+
+
+def test_punctuation_alone_is_never_an_anchor():
+ # As an anchor the dash would hand the whole overlap to the right chunk.
+ left = [word("-", 9.50, 9.55, 0), word("yes", 10.4, 10.6, 0)]
+ right = [word("-", 9.52, 9.55, 8), word("no", 10.4, 10.6, 8)]
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
+ assert words[0] is left[0] and texts(words) == ["-", "no"]
--
2.39.5
+162
View File
@@ -0,0 +1,162 @@
# Scriberr local patches — contract
We build Scriberr from source (no upstream sm_120 image; see `../README.md`), so
we can carry patches on that build. This directory holds them, and
`scripts/scriberr-rebuild` applies them to a pinned upstream sha, builds a
distinctly tagged image, and proves it before anyone deploys it.
| patch | against | status |
|---|---|---|
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | carried; upstream PR **prepared, not opened** (`upstream-pr/`) |
Ruling: Prime, 2026-09-30, "build the slicer". Opening the upstream PR is
outward-facing and needs Prime's explicit yes.
## 0001 — pause-aware Parakeet slicer
### What it changes
One file of product code, `internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py`,
plus one new test file beside it (`tests/test_parakeet_slicing.py`). No Go change.
Upstream cuts long audio at fixed `--chunk-len` marks with no overlap, so a
word that straddles a mark is chopped in two, lost, or transcribed twice. The
patch:
1. **Overlaps adjacent chunks** by `--overlap` seconds (default **4**), half on
each side of the cut, counted *inside* `--chunk-len`.
2. **Hands over at an agreed word.** In each overlap, the chunks switch at the
word nearest the cut that both transcribed alike: the same text after
lowercasing and stripping punctuation (punctuation alone never counts), with
start times within 0.5 s. The left chunk keeps the words before it and the
right chunk keeps the rest, the anchor taken from whichever chunk keeps the
words in time order. With no agreed word, both split at the cut by start time.
Segments are trimmed to the words their chunk keeps, and `transcription` is
the stitched words joined by spaces (upstream's text already equals that).
3. **Optional pause-aware cuts** (`--pause-search N`, default **off**): each cut
moves back to the middle of the quietest 0.3 s within the last N seconds
before the limit. It measured neutral once the stitch was right, so it is not
the default; Go never passes the flag.
4. **Imports NeMo inside `transcribe_buffered()`** so the pure helpers
(`plan_slices`, `stitch_slices`) import and test without a GPU or NeMo.
`--overlap 0` (with pause search off, the default) reproduces upstream's output
exactly: words, segments and text were byte-identical on all four test recordings.
Why the handover is by agreed word and not simply "each word goes to the chunk
its start time falls in" (the first design): at a quarter of the stitches the
two chunks put the *same* word on opposite sides of the cut, one frame apart,
so it was kept twice. Parakeet timestamps a word that follows a pause anywhere
inside the pause. Details in the bench doc.
### The seam it must keep (Go ↔ Python)
`parakeet_adapter.go` is not patched, so the script's CLI and JSON are frozen:
- **Invocation** (Go, `buildBufferedArgs`):
`uv run --native-tls --project <env> python parakeet_transcribe_buffered.py <audio> --output <json> --chunk-len <PARAKEET_CHUNK_THRESHOLD_SECS>`.
Go never passes `--overlap` or `--pause-search`, so **their defaults are the
shipped behaviour**. New flags must stay optional.
- **JSON** (Go, `parseResult`): `transcription` (str), `language` (str),
`word_timestamps` [{`word` str, `start_offset` **int**, `end_offset` **int**,
`start` float, `end` float}], `segment_timestamps` (same, with `segment`),
`audio_file`, `model`, `buffered`, `chunk_duration_secs`, `num_chunks`. An int
field that becomes a float fails the Go unmarshal. Extra keys are fine; the
patch adds `overlap_secs`, `pause_search_secs` and `cut_times`.
- `start_offset`/`end_offset` stay chunk-relative frame indices, as upstream
leaves them; Go does not read them.
- `scripts/scriberr-seam-check.py` asserts all of the above plus the stitch
invariants (word starts never go backwards; segments tile the words once).
### The memory bound
Scriberr shares fv-ml1 GPU 1 with intern-decision. **Parakeet's per-process
peak must stay ≤ 5,496 MiB** (nvidia-smi used_memory, sampled every 0.2 s), the
measured peak of upstream's 120 s slicer with `expandable_segments`. The patch
keeps it because the overlap counts **inside** `--chunk-len`: consecutive cuts
are at most `chunk-len − overlap` apart, so no chunk ever exceeds `--chunk-len`
(120 s in our compose). The pause search runs on the CPU copy of the waveform.
### Measured (2026-09-30, `docs/pfi/scriberr-slicer-bench-2026-09-30.md`)
Four recordings, 118 min in total: Prime's two uploads (private, metrics only),
the first 30 min of a Supreme Court oral argument, and a LibriVox dramatic
reading. Each was scored against a no-cut whole-file transcript, with every
variant at three cut placements.
| | upstream (fixed 120 s) | **patch (overlap 4 s, agreed-word handover)** |
|---|---|---|
| cuts with an error within ±3 s | 52 % (93/179) | **22 % (41/184)** |
| background: same test midway between cuts | 19 % | 18 % |
| near-cut error events | 102 | 47 |
| words duplicated at cuts | 17 | 2 |
| GPU peak, 35-min file, n=3 | 5,496 MiB | **5,496 MiB** (budget 5,496) |
Floor: ±0.08 on a pooled damaged-cut rate (2 SE). Pause-only, overlap-only with
the handover, and pause + overlap at 4 or 8 s all land inside that floor of
each other; the default is the simplest of them.
⚠ **Separate finding, not fixed by this patch:** Parakeet sometimes skips a run
of ≥10 consecutive words mid-chunk (12–17 runs and 500–720 words per 12
transcripts, for upstream's slicer too). See the bench doc.
### Upgrading upstream
```bash
# on nh3-dev, from this repo
scripts/scriberr-rebuild --sha <full 40-char upstream sha> --suffix slicer1
```
It clones that sha into a new `/opt/docker/src/scriberr-<sha7>-<suffix>`,
`git apply --check`s each patch (a conflict stops the run and names the patch),
builds `scriberr:local-blackwell-<sha7>-<suffix>` without touching older tags,
then runs the embed, unit, seam and memory stages and prints a PASS/FAIL table.
Memory runs on GPU 3 and refuses a GPU that is not idle. A conflict means the
patch needs rebasing: in a checkout of the new sha, `git am -3` the old patch,
resolve, run the unit tests, then `git format-patch -1 --stdout >
0001-parakeet-pause-aware-slicer.patch` and re-run the rebuild.
Also re-measure after an upgrade that touches NeMo, torch or the slicer
(`docs/pfi/scriberr-slicer-bench-2026-09-30.md` has the harness).
**Disk.** Everything a rebuild writes lands on fv-ml1's root pool (zroot), not
`/tank`: the build dir under `/opt/docker/src` (~125 MB) and the image, which
shares ~6.1 GB of layers with the other Scriberr images and adds ~120 MB of its
own (measured 2026-09-30). The script refuses to build with less than 20 GB free
under Docker's root dir. Once a deploy has soaked, remove superseded builds by
their literal names: `docker rmi scriberr:local-blackwell-<sha7>-<suffix>` and
`sudo -n rm -rf /opt/docker/src/scriberr-<sha7>-<suffix>`, keeping the running
tag and the one before it for rollback.
⚠ `Dockerfile.cuda.12.9` installs the **latest** `uv`, `yt-dlp` and `deno` at
build time, so a rebuild changes those too, not just our patch. The seam stage
is what catches a `uv run` behaviour change.
### Deploy (manual; the rebuild script never does this)
`SCRIBERR_IMAGE` in `/opt/docker/compose/scriberr/.env` on fv-ml1 selects the
image (`.env.example` documents it). That `.env` is lkraven-owned mode 600,
kept out of git.
```bash
ssh infra-ops@10.251.50.54
cd /opt/docker/compose/scriberr
sudo -n cp -p .env .env.bak-$(date +%Y%m%d-%H%M) | cat # backup first
sudo -n grep -n '^SCRIBERR_IMAGE=' .env | cat # current value
# edit SCRIBERR_IMAGE=scriberr:local-blackwell-<sha7>-<suffix> with sudo -n
docker logs --since 2m scriberr 2>&1 | grep -c "Processing single-track job" # must be 0
sudo -n docker compose config >/dev/null | cat && sudo -n docker compose up -d scriberr | cat
```
Then prove the embed path live: the env's rewritten copy must match the patch.
```bash
docker exec scriberr sha256sum /app/whisperx-env/parakeet/parakeet_transcribe_buffered.py
sha256sum /opt/docker/src/scriberr-<sha7>-<suffix>/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
```
Record the deploy with `scripts/ops-log record`.
**Rollback:** set `SCRIBERR_IMAGE` back to the previous tag (or restore the
`.env` backup) and `docker compose up -d scriberr`. The old image is never
deleted by the rebuild.
+113
View File
@@ -0,0 +1,113 @@
# Upstream PR — prepared, NOT opened
**Status:** ready to submit to `rishikanthc/Scriberr`, **held for Prime's explicit
yes** (opening a PR is outward-facing). Nothing has been pushed to GitHub.
- **Diff:** `../0001-parakeet-pause-aware-slicer.patch`, a `git format-patch` of one
commit authored by Vuong Hoang against upstream `a353078` (current HEAD,
2026-09-20). It is the same file we carry, so the PR and our build cannot drift.
- **To open it (after the yes):** fork on GitHub, then
`git clone <fork> && cd Scriberr && git checkout -b parakeet-overlap-stitch a353078 &&
git am <path>/0001-parakeet-pause-aware-slicer.patch && git push -u origin parakeet-overlap-stitch`,
and open the PR with the title and body below.
- **Before sending, decide:** keep the opt-in `--pause-search` in the upstream
version, or drop it for a smaller diff (about 40 lines with its tests). It measured
neutral; see the body's last paragraph.
---
## Title
Parakeet buffered transcription: overlap chunks and stitch at an agreed word
## Body
### Problem
For audio longer than `PARAKEET_CHUNK_THRESHOLD_SECS`, `parakeet_transcribe_buffered.py`
cuts the file at fixed `--chunk-len` marks with no overlap. A word that straddles a
mark is chopped in two, lost, or transcribed twice. We measured it on four
recordings (118 minutes, see below): **52 % of the cuts had a transcription error
within ±3 s of them, against 19 % at points midway between cuts.**
Lowering `PARAKEET_CHUNK_THRESHOLD_SECS` to save GPU memory, which is what the
knob is for on smaller cards, makes this worse, because there are more cuts.
### Change
One Python file, plus a new test file. No Go changes; the CLI and JSON that
`parakeet_adapter.go` reads are unchanged.
1. **Overlap.** Adjacent chunks share `--overlap` seconds (default 4), half on
each side of the cut. The overlap counts *inside* `--chunk-len`, so no chunk
gets longer and peak GPU memory is unchanged.
2. **Stitch at an agreed word.** In each overlap, the two chunks hand over at the
word nearest the cut that both transcribed alike: same text after lowercasing
and stripping punctuation, start times within 0.5 s. The left chunk keeps the
words before it and the right chunk the rest, with the anchor taken from
whichever chunk keeps word starts in time order. With no agreed word, both
split at the cut. Segments are trimmed to match, and `transcription` is the
stitched words joined by spaces (which is what it already equals for Parakeet).
The obvious rule, "keep each word from the chunk whose half its start time falls
in", is not enough. A word that follows a pause can be timestamped anywhere in
the pause, so the two chunks often put the *same* word on opposite sides of the
cut, one frame apart, and it comes out twice (or not at all).
3. `--pause-search N` (opt-in, off by default): move each cut back to the
quietest 0.3 s within the last N seconds before the limit.
4. NeMo is imported inside `transcribe_buffered()` so the pure helpers
(`plan_slices`, `stitch_slices`) can be unit-tested without a GPU.
New flags are optional with defaults, because the Go side does not pass them. The
JSON gains `overlap_secs`, `pause_search_secs` and `cut_times`. `--overlap 0`
reproduces the previous output exactly (words, segments and text were
byte-identical on all four test files).
### Measurements
Setup: RTX PRO 6000 Blackwell, the `Dockerfile.cuda.12.9` image, parakeet-tdt-0.6b-v3,
16 kHz mono input. Each run was scored against a **no-cut reference**: the same
model over the whole file in one pass with local attention (`parakeet_transcribe.py
--context-left 255 --context-right 255`). Words were aligned after lowercasing and
stripping punctuation. A cut counts as damaged if any error lies within ±3 s of it,
and the same test at points midway between cuts gives the background rate. Each
variant ran at three chunk lengths (120, 110 and 100 s) so the cuts land in
different places. Decoding is deterministic, so repeats are identical.
Recordings: the first 30 minutes of a U.S. Supreme Court oral argument (No. 22-451,
public domain), section 1 of the LibriVox dramatic reading *The Trial of Oscar
Wilde* (public domain), and two private conversational recordings (22 and 35 min;
numbers only).
| slicer | damaged cuts | background | error events near cuts | words duplicated at cuts |
|---|---|---|---|---|
| current (fixed, no overlap) | 93/179 = **52 %** | 19 % | 102 | 17 |
| overlap, start-time stitch | 52/184 = 28 % | 18 % | 60 | 18 |
| **overlap, agreed-word stitch (this PR)** | 41/184 = **22 %** | 18 % | 47 | 2 |
| pause-aware cut, no overlap | 53/201 = 26 % | 18 % | 59 | 0 |
| pause-aware + overlap, agreed-word stitch | 52/207 = 25 % | 18 % | 54 | 2 |
The 2-standard-error band on a difference of these rates is about ±0.08, so the
last three rows are statistically tied and all beat the current slicer by a wide
margin. The public files alone tell the same story: current 52/86 damaged cuts;
this PR 22/89. Peak GPU memory for a 35-minute file with `--chunk-len 120` is
unchanged (5,496 MiB before and after, n=3, `expandable_segments:True`).
Pause-aware cutting is included but off by default: once the stitch was right, it
did not measurably help, and without an overlap it drops words just before its
cuts. Happy to drop it from this PR if you prefer the smaller diff.
### Tests
```
cd internal/transcription/adapters/py/nvidia
python -m pytest tests/test_parakeet_slicing.py # needs numpy, librosa, soundfile; no GPU
```
The tests cover synthetic waveforms with known pauses, audio shorter than one chunk,
a pause at the very start or end, audio with no pause, the chunk limit with the
overlap included, and stitching with known word lists, including a word the two
chunks timestamp on either side of the cut, a word one chunk missed, a longer text
match at a different time, an anchor that would reverse time order, and
punctuation-only tokens. The existing `test_parakeet_transcribe_buffered.py`
still passes, since its `--chunk-len 10` run now also exercises the overlap.