diff --git a/docs/pfi/scriberr-slicer-bench-2026-09-30.md b/docs/pfi/scriberr-slicer-bench-2026-09-30.md new file mode 100644 index 0000000..e3cc92f --- /dev/null +++ b/docs/pfi/scriberr-slicer-bench-2026-09-30.md @@ -0,0 +1,220 @@ +# Scriberr Parakeet slicer bench (2026-09-30) + +Prime's ruling, 2026-09-30: "build the slicer". This document records how the +pause-aware slicer patch (`stacks/scriberr/patches/0001-parakeet-pause-aware-slicer.patch`) +was measured and why the shipped variant was chosen. **Metrics only**: two of +the recordings are Prime's and private, so no transcript text appears here or +anywhere in git. Those recordings, and every transcript made from them, stay on +fv-ml1 in `/tank/spikes/scriberr-slicer/private/` (mode 700). + +## Recordings (4 files, 118 minutes) + +All four were converted to 16 kHz mono WAV with the image's ffmpeg, which is +what Scriberr feeds Parakeet. + +| id | length | what | provenance / licence | +|---|---|---|---| +| p1 | 35.3 min | Prime's upload, conversational | private | +| p2 | 22.3 min | Prime's upload, conversational | private | +| scotus | 30.0 min (first half hour) | U.S. Supreme Court oral argument, *Loper Bright Enterprises v. Raimondo*, No. 22-451, argued 2024-01-17; spontaneous multi-speaker speech with interruptions | `https://www.supremecourt.gov/media/audio/mp3files/22-451.mp3` (sha256 `7e704b1f…5dc2`); U.S. Government work, public domain (17 U.S.C. §105) | +| wilde | 24.4 min | LibriVox *The Trial of Oscar Wilde* (dramatic reading), section 1; several readers, courtroom dialogue | `https://archive.org/download/trialofoscarwilde_1601_librivox/trialofoscarwilde_01_anon_64kb.mp3` (sha256 `492f4c56…3420`); public domain (LibriVox, PD Mark 1.0) | + +## Harness + +- **Image and env:** `scriberr:local-blackwell` (the live image), the live + `/tank/scriberr/whisperx-env` mounted **read-only**, `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`, + invoked exactly as Scriberr does (`uv run --native-tls --project /app/whisperx-env/parakeet python …`). + It runs as uid 1002 with `USER` set, not as appuser 10001; that changes file + ownership only. +- **GPU:** transient containers on **GPU 3 only** (verified: the container sees + one card, UUID `GPU-186dacf4…`). GPU 3 was at 2 MiB before and after. +- **Reference:** upstream's standard script on the whole file in one pass with + local attention (`--context-left 255 --context-right 255`), which has no cuts. + It needs >16 GB, which is why it cannot run on GPU 1. +- **Variants** run through the real `transcribe_buffered()`; the only harness + change is that the model is loaded once per process instead of once per call. + The CLI memory runs below reproduce the harness output byte for byte. +- **Placements:** every variant at three maximum lengths, 120, 110 and 100 s. + Parakeet is deterministic, so a repeat run cannot supply variance; moving the + cuts can. Each (file, variant) therefore has n = 3 different cut placements. + +## Metric + +Each variant's words are aligned against the reference, after lowercasing and +stripping punctuation (difflib opcodes, then exact Levenshtein inside each +non-matching block). Each error gets a time: the hypothesis word's start for a +substitution or insertion, the reference word's start for a deletion. + +- **near-cut / elsewhere word errors (S/I/D):** near-cut means within ±3 s of + any of that variant's cuts (for overlap variants, the cut is the stitch point + at the middle of the overlap, and both chunk edges lie within ±3 s of it). +- **dropped / duplicated at cuts:** deletions near a cut, and insertions near a + cut that repeat an adjacent word. +- **error events:** errors clustered with gaps ≤ 1 s; rate per minute + elsewhere. +- **damaged cuts (the decision metric):** cuts with at least one error within + ±3 s, against **damaged phantoms**, the same test at points midway between + the variant's own cuts (same count, same spacing, as far from any cut as the + audio gets). The phantom rate is the background any slicer's cuts are judged + against. + +### Why the decision metric is not raw word counts + +The negative control caught it. Word errors elsewhere varied by up to ±50 % +between slicers of the same length (p1: 131 to 215). Most of that comes from a +few **unstable stretches**, the same stretches for every variant (p1: 976–992 s, +1291–1298 s, 928–934 s), where the reference and any slicing disagree by 20–60 +words depending on context, at arbitrary distances from any cut. That is heavy-tailed +noise, not cut damage. Counting events, and asking per cut whether anything near +it went wrong, is robust to it; the word counts are still reported. + +## Variants + +| name | overlap | pause search | stitch | +|---|---|---|---| +| fixed | 0 | off | none: upstream's slicer (patched code with both off reproduces it byte for byte) | +| pause | 0 | 25 s | none needed | +| overlap | 4 s | off | v1: each word kept by the chunk whose cut-to-cut range holds its start time (the brief's rule) | +| both | 4 s | 25 s | v1 | +| **overlap2** | 4 s | off | **v2: hand over at the nearest word both chunks transcribed alike, within 0.5 s** (shipped as v3, identical output) | +| both2 | 4 s | 25 s | v2 | +| both8 | 8 s | 25 s | v2 | + +**v3** is v2 after code review, and it is what ships. Anchors are paired by +time first (same text *and* within 0.5 s, so a longer text match elsewhere in +the overlap cannot crowd out the true anchor), punctuation alone never anchors, +and the anchor's copy comes from whichever chunk keeps word starts in time order. +Re-run on all 12 overlap runs (4 files × 3 placements), **v3's output is +byte-identical to v2's** (and "both3" to "both2"): the reviewer's cases did not +occur in this audio, so every v2 number below holds for the shipped code. + +v2 exists because v1 failed its own goal. At 26 of 108 "both" stitches the same +word sat on both sides of the cut, the two copies' start times 0.00–0.09 s apart +(one encoder frame): a word that follows a pause is timestamped anywhere in the +pause, and a pause is exactly where a pause-aware cut lands. Start time at the +midpoint is the worst possible stitch rule for pause cuts. + +## Results + +### Pooled (4 files × 3 placements) + +| variant | damaged cuts | damaged phantoms | excess | by length 120 / 110 / 100 | near-cut words | dropped | duplicated | near-cut events | events elsewhere /min | +|---|---|---|---|---|---|---|---|---|---| +| fixed (upstream) | 93/179 = 52 % | 34/179 = 19 % | +0.33 | 0.48 / 0.52 / 0.55 | 223 | 56 | 17 | 102 | 2.08 | +| pause | 53/201 = 26 % | 36/201 = 18 % | +0.085 | 0.33 / 0.20 / 0.27 | 113 | 52 | 0 | 59 | 2.23 | +| overlap (v1) | 52/184 = 28 % | 33/184 = 18 % | +0.10 | 0.30 / 0.28 / 0.27 | 117 | 25 | 18 | 60 | 2.00 | +| both (v1) | 75/207 = 36 % | 38/207 = 18 % | +0.18 | 0.39 / 0.35 / 0.36 | 126 | 21 | 40 | 87 | 2.35 | +| **overlap2** | **41/184 = 22 %** | 33/184 = 18 % | **+0.04** | 0.21 / 0.22 / 0.24 | 101 | 24 | 2 | **47** | 2.00 | +| both2 | 52/207 = 25 % | 38/207 = 18 % | +0.07 | 0.32 / 0.20 / 0.24 | 89 | 22 | 2 | 54 | 2.35 | +| both8 | 48/222 = 22 % | 46/222 = 21 % | +0.01 | 0.23 / 0.23 / 0.19 | 120 | 42 | 0 | 54 | 2.18 | + +### Per file (summed over the 3 placements) + +| file | variant | damaged cuts | phantom | near S/I/D | near words | else words/min | drop | dup | events near | events else/min | +|---|---|---|---|---|---|---|---|---|---|---| +| p1 | fixed | 27/57 | 11/57 | 21/27/9 | 57 | 4.16 | 9 | 10 | 29 | 2.15 | +| p1 | pause | 19/64 | 8/64 | 20/3/25 | 48 | 5.62 | 25 | 0 | 20 | 2.39 | +| p1 | overlap | 20/59 | 16/59 | 16/26/4 | 46 | 4.66 | 4 | 8 | 23 | 2.07 | +| p1 | both | 29/67 | 13/67 | 23/25/4 | 52 | 5.70 | 4 | 17 | 35 | 2.56 | +| p1 | overlap2 | 15/59 | 16/59 | 16/19/4 | 39 | 4.66 | 4 | 0 | 17 | 2.07 | +| p1 | both2 | 21/67 | 13/67 | 23/8/5 | 36 | 5.70 | 5 | 0 | 21 | 2.56 | +| p1 | both8 | 18/70 | 15/70 | 16/2/8 | 26 | 6.22 | 8 | 0 | 18 | 2.40 | +| p2 | fixed | 14/36 | 4/36 | 10/7/14 | 31 | 4.45 | 14 | 2 | 16 | 1.39 | +| p2 | pause | 10/40 | 6/40 | 11/1/1 | 13 | 2.71 | 1 | 0 | 10 | 1.56 | +| p2 | overlap | 6/36 | 6/36 | 4/2/0 | 6 | 3.56 | 0 | 2 | 6 | 1.47 | +| p2 | both | 16/41 | 6/41 | 13/10/0 | 23 | 3.11 | 0 | 8 | 18 | 1.64 | +| p2 | overlap2 | 4/36 | 6/36 | 4/0/0 | 4 | 3.56 | 0 | 0 | 4 | 1.47 | +| p2 | both2 | 10/41 | 6/41 | 13/2/0 | 15 | 3.11 | 0 | 0 | 11 | 1.64 | +| p2 | both8 | 9/43 | 7/43 | 9/1/1 | 11 | 3.17 | 1 | 0 | 9 | 1.58 | +| scotus | fixed | 27/47 | 15/47 | 14/49/21 | 84 | 9.47 | 21 | 4 | 30 | 3.34 | +| scotus | pause | 18/54 | 17/54 | 13/9/10 | 32 | 8.63 | 10 | 0 | 22 | 3.36 | +| scotus | overlap | 20/49 | 6/49 | 16/23/11 | 50 | 9.43 | 11 | 5 | 24 | 3.01 | +| scotus | both | 19/55 | 14/55 | 9/13/5 | 27 | 9.46 | 5 | 8 | 21 | 3.39 | +| scotus | overlap2 | 18/49 | 6/49 | 15/20/10 | 45 | 9.43 | 10 | 2 | 21 | 3.01 | +| scotus | both2 | 15/55 | 14/55 | 9/7/5 | 21 | 9.46 | 5 | 2 | 15 | 3.39 | +| scotus | both8 | 14/61 | 19/61 | 14/31/10 | 55 | 8.19 | 10 | 0 | 19 | 3.30 | +| wilde | fixed | 25/39 | 4/39 | 9/30/12 | 51 | 6.91 | 12 | 1 | 27 | 1.03 | +| wilde | pause | 6/43 | 5/43 | 4/0/16 | 20 | 6.95 | 16 | 0 | 7 | 1.23 | +| wilde | overlap | 6/40 | 5/40 | 2/3/10 | 15 | 6.74 | 10 | 3 | 7 | 1.13 | +| wilde | both | 11/44 | 5/44 | 4/8/12 | 24 | 9.43 | 12 | 7 | 13 | 1.43 | +| wilde | overlap2 | 4/40 | 5/40 | 3/0/10 | 13 | 6.74 | 10 | 0 | 5 | 1.13 | +| wilde | both2 | 6/44 | 5/44 | 4/1/12 | 17 | 9.43 | 12 | 0 | 7 | 1.43 | +| wilde | both8 | 7/48 | 5/48 | 5/0/23 | 28 | 6.80 | 23 | 0 | 8 | 1.04 | + +### Where the remaining near-cut errors sit + +Signed distance from the cut, pooled over files: upstream's errors pile up +within ±0.5 s (chopped words; at 100 s, 24 insertions in [−0.5, 0) alone). +Pause-only leaves deletions 1–3 s **before** its cuts, where the left chunk ends +(8/9, 3/4, 7/8 in [−3, −1) at the three lengths). overlap2's residue is spread +evenly over ±3 s, which is what background looks like. both2 and both8 are clean +at the handover but keep some of pause-only's pre-cut deletions. + +### Controls + +- **A-vs-A:** the same variant twice gives byte-identical words, segments and + text (fixed-120, both-120, both2-120 on all 4 files, in one process), and the + shipped script run three times as separate CLI processes matches the harness + output byte for byte (p1 and scotus). Decoding is deterministic, so run-to-run + variance is zero; the variance that matters comes from cut placement. +- **Harness validity:** patched code with overlap and pause search off equals + upstream's unmodified script on all 4 files; v2 without overlap equals v1. +- **Positive control:** upstream's fixed cutter damages 52 % of its cuts + against a 19 % background (+0.33, about 7 standard errors), per file 47–64 % + against 11–25 %. The instrument sees cut damage. +- **Negative control:** the phantom (background) rate is 18–21 % for every + variant, and error events away from cuts run at 2.0–2.35 per minute for all + of them. **Word counts away from cuts do not agree across slicers** (see + "Parakeet drops stretches" below), which is why they are not the decision metric. +- **Sensitivity floor:** at ~180–220 cuts per variant the 2-standard-error band + on a difference of damaged-cut rates is **±0.08** pooled and **±0.17** for + one file. No slicer can be measured below the background (~18–21 %). Word-count + differences under ~60 words are unresolvable, because a single dropped stretch + (next section) is 10–60 words. + +### Verdict + +Every variant that overlaps with the v2 handover, and pause-only, beats +upstream by 0.26–0.30 in damaged-cut rate, far outside the ±0.08 floor. Among +them the differences are **inside** the floor. overlap2 ships: tied best on +damaged cuts, fewest near-cut error events (47 vs 102), near-zero duplicates, +nothing concentrated at the handover, and the simplest mechanism. Pause search +measured neutral once the stitch was fixed, so it stays in the patch as an +opt-in `--pause-search`, off by default. + +## Memory (GPU 3, production invocation, nvidia-smi every 0.2 s) + +| run | n | per-process peak | +|---|---|---| +| upstream script, 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB (reproduces the 2026-09-30 budget figure) | +| shipped script (v2), 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB | +| shipped script (v3, final), 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB; all 3 CLI outputs byte-identical to the harness | +| shipped script, 120 s, scotus (public) | 3 | 5,496 / 5,496 / 5,496 MiB | +| positive control: shipped script, **300 s**, scotus | 1 | 7,056 MiB (deterministic; 300 s was n=3 in the earlier table) | + +Zero spread; the peak is set by the 120 s maximum, and the overlap sits inside +it. The scotus file reads the same peak as p1, so the rebuild script uses it +as its public memory fixture. A spike shorter than the 0.2 s sample period +could be missed. + +## Parakeet drops stretches of speech (separate finding, not the slicer) + +Every chunked variant, **upstream's included**, sometimes skips a run of ≥10 +consecutive words in the middle of a chunk, and which runs it skips changes +chaotically with the cut placement. Wilde at 120 s loses nothing, at 110 s +171 words, and overlap2 at 110 s loses one 60-second stretch. Across 4 files × 3 +placements: + +| variant | runs of ≥10 words lost | words lost | +|---|---|---| +| fixed (upstream) | 15 | 599 | +| pause | 17 | 511 | +| overlap2 | 12 | 656 | +| both2 | 14 | 719 | +| both8 | 12 | 560 | + +The whole-file local-attention reference does the same: runs where every +slicer has words the reference lacks (13–16 per 12 comparisons, 340–400 words). +Today's production setting (upstream, 120 s) lost 85 words in 2 runs on p2. The +slicer neither causes nor cures it; it needs its own investigation (decoder +settings, chunk length, or model), which was out of scope here. diff --git a/scripts/scriberr-rebuild b/scripts/scriberr-rebuild new file mode 100755 index 0000000..fe96287 --- /dev/null +++ b/scripts/scriberr-rebuild @@ -0,0 +1,263 @@ +#!/usr/bin/env bash +# scriberr-rebuild — rebuild Scriberr's Blackwell image at a PINNED upstream sha +# with our patches applied, then prove the result before anyone deploys it. +# +# Runs on nh3-dev and drives fv-ml1 over ssh. Deploy is a SEPARATE manual step +# (stacks/scriberr/patches/README.md § Deploy); this script never touches the +# live container, its .env, or GPU 1. +# +# Why this exists: upstream publishes no sm_120 image, so we build from source, +# and we carry a patch to the Parakeet slicer (stacks/scriberr/patches/). Upstream +# moves slowly, so an upgrade should be one command plus a verdict. +# +# Stages (each prints PASS/FAIL; the first FAIL stops the run): +# clone clean shallow clone of upstream at the pinned sha, in a NEW dir +# /opt/docker/src/scriberr-- (reused only if it already +# holds that sha with every patch applied) +# patch `git apply --check` then `git apply`, patch by patch; a conflict +# stops the run loudly and names the patch +# build docker build -f Dockerfile.cuda.12.9 -t scriberr:local-blackwell-- +# — a DISTINCT tag, so the running image is never overwritten +# embed the patched script's exact bytes are inside the new Go binary +# (Scriberr rewrites the env's copy from this embed on every start) +# unit the slicer's pure-function tests, under the live env's numpy/librosa +# seam patched script, production invocation, short fixture at +# --chunk-len 10, JSON validated against the Go struct +# memory same on a long recording at --chunk-len 120 on an IDLE GPU, +# nvidia-smi sampled every 0.2 s; per-process peak <= the budget +# +# Usage: +# scripts/scriberr-rebuild [--sha SHA40] [--suffix NAME] [--gpu N] +# [--budget MIB] [--memory-audio PATH] [--reuse-image] +# +# Defaults: --sha PINNED_SHA below, --suffix slicer1, --gpu 3, --budget 5496, +# --memory-audio the public 30-min SCOTUS fixture. The GPU must be idle +# (< 100 MiB used), which in practice means GPU 3; GPUs 0-2 run live seats. +set -euo pipefail + +PINNED_SHA=a353078fd96b8aca4002681813524b7397c90df1 # upstream HEAD 2026-09-20 +UPSTREAM=https://github.com/rishikanthc/Scriberr.git +HOST=${SCRIBERR_REBUILD_HOST:-infra-ops@10.251.50.54} +ENV_DIR=/tank/scriberr/whisperx-env # live env, always mounted READ-ONLY +TOOLS=/opt/docker/src/scriberr-rebuild # fixtures + seam checker on fv-ml1 +SCRIPT_REL=internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py +TEST_REL=internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py +SEAM_AUDIO_REL=tests/data/AMI-Corpus-IB4002.Mix-Headset-clip.wav + +SHA=$PINNED_SHA SUFFIX=slicer1 GPU=3 BUDGET=5496 REUSE_IMAGE=0 +MEM_AUDIO=$TOOLS/fixtures/scotus-22-451-first30m.wav +while [ $# -gt 0 ]; do + case $1 in + --sha) SHA=$2; shift 2 ;; + --suffix) SUFFIX=$2; shift 2 ;; + --gpu) GPU=$2; shift 2 ;; + --budget) BUDGET=$2; shift 2 ;; + --memory-audio) MEM_AUDIO=$2; shift 2 ;; + --reuse-image) REUSE_IMAGE=1; shift ;; + -h|--help) sed -n '2,/^set -euo/p' "$0" | sed '$d; s/^# \{0,1\}//'; exit 0 ;; + *) echo "unknown argument: $1 (see --help)" >&2; exit 2 ;; + esac +done +[[ $SHA =~ ^[0-9a-f]{40}$ ]] || { echo "--sha must be a full 40-char sha (GitHub fetches by full sha)" >&2; exit 2; } +[[ $SUFFIX =~ ^[a-z0-9][a-z0-9.-]*$ ]] || { echo "--suffix must be [a-z0-9.-]" >&2; exit 2; } +[[ $GPU =~ ^[0-9]+$ && $BUDGET =~ ^[0-9]+$ ]] || { echo "--gpu and --budget must be integers" >&2; exit 2; } + +REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +PATCH_DIR=$REPO_ROOT/stacks/scriberr/patches +mapfile -t PATCHES < <(find "$PATCH_DIR" -maxdepth 1 -name '*.patch' | sort) +[ ${#PATCHES[@]} -gt 0 ] || { echo "no patches in $PATCH_DIR" >&2; exit 2; } +# The only paths a reused build dir may differ from upstream in. +mapfile -t PATCHED_PATHS < <(sed -n 's#^+++ b/##p' "${PATCHES[@]}" | sort -u) +SHA7=${SHA:0:7} +TAG=scriberr:local-blackwell-$SHA7-$SUFFIX +BUILD_DIR=/opt/docker/src/scriberr-$SHA7-$SUFFIX +PATCH_SUM=$(cat "${PATCHES[@]}" | sha256sum | cut -c1-16) +CNAME=scriberr-rebuild-$SHA7-$SUFFIX # every GPU container, so cleanup can find it +SAMPLES=/tmp/$CNAME.mem.csv +BUILD_LOG=/tmp/$CNAME.build.log +MIN_FREE_GB=20 # under Docker's root dir (zroot); an image adds ~0.1-6 GB + +# One multiplexed ssh connection serves the ~15 remote calls of a run. +CM=(-o BatchMode=yes -o ControlMaster=auto -o "ControlPath=$HOME/.ssh/cm-%C" -o ControlPersist=120) +SSH=(ssh "${CM[@]}" "$HOST") +SCP=(scp -q "${CM[@]}") + +RESULTS=() +pass() { RESULTS+=("PASS $1 $2"); echo "== PASS $1: $2"; } +fail() { + RESULTS+=("FAIL $1 $2"); echo "== FAIL $1: $2" >&2 + summary; exit 1 +} +summary() { + echo; echo "scriberr-rebuild upstream=$SHA7 patches=$PATCH_SUM tag=$TAG" + printf ' %s\n' "${RESULTS[@]}" +} +record() { # best effort: the audit trail must not become a failure point + "$REPO_ROOT/scripts/ops-log" record --host fv-ml1 --action "$1" --target "$2" \ + --detail "$3" >/dev/null 2>&1 || echo "(ops-log record failed; continuing)" >&2 +} +remote() { "${SSH[@]}" bash -s -- "$@"; } + +# Whatever happens (a FAIL, Ctrl-C, a dropped connection), never leave the +# nvidia-smi sampler or a transient GPU container behind on fv-ml1. +cleanup() { + "${SSH[@]}" "[ -s $SAMPLES.pid ] && kill \$(cat $SAMPLES.pid) 2>/dev/null; : >$SAMPLES.pid; \ + docker rm -f $CNAME >/dev/null 2>&1; true" 2>/dev/null || true +} +trap cleanup EXIT +trap 'exit 130' INT TERM +# An unguarded remote call that fails must still say so and print the table. +set -E +trap 'echo "== ABORT: unexpected failure at line $LINENO (see output above)" >&2; summary' ERR + +# Every container run: the new image, the live env READ-ONLY, the build tree +# read-only, and nothing else writable but the container's own /tmp. +DOCKER_RUN="docker run --rm --user 1002:1003 -e HOME=/tmp -e USER=infra-ops -e LOGNAME=infra-ops \ + -e PYTHONDONTWRITEBYTECODE=1 -e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True -e UV_LINK_MODE=copy \ + -v $ENV_DIR:/app/whisperx-env:ro -v $BUILD_DIR:/src:ro -v $TOOLS:/tools:ro --entrypoint bash" +UVRUN="uv run --native-tls --project /app/whisperx-env/parakeet" + +# ── clone ────────────────────────────────────────────────────────────────── +if out=$(remote "$BUILD_DIR" "$UPSTREAM" "$SHA" "$TOOLS" "${PATCHED_PATHS[@]}" 2>&1 <<'EOF' +set -euo pipefail +dir=$1 upstream=$2 sha=$3 tools=$4 +shift 4 +sudo -n install -d -o infra-ops -g infra-ops "$tools" "$tools/fixtures" | cat +if [ -e "$dir" ]; then + [ "$(git -C "$dir" rev-parse HEAD 2>/dev/null)" = "$sha" ] \ + || { echo "$dir exists but is not a checkout of $sha; remove it by hand (sudo -n rm -rf $dir) or pick --suffix"; exit 1; } + extra=$({ git -C "$dir" diff --name-only HEAD; git -C "$dir" ls-files --others --exclude-standard; } \ + | sort -u | grep -vxF -f <(printf '%s\n' "$@") || true) + [ -z "$extra" ] || { echo "$dir has changes outside the patches ($extra); remove it by hand or pick --suffix"; exit 1; } + echo "reusing $dir" + exit 0 +fi +sudo -n install -d -o infra-ops -g infra-ops "$dir" | cat +cd "$dir" +git init -q +git remote add origin "$upstream" +git fetch -q --depth 1 origin "$sha" \ + || { echo "fetch of $sha failed; $dir is left empty, remove it by hand (sudo -n rm -rf $dir)"; exit 1; } +git checkout -q --detach FETCH_HEAD +[ "$(git rev-parse HEAD)" = "$sha" ] || { echo "checked out $(git rev-parse HEAD), wanted $sha"; exit 1; } +echo "cloned $sha into $dir" +EOF +); then pass clone "$out"; else fail clone "$out"; fi +if [[ $out == reusing* ]]; then REUSED=1; else REUSED=0; record create "$BUILD_DIR" "scriberr-rebuild: clean clone of upstream $SHA7"; fi + +# ── patch ────────────────────────────────────────────────────────────────── +for p in "${PATCHES[@]}"; do + name=$(basename "$p") + # Already applied (a reused dir)? `apply --reverse --check` succeeds only then. + if "${SSH[@]}" "cd $BUILD_DIR && git apply --reverse --check -" <"$p" >/dev/null 2>&1; then + pass patch "$name already applied" + continue + fi + if ! out=$("${SSH[@]}" "cd $BUILD_DIR && git apply --check -" <"$p" 2>&1); then + if [ "$REUSED" = 1 ]; then + fail patch "$name does not apply to the REUSED $BUILD_DIR, which holds an older state of the patches; pick a new --suffix or remove the dir by hand (sudo -n rm -rf $BUILD_DIR): +$out" + fi + fail patch "$name DOES NOT APPLY to upstream $SHA7 — rebase the patch before upgrading: +$out" + fi + "${SSH[@]}" "cd $BUILD_DIR && git apply -" <"$p" || fail patch "$name: git apply failed after a clean --check" + record patch "$BUILD_DIR" "scriberr-rebuild: git apply $name" + pass patch "$name applied" +done + +# ── build ────────────────────────────────────────────────────────────────── +if "${SSH[@]}" "docker image inspect $TAG >/dev/null 2>&1"; then + [ "$REUSE_IMAGE" = 1 ] || fail build "$TAG already exists; pass --reuse-image to test it, or pick a new --suffix" + pass build "reusing existing $TAG" +else + free=$("${SSH[@]}" "df -BG --output=avail \$(docker info -f '{{.DockerRootDir}}') | tail -1") \ + || fail build "could not read free space on fv-ml1" + free=${free//[!0-9]/} + [ "${free:-0}" -ge "$MIN_FREE_GB" ] \ + || fail build "only ${free:-?} GB free under Docker's root dir, need $MIN_FREE_GB; remove superseded scriberr tags first (patches/README.md)" + echo "building $TAG (several minutes; log on fv-ml1 at $BUILD_LOG)" + if "${SSH[@]}" "docker build -f $BUILD_DIR/Dockerfile.cuda.12.9 -t $TAG \ + --label org.phasefinal.scriberr.upstream=$SHA --label org.phasefinal.scriberr.patches=$PATCH_SUM \ + $BUILD_DIR >$BUILD_LOG 2>&1"; then + record build "$TAG" "scriberr-rebuild: upstream $SHA7 + patches $PATCH_SUM; old images kept" + pass build "$TAG ($free GB was free)" + else + "${SSH[@]}" "tail -25 $BUILD_LOG" >&2 || true + fail build "docker build failed (log tail above)" + fi +fi + +# ── embed ────────────────────────────────────────────────────────────────── +if out=$(remote "$TAG" "$BUILD_DIR" "$SCRIPT_REL" 2>&1 <<'EOF' +docker run --rm -v "$2":/src:ro --entrypoint python3 "$1" -c " +import sys +script = open('/src/$3', 'rb').read() +sys.exit(0 if script in open('/app/scriberr', 'rb').read() else 1)" +EOF +); then + pass embed "patched $(basename "$SCRIPT_REL") is byte-identical inside /app/scriberr" +else + fail embed "the binary does not embed the patched script ${out:+($out)}" +fi + +# ── unit ─────────────────────────────────────────────────────────────────── +if out=$("${SSH[@]}" "$DOCKER_RUN $TAG -c 'cd /tmp && $UVRUN --with pytest \ + python -m pytest -q -p no:cacheprovider /src/$TEST_REL 2>&1 | tail -3'" 2>&1) \ + && grep -q ' passed' <<<"$out" && ! grep -Eq 'failed|error' <<<"$out"; then + pass unit "$(tail -1 <<<"$out")" +else + fail unit "$out" +fi + +# The checker travels with this script, so the host copy is refreshed each run. +"${SCP[@]}" "$REPO_ROOT/scripts/scriberr-seam-check.py" "$HOST:$TOOLS/seam-check.py" \ + || fail seam "could not copy the seam checker to fv-ml1:$TOOLS" +record update "$TOOLS/seam-check.py" "scriberr-rebuild: refreshed the seam checker" + +# One GPU run of the patched script under the production invocation, output +# kept inside the container, validated there. Prints the seam checker's line. +gpu_run() { # $1 audio path on host, $2 --chunk-len, $3 --min-chunks + local audio_dir; audio_dir=$(dirname "$1") + "${SSH[@]}" "$DOCKER_RUN --name $CNAME --gpus '\"device=$GPU\"' -e NVIDIA_VISIBLE_DEVICES=$GPU \ + -v $audio_dir:/audio:ro $TAG -c 'cd /tmp && $UVRUN python /src/$SCRIPT_REL /audio/$(basename "$1") \ + --output /tmp/out.json --chunk-len $2 >/tmp/run.log 2>&1 || { tail -5 /tmp/run.log; exit 1; }; \ + python3 /tools/seam-check.py /tmp/out.json --min-chunks $3'" +} +gpu_idle() { + local used + used=$("${SSH[@]}" "nvidia-smi -i $GPU --query-gpu=memory.used --format=csv,noheader,nounits") \ + || fail "$1" "could not read GPU $GPU memory on fv-ml1" + used=${used//[!0-9]/} + [ -n "$used" ] && [ "$used" -lt 100 ] \ + || fail "$1" "GPU $GPU is not idle (${used:-?} MiB used); refusing to share a live card" +} + +# ── seam ─────────────────────────────────────────────────────────────────── +gpu_idle seam +if out=$(gpu_run "$BUILD_DIR/$SEAM_AUDIO_REL" 10 2 2>&1); then pass seam "$(tail -1 <<<"$out")" +else fail seam "$out"; fi + +# ── memory ───────────────────────────────────────────────────────────────── +"${SSH[@]}" "test -s $MEM_AUDIO" || fail memory "memory audio $MEM_AUDIO not found on fv-ml1" +gpu_idle memory +"${SSH[@]}" "nohup nvidia-smi -i $GPU --query-compute-apps=pid,used_memory \ + --format=csv,noheader,nounits -lms 200 $SAMPLES 2>/dev/null & echo \$! >$SAMPLES.pid" \ + || fail memory "could not start the nvidia-smi sampler" +record run "$CNAME" "transient $TAG on GPU $GPU, --chunk-len 120, env ro; removed on exit" +rc=0; out=$(gpu_run "$MEM_AUDIO" 120 2 2>&1) || rc=$? # `||`, not set +e: keeps the ERR trap quiet +"${SSH[@]}" "kill \$(cat $SAMPLES.pid) 2>/dev/null; : >$SAMPLES.pid" || true +read -r pids peak < <("${SSH[@]}" \ + "awk -F', *' 'NF==2 {if (!(\$1 in p)) {p[\$1]=1; n++}; if (\$2+0>m) m=\$2+0} END {print n+0, m+0}' $SAMPLES") \ + || fail memory "no GPU samples could be read back from $SAMPLES" +[ "$rc" = 0 ] || fail memory "memory run failed: $out" +[ "$pids" = 1 ] || fail memory "saw $pids processes on GPU $GPU during the run; the peak is not attributable" +if [ "$peak" -le "$BUDGET" ]; then + pass memory "peak $peak MiB <= budget $BUDGET MiB (0.2 s samples, GPU $GPU); $(tail -1 <<<"$out")" +else + fail memory "peak $peak MiB > budget $BUDGET MiB — do NOT deploy beside intern-decision" +fi + +summary +echo +echo "VERDICT: PASS. Deploy is manual: stacks/scriberr/patches/README.md § Deploy (SCRIBERR_IMAGE=$TAG)." diff --git a/scripts/scriberr-seam-check.py b/scripts/scriberr-seam-check.py new file mode 100755 index 0000000..18ef4ef --- /dev/null +++ b/scripts/scriberr-seam-check.py @@ -0,0 +1,86 @@ +#!/usr/bin/env python3 +"""Validate a parakeet_transcribe_buffered.py result against the Go seam. + +Scriberr's parakeet_adapter.go (parseResult) unmarshals this JSON into a struct +with typed fields; a float where Go expects an int, or a missing key, fails the +job. This checks the shape Go reads plus the stitching invariants the slicer +patch promises. Stdlib only, so it runs under any python3. Prints counts, never +transcript text. + +usage: scriberr-seam-check.py RESULT.json [--min-chunks N] +""" +import argparse +import json +import sys + +NUMBER = (int, float) + + +def fail(msg): + print(f"SEAM FAIL: {msg}") + sys.exit(1) + + +def check_items(items, text_key, name): + for i, item in enumerate(items): + if not isinstance(item, dict): + fail(f"{name}[{i}] is not an object") + if not isinstance(item.get(text_key), str): + fail(f"{name}[{i}].{text_key} is not a string") + for key in ("start_offset", "end_offset"): + if type(item.get(key)) is not int: + fail(f"{name}[{i}].{key} is not an integer (Go field is int)") + for key in ("start", "end"): + if not isinstance(item.get(key), NUMBER) or isinstance(item.get(key), bool): + fail(f"{name}[{i}].{key} is not a number") + if item["start"] > item["end"]: + fail(f"{name}[{i}] starts after it ends") + + +def main(): + parser = argparse.ArgumentParser(description="Validate a buffered Parakeet result for Go.") + parser.add_argument("result", help="result JSON written by parakeet_transcribe_buffered.py") + parser.add_argument("--min-chunks", type=int, default=1, + help="fail unless the run used at least this many chunks") + args = parser.parse_args() + min_chunks = args.min_chunks + try: + data = json.load(open(args.result, encoding="utf-8")) + except (OSError, ValueError) as e: + fail(f"cannot read {args.result}: {e}") + if not isinstance(data, dict): + fail("the result is not a JSON object") + + required = {"transcription": str, "language": str, "word_timestamps": list, + "segment_timestamps": list, "audio_file": str, "model": str} + for key, kind in required.items(): + if not isinstance(data.get(key), kind): + fail(f"'{key}' missing or not {kind.__name__}") + if data.get("buffered") is not True: + fail("'buffered' is not true") + if not isinstance(data.get("chunk_duration_secs"), NUMBER): + fail("'chunk_duration_secs' is not a number") + if type(data.get("num_chunks")) is not int or data["num_chunks"] < min_chunks: + fail(f"'num_chunks' is not an integer >= {min_chunks}") + + words, segments = data["word_timestamps"], data["segment_timestamps"] + if not words or not data["transcription"].strip(): + fail("empty transcript") + check_items(words, "word", "word_timestamps") + check_items(segments, "segment", "segment_timestamps") + + starts = [w["start"] for w in words] + if starts != sorted(starts): + fail("word start times go backwards (a stitch repeated or reordered words)") + joined = " ".join(w["word"] for w in words) + if data["transcription"] != joined: + fail("'transcription' is not the stitched words joined by spaces") + if " ".join(s["segment"] for s in segments) != joined: + fail("segments do not cover the stitched words exactly once, in order") + + print(f"SEAM OK: {len(words)} words, {len(segments)} segments, " + f"{data['num_chunks']} chunks, cuts at {len(data.get('cut_times', []))} points") + + +if __name__ == "__main__": + main() diff --git a/stacks/scriberr/.env.example b/stacks/scriberr/.env.example index 5e24d73..3af4825 100644 --- a/stacks/scriberr/.env.example +++ b/stacks/scriberr/.env.example @@ -4,6 +4,11 @@ # ── Image ──────────────────────────────────────────────────────────────── # Built locally from Dockerfile.cuda.12.9 — see the compose header for why # the published scriberr-cuda image is NOT usable on these Blackwell cards. +# Patched builds come from scripts/scriberr-rebuild and are tagged +# scriberr:local-blackwell-- (e.g. -a353078-slicer1) +# so every build keeps its own tag. Deploying = pointing this at a new tag; +# rollback = pointing it back. See stacks/scriberr/patches/README.md. +# (Unset falls back to the original unpatched scriberr:local-blackwell.) SCRIBERR_IMAGE=scriberr:local-blackwell # ── Network ────────────────────────────────────────────────────────────── diff --git a/stacks/scriberr/README.md b/stacks/scriberr/README.md index 9b1a485..ad50b10 100644 --- a/stacks/scriberr/README.md +++ b/stacks/scriberr/README.md @@ -27,15 +27,21 @@ image** — it will fail on these cards or quietly fall back to CPU. ### Rebuilding +We carry local patches (`patches/`, currently the pause-aware Parakeet +slicer), so a rebuild is one command from nh3-dev, pinned to an upstream sha: + ```bash -ssh fv-ml1 -cd /tank/scriberr/src/Scriberr -git pull -docker build -f Dockerfile.cuda.12.9 -t scriberr:local-blackwell . -cd /opt/docker/compose/scriberr && docker compose up -d +scripts/scriberr-rebuild --sha --suffix slicer1 ``` -Source checkout lives on `/tank`, not the root pool — see storage below. +It makes a clean clone in `/opt/docker/src/scriberr--` on +fv-ml1, `git apply --check`s the patches (a conflict stops it), builds +`scriberr:local-blackwell--` beside the old images, and checks +the embed, the unit tests, the Go↔Python JSON seam, and the GPU memory budget. +Deploying it is a separate manual step: `patches/README.md` § Deploy. + +The old checkout at `/tank/scriberr/src/Scriberr` (lkraven-owned, shallow) +built the original `scriberr:local-blackwell` and is left as it was. ## Deploy @@ -62,7 +68,7 @@ mounts are bind-mounted onto `/tank` (4+ TB) instead of named volumes: |---|---|---| | `/tank/scriberr/data` | `/app/data` | SQLite DB, uploads, transcripts | | `/tank/scriberr/whisperx-env` | `/app/whisperx-env` | Python env + model weights | -| `/tank/scriberr/src/Scriberr` | — | build checkout | +| `/tank/scriberr/src/Scriberr` | — | original build checkout (patched builds: `/opt/docker/src/scriberr--`) | Both are owned by uid/gid 1000 to match `PUID`/`PGID`. @@ -164,3 +170,11 @@ Peak GPU memory on a 35-minute file: - The two env knobs are read by upstream's Go code (`parakeet_adapter.go`), so they survive image upgrades for as long as upstream keeps them. Re-measure the peak after any upgrade. +- **The slicer itself is patched** (`patches/0001-parakeet-pause-aware-slicer.patch`, + 2026-09-30). Adjacent 120 s slices now overlap by 4 s and are stitched at a word + both transcribed, which cut the share of cuts with an error nearby from 52 % to + 22 % against a 19 % background. The overlap sits *inside* the 120 s, so the + peak is unchanged (5,496 MiB, n=3). See `patches/README.md` and + `docs/pfi/scriberr-slicer-bench-2026-09-30.md`. That bench also found that + Parakeet sometimes skips stretches of ≥10 words mid-slice, with or without the + patch; that is still open. diff --git a/stacks/scriberr/patches/0001-parakeet-pause-aware-slicer.patch b/stacks/scriberr/patches/0001-parakeet-pause-aware-slicer.patch new file mode 100644 index 0000000..20bbf1c --- /dev/null +++ b/stacks/scriberr/patches/0001-parakeet-pause-aware-slicer.patch @@ -0,0 +1,686 @@ +From 2dafe7ce81c217609ff2c8616e43b0d72255e170 Mon Sep 17 00:00:00 2001 +From: Vuong Hoang +Date: Wed, 30 Sep 2026 11:41:07 -0700 +Subject: [PATCH] fix(parakeet): overlap buffered chunks and stitch at an + agreed word + +parakeet_transcribe_buffered.py cut long audio at fixed --chunk-len marks +with no overlap, so a word straddling a mark was chopped, dropped or +transcribed twice. Adjacent chunks now overlap by --overlap seconds +(default 4, counted inside --chunk-len so no chunk grows), and in each +overlap the chunks hand over at the word nearest the cut that both +transcribed with the same text at nearly the same time (within 0.5 s), +keeping whichever copy leaves the words in time order. With no such word +they split at the cut. Splitting both chunks at the cut by word start time +is not enough on its own: a word after a pause can be timestamped anywhere +in the pause, so the two chunks may place it on opposite sides of the cut. + +--pause-search N (opt-in) also moves each cut back to the quietest 0.3 s +in the last N seconds before the limit. Measured neutral on top of the +overlap, so it is off by default. + +The CLI and JSON the Go adapter reads are unchanged; the new flags are +optional, and the JSON gains overlap_secs, pause_search_secs and +cut_times. --overlap 0 reproduces the previous output exactly. NeMo is +now imported inside transcribe_buffered() so the slicing and stitching +helpers can be unit-tested without a GPU. If a chunk ever returns text +without word timestamps, its text is kept rather than dropped. +--- + .../py/nvidia/parakeet_transcribe_buffered.py | 221 ++++++++++-- + .../py/nvidia/tests/test_parakeet_slicing.py | 329 ++++++++++++++++++ + 2 files changed, 525 insertions(+), 25 deletions(-) + create mode 100644 internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py + +diff --git a/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py b/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py +index 29d5047..ba755c1 100644 +--- a/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py ++++ b/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py +@@ -2,6 +2,10 @@ + """ + NVIDIA Parakeet buffered inference for long audio files. + Splits audio into chunks to avoid GPU memory issues. ++ ++Adjacent chunks overlap slightly, and in each overlap the chunks hand over at ++a word both transcribed alike, so a word near a cut is neither chopped, dropped ++nor repeated. Optionally (--pause-search) each cut also moves into a pause. + """ + + import argparse +@@ -12,37 +16,178 @@ import librosa + import soundfile as sf + import numpy as np + from pathlib import Path +-import nemo.collections.asr as nemo_asr + ++DEFAULT_OVERLAP_SECS = 4.0 ++DEFAULT_PAUSE_SEARCH_SECS = 0.0 # opt-in; measured no gain on top of the overlap ++QUIET_WINDOW_SECS = 0.3 ++SAME_WORD_SECS = 0.5 ++ ++ ++def plan_slices(audio, sr, max_chunk_secs, overlap_secs=0.0, search_secs=0.0): ++ """Choose where to cut `audio` so no chunk exceeds `max_chunk_secs`. + +-def split_audio_file(audio_path, chunk_duration_secs=300): +- """Split audio file into chunks of specified duration.""" ++ Returns (spans, cuts): `cuts` are the sample indices where one chunk's ++ share of the audio ends and the next one's begins; `spans` are the ++ (start, end) samples actually transcribed, each cut-to-cut range widened ++ by half the overlap on both sides. The overlap counts towards the limit. ++ ++ With `search_secs` > 0, each cut moves back from the limit to the middle ++ of the quietest QUIET_WINDOW_SECS window within the last `search_secs`. ++ """ ++ overlap_secs = max(0.0, min(overlap_secs, max_chunk_secs / 4)) ++ step = int((max_chunk_secs - overlap_secs) * sr) ++ if step < 1: ++ raise ValueError(f"chunk length must be positive, got {max_chunk_secs}s") ++ search = min(int(search_secs * sr), step // 2) ++ window = max(1, int(QUIET_WINDOW_SECS * sr)) ++ ++ cuts = [] ++ position = 0 ++ while len(audio) - position > step: ++ cut = position + step ++ if search > window: ++ cut = _quietest_point(audio, cut - search, cut, window) ++ cuts.append(cut) ++ position = cut ++ ++ pad = int(overlap_secs * sr) // 2 ++ edges = [0] + cuts + [len(audio)] ++ spans = [(max(0, start - pad), min(len(audio), end + pad)) ++ for start, end in zip(edges, edges[1:])] ++ return spans, cuts ++ ++ ++def _quietest_point(audio, start, end, window): ++ """Middle of the lowest-energy `window` samples within audio[start:end]. ++ ++ Ties go to the latest window, which keeps chunks as long as allowed. ++ """ ++ x = audio[start:end].astype(np.float64) ++ cumulative = np.concatenate(([0.0], np.cumsum(x * x))) ++ energy = cumulative[window:] - cumulative[:-window] ++ latest_min = len(energy) - 1 - int(np.argmin(energy[::-1])) ++ return start + latest_min + window // 2 ++ ++ ++def stitch_slices(slice_results, cut_times, chunk_spans=None): ++ """Merge per-chunk (words, segments), already shifted to absolute time. ++ ++ `chunk_spans` gives each chunk's (start, end) in seconds; omit it when the ++ chunks do not overlap. Each chunk contributes the words between its two ++ handovers. A handover is at the cut, unless the chunks overlap: then it ++ moves to the nearest word in the overlap that both chunks transcribed ++ alike, at the same time, and the left chunk keeps the words before it, ++ the right chunk that word and the ones after. (Splitting both chunks at ++ the cut is fragile: a word that follows a pause can be timestamped ++ anywhere in the pause, so the two chunks may put it on opposite sides of ++ the cut and keep it twice, or not at all.) ++ ++ Segments are trimmed to the words their chunk keeps, and dropped if none. ++ """ ++ first = [0] * len(slice_results) ++ last = [len(chunk_words) for chunk_words, _ in slice_results] ++ for k, cut in enumerate(cut_times): ++ overlap = (chunk_spans[k + 1][0], chunk_spans[k][1]) if chunk_spans else (cut, cut) ++ last[k], first[k + 1] = _handover( ++ slice_results[k][0], slice_results[k + 1][0], cut, overlap ++ ) ++ ++ words, segments = [], [] ++ for (chunk_words, chunk_segments), lo, hi in zip(slice_results, first, last): ++ words.extend(chunk_words[lo:hi]) ++ for seg, (start, stop) in zip(chunk_segments, _segment_ranges(chunk_words, chunk_segments)): ++ kept = chunk_words[max(start, lo):min(stop, hi)] ++ if start == stop: # a segment without words: keep it where its chunk does ++ if lo <= start < hi: ++ segments.append(seg) ++ elif len(kept) == stop - start: ++ segments.append(seg) ++ elif kept: ++ segments.append({ ++ **seg, ++ "segment": " ".join(w["word"] for w in kept), ++ "start_offset": kept[0]["start_offset"], ++ "end_offset": kept[-1]["end_offset"], ++ "start": kept[0]["start"], ++ "end": kept[-1]["end"], ++ }) ++ return words, segments ++ ++ ++def _handover(left, right, cut, overlap): ++ """(i, j): the left chunk keeps left[:i] and the right chunk right[j:]. ++ ++ Anchors are words in the overlap that both chunks transcribed with the same ++ text at nearly the same time. At the anchor nearest the cut, the right ++ chunk's copy is kept, or the left chunk's if that is what keeps the words ++ in time order. Without an anchor, both chunks split at the cut. ++ """ ++ i = sum(1 for w in left if w["start"] < cut) ++ j = sum(1 for w in right if w["start"] < cut) ++ in_order = lambda a, b: a == 0 or b == len(right) or left[a - 1]["start"] <= right[b]["start"] ++ anchors = [] ++ for p, lw in enumerate(left): ++ text = _normalize(lw["word"]) ++ if not text or not overlap[0] <= lw["start"] < overlap[1]: ++ continue ++ partners = [q for q, rw in enumerate(right) if _normalize(rw["word"]) == text ++ and abs(rw["start"] - lw["start"]) <= SAME_WORD_SECS] ++ if partners: ++ q = min(partners, key=lambda q: abs(right[q]["start"] - lw["start"])) ++ options = [h for h in ((p, q), (p + 1, q + 1)) if in_order(*h)] ++ if options: ++ anchors.append((abs(lw["start"] + right[q]["start"] - 2 * cut), options[0])) ++ if anchors: ++ i, j = min(anchors)[1] ++ return i, j ++ ++ ++def _normalize(word): ++ return "".join(c for c in word.lower() if c.isalnum() or c == "'") ++ ++ ++def _segment_ranges(words, segments): ++ """[start, stop) word indices of each segment, matched in order by frame offsets.""" ++ ranges, i = [], 0 ++ for seg in segments: ++ while i < len(words) and words[i]["start_offset"] < seg["start_offset"]: ++ i += 1 ++ start = i ++ while i < len(words) and words[i]["end_offset"] <= seg["end_offset"]: ++ i += 1 ++ ranges.append((start, i)) ++ return ranges ++ ++ ++def split_audio_file(audio_path, chunk_duration_secs=300, overlap_secs=0.0, search_secs=0.0): ++ """Split audio file into chunks of at most chunk_duration_secs.""" + audio, sr = librosa.load(audio_path, sr=None, mono=True) +- total_duration = len(audio) / sr +- chunk_samples = int(chunk_duration_secs * sr) ++ spans, cuts = plan_slices(audio, sr, chunk_duration_secs, overlap_secs, search_secs) + + chunks = [] +- for start_sample in range(0, len(audio), chunk_samples): +- end_sample = min(start_sample + chunk_samples, len(audio)) ++ for start_sample, end_sample in spans: + chunk_audio = audio[start_sample:end_sample] +- start_time = start_sample / sr + chunks.append({ + 'audio': chunk_audio, +- 'start_time': start_time, ++ 'start_time': start_sample / sr, + 'duration': len(chunk_audio) / sr + }) + +- return chunks, sr ++ return chunks, sr, [cut / sr for cut in cuts] + + + def transcribe_buffered( + audio_path: str, + output_file: str = None, + chunk_duration_secs: float = 300, # 5 minutes default ++ overlap_secs: float = DEFAULT_OVERLAP_SECS, ++ pause_search_secs: float = DEFAULT_PAUSE_SEARCH_SECS, + ): + """ + Transcribe long audio by splitting into chunks and merging results. + """ ++ import nemo.collections.asr as nemo_asr ++ + # Determine model path + model_filename = "parakeet-tdt-0.6b-v3.nemo" + model_path = None +@@ -79,13 +224,15 @@ def transcribe_buffered( + asr_model.change_decoding_strategy(dec_cfg) + print("✓ CUDA graphs disabled successfully") + +- print(f"Splitting audio into {chunk_duration_secs}s chunks...") +- chunks, sr = split_audio_file(audio_path, chunk_duration_secs) ++ print(f"Splitting audio into chunks of at most {chunk_duration_secs}s " ++ f"(overlap {overlap_secs}s, pause search {pause_search_secs}s)...") ++ chunks, sr, cut_times = split_audio_file( ++ audio_path, chunk_duration_secs, overlap_secs, pause_search_secs ++ ) + print(f"Created {len(chunks)} chunks") + +- all_words = [] +- all_segments = [] +- full_text = [] ++ slice_results = [] ++ chunk_texts = [] + + for i, chunk_info in enumerate(chunks): + print(f"Transcribing chunk {i+1}/{len(chunks)} (duration: {chunk_info['duration']:.1f}s)...") +@@ -104,25 +251,26 @@ def transcribe_buffered( + + result_data = output[0] + chunk_text = result_data.text +- full_text.append(chunk_text) ++ chunk_texts.append(chunk_text) ++ chunk_words = [] ++ chunk_segments = [] + + # Extract and adjust timestamps + if hasattr(result_data, 'timestamp') and result_data.timestamp: +- chunk_words = result_data.timestamp.get("word", []) +- chunk_segments = result_data.timestamp.get("segment", []) +- + # Adjust timestamps by chunk start time +- for word in chunk_words: ++ for word in result_data.timestamp.get("word", []): + word_copy = dict(word) + word_copy['start'] += chunk_info['start_time'] + word_copy['end'] += chunk_info['start_time'] +- all_words.append(word_copy) ++ chunk_words.append(word_copy) + +- for segment in chunk_segments: ++ for segment in result_data.timestamp.get("segment", []): + seg_copy = dict(segment) + seg_copy['start'] += chunk_info['start_time'] + seg_copy['end'] += chunk_info['start_time'] +- all_segments.append(seg_copy) ++ chunk_segments.append(seg_copy) ++ ++ slice_results.append((chunk_words, chunk_segments)) + + print(f"Chunk {i+1} complete: {len(chunk_text)} characters") + +@@ -131,7 +279,15 @@ def transcribe_buffered( + if os.path.exists(chunk_path): + os.remove(chunk_path) + +- final_text = " ".join(full_text) ++ chunk_spans = [(c['start_time'], c['start_time'] + c['duration']) for c in chunks] ++ all_words, all_segments = stitch_slices(slice_results, cut_times, chunk_spans) ++ if any(text.strip() and not words for (words, _), text in zip(slice_results, chunk_texts)): ++ # A chunk came back without word timestamps, so there is nothing to ++ # stitch it by; keep its text rather than lose it. ++ print("Warning: a chunk has text but no word timestamps; joining chunk texts") ++ final_text = " ".join(chunk_texts) ++ else: ++ final_text = " ".join(w["word"] for w in all_words) + print(f"Transcription complete: {len(final_text)} characters total") + + output_data = { +@@ -144,6 +300,9 @@ def transcribe_buffered( + "buffered": True, + "chunk_duration_secs": chunk_duration_secs, + "num_chunks": len(chunks), ++ "overlap_secs": overlap_secs, ++ "pause_search_secs": pause_search_secs, ++ "cut_times": cut_times, + } + + if output_file: +@@ -162,7 +321,17 @@ def main(): + parser.add_argument("--output", "-o", help="Output file path", required=True) + parser.add_argument( + "--chunk-len", type=float, default=300, +- help="Chunk duration in seconds (default: 300 = 5 minutes)" ++ help="Maximum chunk duration in seconds, overlap included (default: 300 = 5 minutes)" ++ ) ++ parser.add_argument( ++ "--overlap", type=float, default=DEFAULT_OVERLAP_SECS, ++ help=f"Seconds shared by adjacent chunks, capped at a quarter of --chunk-len " ++ f"(default: {DEFAULT_OVERLAP_SECS}; 0 disables)" ++ ) ++ parser.add_argument( ++ "--pause-search", type=float, default=DEFAULT_PAUSE_SEARCH_SECS, ++ help=f"Seconds before each chunk limit searched for the quietest point to cut at, " ++ f"e.g. 25 (default: {DEFAULT_PAUSE_SEARCH_SECS}, cut at the limit)" + ) + + args = parser.parse_args() +@@ -175,6 +344,8 @@ def main(): + audio_path=args.audio_file, + output_file=args.output, + chunk_duration_secs=args.chunk_len, ++ overlap_secs=args.overlap, ++ pause_search_secs=args.pause_search, + ) + + +diff --git a/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py b/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py +new file mode 100644 +index 0000000..6a35947 +--- /dev/null ++++ b/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py +@@ -0,0 +1,329 @@ ++"""Unit tests for the slicing and stitching helpers in parakeet_transcribe_buffered.py. ++ ++These are pure functions: they need numpy, librosa and soundfile (imported by the ++script) but no GPU, no NeMo and no model. ++""" ++import sys ++from pathlib import Path ++ ++import numpy as np ++import pytest ++ ++sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) ++from parakeet_transcribe_buffered import plan_slices, stitch_slices # noqa: E402 ++ ++SR = 16000 ++ ++ ++def speech(seconds, seed=0): ++ """Stand-in for continuous speech: broadband noise at a speech-like level.""" ++ rng = np.random.default_rng(seed) ++ return (0.1 * rng.standard_normal(int(seconds * SR))).astype(np.float32) ++ ++ ++def with_pauses(audio, pauses, level=0.001, seed=1): ++ """Replace each (start_s, end_s) span with low-level room noise (or zeros).""" ++ rng = np.random.default_rng(seed) ++ out = audio.copy() ++ for start, end in pauses: ++ a, b = int(start * SR), int(end * SR) ++ out[a:b] = level * rng.standard_normal(b - a) ++ return out ++ ++ ++def assert_valid_plan(spans, cuts, num_samples, max_secs): ++ assert spans[0][0] == 0 and spans[-1][1] == num_samples ++ assert len(spans) == len(cuts) + 1 ++ for start, end in spans: ++ assert 0 < end - start <= max_secs * SR ++ for (a0, a1), (b0, b1), cut in zip(spans, spans[1:], cuts): ++ assert b0 <= cut <= a1, "each cut must lie inside both neighbouring slices" ++ assert a0 < cut < b1 ++ ++ ++# -- plan_slices: cut placement ------------------------------------------------- ++ ++ ++def test_audio_shorter_than_one_slice_is_not_cut(): ++ audio = speech(60) ++ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=4, search_secs=25) ++ assert cuts == [] ++ assert spans == [(0, len(audio))] ++ ++ ++def test_audio_exactly_one_slice_long_is_not_cut(): ++ audio = speech(120) ++ spans, cuts = plan_slices(audio, SR, 120) ++ assert cuts == [] ++ assert spans == [(0, len(audio))] ++ ++ ++def test_without_search_or_overlap_the_legacy_fixed_grid_is_reproduced(): ++ audio = speech(300) ++ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=0, search_secs=0) ++ assert cuts == [120 * SR, 240 * SR] ++ assert spans == [(0, 120 * SR), (120 * SR, 240 * SR), (240 * SR, 300 * SR)] ++ ++ ++def test_cuts_land_in_the_pauses_before_the_limit(): ++ pauses = [(100.0, 100.5), (215.0, 215.5), (330.0, 330.5)] ++ audio = with_pauses(speech(400), pauses) ++ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=0, search_secs=25) ++ assert len(cuts) == 3 ++ for cut, (start, end) in zip(cuts, pauses): ++ assert start <= cut / SR <= end ++ assert_valid_plan(spans, cuts, len(audio), 120) ++ ++ ++def test_the_quietest_pause_wins(): ++ # Two pauses inside the same search window; the later one is louder. ++ audio = with_pauses(speech(200), [(100.0, 100.6)], level=0.0) ++ audio = with_pauses(audio, [(115.0, 115.6)], level=0.01) ++ _, cuts = plan_slices(audio, SR, 120, search_secs=25) ++ assert 100.0 <= cuts[0] / SR <= 100.6 ++ ++ ++def test_audio_with_no_pause_is_still_cut_within_the_limit(): ++ audio = speech(400) ++ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=0, search_secs=25) ++ assert_valid_plan(spans, cuts, len(audio), 120) ++ edges = [0] + cuts ++ for prev, cut in zip(edges, cuts): ++ assert 95 * SR <= cut - prev <= 120 * SR, "cut must fall inside its search window" ++ ++ ++def test_a_pause_at_the_very_start_is_never_a_cut(): ++ audio = with_pauses(speech(200), [(0.0, 5.0)], level=0.0) ++ spans, cuts = plan_slices(audio, SR, 120, search_secs=25) ++ assert len(cuts) == 1 and 95 <= cuts[0] / SR <= 120 ++ assert_valid_plan(spans, cuts, len(audio), 120) ++ ++ ++def test_a_pause_at_the_very_end_leaves_no_empty_slice(): ++ # First cut ~100.35 s, so the second search window is ~[195, 220] s and ++ # holds the start of the trailing silence (217-222 s). ++ audio = with_pauses(speech(222), [(100.0, 100.5), (217.0, 222.0)], level=0.0) ++ spans, cuts = plan_slices(audio, SR, 120, search_secs=25) ++ assert len(cuts) == 2 ++ assert 217.0 <= cuts[1] / SR < 222.0 ++ assert_valid_plan(spans, cuts, len(audio), 120) ++ ++ ++# -- plan_slices: overlap ------------------------------------------------------- ++ ++ ++def test_overlap_is_included_in_the_slice_limit_and_centred_on_each_cut(): ++ pauses = [(100.0, 100.5), (215.0, 215.5), (330.0, 330.5)] ++ audio = with_pauses(speech(400), pauses) ++ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=4, search_secs=25) ++ assert_valid_plan(spans, cuts, len(audio), 120) ++ for (_, a1), (b0, _), cut in zip(spans, spans[1:], cuts): ++ assert a1 - b0 == 4 * SR ++ assert cut - b0 == a1 - cut ++ for cut, (start, end) in zip(cuts, pauses): ++ assert start <= cut / SR <= end ++ ++ ++def test_overlap_without_pause_search_uses_a_fixed_grid(): ++ audio = speech(300) ++ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=4, search_secs=0) ++ assert cuts == [116 * SR, 232 * SR] ++ assert spans == [(0, 118 * SR), (114 * SR, 234 * SR), (230 * SR, 300 * SR)] ++ ++ ++def test_short_slice_lengths_clamp_overlap_and_search(): ++ # Upstream's own buffered test runs a 19 s clip with --chunk-len 10. ++ audio = speech(19) ++ spans, cuts = plan_slices(audio, SR, 10, overlap_secs=4, search_secs=25) ++ assert len(spans) >= 2 ++ assert_valid_plan(spans, cuts, len(audio), 10) ++ ++ ++# -- stitch_slices ------------------------------------------------------------------ ++ ++FRAME = 0.08 ++OVERLAPPING = [(0.0, 12.0), (8.0, 20.0)] # two chunks sharing 8-12 s, cut at 10 s ++ ++ ++def word(text, start, end, slice_start): ++ return { ++ "word": text, ++ "start_offset": round((start - slice_start) / FRAME), ++ "end_offset": round((end - slice_start) / FRAME), ++ "start": start, ++ "end": end, ++ } ++ ++ ++def segment(words): ++ return { ++ "segment": " ".join(w["word"] for w in words), ++ "start_offset": words[0]["start_offset"], ++ "end_offset": words[-1]["end_offset"], ++ "start": words[0]["start"], ++ "end": words[-1]["end"], ++ } ++ ++ ++def texts(items, key="word"): ++ return [item[key] for item in items] ++ ++ ++def test_no_overlap_stitching_is_plain_concatenation(): ++ left = [word("one", 1.0, 1.4, 0), word("two", 5.0, 5.3, 0)] ++ right = [word("three", 10.5, 10.9, 10), word("four", 14.0, 14.4, 10)] ++ words, segments = stitch_slices( ++ [(left, [segment(left)]), (right, [segment(right)])], [10.0] ++ ) ++ assert words == left + right ++ assert segments == [segment(left), segment(right)] ++ ++ ++def test_overlapping_slices_keep_every_word_exactly_once(): ++ # Slices [0, 12] and [8, 20], cut at 10. Both transcribe the overlap and ++ # their timestamps for the same word differ by a few ms. ++ left = [ ++ word("a", 1.0, 1.3, 0), ++ word("b", 5.0, 5.4, 0), ++ word("c", 9.00, 9.40, 0), ++ word("d", 10.50, 10.90, 0), ++ word("e", 11.50, 11.80, 0), ++ ] ++ right = [ ++ word("c", 9.02, 9.40, 8), ++ word("d", 10.48, 10.90, 8), ++ word("e", 11.52, 11.80, 8), ++ word("f", 15.00, 15.40, 8), ++ ] ++ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING) ++ assert texts(words) == ["a", "b", "c", "d", "e", "f"] ++ assert words[2] is left[2] and words[3] is right[1] ++ ++ ++def test_a_word_straddling_the_cut_is_kept_once(): ++ left = [word("over", 9.90, 10.30, 0), word("the", 10.40, 10.55, 0)] ++ right = [word("over", 9.92, 10.30, 8), word("the", 10.40, 10.55, 8)] ++ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING) ++ assert texts(words) == ["over", "the"] ++ # Both chunks agree on it, so the right chunk takes over from it. ++ assert words[0] is right[0] and words[1] is right[1] ++ ++ ++def test_a_segment_straddling_the_cut_is_trimmed_to_the_words_each_slice_owns(): ++ left_words = [ ++ word("Hello", 8.5, 8.9, 0), ++ word("there", 9.2, 9.6, 0), ++ word("friend.", 10.4, 10.9, 0), ++ ] ++ right_words = [ ++ word("there", 9.21, 9.6, 8), ++ word("friend.", 10.41, 10.9, 8), ++ word("Bye.", 13.0, 13.4, 8), ++ ] ++ words, segments = stitch_slices( ++ [ ++ (left_words, [segment(left_words)]), ++ (right_words, [segment(right_words[:2]), segment(right_words[2:])]), ++ ], ++ [10.0], ++ OVERLAPPING, ++ ) ++ assert texts(words) == ["Hello", "there", "friend.", "Bye."] ++ assert texts(segments, "segment") == ["Hello there", "friend.", "Bye."] ++ assert segments[0]["end"] == 9.6 and segments[1]["start"] == 10.41 ++ # A segment that needed no trimming is passed through untouched. ++ assert segments[2] == segment(right_words[2:]) ++ # Every word appears in exactly one segment, in order. ++ assert " ".join(texts(segments, "segment")) == " ".join(texts(words)) ++ ++ ++def test_a_segment_wholly_inside_the_other_slices_share_is_dropped(): ++ left_words = [word("a", 2.0, 2.3, 0), word("b.", 10.6, 11.0, 0)] ++ right_words = [word("b.", 10.61, 11.0, 8), word("c", 12.0, 12.3, 8)] ++ _, segments = stitch_slices( ++ [ ++ (left_words, [segment(left_words[:1]), segment(left_words[1:])]), ++ (right_words, [segment(right_words[:1]), segment(right_words[1:])]), ++ ], ++ [10.0], ++ OVERLAPPING, ++ ) ++ assert texts(segments, "segment") == ["a", "b.", "c"] ++ assert segments[1]["start"] == 10.61 ++ ++ ++def test_a_non_positive_chunk_length_is_rejected_rather_than_looping(): ++ with pytest.raises(ValueError): ++ plan_slices(speech(5), SR, 0) ++ ++ ++def test_a_word_the_two_chunks_timestamp_either_side_of_the_cut_is_kept_once(): ++ # After a pause TDT may place a word's start anywhere in the pause, so the ++ # two chunks can disagree about which side of the cut it starts on. ++ left = [word("so", 8.2, 8.5, 0), word("then", 9.98, 10.3, 0), word("we", 10.4, 10.6, 0)] ++ right = [word("so", 8.2, 8.5, 8), word("then", 10.03, 10.3, 8), word("we", 10.4, 10.6, 8)] ++ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING) ++ assert texts(words) == ["so", "then", "we"] ++ # ...and the mirror image, where splitting both at the cut would drop it. ++ left = [word("so", 8.2, 8.5, 0), word("then", 10.03, 10.3, 0), word("we", 10.4, 10.6, 0)] ++ right = [word("so", 8.2, 8.5, 8), word("then", 9.98, 10.3, 8), word("we", 10.4, 10.6, 8)] ++ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING) ++ assert texts(words) == ["so", "then", "we"] ++ ++ ++def test_handover_happens_at_the_agreed_word_nearest_the_cut(): ++ # The chunks differ in casing/punctuation and the left one drops "really" ++ # near its end; the right chunk's version of the overlap after the cut wins. ++ left = [word("It", 8.5, 8.7, 0), word("was", 9.6, 9.9, 0), word("good,", 11.0, 11.4, 0)] ++ right = [word("it", 8.5, 8.7, 8), word("was", 9.62, 9.9, 8), word("really", 10.3, 10.7, 8), ++ word("good.", 11.0, 11.4, 8)] ++ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING) ++ assert texts(words) == ["It", "was", "really", "good."] ++ assert words[0] is left[0] and words[1] is right[1] ++ ++ ++def test_without_an_agreed_word_the_split_falls_back_to_the_cut(): ++ left = [word("alpha", 9.0, 9.4, 0), word("beta", 10.5, 10.9, 0)] ++ right = [word("gamma", 9.1, 9.4, 8), word("delta", 10.6, 10.9, 8)] ++ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING) ++ assert texts(words) == ["alpha", "delta"] ++ ++ ++def test_another_occurrence_of_the_word_elsewhere_in_the_overlap_is_not_an_anchor(): ++ # The chunks disagree everywhere except on "the", but the left chunk's ++ # "the" (9.0 s) and the right chunk's (11.0 s) are different words. As an ++ # anchor they would average to the cut and discard the left one. ++ left = [word("the", 9.0, 9.2, 0), word("dog", 10.5, 10.8, 0)] ++ right = [word("cat", 9.3, 9.6, 8), word("the", 11.0, 11.2, 8)] ++ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING) ++ assert [w["start"] for w in words] == [9.0, 11.0] ++ ++ ++def test_the_handover_never_puts_words_out_of_time_order(): ++ # "y" agrees (0.45 s apart), but the right chunk's "y" (9.65) after the ++ # left chunk's "x" (9.70) would run time backwards, so the left chunk's ++ # copy is kept instead. Splitting at the cut would lose "y" altogether. ++ left = [word("x", 9.70, 9.90, 0), word("y", 10.10, 10.30, 0)] ++ right = [word("z", 9.40, 9.60, 8), word("y", 9.65, 9.90, 8), word("w", 10.6, 10.8, 8)] ++ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING) ++ assert texts(words) == ["x", "y", "w"] and words[1] is left[1] ++ starts = [w["start"] for w in words] ++ assert starts == sorted(starts) ++ ++ ++def test_a_co_timed_anchor_is_found_even_when_a_longer_match_lies_elsewhere(): ++ # "x y" recurs later in the right chunk, a longer text match than "z", but ++ # at a different time. Only "z" is the same word in both chunks, and it ++ # straddles the cut, so splitting both at the cut would keep it twice. ++ left = [word("x", 8.2, 8.3, 0), word("y", 8.4, 8.5, 0), word("z", 9.98, 10.2, 0)] ++ right = [word("z", 10.03, 10.2, 8), word("x", 11.0, 11.1, 8), word("y", 11.2, 11.3, 8)] ++ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING) ++ assert texts(words) == ["x", "y", "z", "x", "y"] ++ ++ ++def test_punctuation_alone_is_never_an_anchor(): ++ # As an anchor the dash would hand the whole overlap to the right chunk. ++ left = [word("-", 9.50, 9.55, 0), word("yes", 10.4, 10.6, 0)] ++ right = [word("-", 9.52, 9.55, 8), word("no", 10.4, 10.6, 8)] ++ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING) ++ assert words[0] is left[0] and texts(words) == ["-", "no"] +-- +2.39.5 + diff --git a/stacks/scriberr/patches/README.md b/stacks/scriberr/patches/README.md new file mode 100644 index 0000000..008c99a --- /dev/null +++ b/stacks/scriberr/patches/README.md @@ -0,0 +1,162 @@ +# Scriberr local patches — contract + +We build Scriberr from source (no upstream sm_120 image; see `../README.md`), so +we can carry patches on that build. This directory holds them, and +`scripts/scriberr-rebuild` applies them to a pinned upstream sha, builds a +distinctly tagged image, and proves it before anyone deploys it. + +| patch | against | status | +|---|---|---| +| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | carried; upstream PR **prepared, not opened** (`upstream-pr/`) | + +Ruling: Prime, 2026-09-30, "build the slicer". Opening the upstream PR is +outward-facing and needs Prime's explicit yes. + +## 0001 — pause-aware Parakeet slicer + +### What it changes + +One file of product code, `internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py`, +plus one new test file beside it (`tests/test_parakeet_slicing.py`). No Go change. + +Upstream cuts long audio at fixed `--chunk-len` marks with no overlap, so a +word that straddles a mark is chopped in two, lost, or transcribed twice. The +patch: + +1. **Overlaps adjacent chunks** by `--overlap` seconds (default **4**), half on + each side of the cut, counted *inside* `--chunk-len`. +2. **Hands over at an agreed word.** In each overlap, the chunks switch at the + word nearest the cut that both transcribed alike: the same text after + lowercasing and stripping punctuation (punctuation alone never counts), with + start times within 0.5 s. The left chunk keeps the words before it and the + right chunk keeps the rest, the anchor taken from whichever chunk keeps the + words in time order. With no agreed word, both split at the cut by start time. + Segments are trimmed to the words their chunk keeps, and `transcription` is + the stitched words joined by spaces (upstream's text already equals that). +3. **Optional pause-aware cuts** (`--pause-search N`, default **off**): each cut + moves back to the middle of the quietest 0.3 s within the last N seconds + before the limit. It measured neutral once the stitch was right, so it is not + the default; Go never passes the flag. +4. **Imports NeMo inside `transcribe_buffered()`** so the pure helpers + (`plan_slices`, `stitch_slices`) import and test without a GPU or NeMo. + +`--overlap 0` (with pause search off, the default) reproduces upstream's output +exactly: words, segments and text were byte-identical on all four test recordings. + +Why the handover is by agreed word and not simply "each word goes to the chunk +its start time falls in" (the first design): at a quarter of the stitches the +two chunks put the *same* word on opposite sides of the cut, one frame apart, +so it was kept twice. Parakeet timestamps a word that follows a pause anywhere +inside the pause. Details in the bench doc. + +### The seam it must keep (Go ↔ Python) + +`parakeet_adapter.go` is not patched, so the script's CLI and JSON are frozen: + +- **Invocation** (Go, `buildBufferedArgs`): + `uv run --native-tls --project python parakeet_transcribe_buffered.py