feat(scriberr): overlap-and-stitch Parakeet slicer patch, rebuild script, bench
Carry patches/0001 on our Scriberr build (upstream a353078): adjacent buffered chunks overlap by 4 s inside --chunk-len and hand over at a word both chunks transcribed alike, instead of cutting at fixed marks with no overlap. Pause-aware cutting is included as an opt-in (--pause-search); it measured neutral once the stitch was right. The Go<->Python CLI and JSON seam is unchanged. Bench (4 recordings, 118 min, 3 cut placements each, against a no-cut whole-file reference; metrics only, private audio stays on fv-ml1): cuts with an error within +-3 s fall from 52% (93/179) to 22% (41/184) against a 19% background; floor +-0.08. Positive control: upstream's cutter +0.33 over background. A-vs-A byte-identical in-process and across CLI processes. Peak GPU memory unchanged at 5,496 MiB (n=3). Also found: Parakeet skips runs of >=10 words mid-chunk with any slicer, upstream's included; not addressed here. scripts/scriberr-rebuild clones a pinned upstream sha into a new /opt/docker/src dir, git-apply-checks the patches, builds a distinct tag, and checks embed, unit tests, the JSON seam (scriberr-seam-check.py) and the memory budget on idle GPU 3. Deploy stays manual. The upstream PR is prepared under patches/upstream-pr/ and not opened.
This commit is contained in:
@@ -0,0 +1,220 @@
|
|||||||
|
# Scriberr Parakeet slicer bench (2026-09-30)
|
||||||
|
|
||||||
|
Prime's ruling, 2026-09-30: "build the slicer". This document records how the
|
||||||
|
pause-aware slicer patch (`stacks/scriberr/patches/0001-parakeet-pause-aware-slicer.patch`)
|
||||||
|
was measured and why the shipped variant was chosen. **Metrics only**: two of
|
||||||
|
the recordings are Prime's and private, so no transcript text appears here or
|
||||||
|
anywhere in git. Those recordings, and every transcript made from them, stay on
|
||||||
|
fv-ml1 in `/tank/spikes/scriberr-slicer/private/` (mode 700).
|
||||||
|
|
||||||
|
## Recordings (4 files, 118 minutes)
|
||||||
|
|
||||||
|
All four were converted to 16 kHz mono WAV with the image's ffmpeg, which is
|
||||||
|
what Scriberr feeds Parakeet.
|
||||||
|
|
||||||
|
| id | length | what | provenance / licence |
|
||||||
|
|---|---|---|---|
|
||||||
|
| p1 | 35.3 min | Prime's upload, conversational | private |
|
||||||
|
| p2 | 22.3 min | Prime's upload, conversational | private |
|
||||||
|
| scotus | 30.0 min (first half hour) | U.S. Supreme Court oral argument, *Loper Bright Enterprises v. Raimondo*, No. 22-451, argued 2024-01-17; spontaneous multi-speaker speech with interruptions | `https://www.supremecourt.gov/media/audio/mp3files/22-451.mp3` (sha256 `7e704b1f…5dc2`); U.S. Government work, public domain (17 U.S.C. §105) |
|
||||||
|
| wilde | 24.4 min | LibriVox *The Trial of Oscar Wilde* (dramatic reading), section 1; several readers, courtroom dialogue | `https://archive.org/download/trialofoscarwilde_1601_librivox/trialofoscarwilde_01_anon_64kb.mp3` (sha256 `492f4c56…3420`); public domain (LibriVox, PD Mark 1.0) |
|
||||||
|
|
||||||
|
## Harness
|
||||||
|
|
||||||
|
- **Image and env:** `scriberr:local-blackwell` (the live image), the live
|
||||||
|
`/tank/scriberr/whisperx-env` mounted **read-only**, `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`,
|
||||||
|
invoked exactly as Scriberr does (`uv run --native-tls --project /app/whisperx-env/parakeet python …`).
|
||||||
|
It runs as uid 1002 with `USER` set, not as appuser 10001; that changes file
|
||||||
|
ownership only.
|
||||||
|
- **GPU:** transient containers on **GPU 3 only** (verified: the container sees
|
||||||
|
one card, UUID `GPU-186dacf4…`). GPU 3 was at 2 MiB before and after.
|
||||||
|
- **Reference:** upstream's standard script on the whole file in one pass with
|
||||||
|
local attention (`--context-left 255 --context-right 255`), which has no cuts.
|
||||||
|
It needs >16 GB, which is why it cannot run on GPU 1.
|
||||||
|
- **Variants** run through the real `transcribe_buffered()`; the only harness
|
||||||
|
change is that the model is loaded once per process instead of once per call.
|
||||||
|
The CLI memory runs below reproduce the harness output byte for byte.
|
||||||
|
- **Placements:** every variant at three maximum lengths, 120, 110 and 100 s.
|
||||||
|
Parakeet is deterministic, so a repeat run cannot supply variance; moving the
|
||||||
|
cuts can. Each (file, variant) therefore has n = 3 different cut placements.
|
||||||
|
|
||||||
|
## Metric
|
||||||
|
|
||||||
|
Each variant's words are aligned against the reference, after lowercasing and
|
||||||
|
stripping punctuation (difflib opcodes, then exact Levenshtein inside each
|
||||||
|
non-matching block). Each error gets a time: the hypothesis word's start for a
|
||||||
|
substitution or insertion, the reference word's start for a deletion.
|
||||||
|
|
||||||
|
- **near-cut / elsewhere word errors (S/I/D):** near-cut means within ±3 s of
|
||||||
|
any of that variant's cuts (for overlap variants, the cut is the stitch point
|
||||||
|
at the middle of the overlap, and both chunk edges lie within ±3 s of it).
|
||||||
|
- **dropped / duplicated at cuts:** deletions near a cut, and insertions near a
|
||||||
|
cut that repeat an adjacent word.
|
||||||
|
- **error events:** errors clustered with gaps ≤ 1 s; rate per minute
|
||||||
|
elsewhere.
|
||||||
|
- **damaged cuts (the decision metric):** cuts with at least one error within
|
||||||
|
±3 s, against **damaged phantoms**, the same test at points midway between
|
||||||
|
the variant's own cuts (same count, same spacing, as far from any cut as the
|
||||||
|
audio gets). The phantom rate is the background any slicer's cuts are judged
|
||||||
|
against.
|
||||||
|
|
||||||
|
### Why the decision metric is not raw word counts
|
||||||
|
|
||||||
|
The negative control caught it. Word errors elsewhere varied by up to ±50 %
|
||||||
|
between slicers of the same length (p1: 131 to 215). Most of that comes from a
|
||||||
|
few **unstable stretches**, the same stretches for every variant (p1: 976–992 s,
|
||||||
|
1291–1298 s, 928–934 s), where the reference and any slicing disagree by 20–60
|
||||||
|
words depending on context, at arbitrary distances from any cut. That is heavy-tailed
|
||||||
|
noise, not cut damage. Counting events, and asking per cut whether anything near
|
||||||
|
it went wrong, is robust to it; the word counts are still reported.
|
||||||
|
|
||||||
|
## Variants
|
||||||
|
|
||||||
|
| name | overlap | pause search | stitch |
|
||||||
|
|---|---|---|---|
|
||||||
|
| fixed | 0 | off | none: upstream's slicer (patched code with both off reproduces it byte for byte) |
|
||||||
|
| pause | 0 | 25 s | none needed |
|
||||||
|
| overlap | 4 s | off | v1: each word kept by the chunk whose cut-to-cut range holds its start time (the brief's rule) |
|
||||||
|
| both | 4 s | 25 s | v1 |
|
||||||
|
| **overlap2** | 4 s | off | **v2: hand over at the nearest word both chunks transcribed alike, within 0.5 s** (shipped as v3, identical output) |
|
||||||
|
| both2 | 4 s | 25 s | v2 |
|
||||||
|
| both8 | 8 s | 25 s | v2 |
|
||||||
|
|
||||||
|
**v3** is v2 after code review, and it is what ships. Anchors are paired by
|
||||||
|
time first (same text *and* within 0.5 s, so a longer text match elsewhere in
|
||||||
|
the overlap cannot crowd out the true anchor), punctuation alone never anchors,
|
||||||
|
and the anchor's copy comes from whichever chunk keeps word starts in time order.
|
||||||
|
Re-run on all 12 overlap runs (4 files × 3 placements), **v3's output is
|
||||||
|
byte-identical to v2's** (and "both3" to "both2"): the reviewer's cases did not
|
||||||
|
occur in this audio, so every v2 number below holds for the shipped code.
|
||||||
|
|
||||||
|
v2 exists because v1 failed its own goal. At 26 of 108 "both" stitches the same
|
||||||
|
word sat on both sides of the cut, the two copies' start times 0.00–0.09 s apart
|
||||||
|
(one encoder frame): a word that follows a pause is timestamped anywhere in the
|
||||||
|
pause, and a pause is exactly where a pause-aware cut lands. Start time at the
|
||||||
|
midpoint is the worst possible stitch rule for pause cuts.
|
||||||
|
|
||||||
|
## Results
|
||||||
|
|
||||||
|
### Pooled (4 files × 3 placements)
|
||||||
|
|
||||||
|
| variant | damaged cuts | damaged phantoms | excess | by length 120 / 110 / 100 | near-cut words | dropped | duplicated | near-cut events | events elsewhere /min |
|
||||||
|
|---|---|---|---|---|---|---|---|---|---|
|
||||||
|
| fixed (upstream) | 93/179 = 52 % | 34/179 = 19 % | +0.33 | 0.48 / 0.52 / 0.55 | 223 | 56 | 17 | 102 | 2.08 |
|
||||||
|
| pause | 53/201 = 26 % | 36/201 = 18 % | +0.085 | 0.33 / 0.20 / 0.27 | 113 | 52 | 0 | 59 | 2.23 |
|
||||||
|
| overlap (v1) | 52/184 = 28 % | 33/184 = 18 % | +0.10 | 0.30 / 0.28 / 0.27 | 117 | 25 | 18 | 60 | 2.00 |
|
||||||
|
| both (v1) | 75/207 = 36 % | 38/207 = 18 % | +0.18 | 0.39 / 0.35 / 0.36 | 126 | 21 | 40 | 87 | 2.35 |
|
||||||
|
| **overlap2** | **41/184 = 22 %** | 33/184 = 18 % | **+0.04** | 0.21 / 0.22 / 0.24 | 101 | 24 | 2 | **47** | 2.00 |
|
||||||
|
| both2 | 52/207 = 25 % | 38/207 = 18 % | +0.07 | 0.32 / 0.20 / 0.24 | 89 | 22 | 2 | 54 | 2.35 |
|
||||||
|
| both8 | 48/222 = 22 % | 46/222 = 21 % | +0.01 | 0.23 / 0.23 / 0.19 | 120 | 42 | 0 | 54 | 2.18 |
|
||||||
|
|
||||||
|
### Per file (summed over the 3 placements)
|
||||||
|
|
||||||
|
| file | variant | damaged cuts | phantom | near S/I/D | near words | else words/min | drop | dup | events near | events else/min |
|
||||||
|
|---|---|---|---|---|---|---|---|---|---|---|
|
||||||
|
| p1 | fixed | 27/57 | 11/57 | 21/27/9 | 57 | 4.16 | 9 | 10 | 29 | 2.15 |
|
||||||
|
| p1 | pause | 19/64 | 8/64 | 20/3/25 | 48 | 5.62 | 25 | 0 | 20 | 2.39 |
|
||||||
|
| p1 | overlap | 20/59 | 16/59 | 16/26/4 | 46 | 4.66 | 4 | 8 | 23 | 2.07 |
|
||||||
|
| p1 | both | 29/67 | 13/67 | 23/25/4 | 52 | 5.70 | 4 | 17 | 35 | 2.56 |
|
||||||
|
| p1 | overlap2 | 15/59 | 16/59 | 16/19/4 | 39 | 4.66 | 4 | 0 | 17 | 2.07 |
|
||||||
|
| p1 | both2 | 21/67 | 13/67 | 23/8/5 | 36 | 5.70 | 5 | 0 | 21 | 2.56 |
|
||||||
|
| p1 | both8 | 18/70 | 15/70 | 16/2/8 | 26 | 6.22 | 8 | 0 | 18 | 2.40 |
|
||||||
|
| p2 | fixed | 14/36 | 4/36 | 10/7/14 | 31 | 4.45 | 14 | 2 | 16 | 1.39 |
|
||||||
|
| p2 | pause | 10/40 | 6/40 | 11/1/1 | 13 | 2.71 | 1 | 0 | 10 | 1.56 |
|
||||||
|
| p2 | overlap | 6/36 | 6/36 | 4/2/0 | 6 | 3.56 | 0 | 2 | 6 | 1.47 |
|
||||||
|
| p2 | both | 16/41 | 6/41 | 13/10/0 | 23 | 3.11 | 0 | 8 | 18 | 1.64 |
|
||||||
|
| p2 | overlap2 | 4/36 | 6/36 | 4/0/0 | 4 | 3.56 | 0 | 0 | 4 | 1.47 |
|
||||||
|
| p2 | both2 | 10/41 | 6/41 | 13/2/0 | 15 | 3.11 | 0 | 0 | 11 | 1.64 |
|
||||||
|
| p2 | both8 | 9/43 | 7/43 | 9/1/1 | 11 | 3.17 | 1 | 0 | 9 | 1.58 |
|
||||||
|
| scotus | fixed | 27/47 | 15/47 | 14/49/21 | 84 | 9.47 | 21 | 4 | 30 | 3.34 |
|
||||||
|
| scotus | pause | 18/54 | 17/54 | 13/9/10 | 32 | 8.63 | 10 | 0 | 22 | 3.36 |
|
||||||
|
| scotus | overlap | 20/49 | 6/49 | 16/23/11 | 50 | 9.43 | 11 | 5 | 24 | 3.01 |
|
||||||
|
| scotus | both | 19/55 | 14/55 | 9/13/5 | 27 | 9.46 | 5 | 8 | 21 | 3.39 |
|
||||||
|
| scotus | overlap2 | 18/49 | 6/49 | 15/20/10 | 45 | 9.43 | 10 | 2 | 21 | 3.01 |
|
||||||
|
| scotus | both2 | 15/55 | 14/55 | 9/7/5 | 21 | 9.46 | 5 | 2 | 15 | 3.39 |
|
||||||
|
| scotus | both8 | 14/61 | 19/61 | 14/31/10 | 55 | 8.19 | 10 | 0 | 19 | 3.30 |
|
||||||
|
| wilde | fixed | 25/39 | 4/39 | 9/30/12 | 51 | 6.91 | 12 | 1 | 27 | 1.03 |
|
||||||
|
| wilde | pause | 6/43 | 5/43 | 4/0/16 | 20 | 6.95 | 16 | 0 | 7 | 1.23 |
|
||||||
|
| wilde | overlap | 6/40 | 5/40 | 2/3/10 | 15 | 6.74 | 10 | 3 | 7 | 1.13 |
|
||||||
|
| wilde | both | 11/44 | 5/44 | 4/8/12 | 24 | 9.43 | 12 | 7 | 13 | 1.43 |
|
||||||
|
| wilde | overlap2 | 4/40 | 5/40 | 3/0/10 | 13 | 6.74 | 10 | 0 | 5 | 1.13 |
|
||||||
|
| wilde | both2 | 6/44 | 5/44 | 4/1/12 | 17 | 9.43 | 12 | 0 | 7 | 1.43 |
|
||||||
|
| wilde | both8 | 7/48 | 5/48 | 5/0/23 | 28 | 6.80 | 23 | 0 | 8 | 1.04 |
|
||||||
|
|
||||||
|
### Where the remaining near-cut errors sit
|
||||||
|
|
||||||
|
Signed distance from the cut, pooled over files: upstream's errors pile up
|
||||||
|
within ±0.5 s (chopped words; at 100 s, 24 insertions in [−0.5, 0) alone).
|
||||||
|
Pause-only leaves deletions 1–3 s **before** its cuts, where the left chunk ends
|
||||||
|
(8/9, 3/4, 7/8 in [−3, −1) at the three lengths). overlap2's residue is spread
|
||||||
|
evenly over ±3 s, which is what background looks like. both2 and both8 are clean
|
||||||
|
at the handover but keep some of pause-only's pre-cut deletions.
|
||||||
|
|
||||||
|
### Controls
|
||||||
|
|
||||||
|
- **A-vs-A:** the same variant twice gives byte-identical words, segments and
|
||||||
|
text (fixed-120, both-120, both2-120 on all 4 files, in one process), and the
|
||||||
|
shipped script run three times as separate CLI processes matches the harness
|
||||||
|
output byte for byte (p1 and scotus). Decoding is deterministic, so run-to-run
|
||||||
|
variance is zero; the variance that matters comes from cut placement.
|
||||||
|
- **Harness validity:** patched code with overlap and pause search off equals
|
||||||
|
upstream's unmodified script on all 4 files; v2 without overlap equals v1.
|
||||||
|
- **Positive control:** upstream's fixed cutter damages 52 % of its cuts
|
||||||
|
against a 19 % background (+0.33, about 7 standard errors), per file 47–64 %
|
||||||
|
against 11–25 %. The instrument sees cut damage.
|
||||||
|
- **Negative control:** the phantom (background) rate is 18–21 % for every
|
||||||
|
variant, and error events away from cuts run at 2.0–2.35 per minute for all
|
||||||
|
of them. **Word counts away from cuts do not agree across slicers** (see
|
||||||
|
"Parakeet drops stretches" below), which is why they are not the decision metric.
|
||||||
|
- **Sensitivity floor:** at ~180–220 cuts per variant the 2-standard-error band
|
||||||
|
on a difference of damaged-cut rates is **±0.08** pooled and **±0.17** for
|
||||||
|
one file. No slicer can be measured below the background (~18–21 %). Word-count
|
||||||
|
differences under ~60 words are unresolvable, because a single dropped stretch
|
||||||
|
(next section) is 10–60 words.
|
||||||
|
|
||||||
|
### Verdict
|
||||||
|
|
||||||
|
Every variant that overlaps with the v2 handover, and pause-only, beats
|
||||||
|
upstream by 0.26–0.30 in damaged-cut rate, far outside the ±0.08 floor. Among
|
||||||
|
them the differences are **inside** the floor. overlap2 ships: tied best on
|
||||||
|
damaged cuts, fewest near-cut error events (47 vs 102), near-zero duplicates,
|
||||||
|
nothing concentrated at the handover, and the simplest mechanism. Pause search
|
||||||
|
measured neutral once the stitch was fixed, so it stays in the patch as an
|
||||||
|
opt-in `--pause-search`, off by default.
|
||||||
|
|
||||||
|
## Memory (GPU 3, production invocation, nvidia-smi every 0.2 s)
|
||||||
|
|
||||||
|
| run | n | per-process peak |
|
||||||
|
|---|---|---|
|
||||||
|
| upstream script, 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB (reproduces the 2026-09-30 budget figure) |
|
||||||
|
| shipped script (v2), 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB |
|
||||||
|
| shipped script (v3, final), 120 s, p1 | 3 | 5,496 / 5,496 / 5,496 MiB; all 3 CLI outputs byte-identical to the harness |
|
||||||
|
| shipped script, 120 s, scotus (public) | 3 | 5,496 / 5,496 / 5,496 MiB |
|
||||||
|
| positive control: shipped script, **300 s**, scotus | 1 | 7,056 MiB (deterministic; 300 s was n=3 in the earlier table) |
|
||||||
|
|
||||||
|
Zero spread; the peak is set by the 120 s maximum, and the overlap sits inside
|
||||||
|
it. The scotus file reads the same peak as p1, so the rebuild script uses it
|
||||||
|
as its public memory fixture. A spike shorter than the 0.2 s sample period
|
||||||
|
could be missed.
|
||||||
|
|
||||||
|
## Parakeet drops stretches of speech (separate finding, not the slicer)
|
||||||
|
|
||||||
|
Every chunked variant, **upstream's included**, sometimes skips a run of ≥10
|
||||||
|
consecutive words in the middle of a chunk, and which runs it skips changes
|
||||||
|
chaotically with the cut placement. Wilde at 120 s loses nothing, at 110 s
|
||||||
|
171 words, and overlap2 at 110 s loses one 60-second stretch. Across 4 files × 3
|
||||||
|
placements:
|
||||||
|
|
||||||
|
| variant | runs of ≥10 words lost | words lost |
|
||||||
|
|---|---|---|
|
||||||
|
| fixed (upstream) | 15 | 599 |
|
||||||
|
| pause | 17 | 511 |
|
||||||
|
| overlap2 | 12 | 656 |
|
||||||
|
| both2 | 14 | 719 |
|
||||||
|
| both8 | 12 | 560 |
|
||||||
|
|
||||||
|
The whole-file local-attention reference does the same: runs where every
|
||||||
|
slicer has words the reference lacks (13–16 per 12 comparisons, 340–400 words).
|
||||||
|
Today's production setting (upstream, 120 s) lost 85 words in 2 runs on p2. The
|
||||||
|
slicer neither causes nor cures it; it needs its own investigation (decoder
|
||||||
|
settings, chunk length, or model), which was out of scope here.
|
||||||
Executable
+263
@@ -0,0 +1,263 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# scriberr-rebuild — rebuild Scriberr's Blackwell image at a PINNED upstream sha
|
||||||
|
# with our patches applied, then prove the result before anyone deploys it.
|
||||||
|
#
|
||||||
|
# Runs on nh3-dev and drives fv-ml1 over ssh. Deploy is a SEPARATE manual step
|
||||||
|
# (stacks/scriberr/patches/README.md § Deploy); this script never touches the
|
||||||
|
# live container, its .env, or GPU 1.
|
||||||
|
#
|
||||||
|
# Why this exists: upstream publishes no sm_120 image, so we build from source,
|
||||||
|
# and we carry a patch to the Parakeet slicer (stacks/scriberr/patches/). Upstream
|
||||||
|
# moves slowly, so an upgrade should be one command plus a verdict.
|
||||||
|
#
|
||||||
|
# Stages (each prints PASS/FAIL; the first FAIL stops the run):
|
||||||
|
# clone clean shallow clone of upstream at the pinned sha, in a NEW dir
|
||||||
|
# /opt/docker/src/scriberr-<sha7>-<suffix> (reused only if it already
|
||||||
|
# holds that sha with every patch applied)
|
||||||
|
# patch `git apply --check` then `git apply`, patch by patch; a conflict
|
||||||
|
# stops the run loudly and names the patch
|
||||||
|
# build docker build -f Dockerfile.cuda.12.9 -t scriberr:local-blackwell-<sha7>-<suffix>
|
||||||
|
# — a DISTINCT tag, so the running image is never overwritten
|
||||||
|
# embed the patched script's exact bytes are inside the new Go binary
|
||||||
|
# (Scriberr rewrites the env's copy from this embed on every start)
|
||||||
|
# unit the slicer's pure-function tests, under the live env's numpy/librosa
|
||||||
|
# seam patched script, production invocation, short fixture at
|
||||||
|
# --chunk-len 10, JSON validated against the Go struct
|
||||||
|
# memory same on a long recording at --chunk-len 120 on an IDLE GPU,
|
||||||
|
# nvidia-smi sampled every 0.2 s; per-process peak <= the budget
|
||||||
|
#
|
||||||
|
# Usage:
|
||||||
|
# scripts/scriberr-rebuild [--sha SHA40] [--suffix NAME] [--gpu N]
|
||||||
|
# [--budget MIB] [--memory-audio PATH] [--reuse-image]
|
||||||
|
#
|
||||||
|
# Defaults: --sha PINNED_SHA below, --suffix slicer1, --gpu 3, --budget 5496,
|
||||||
|
# --memory-audio the public 30-min SCOTUS fixture. The GPU must be idle
|
||||||
|
# (< 100 MiB used), which in practice means GPU 3; GPUs 0-2 run live seats.
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
PINNED_SHA=a353078fd96b8aca4002681813524b7397c90df1 # upstream HEAD 2026-09-20
|
||||||
|
UPSTREAM=https://github.com/rishikanthc/Scriberr.git
|
||||||
|
HOST=${SCRIBERR_REBUILD_HOST:-infra-ops@10.251.50.54}
|
||||||
|
ENV_DIR=/tank/scriberr/whisperx-env # live env, always mounted READ-ONLY
|
||||||
|
TOOLS=/opt/docker/src/scriberr-rebuild # fixtures + seam checker on fv-ml1
|
||||||
|
SCRIPT_REL=internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
|
||||||
|
TEST_REL=internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
|
||||||
|
SEAM_AUDIO_REL=tests/data/AMI-Corpus-IB4002.Mix-Headset-clip.wav
|
||||||
|
|
||||||
|
SHA=$PINNED_SHA SUFFIX=slicer1 GPU=3 BUDGET=5496 REUSE_IMAGE=0
|
||||||
|
MEM_AUDIO=$TOOLS/fixtures/scotus-22-451-first30m.wav
|
||||||
|
while [ $# -gt 0 ]; do
|
||||||
|
case $1 in
|
||||||
|
--sha) SHA=$2; shift 2 ;;
|
||||||
|
--suffix) SUFFIX=$2; shift 2 ;;
|
||||||
|
--gpu) GPU=$2; shift 2 ;;
|
||||||
|
--budget) BUDGET=$2; shift 2 ;;
|
||||||
|
--memory-audio) MEM_AUDIO=$2; shift 2 ;;
|
||||||
|
--reuse-image) REUSE_IMAGE=1; shift ;;
|
||||||
|
-h|--help) sed -n '2,/^set -euo/p' "$0" | sed '$d; s/^# \{0,1\}//'; exit 0 ;;
|
||||||
|
*) echo "unknown argument: $1 (see --help)" >&2; exit 2 ;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
[[ $SHA =~ ^[0-9a-f]{40}$ ]] || { echo "--sha must be a full 40-char sha (GitHub fetches by full sha)" >&2; exit 2; }
|
||||||
|
[[ $SUFFIX =~ ^[a-z0-9][a-z0-9.-]*$ ]] || { echo "--suffix must be [a-z0-9.-]" >&2; exit 2; }
|
||||||
|
[[ $GPU =~ ^[0-9]+$ && $BUDGET =~ ^[0-9]+$ ]] || { echo "--gpu and --budget must be integers" >&2; exit 2; }
|
||||||
|
|
||||||
|
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||||
|
PATCH_DIR=$REPO_ROOT/stacks/scriberr/patches
|
||||||
|
mapfile -t PATCHES < <(find "$PATCH_DIR" -maxdepth 1 -name '*.patch' | sort)
|
||||||
|
[ ${#PATCHES[@]} -gt 0 ] || { echo "no patches in $PATCH_DIR" >&2; exit 2; }
|
||||||
|
# The only paths a reused build dir may differ from upstream in.
|
||||||
|
mapfile -t PATCHED_PATHS < <(sed -n 's#^+++ b/##p' "${PATCHES[@]}" | sort -u)
|
||||||
|
SHA7=${SHA:0:7}
|
||||||
|
TAG=scriberr:local-blackwell-$SHA7-$SUFFIX
|
||||||
|
BUILD_DIR=/opt/docker/src/scriberr-$SHA7-$SUFFIX
|
||||||
|
PATCH_SUM=$(cat "${PATCHES[@]}" | sha256sum | cut -c1-16)
|
||||||
|
CNAME=scriberr-rebuild-$SHA7-$SUFFIX # every GPU container, so cleanup can find it
|
||||||
|
SAMPLES=/tmp/$CNAME.mem.csv
|
||||||
|
BUILD_LOG=/tmp/$CNAME.build.log
|
||||||
|
MIN_FREE_GB=20 # under Docker's root dir (zroot); an image adds ~0.1-6 GB
|
||||||
|
|
||||||
|
# One multiplexed ssh connection serves the ~15 remote calls of a run.
|
||||||
|
CM=(-o BatchMode=yes -o ControlMaster=auto -o "ControlPath=$HOME/.ssh/cm-%C" -o ControlPersist=120)
|
||||||
|
SSH=(ssh "${CM[@]}" "$HOST")
|
||||||
|
SCP=(scp -q "${CM[@]}")
|
||||||
|
|
||||||
|
RESULTS=()
|
||||||
|
pass() { RESULTS+=("PASS $1 $2"); echo "== PASS $1: $2"; }
|
||||||
|
fail() {
|
||||||
|
RESULTS+=("FAIL $1 $2"); echo "== FAIL $1: $2" >&2
|
||||||
|
summary; exit 1
|
||||||
|
}
|
||||||
|
summary() {
|
||||||
|
echo; echo "scriberr-rebuild upstream=$SHA7 patches=$PATCH_SUM tag=$TAG"
|
||||||
|
printf ' %s\n' "${RESULTS[@]}"
|
||||||
|
}
|
||||||
|
record() { # best effort: the audit trail must not become a failure point
|
||||||
|
"$REPO_ROOT/scripts/ops-log" record --host fv-ml1 --action "$1" --target "$2" \
|
||||||
|
--detail "$3" >/dev/null 2>&1 || echo "(ops-log record failed; continuing)" >&2
|
||||||
|
}
|
||||||
|
remote() { "${SSH[@]}" bash -s -- "$@"; }
|
||||||
|
|
||||||
|
# Whatever happens (a FAIL, Ctrl-C, a dropped connection), never leave the
|
||||||
|
# nvidia-smi sampler or a transient GPU container behind on fv-ml1.
|
||||||
|
cleanup() {
|
||||||
|
"${SSH[@]}" "[ -s $SAMPLES.pid ] && kill \$(cat $SAMPLES.pid) 2>/dev/null; : >$SAMPLES.pid; \
|
||||||
|
docker rm -f $CNAME >/dev/null 2>&1; true" 2>/dev/null || true
|
||||||
|
}
|
||||||
|
trap cleanup EXIT
|
||||||
|
trap 'exit 130' INT TERM
|
||||||
|
# An unguarded remote call that fails must still say so and print the table.
|
||||||
|
set -E
|
||||||
|
trap 'echo "== ABORT: unexpected failure at line $LINENO (see output above)" >&2; summary' ERR
|
||||||
|
|
||||||
|
# Every container run: the new image, the live env READ-ONLY, the build tree
|
||||||
|
# read-only, and nothing else writable but the container's own /tmp.
|
||||||
|
DOCKER_RUN="docker run --rm --user 1002:1003 -e HOME=/tmp -e USER=infra-ops -e LOGNAME=infra-ops \
|
||||||
|
-e PYTHONDONTWRITEBYTECODE=1 -e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True -e UV_LINK_MODE=copy \
|
||||||
|
-v $ENV_DIR:/app/whisperx-env:ro -v $BUILD_DIR:/src:ro -v $TOOLS:/tools:ro --entrypoint bash"
|
||||||
|
UVRUN="uv run --native-tls --project /app/whisperx-env/parakeet"
|
||||||
|
|
||||||
|
# ── clone ──────────────────────────────────────────────────────────────────
|
||||||
|
if out=$(remote "$BUILD_DIR" "$UPSTREAM" "$SHA" "$TOOLS" "${PATCHED_PATHS[@]}" 2>&1 <<'EOF'
|
||||||
|
set -euo pipefail
|
||||||
|
dir=$1 upstream=$2 sha=$3 tools=$4
|
||||||
|
shift 4
|
||||||
|
sudo -n install -d -o infra-ops -g infra-ops "$tools" "$tools/fixtures" | cat
|
||||||
|
if [ -e "$dir" ]; then
|
||||||
|
[ "$(git -C "$dir" rev-parse HEAD 2>/dev/null)" = "$sha" ] \
|
||||||
|
|| { echo "$dir exists but is not a checkout of $sha; remove it by hand (sudo -n rm -rf $dir) or pick --suffix"; exit 1; }
|
||||||
|
extra=$({ git -C "$dir" diff --name-only HEAD; git -C "$dir" ls-files --others --exclude-standard; } \
|
||||||
|
| sort -u | grep -vxF -f <(printf '%s\n' "$@") || true)
|
||||||
|
[ -z "$extra" ] || { echo "$dir has changes outside the patches ($extra); remove it by hand or pick --suffix"; exit 1; }
|
||||||
|
echo "reusing $dir"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
sudo -n install -d -o infra-ops -g infra-ops "$dir" | cat
|
||||||
|
cd "$dir"
|
||||||
|
git init -q
|
||||||
|
git remote add origin "$upstream"
|
||||||
|
git fetch -q --depth 1 origin "$sha" \
|
||||||
|
|| { echo "fetch of $sha failed; $dir is left empty, remove it by hand (sudo -n rm -rf $dir)"; exit 1; }
|
||||||
|
git checkout -q --detach FETCH_HEAD
|
||||||
|
[ "$(git rev-parse HEAD)" = "$sha" ] || { echo "checked out $(git rev-parse HEAD), wanted $sha"; exit 1; }
|
||||||
|
echo "cloned $sha into $dir"
|
||||||
|
EOF
|
||||||
|
); then pass clone "$out"; else fail clone "$out"; fi
|
||||||
|
if [[ $out == reusing* ]]; then REUSED=1; else REUSED=0; record create "$BUILD_DIR" "scriberr-rebuild: clean clone of upstream $SHA7"; fi
|
||||||
|
|
||||||
|
# ── patch ──────────────────────────────────────────────────────────────────
|
||||||
|
for p in "${PATCHES[@]}"; do
|
||||||
|
name=$(basename "$p")
|
||||||
|
# Already applied (a reused dir)? `apply --reverse --check` succeeds only then.
|
||||||
|
if "${SSH[@]}" "cd $BUILD_DIR && git apply --reverse --check -" <"$p" >/dev/null 2>&1; then
|
||||||
|
pass patch "$name already applied"
|
||||||
|
continue
|
||||||
|
fi
|
||||||
|
if ! out=$("${SSH[@]}" "cd $BUILD_DIR && git apply --check -" <"$p" 2>&1); then
|
||||||
|
if [ "$REUSED" = 1 ]; then
|
||||||
|
fail patch "$name does not apply to the REUSED $BUILD_DIR, which holds an older state of the patches; pick a new --suffix or remove the dir by hand (sudo -n rm -rf $BUILD_DIR):
|
||||||
|
$out"
|
||||||
|
fi
|
||||||
|
fail patch "$name DOES NOT APPLY to upstream $SHA7 — rebase the patch before upgrading:
|
||||||
|
$out"
|
||||||
|
fi
|
||||||
|
"${SSH[@]}" "cd $BUILD_DIR && git apply -" <"$p" || fail patch "$name: git apply failed after a clean --check"
|
||||||
|
record patch "$BUILD_DIR" "scriberr-rebuild: git apply $name"
|
||||||
|
pass patch "$name applied"
|
||||||
|
done
|
||||||
|
|
||||||
|
# ── build ──────────────────────────────────────────────────────────────────
|
||||||
|
if "${SSH[@]}" "docker image inspect $TAG >/dev/null 2>&1"; then
|
||||||
|
[ "$REUSE_IMAGE" = 1 ] || fail build "$TAG already exists; pass --reuse-image to test it, or pick a new --suffix"
|
||||||
|
pass build "reusing existing $TAG"
|
||||||
|
else
|
||||||
|
free=$("${SSH[@]}" "df -BG --output=avail \$(docker info -f '{{.DockerRootDir}}') | tail -1") \
|
||||||
|
|| fail build "could not read free space on fv-ml1"
|
||||||
|
free=${free//[!0-9]/}
|
||||||
|
[ "${free:-0}" -ge "$MIN_FREE_GB" ] \
|
||||||
|
|| fail build "only ${free:-?} GB free under Docker's root dir, need $MIN_FREE_GB; remove superseded scriberr tags first (patches/README.md)"
|
||||||
|
echo "building $TAG (several minutes; log on fv-ml1 at $BUILD_LOG)"
|
||||||
|
if "${SSH[@]}" "docker build -f $BUILD_DIR/Dockerfile.cuda.12.9 -t $TAG \
|
||||||
|
--label org.phasefinal.scriberr.upstream=$SHA --label org.phasefinal.scriberr.patches=$PATCH_SUM \
|
||||||
|
$BUILD_DIR >$BUILD_LOG 2>&1"; then
|
||||||
|
record build "$TAG" "scriberr-rebuild: upstream $SHA7 + patches $PATCH_SUM; old images kept"
|
||||||
|
pass build "$TAG ($free GB was free)"
|
||||||
|
else
|
||||||
|
"${SSH[@]}" "tail -25 $BUILD_LOG" >&2 || true
|
||||||
|
fail build "docker build failed (log tail above)"
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── embed ──────────────────────────────────────────────────────────────────
|
||||||
|
if out=$(remote "$TAG" "$BUILD_DIR" "$SCRIPT_REL" 2>&1 <<'EOF'
|
||||||
|
docker run --rm -v "$2":/src:ro --entrypoint python3 "$1" -c "
|
||||||
|
import sys
|
||||||
|
script = open('/src/$3', 'rb').read()
|
||||||
|
sys.exit(0 if script in open('/app/scriberr', 'rb').read() else 1)"
|
||||||
|
EOF
|
||||||
|
); then
|
||||||
|
pass embed "patched $(basename "$SCRIPT_REL") is byte-identical inside /app/scriberr"
|
||||||
|
else
|
||||||
|
fail embed "the binary does not embed the patched script ${out:+($out)}"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ── unit ───────────────────────────────────────────────────────────────────
|
||||||
|
if out=$("${SSH[@]}" "$DOCKER_RUN $TAG -c 'cd /tmp && $UVRUN --with pytest \
|
||||||
|
python -m pytest -q -p no:cacheprovider /src/$TEST_REL 2>&1 | tail -3'" 2>&1) \
|
||||||
|
&& grep -q ' passed' <<<"$out" && ! grep -Eq 'failed|error' <<<"$out"; then
|
||||||
|
pass unit "$(tail -1 <<<"$out")"
|
||||||
|
else
|
||||||
|
fail unit "$out"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# The checker travels with this script, so the host copy is refreshed each run.
|
||||||
|
"${SCP[@]}" "$REPO_ROOT/scripts/scriberr-seam-check.py" "$HOST:$TOOLS/seam-check.py" \
|
||||||
|
|| fail seam "could not copy the seam checker to fv-ml1:$TOOLS"
|
||||||
|
record update "$TOOLS/seam-check.py" "scriberr-rebuild: refreshed the seam checker"
|
||||||
|
|
||||||
|
# One GPU run of the patched script under the production invocation, output
|
||||||
|
# kept inside the container, validated there. Prints the seam checker's line.
|
||||||
|
gpu_run() { # $1 audio path on host, $2 --chunk-len, $3 --min-chunks
|
||||||
|
local audio_dir; audio_dir=$(dirname "$1")
|
||||||
|
"${SSH[@]}" "$DOCKER_RUN --name $CNAME --gpus '\"device=$GPU\"' -e NVIDIA_VISIBLE_DEVICES=$GPU \
|
||||||
|
-v $audio_dir:/audio:ro $TAG -c 'cd /tmp && $UVRUN python /src/$SCRIPT_REL /audio/$(basename "$1") \
|
||||||
|
--output /tmp/out.json --chunk-len $2 >/tmp/run.log 2>&1 || { tail -5 /tmp/run.log; exit 1; }; \
|
||||||
|
python3 /tools/seam-check.py /tmp/out.json --min-chunks $3'"
|
||||||
|
}
|
||||||
|
gpu_idle() {
|
||||||
|
local used
|
||||||
|
used=$("${SSH[@]}" "nvidia-smi -i $GPU --query-gpu=memory.used --format=csv,noheader,nounits") \
|
||||||
|
|| fail "$1" "could not read GPU $GPU memory on fv-ml1"
|
||||||
|
used=${used//[!0-9]/}
|
||||||
|
[ -n "$used" ] && [ "$used" -lt 100 ] \
|
||||||
|
|| fail "$1" "GPU $GPU is not idle (${used:-?} MiB used); refusing to share a live card"
|
||||||
|
}
|
||||||
|
|
||||||
|
# ── seam ───────────────────────────────────────────────────────────────────
|
||||||
|
gpu_idle seam
|
||||||
|
if out=$(gpu_run "$BUILD_DIR/$SEAM_AUDIO_REL" 10 2 2>&1); then pass seam "$(tail -1 <<<"$out")"
|
||||||
|
else fail seam "$out"; fi
|
||||||
|
|
||||||
|
# ── memory ─────────────────────────────────────────────────────────────────
|
||||||
|
"${SSH[@]}" "test -s $MEM_AUDIO" || fail memory "memory audio $MEM_AUDIO not found on fv-ml1"
|
||||||
|
gpu_idle memory
|
||||||
|
"${SSH[@]}" "nohup nvidia-smi -i $GPU --query-compute-apps=pid,used_memory \
|
||||||
|
--format=csv,noheader,nounits -lms 200 </dev/null >$SAMPLES 2>/dev/null & echo \$! >$SAMPLES.pid" \
|
||||||
|
|| fail memory "could not start the nvidia-smi sampler"
|
||||||
|
record run "$CNAME" "transient $TAG on GPU $GPU, --chunk-len 120, env ro; removed on exit"
|
||||||
|
rc=0; out=$(gpu_run "$MEM_AUDIO" 120 2 2>&1) || rc=$? # `||`, not set +e: keeps the ERR trap quiet
|
||||||
|
"${SSH[@]}" "kill \$(cat $SAMPLES.pid) 2>/dev/null; : >$SAMPLES.pid" || true
|
||||||
|
read -r pids peak < <("${SSH[@]}" \
|
||||||
|
"awk -F', *' 'NF==2 {if (!(\$1 in p)) {p[\$1]=1; n++}; if (\$2+0>m) m=\$2+0} END {print n+0, m+0}' $SAMPLES") \
|
||||||
|
|| fail memory "no GPU samples could be read back from $SAMPLES"
|
||||||
|
[ "$rc" = 0 ] || fail memory "memory run failed: $out"
|
||||||
|
[ "$pids" = 1 ] || fail memory "saw $pids processes on GPU $GPU during the run; the peak is not attributable"
|
||||||
|
if [ "$peak" -le "$BUDGET" ]; then
|
||||||
|
pass memory "peak $peak MiB <= budget $BUDGET MiB (0.2 s samples, GPU $GPU); $(tail -1 <<<"$out")"
|
||||||
|
else
|
||||||
|
fail memory "peak $peak MiB > budget $BUDGET MiB — do NOT deploy beside intern-decision"
|
||||||
|
fi
|
||||||
|
|
||||||
|
summary
|
||||||
|
echo
|
||||||
|
echo "VERDICT: PASS. Deploy is manual: stacks/scriberr/patches/README.md § Deploy (SCRIBERR_IMAGE=$TAG)."
|
||||||
Executable
+86
@@ -0,0 +1,86 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Validate a parakeet_transcribe_buffered.py result against the Go seam.
|
||||||
|
|
||||||
|
Scriberr's parakeet_adapter.go (parseResult) unmarshals this JSON into a struct
|
||||||
|
with typed fields; a float where Go expects an int, or a missing key, fails the
|
||||||
|
job. This checks the shape Go reads plus the stitching invariants the slicer
|
||||||
|
patch promises. Stdlib only, so it runs under any python3. Prints counts, never
|
||||||
|
transcript text.
|
||||||
|
|
||||||
|
usage: scriberr-seam-check.py RESULT.json [--min-chunks N]
|
||||||
|
"""
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import sys
|
||||||
|
|
||||||
|
NUMBER = (int, float)
|
||||||
|
|
||||||
|
|
||||||
|
def fail(msg):
|
||||||
|
print(f"SEAM FAIL: {msg}")
|
||||||
|
sys.exit(1)
|
||||||
|
|
||||||
|
|
||||||
|
def check_items(items, text_key, name):
|
||||||
|
for i, item in enumerate(items):
|
||||||
|
if not isinstance(item, dict):
|
||||||
|
fail(f"{name}[{i}] is not an object")
|
||||||
|
if not isinstance(item.get(text_key), str):
|
||||||
|
fail(f"{name}[{i}].{text_key} is not a string")
|
||||||
|
for key in ("start_offset", "end_offset"):
|
||||||
|
if type(item.get(key)) is not int:
|
||||||
|
fail(f"{name}[{i}].{key} is not an integer (Go field is int)")
|
||||||
|
for key in ("start", "end"):
|
||||||
|
if not isinstance(item.get(key), NUMBER) or isinstance(item.get(key), bool):
|
||||||
|
fail(f"{name}[{i}].{key} is not a number")
|
||||||
|
if item["start"] > item["end"]:
|
||||||
|
fail(f"{name}[{i}] starts after it ends")
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
parser = argparse.ArgumentParser(description="Validate a buffered Parakeet result for Go.")
|
||||||
|
parser.add_argument("result", help="result JSON written by parakeet_transcribe_buffered.py")
|
||||||
|
parser.add_argument("--min-chunks", type=int, default=1,
|
||||||
|
help="fail unless the run used at least this many chunks")
|
||||||
|
args = parser.parse_args()
|
||||||
|
min_chunks = args.min_chunks
|
||||||
|
try:
|
||||||
|
data = json.load(open(args.result, encoding="utf-8"))
|
||||||
|
except (OSError, ValueError) as e:
|
||||||
|
fail(f"cannot read {args.result}: {e}")
|
||||||
|
if not isinstance(data, dict):
|
||||||
|
fail("the result is not a JSON object")
|
||||||
|
|
||||||
|
required = {"transcription": str, "language": str, "word_timestamps": list,
|
||||||
|
"segment_timestamps": list, "audio_file": str, "model": str}
|
||||||
|
for key, kind in required.items():
|
||||||
|
if not isinstance(data.get(key), kind):
|
||||||
|
fail(f"'{key}' missing or not {kind.__name__}")
|
||||||
|
if data.get("buffered") is not True:
|
||||||
|
fail("'buffered' is not true")
|
||||||
|
if not isinstance(data.get("chunk_duration_secs"), NUMBER):
|
||||||
|
fail("'chunk_duration_secs' is not a number")
|
||||||
|
if type(data.get("num_chunks")) is not int or data["num_chunks"] < min_chunks:
|
||||||
|
fail(f"'num_chunks' is not an integer >= {min_chunks}")
|
||||||
|
|
||||||
|
words, segments = data["word_timestamps"], data["segment_timestamps"]
|
||||||
|
if not words or not data["transcription"].strip():
|
||||||
|
fail("empty transcript")
|
||||||
|
check_items(words, "word", "word_timestamps")
|
||||||
|
check_items(segments, "segment", "segment_timestamps")
|
||||||
|
|
||||||
|
starts = [w["start"] for w in words]
|
||||||
|
if starts != sorted(starts):
|
||||||
|
fail("word start times go backwards (a stitch repeated or reordered words)")
|
||||||
|
joined = " ".join(w["word"] for w in words)
|
||||||
|
if data["transcription"] != joined:
|
||||||
|
fail("'transcription' is not the stitched words joined by spaces")
|
||||||
|
if " ".join(s["segment"] for s in segments) != joined:
|
||||||
|
fail("segments do not cover the stitched words exactly once, in order")
|
||||||
|
|
||||||
|
print(f"SEAM OK: {len(words)} words, {len(segments)} segments, "
|
||||||
|
f"{data['num_chunks']} chunks, cuts at {len(data.get('cut_times', []))} points")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -4,6 +4,11 @@
|
|||||||
# ── Image ────────────────────────────────────────────────────────────────
|
# ── Image ────────────────────────────────────────────────────────────────
|
||||||
# Built locally from Dockerfile.cuda.12.9 — see the compose header for why
|
# Built locally from Dockerfile.cuda.12.9 — see the compose header for why
|
||||||
# the published scriberr-cuda image is NOT usable on these Blackwell cards.
|
# the published scriberr-cuda image is NOT usable on these Blackwell cards.
|
||||||
|
# Patched builds come from scripts/scriberr-rebuild and are tagged
|
||||||
|
# scriberr:local-blackwell-<upstream sha7>-<suffix> (e.g. -a353078-slicer1)
|
||||||
|
# so every build keeps its own tag. Deploying = pointing this at a new tag;
|
||||||
|
# rollback = pointing it back. See stacks/scriberr/patches/README.md.
|
||||||
|
# (Unset falls back to the original unpatched scriberr:local-blackwell.)
|
||||||
SCRIBERR_IMAGE=scriberr:local-blackwell
|
SCRIBERR_IMAGE=scriberr:local-blackwell
|
||||||
|
|
||||||
# ── Network ──────────────────────────────────────────────────────────────
|
# ── Network ──────────────────────────────────────────────────────────────
|
||||||
|
|||||||
@@ -27,15 +27,21 @@ image** — it will fail on these cards or quietly fall back to CPU.
|
|||||||
|
|
||||||
### Rebuilding
|
### Rebuilding
|
||||||
|
|
||||||
|
We carry local patches (`patches/`, currently the pause-aware Parakeet
|
||||||
|
slicer), so a rebuild is one command from nh3-dev, pinned to an upstream sha:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
ssh fv-ml1
|
scripts/scriberr-rebuild --sha <full upstream sha> --suffix slicer1
|
||||||
cd /tank/scriberr/src/Scriberr
|
|
||||||
git pull
|
|
||||||
docker build -f Dockerfile.cuda.12.9 -t scriberr:local-blackwell .
|
|
||||||
cd /opt/docker/compose/scriberr && docker compose up -d
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Source checkout lives on `/tank`, not the root pool — see storage below.
|
It makes a clean clone in `/opt/docker/src/scriberr-<sha7>-<suffix>` on
|
||||||
|
fv-ml1, `git apply --check`s the patches (a conflict stops it), builds
|
||||||
|
`scriberr:local-blackwell-<sha7>-<suffix>` beside the old images, and checks
|
||||||
|
the embed, the unit tests, the Go↔Python JSON seam, and the GPU memory budget.
|
||||||
|
Deploying it is a separate manual step: `patches/README.md` § Deploy.
|
||||||
|
|
||||||
|
The old checkout at `/tank/scriberr/src/Scriberr` (lkraven-owned, shallow)
|
||||||
|
built the original `scriberr:local-blackwell` and is left as it was.
|
||||||
|
|
||||||
## Deploy
|
## Deploy
|
||||||
|
|
||||||
@@ -62,7 +68,7 @@ mounts are bind-mounted onto `/tank` (4+ TB) instead of named volumes:
|
|||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `/tank/scriberr/data` | `/app/data` | SQLite DB, uploads, transcripts |
|
| `/tank/scriberr/data` | `/app/data` | SQLite DB, uploads, transcripts |
|
||||||
| `/tank/scriberr/whisperx-env` | `/app/whisperx-env` | Python env + model weights |
|
| `/tank/scriberr/whisperx-env` | `/app/whisperx-env` | Python env + model weights |
|
||||||
| `/tank/scriberr/src/Scriberr` | — | build checkout |
|
| `/tank/scriberr/src/Scriberr` | — | original build checkout (patched builds: `/opt/docker/src/scriberr-<sha7>-<suffix>`) |
|
||||||
|
|
||||||
Both are owned by uid/gid 1000 to match `PUID`/`PGID`.
|
Both are owned by uid/gid 1000 to match `PUID`/`PGID`.
|
||||||
|
|
||||||
@@ -164,3 +170,11 @@ Peak GPU memory on a 35-minute file:
|
|||||||
- The two env knobs are read by upstream's Go code (`parakeet_adapter.go`), so
|
- The two env knobs are read by upstream's Go code (`parakeet_adapter.go`), so
|
||||||
they survive image upgrades for as long as upstream keeps them. Re-measure the
|
they survive image upgrades for as long as upstream keeps them. Re-measure the
|
||||||
peak after any upgrade.
|
peak after any upgrade.
|
||||||
|
- **The slicer itself is patched** (`patches/0001-parakeet-pause-aware-slicer.patch`,
|
||||||
|
2026-09-30). Adjacent 120 s slices now overlap by 4 s and are stitched at a word
|
||||||
|
both transcribed, which cut the share of cuts with an error nearby from 52 % to
|
||||||
|
22 % against a 19 % background. The overlap sits *inside* the 120 s, so the
|
||||||
|
peak is unchanged (5,496 MiB, n=3). See `patches/README.md` and
|
||||||
|
`docs/pfi/scriberr-slicer-bench-2026-09-30.md`. That bench also found that
|
||||||
|
Parakeet sometimes skips stretches of ≥10 words mid-slice, with or without the
|
||||||
|
patch; that is still open.
|
||||||
|
|||||||
@@ -0,0 +1,686 @@
|
|||||||
|
From 2dafe7ce81c217609ff2c8616e43b0d72255e170 Mon Sep 17 00:00:00 2001
|
||||||
|
From: Vuong Hoang <vh@phasefinal.com>
|
||||||
|
Date: Wed, 30 Sep 2026 11:41:07 -0700
|
||||||
|
Subject: [PATCH] fix(parakeet): overlap buffered chunks and stitch at an
|
||||||
|
agreed word
|
||||||
|
|
||||||
|
parakeet_transcribe_buffered.py cut long audio at fixed --chunk-len marks
|
||||||
|
with no overlap, so a word straddling a mark was chopped, dropped or
|
||||||
|
transcribed twice. Adjacent chunks now overlap by --overlap seconds
|
||||||
|
(default 4, counted inside --chunk-len so no chunk grows), and in each
|
||||||
|
overlap the chunks hand over at the word nearest the cut that both
|
||||||
|
transcribed with the same text at nearly the same time (within 0.5 s),
|
||||||
|
keeping whichever copy leaves the words in time order. With no such word
|
||||||
|
they split at the cut. Splitting both chunks at the cut by word start time
|
||||||
|
is not enough on its own: a word after a pause can be timestamped anywhere
|
||||||
|
in the pause, so the two chunks may place it on opposite sides of the cut.
|
||||||
|
|
||||||
|
--pause-search N (opt-in) also moves each cut back to the quietest 0.3 s
|
||||||
|
in the last N seconds before the limit. Measured neutral on top of the
|
||||||
|
overlap, so it is off by default.
|
||||||
|
|
||||||
|
The CLI and JSON the Go adapter reads are unchanged; the new flags are
|
||||||
|
optional, and the JSON gains overlap_secs, pause_search_secs and
|
||||||
|
cut_times. --overlap 0 reproduces the previous output exactly. NeMo is
|
||||||
|
now imported inside transcribe_buffered() so the slicing and stitching
|
||||||
|
helpers can be unit-tested without a GPU. If a chunk ever returns text
|
||||||
|
without word timestamps, its text is kept rather than dropped.
|
||||||
|
---
|
||||||
|
.../py/nvidia/parakeet_transcribe_buffered.py | 221 ++++++++++--
|
||||||
|
.../py/nvidia/tests/test_parakeet_slicing.py | 329 ++++++++++++++++++
|
||||||
|
2 files changed, 525 insertions(+), 25 deletions(-)
|
||||||
|
create mode 100644 internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
|
||||||
|
|
||||||
|
diff --git a/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py b/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
|
||||||
|
index 29d5047..ba755c1 100644
|
||||||
|
--- a/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
|
||||||
|
+++ b/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
|
||||||
|
@@ -2,6 +2,10 @@
|
||||||
|
"""
|
||||||
|
NVIDIA Parakeet buffered inference for long audio files.
|
||||||
|
Splits audio into chunks to avoid GPU memory issues.
|
||||||
|
+
|
||||||
|
+Adjacent chunks overlap slightly, and in each overlap the chunks hand over at
|
||||||
|
+a word both transcribed alike, so a word near a cut is neither chopped, dropped
|
||||||
|
+nor repeated. Optionally (--pause-search) each cut also moves into a pause.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
@@ -12,37 +16,178 @@ import librosa
|
||||||
|
import soundfile as sf
|
||||||
|
import numpy as np
|
||||||
|
from pathlib import Path
|
||||||
|
-import nemo.collections.asr as nemo_asr
|
||||||
|
|
||||||
|
+DEFAULT_OVERLAP_SECS = 4.0
|
||||||
|
+DEFAULT_PAUSE_SEARCH_SECS = 0.0 # opt-in; measured no gain on top of the overlap
|
||||||
|
+QUIET_WINDOW_SECS = 0.3
|
||||||
|
+SAME_WORD_SECS = 0.5
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def plan_slices(audio, sr, max_chunk_secs, overlap_secs=0.0, search_secs=0.0):
|
||||||
|
+ """Choose where to cut `audio` so no chunk exceeds `max_chunk_secs`.
|
||||||
|
|
||||||
|
-def split_audio_file(audio_path, chunk_duration_secs=300):
|
||||||
|
- """Split audio file into chunks of specified duration."""
|
||||||
|
+ Returns (spans, cuts): `cuts` are the sample indices where one chunk's
|
||||||
|
+ share of the audio ends and the next one's begins; `spans` are the
|
||||||
|
+ (start, end) samples actually transcribed, each cut-to-cut range widened
|
||||||
|
+ by half the overlap on both sides. The overlap counts towards the limit.
|
||||||
|
+
|
||||||
|
+ With `search_secs` > 0, each cut moves back from the limit to the middle
|
||||||
|
+ of the quietest QUIET_WINDOW_SECS window within the last `search_secs`.
|
||||||
|
+ """
|
||||||
|
+ overlap_secs = max(0.0, min(overlap_secs, max_chunk_secs / 4))
|
||||||
|
+ step = int((max_chunk_secs - overlap_secs) * sr)
|
||||||
|
+ if step < 1:
|
||||||
|
+ raise ValueError(f"chunk length must be positive, got {max_chunk_secs}s")
|
||||||
|
+ search = min(int(search_secs * sr), step // 2)
|
||||||
|
+ window = max(1, int(QUIET_WINDOW_SECS * sr))
|
||||||
|
+
|
||||||
|
+ cuts = []
|
||||||
|
+ position = 0
|
||||||
|
+ while len(audio) - position > step:
|
||||||
|
+ cut = position + step
|
||||||
|
+ if search > window:
|
||||||
|
+ cut = _quietest_point(audio, cut - search, cut, window)
|
||||||
|
+ cuts.append(cut)
|
||||||
|
+ position = cut
|
||||||
|
+
|
||||||
|
+ pad = int(overlap_secs * sr) // 2
|
||||||
|
+ edges = [0] + cuts + [len(audio)]
|
||||||
|
+ spans = [(max(0, start - pad), min(len(audio), end + pad))
|
||||||
|
+ for start, end in zip(edges, edges[1:])]
|
||||||
|
+ return spans, cuts
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def _quietest_point(audio, start, end, window):
|
||||||
|
+ """Middle of the lowest-energy `window` samples within audio[start:end].
|
||||||
|
+
|
||||||
|
+ Ties go to the latest window, which keeps chunks as long as allowed.
|
||||||
|
+ """
|
||||||
|
+ x = audio[start:end].astype(np.float64)
|
||||||
|
+ cumulative = np.concatenate(([0.0], np.cumsum(x * x)))
|
||||||
|
+ energy = cumulative[window:] - cumulative[:-window]
|
||||||
|
+ latest_min = len(energy) - 1 - int(np.argmin(energy[::-1]))
|
||||||
|
+ return start + latest_min + window // 2
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def stitch_slices(slice_results, cut_times, chunk_spans=None):
|
||||||
|
+ """Merge per-chunk (words, segments), already shifted to absolute time.
|
||||||
|
+
|
||||||
|
+ `chunk_spans` gives each chunk's (start, end) in seconds; omit it when the
|
||||||
|
+ chunks do not overlap. Each chunk contributes the words between its two
|
||||||
|
+ handovers. A handover is at the cut, unless the chunks overlap: then it
|
||||||
|
+ moves to the nearest word in the overlap that both chunks transcribed
|
||||||
|
+ alike, at the same time, and the left chunk keeps the words before it,
|
||||||
|
+ the right chunk that word and the ones after. (Splitting both chunks at
|
||||||
|
+ the cut is fragile: a word that follows a pause can be timestamped
|
||||||
|
+ anywhere in the pause, so the two chunks may put it on opposite sides of
|
||||||
|
+ the cut and keep it twice, or not at all.)
|
||||||
|
+
|
||||||
|
+ Segments are trimmed to the words their chunk keeps, and dropped if none.
|
||||||
|
+ """
|
||||||
|
+ first = [0] * len(slice_results)
|
||||||
|
+ last = [len(chunk_words) for chunk_words, _ in slice_results]
|
||||||
|
+ for k, cut in enumerate(cut_times):
|
||||||
|
+ overlap = (chunk_spans[k + 1][0], chunk_spans[k][1]) if chunk_spans else (cut, cut)
|
||||||
|
+ last[k], first[k + 1] = _handover(
|
||||||
|
+ slice_results[k][0], slice_results[k + 1][0], cut, overlap
|
||||||
|
+ )
|
||||||
|
+
|
||||||
|
+ words, segments = [], []
|
||||||
|
+ for (chunk_words, chunk_segments), lo, hi in zip(slice_results, first, last):
|
||||||
|
+ words.extend(chunk_words[lo:hi])
|
||||||
|
+ for seg, (start, stop) in zip(chunk_segments, _segment_ranges(chunk_words, chunk_segments)):
|
||||||
|
+ kept = chunk_words[max(start, lo):min(stop, hi)]
|
||||||
|
+ if start == stop: # a segment without words: keep it where its chunk does
|
||||||
|
+ if lo <= start < hi:
|
||||||
|
+ segments.append(seg)
|
||||||
|
+ elif len(kept) == stop - start:
|
||||||
|
+ segments.append(seg)
|
||||||
|
+ elif kept:
|
||||||
|
+ segments.append({
|
||||||
|
+ **seg,
|
||||||
|
+ "segment": " ".join(w["word"] for w in kept),
|
||||||
|
+ "start_offset": kept[0]["start_offset"],
|
||||||
|
+ "end_offset": kept[-1]["end_offset"],
|
||||||
|
+ "start": kept[0]["start"],
|
||||||
|
+ "end": kept[-1]["end"],
|
||||||
|
+ })
|
||||||
|
+ return words, segments
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def _handover(left, right, cut, overlap):
|
||||||
|
+ """(i, j): the left chunk keeps left[:i] and the right chunk right[j:].
|
||||||
|
+
|
||||||
|
+ Anchors are words in the overlap that both chunks transcribed with the same
|
||||||
|
+ text at nearly the same time. At the anchor nearest the cut, the right
|
||||||
|
+ chunk's copy is kept, or the left chunk's if that is what keeps the words
|
||||||
|
+ in time order. Without an anchor, both chunks split at the cut.
|
||||||
|
+ """
|
||||||
|
+ i = sum(1 for w in left if w["start"] < cut)
|
||||||
|
+ j = sum(1 for w in right if w["start"] < cut)
|
||||||
|
+ in_order = lambda a, b: a == 0 or b == len(right) or left[a - 1]["start"] <= right[b]["start"]
|
||||||
|
+ anchors = []
|
||||||
|
+ for p, lw in enumerate(left):
|
||||||
|
+ text = _normalize(lw["word"])
|
||||||
|
+ if not text or not overlap[0] <= lw["start"] < overlap[1]:
|
||||||
|
+ continue
|
||||||
|
+ partners = [q for q, rw in enumerate(right) if _normalize(rw["word"]) == text
|
||||||
|
+ and abs(rw["start"] - lw["start"]) <= SAME_WORD_SECS]
|
||||||
|
+ if partners:
|
||||||
|
+ q = min(partners, key=lambda q: abs(right[q]["start"] - lw["start"]))
|
||||||
|
+ options = [h for h in ((p, q), (p + 1, q + 1)) if in_order(*h)]
|
||||||
|
+ if options:
|
||||||
|
+ anchors.append((abs(lw["start"] + right[q]["start"] - 2 * cut), options[0]))
|
||||||
|
+ if anchors:
|
||||||
|
+ i, j = min(anchors)[1]
|
||||||
|
+ return i, j
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def _normalize(word):
|
||||||
|
+ return "".join(c for c in word.lower() if c.isalnum() or c == "'")
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def _segment_ranges(words, segments):
|
||||||
|
+ """[start, stop) word indices of each segment, matched in order by frame offsets."""
|
||||||
|
+ ranges, i = [], 0
|
||||||
|
+ for seg in segments:
|
||||||
|
+ while i < len(words) and words[i]["start_offset"] < seg["start_offset"]:
|
||||||
|
+ i += 1
|
||||||
|
+ start = i
|
||||||
|
+ while i < len(words) and words[i]["end_offset"] <= seg["end_offset"]:
|
||||||
|
+ i += 1
|
||||||
|
+ ranges.append((start, i))
|
||||||
|
+ return ranges
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def split_audio_file(audio_path, chunk_duration_secs=300, overlap_secs=0.0, search_secs=0.0):
|
||||||
|
+ """Split audio file into chunks of at most chunk_duration_secs."""
|
||||||
|
audio, sr = librosa.load(audio_path, sr=None, mono=True)
|
||||||
|
- total_duration = len(audio) / sr
|
||||||
|
- chunk_samples = int(chunk_duration_secs * sr)
|
||||||
|
+ spans, cuts = plan_slices(audio, sr, chunk_duration_secs, overlap_secs, search_secs)
|
||||||
|
|
||||||
|
chunks = []
|
||||||
|
- for start_sample in range(0, len(audio), chunk_samples):
|
||||||
|
- end_sample = min(start_sample + chunk_samples, len(audio))
|
||||||
|
+ for start_sample, end_sample in spans:
|
||||||
|
chunk_audio = audio[start_sample:end_sample]
|
||||||
|
- start_time = start_sample / sr
|
||||||
|
chunks.append({
|
||||||
|
'audio': chunk_audio,
|
||||||
|
- 'start_time': start_time,
|
||||||
|
+ 'start_time': start_sample / sr,
|
||||||
|
'duration': len(chunk_audio) / sr
|
||||||
|
})
|
||||||
|
|
||||||
|
- return chunks, sr
|
||||||
|
+ return chunks, sr, [cut / sr for cut in cuts]
|
||||||
|
|
||||||
|
|
||||||
|
def transcribe_buffered(
|
||||||
|
audio_path: str,
|
||||||
|
output_file: str = None,
|
||||||
|
chunk_duration_secs: float = 300, # 5 minutes default
|
||||||
|
+ overlap_secs: float = DEFAULT_OVERLAP_SECS,
|
||||||
|
+ pause_search_secs: float = DEFAULT_PAUSE_SEARCH_SECS,
|
||||||
|
):
|
||||||
|
"""
|
||||||
|
Transcribe long audio by splitting into chunks and merging results.
|
||||||
|
"""
|
||||||
|
+ import nemo.collections.asr as nemo_asr
|
||||||
|
+
|
||||||
|
# Determine model path
|
||||||
|
model_filename = "parakeet-tdt-0.6b-v3.nemo"
|
||||||
|
model_path = None
|
||||||
|
@@ -79,13 +224,15 @@ def transcribe_buffered(
|
||||||
|
asr_model.change_decoding_strategy(dec_cfg)
|
||||||
|
print("✓ CUDA graphs disabled successfully")
|
||||||
|
|
||||||
|
- print(f"Splitting audio into {chunk_duration_secs}s chunks...")
|
||||||
|
- chunks, sr = split_audio_file(audio_path, chunk_duration_secs)
|
||||||
|
+ print(f"Splitting audio into chunks of at most {chunk_duration_secs}s "
|
||||||
|
+ f"(overlap {overlap_secs}s, pause search {pause_search_secs}s)...")
|
||||||
|
+ chunks, sr, cut_times = split_audio_file(
|
||||||
|
+ audio_path, chunk_duration_secs, overlap_secs, pause_search_secs
|
||||||
|
+ )
|
||||||
|
print(f"Created {len(chunks)} chunks")
|
||||||
|
|
||||||
|
- all_words = []
|
||||||
|
- all_segments = []
|
||||||
|
- full_text = []
|
||||||
|
+ slice_results = []
|
||||||
|
+ chunk_texts = []
|
||||||
|
|
||||||
|
for i, chunk_info in enumerate(chunks):
|
||||||
|
print(f"Transcribing chunk {i+1}/{len(chunks)} (duration: {chunk_info['duration']:.1f}s)...")
|
||||||
|
@@ -104,25 +251,26 @@ def transcribe_buffered(
|
||||||
|
|
||||||
|
result_data = output[0]
|
||||||
|
chunk_text = result_data.text
|
||||||
|
- full_text.append(chunk_text)
|
||||||
|
+ chunk_texts.append(chunk_text)
|
||||||
|
+ chunk_words = []
|
||||||
|
+ chunk_segments = []
|
||||||
|
|
||||||
|
# Extract and adjust timestamps
|
||||||
|
if hasattr(result_data, 'timestamp') and result_data.timestamp:
|
||||||
|
- chunk_words = result_data.timestamp.get("word", [])
|
||||||
|
- chunk_segments = result_data.timestamp.get("segment", [])
|
||||||
|
-
|
||||||
|
# Adjust timestamps by chunk start time
|
||||||
|
- for word in chunk_words:
|
||||||
|
+ for word in result_data.timestamp.get("word", []):
|
||||||
|
word_copy = dict(word)
|
||||||
|
word_copy['start'] += chunk_info['start_time']
|
||||||
|
word_copy['end'] += chunk_info['start_time']
|
||||||
|
- all_words.append(word_copy)
|
||||||
|
+ chunk_words.append(word_copy)
|
||||||
|
|
||||||
|
- for segment in chunk_segments:
|
||||||
|
+ for segment in result_data.timestamp.get("segment", []):
|
||||||
|
seg_copy = dict(segment)
|
||||||
|
seg_copy['start'] += chunk_info['start_time']
|
||||||
|
seg_copy['end'] += chunk_info['start_time']
|
||||||
|
- all_segments.append(seg_copy)
|
||||||
|
+ chunk_segments.append(seg_copy)
|
||||||
|
+
|
||||||
|
+ slice_results.append((chunk_words, chunk_segments))
|
||||||
|
|
||||||
|
print(f"Chunk {i+1} complete: {len(chunk_text)} characters")
|
||||||
|
|
||||||
|
@@ -131,7 +279,15 @@ def transcribe_buffered(
|
||||||
|
if os.path.exists(chunk_path):
|
||||||
|
os.remove(chunk_path)
|
||||||
|
|
||||||
|
- final_text = " ".join(full_text)
|
||||||
|
+ chunk_spans = [(c['start_time'], c['start_time'] + c['duration']) for c in chunks]
|
||||||
|
+ all_words, all_segments = stitch_slices(slice_results, cut_times, chunk_spans)
|
||||||
|
+ if any(text.strip() and not words for (words, _), text in zip(slice_results, chunk_texts)):
|
||||||
|
+ # A chunk came back without word timestamps, so there is nothing to
|
||||||
|
+ # stitch it by; keep its text rather than lose it.
|
||||||
|
+ print("Warning: a chunk has text but no word timestamps; joining chunk texts")
|
||||||
|
+ final_text = " ".join(chunk_texts)
|
||||||
|
+ else:
|
||||||
|
+ final_text = " ".join(w["word"] for w in all_words)
|
||||||
|
print(f"Transcription complete: {len(final_text)} characters total")
|
||||||
|
|
||||||
|
output_data = {
|
||||||
|
@@ -144,6 +300,9 @@ def transcribe_buffered(
|
||||||
|
"buffered": True,
|
||||||
|
"chunk_duration_secs": chunk_duration_secs,
|
||||||
|
"num_chunks": len(chunks),
|
||||||
|
+ "overlap_secs": overlap_secs,
|
||||||
|
+ "pause_search_secs": pause_search_secs,
|
||||||
|
+ "cut_times": cut_times,
|
||||||
|
}
|
||||||
|
|
||||||
|
if output_file:
|
||||||
|
@@ -162,7 +321,17 @@ def main():
|
||||||
|
parser.add_argument("--output", "-o", help="Output file path", required=True)
|
||||||
|
parser.add_argument(
|
||||||
|
"--chunk-len", type=float, default=300,
|
||||||
|
- help="Chunk duration in seconds (default: 300 = 5 minutes)"
|
||||||
|
+ help="Maximum chunk duration in seconds, overlap included (default: 300 = 5 minutes)"
|
||||||
|
+ )
|
||||||
|
+ parser.add_argument(
|
||||||
|
+ "--overlap", type=float, default=DEFAULT_OVERLAP_SECS,
|
||||||
|
+ help=f"Seconds shared by adjacent chunks, capped at a quarter of --chunk-len "
|
||||||
|
+ f"(default: {DEFAULT_OVERLAP_SECS}; 0 disables)"
|
||||||
|
+ )
|
||||||
|
+ parser.add_argument(
|
||||||
|
+ "--pause-search", type=float, default=DEFAULT_PAUSE_SEARCH_SECS,
|
||||||
|
+ help=f"Seconds before each chunk limit searched for the quietest point to cut at, "
|
||||||
|
+ f"e.g. 25 (default: {DEFAULT_PAUSE_SEARCH_SECS}, cut at the limit)"
|
||||||
|
)
|
||||||
|
|
||||||
|
args = parser.parse_args()
|
||||||
|
@@ -175,6 +344,8 @@ def main():
|
||||||
|
audio_path=args.audio_file,
|
||||||
|
output_file=args.output,
|
||||||
|
chunk_duration_secs=args.chunk_len,
|
||||||
|
+ overlap_secs=args.overlap,
|
||||||
|
+ pause_search_secs=args.pause_search,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
diff --git a/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py b/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
|
||||||
|
new file mode 100644
|
||||||
|
index 0000000..6a35947
|
||||||
|
--- /dev/null
|
||||||
|
+++ b/internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
|
||||||
|
@@ -0,0 +1,329 @@
|
||||||
|
+"""Unit tests for the slicing and stitching helpers in parakeet_transcribe_buffered.py.
|
||||||
|
+
|
||||||
|
+These are pure functions: they need numpy, librosa and soundfile (imported by the
|
||||||
|
+script) but no GPU, no NeMo and no model.
|
||||||
|
+"""
|
||||||
|
+import sys
|
||||||
|
+from pathlib import Path
|
||||||
|
+
|
||||||
|
+import numpy as np
|
||||||
|
+import pytest
|
||||||
|
+
|
||||||
|
+sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
||||||
|
+from parakeet_transcribe_buffered import plan_slices, stitch_slices # noqa: E402
|
||||||
|
+
|
||||||
|
+SR = 16000
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def speech(seconds, seed=0):
|
||||||
|
+ """Stand-in for continuous speech: broadband noise at a speech-like level."""
|
||||||
|
+ rng = np.random.default_rng(seed)
|
||||||
|
+ return (0.1 * rng.standard_normal(int(seconds * SR))).astype(np.float32)
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def with_pauses(audio, pauses, level=0.001, seed=1):
|
||||||
|
+ """Replace each (start_s, end_s) span with low-level room noise (or zeros)."""
|
||||||
|
+ rng = np.random.default_rng(seed)
|
||||||
|
+ out = audio.copy()
|
||||||
|
+ for start, end in pauses:
|
||||||
|
+ a, b = int(start * SR), int(end * SR)
|
||||||
|
+ out[a:b] = level * rng.standard_normal(b - a)
|
||||||
|
+ return out
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def assert_valid_plan(spans, cuts, num_samples, max_secs):
|
||||||
|
+ assert spans[0][0] == 0 and spans[-1][1] == num_samples
|
||||||
|
+ assert len(spans) == len(cuts) + 1
|
||||||
|
+ for start, end in spans:
|
||||||
|
+ assert 0 < end - start <= max_secs * SR
|
||||||
|
+ for (a0, a1), (b0, b1), cut in zip(spans, spans[1:], cuts):
|
||||||
|
+ assert b0 <= cut <= a1, "each cut must lie inside both neighbouring slices"
|
||||||
|
+ assert a0 < cut < b1
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+# -- plan_slices: cut placement -------------------------------------------------
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_audio_shorter_than_one_slice_is_not_cut():
|
||||||
|
+ audio = speech(60)
|
||||||
|
+ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=4, search_secs=25)
|
||||||
|
+ assert cuts == []
|
||||||
|
+ assert spans == [(0, len(audio))]
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_audio_exactly_one_slice_long_is_not_cut():
|
||||||
|
+ audio = speech(120)
|
||||||
|
+ spans, cuts = plan_slices(audio, SR, 120)
|
||||||
|
+ assert cuts == []
|
||||||
|
+ assert spans == [(0, len(audio))]
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_without_search_or_overlap_the_legacy_fixed_grid_is_reproduced():
|
||||||
|
+ audio = speech(300)
|
||||||
|
+ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=0, search_secs=0)
|
||||||
|
+ assert cuts == [120 * SR, 240 * SR]
|
||||||
|
+ assert spans == [(0, 120 * SR), (120 * SR, 240 * SR), (240 * SR, 300 * SR)]
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_cuts_land_in_the_pauses_before_the_limit():
|
||||||
|
+ pauses = [(100.0, 100.5), (215.0, 215.5), (330.0, 330.5)]
|
||||||
|
+ audio = with_pauses(speech(400), pauses)
|
||||||
|
+ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=0, search_secs=25)
|
||||||
|
+ assert len(cuts) == 3
|
||||||
|
+ for cut, (start, end) in zip(cuts, pauses):
|
||||||
|
+ assert start <= cut / SR <= end
|
||||||
|
+ assert_valid_plan(spans, cuts, len(audio), 120)
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_the_quietest_pause_wins():
|
||||||
|
+ # Two pauses inside the same search window; the later one is louder.
|
||||||
|
+ audio = with_pauses(speech(200), [(100.0, 100.6)], level=0.0)
|
||||||
|
+ audio = with_pauses(audio, [(115.0, 115.6)], level=0.01)
|
||||||
|
+ _, cuts = plan_slices(audio, SR, 120, search_secs=25)
|
||||||
|
+ assert 100.0 <= cuts[0] / SR <= 100.6
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_audio_with_no_pause_is_still_cut_within_the_limit():
|
||||||
|
+ audio = speech(400)
|
||||||
|
+ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=0, search_secs=25)
|
||||||
|
+ assert_valid_plan(spans, cuts, len(audio), 120)
|
||||||
|
+ edges = [0] + cuts
|
||||||
|
+ for prev, cut in zip(edges, cuts):
|
||||||
|
+ assert 95 * SR <= cut - prev <= 120 * SR, "cut must fall inside its search window"
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_a_pause_at_the_very_start_is_never_a_cut():
|
||||||
|
+ audio = with_pauses(speech(200), [(0.0, 5.0)], level=0.0)
|
||||||
|
+ spans, cuts = plan_slices(audio, SR, 120, search_secs=25)
|
||||||
|
+ assert len(cuts) == 1 and 95 <= cuts[0] / SR <= 120
|
||||||
|
+ assert_valid_plan(spans, cuts, len(audio), 120)
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_a_pause_at_the_very_end_leaves_no_empty_slice():
|
||||||
|
+ # First cut ~100.35 s, so the second search window is ~[195, 220] s and
|
||||||
|
+ # holds the start of the trailing silence (217-222 s).
|
||||||
|
+ audio = with_pauses(speech(222), [(100.0, 100.5), (217.0, 222.0)], level=0.0)
|
||||||
|
+ spans, cuts = plan_slices(audio, SR, 120, search_secs=25)
|
||||||
|
+ assert len(cuts) == 2
|
||||||
|
+ assert 217.0 <= cuts[1] / SR < 222.0
|
||||||
|
+ assert_valid_plan(spans, cuts, len(audio), 120)
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+# -- plan_slices: overlap -------------------------------------------------------
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_overlap_is_included_in_the_slice_limit_and_centred_on_each_cut():
|
||||||
|
+ pauses = [(100.0, 100.5), (215.0, 215.5), (330.0, 330.5)]
|
||||||
|
+ audio = with_pauses(speech(400), pauses)
|
||||||
|
+ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=4, search_secs=25)
|
||||||
|
+ assert_valid_plan(spans, cuts, len(audio), 120)
|
||||||
|
+ for (_, a1), (b0, _), cut in zip(spans, spans[1:], cuts):
|
||||||
|
+ assert a1 - b0 == 4 * SR
|
||||||
|
+ assert cut - b0 == a1 - cut
|
||||||
|
+ for cut, (start, end) in zip(cuts, pauses):
|
||||||
|
+ assert start <= cut / SR <= end
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_overlap_without_pause_search_uses_a_fixed_grid():
|
||||||
|
+ audio = speech(300)
|
||||||
|
+ spans, cuts = plan_slices(audio, SR, 120, overlap_secs=4, search_secs=0)
|
||||||
|
+ assert cuts == [116 * SR, 232 * SR]
|
||||||
|
+ assert spans == [(0, 118 * SR), (114 * SR, 234 * SR), (230 * SR, 300 * SR)]
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_short_slice_lengths_clamp_overlap_and_search():
|
||||||
|
+ # Upstream's own buffered test runs a 19 s clip with --chunk-len 10.
|
||||||
|
+ audio = speech(19)
|
||||||
|
+ spans, cuts = plan_slices(audio, SR, 10, overlap_secs=4, search_secs=25)
|
||||||
|
+ assert len(spans) >= 2
|
||||||
|
+ assert_valid_plan(spans, cuts, len(audio), 10)
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+# -- stitch_slices ------------------------------------------------------------------
|
||||||
|
+
|
||||||
|
+FRAME = 0.08
|
||||||
|
+OVERLAPPING = [(0.0, 12.0), (8.0, 20.0)] # two chunks sharing 8-12 s, cut at 10 s
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def word(text, start, end, slice_start):
|
||||||
|
+ return {
|
||||||
|
+ "word": text,
|
||||||
|
+ "start_offset": round((start - slice_start) / FRAME),
|
||||||
|
+ "end_offset": round((end - slice_start) / FRAME),
|
||||||
|
+ "start": start,
|
||||||
|
+ "end": end,
|
||||||
|
+ }
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def segment(words):
|
||||||
|
+ return {
|
||||||
|
+ "segment": " ".join(w["word"] for w in words),
|
||||||
|
+ "start_offset": words[0]["start_offset"],
|
||||||
|
+ "end_offset": words[-1]["end_offset"],
|
||||||
|
+ "start": words[0]["start"],
|
||||||
|
+ "end": words[-1]["end"],
|
||||||
|
+ }
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def texts(items, key="word"):
|
||||||
|
+ return [item[key] for item in items]
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_no_overlap_stitching_is_plain_concatenation():
|
||||||
|
+ left = [word("one", 1.0, 1.4, 0), word("two", 5.0, 5.3, 0)]
|
||||||
|
+ right = [word("three", 10.5, 10.9, 10), word("four", 14.0, 14.4, 10)]
|
||||||
|
+ words, segments = stitch_slices(
|
||||||
|
+ [(left, [segment(left)]), (right, [segment(right)])], [10.0]
|
||||||
|
+ )
|
||||||
|
+ assert words == left + right
|
||||||
|
+ assert segments == [segment(left), segment(right)]
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_overlapping_slices_keep_every_word_exactly_once():
|
||||||
|
+ # Slices [0, 12] and [8, 20], cut at 10. Both transcribe the overlap and
|
||||||
|
+ # their timestamps for the same word differ by a few ms.
|
||||||
|
+ left = [
|
||||||
|
+ word("a", 1.0, 1.3, 0),
|
||||||
|
+ word("b", 5.0, 5.4, 0),
|
||||||
|
+ word("c", 9.00, 9.40, 0),
|
||||||
|
+ word("d", 10.50, 10.90, 0),
|
||||||
|
+ word("e", 11.50, 11.80, 0),
|
||||||
|
+ ]
|
||||||
|
+ right = [
|
||||||
|
+ word("c", 9.02, 9.40, 8),
|
||||||
|
+ word("d", 10.48, 10.90, 8),
|
||||||
|
+ word("e", 11.52, 11.80, 8),
|
||||||
|
+ word("f", 15.00, 15.40, 8),
|
||||||
|
+ ]
|
||||||
|
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
|
||||||
|
+ assert texts(words) == ["a", "b", "c", "d", "e", "f"]
|
||||||
|
+ assert words[2] is left[2] and words[3] is right[1]
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_a_word_straddling_the_cut_is_kept_once():
|
||||||
|
+ left = [word("over", 9.90, 10.30, 0), word("the", 10.40, 10.55, 0)]
|
||||||
|
+ right = [word("over", 9.92, 10.30, 8), word("the", 10.40, 10.55, 8)]
|
||||||
|
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
|
||||||
|
+ assert texts(words) == ["over", "the"]
|
||||||
|
+ # Both chunks agree on it, so the right chunk takes over from it.
|
||||||
|
+ assert words[0] is right[0] and words[1] is right[1]
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_a_segment_straddling_the_cut_is_trimmed_to_the_words_each_slice_owns():
|
||||||
|
+ left_words = [
|
||||||
|
+ word("Hello", 8.5, 8.9, 0),
|
||||||
|
+ word("there", 9.2, 9.6, 0),
|
||||||
|
+ word("friend.", 10.4, 10.9, 0),
|
||||||
|
+ ]
|
||||||
|
+ right_words = [
|
||||||
|
+ word("there", 9.21, 9.6, 8),
|
||||||
|
+ word("friend.", 10.41, 10.9, 8),
|
||||||
|
+ word("Bye.", 13.0, 13.4, 8),
|
||||||
|
+ ]
|
||||||
|
+ words, segments = stitch_slices(
|
||||||
|
+ [
|
||||||
|
+ (left_words, [segment(left_words)]),
|
||||||
|
+ (right_words, [segment(right_words[:2]), segment(right_words[2:])]),
|
||||||
|
+ ],
|
||||||
|
+ [10.0],
|
||||||
|
+ OVERLAPPING,
|
||||||
|
+ )
|
||||||
|
+ assert texts(words) == ["Hello", "there", "friend.", "Bye."]
|
||||||
|
+ assert texts(segments, "segment") == ["Hello there", "friend.", "Bye."]
|
||||||
|
+ assert segments[0]["end"] == 9.6 and segments[1]["start"] == 10.41
|
||||||
|
+ # A segment that needed no trimming is passed through untouched.
|
||||||
|
+ assert segments[2] == segment(right_words[2:])
|
||||||
|
+ # Every word appears in exactly one segment, in order.
|
||||||
|
+ assert " ".join(texts(segments, "segment")) == " ".join(texts(words))
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_a_segment_wholly_inside_the_other_slices_share_is_dropped():
|
||||||
|
+ left_words = [word("a", 2.0, 2.3, 0), word("b.", 10.6, 11.0, 0)]
|
||||||
|
+ right_words = [word("b.", 10.61, 11.0, 8), word("c", 12.0, 12.3, 8)]
|
||||||
|
+ _, segments = stitch_slices(
|
||||||
|
+ [
|
||||||
|
+ (left_words, [segment(left_words[:1]), segment(left_words[1:])]),
|
||||||
|
+ (right_words, [segment(right_words[:1]), segment(right_words[1:])]),
|
||||||
|
+ ],
|
||||||
|
+ [10.0],
|
||||||
|
+ OVERLAPPING,
|
||||||
|
+ )
|
||||||
|
+ assert texts(segments, "segment") == ["a", "b.", "c"]
|
||||||
|
+ assert segments[1]["start"] == 10.61
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_a_non_positive_chunk_length_is_rejected_rather_than_looping():
|
||||||
|
+ with pytest.raises(ValueError):
|
||||||
|
+ plan_slices(speech(5), SR, 0)
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_a_word_the_two_chunks_timestamp_either_side_of_the_cut_is_kept_once():
|
||||||
|
+ # After a pause TDT may place a word's start anywhere in the pause, so the
|
||||||
|
+ # two chunks can disagree about which side of the cut it starts on.
|
||||||
|
+ left = [word("so", 8.2, 8.5, 0), word("then", 9.98, 10.3, 0), word("we", 10.4, 10.6, 0)]
|
||||||
|
+ right = [word("so", 8.2, 8.5, 8), word("then", 10.03, 10.3, 8), word("we", 10.4, 10.6, 8)]
|
||||||
|
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
|
||||||
|
+ assert texts(words) == ["so", "then", "we"]
|
||||||
|
+ # ...and the mirror image, where splitting both at the cut would drop it.
|
||||||
|
+ left = [word("so", 8.2, 8.5, 0), word("then", 10.03, 10.3, 0), word("we", 10.4, 10.6, 0)]
|
||||||
|
+ right = [word("so", 8.2, 8.5, 8), word("then", 9.98, 10.3, 8), word("we", 10.4, 10.6, 8)]
|
||||||
|
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
|
||||||
|
+ assert texts(words) == ["so", "then", "we"]
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_handover_happens_at_the_agreed_word_nearest_the_cut():
|
||||||
|
+ # The chunks differ in casing/punctuation and the left one drops "really"
|
||||||
|
+ # near its end; the right chunk's version of the overlap after the cut wins.
|
||||||
|
+ left = [word("It", 8.5, 8.7, 0), word("was", 9.6, 9.9, 0), word("good,", 11.0, 11.4, 0)]
|
||||||
|
+ right = [word("it", 8.5, 8.7, 8), word("was", 9.62, 9.9, 8), word("really", 10.3, 10.7, 8),
|
||||||
|
+ word("good.", 11.0, 11.4, 8)]
|
||||||
|
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
|
||||||
|
+ assert texts(words) == ["It", "was", "really", "good."]
|
||||||
|
+ assert words[0] is left[0] and words[1] is right[1]
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_without_an_agreed_word_the_split_falls_back_to_the_cut():
|
||||||
|
+ left = [word("alpha", 9.0, 9.4, 0), word("beta", 10.5, 10.9, 0)]
|
||||||
|
+ right = [word("gamma", 9.1, 9.4, 8), word("delta", 10.6, 10.9, 8)]
|
||||||
|
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
|
||||||
|
+ assert texts(words) == ["alpha", "delta"]
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_another_occurrence_of_the_word_elsewhere_in_the_overlap_is_not_an_anchor():
|
||||||
|
+ # The chunks disagree everywhere except on "the", but the left chunk's
|
||||||
|
+ # "the" (9.0 s) and the right chunk's (11.0 s) are different words. As an
|
||||||
|
+ # anchor they would average to the cut and discard the left one.
|
||||||
|
+ left = [word("the", 9.0, 9.2, 0), word("dog", 10.5, 10.8, 0)]
|
||||||
|
+ right = [word("cat", 9.3, 9.6, 8), word("the", 11.0, 11.2, 8)]
|
||||||
|
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
|
||||||
|
+ assert [w["start"] for w in words] == [9.0, 11.0]
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_the_handover_never_puts_words_out_of_time_order():
|
||||||
|
+ # "y" agrees (0.45 s apart), but the right chunk's "y" (9.65) after the
|
||||||
|
+ # left chunk's "x" (9.70) would run time backwards, so the left chunk's
|
||||||
|
+ # copy is kept instead. Splitting at the cut would lose "y" altogether.
|
||||||
|
+ left = [word("x", 9.70, 9.90, 0), word("y", 10.10, 10.30, 0)]
|
||||||
|
+ right = [word("z", 9.40, 9.60, 8), word("y", 9.65, 9.90, 8), word("w", 10.6, 10.8, 8)]
|
||||||
|
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
|
||||||
|
+ assert texts(words) == ["x", "y", "w"] and words[1] is left[1]
|
||||||
|
+ starts = [w["start"] for w in words]
|
||||||
|
+ assert starts == sorted(starts)
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_a_co_timed_anchor_is_found_even_when_a_longer_match_lies_elsewhere():
|
||||||
|
+ # "x y" recurs later in the right chunk, a longer text match than "z", but
|
||||||
|
+ # at a different time. Only "z" is the same word in both chunks, and it
|
||||||
|
+ # straddles the cut, so splitting both at the cut would keep it twice.
|
||||||
|
+ left = [word("x", 8.2, 8.3, 0), word("y", 8.4, 8.5, 0), word("z", 9.98, 10.2, 0)]
|
||||||
|
+ right = [word("z", 10.03, 10.2, 8), word("x", 11.0, 11.1, 8), word("y", 11.2, 11.3, 8)]
|
||||||
|
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
|
||||||
|
+ assert texts(words) == ["x", "y", "z", "x", "y"]
|
||||||
|
+
|
||||||
|
+
|
||||||
|
+def test_punctuation_alone_is_never_an_anchor():
|
||||||
|
+ # As an anchor the dash would hand the whole overlap to the right chunk.
|
||||||
|
+ left = [word("-", 9.50, 9.55, 0), word("yes", 10.4, 10.6, 0)]
|
||||||
|
+ right = [word("-", 9.52, 9.55, 8), word("no", 10.4, 10.6, 8)]
|
||||||
|
+ words, _ = stitch_slices([(left, []), (right, [])], [10.0], OVERLAPPING)
|
||||||
|
+ assert words[0] is left[0] and texts(words) == ["-", "no"]
|
||||||
|
--
|
||||||
|
2.39.5
|
||||||
|
|
||||||
@@ -0,0 +1,162 @@
|
|||||||
|
# Scriberr local patches — contract
|
||||||
|
|
||||||
|
We build Scriberr from source (no upstream sm_120 image; see `../README.md`), so
|
||||||
|
we can carry patches on that build. This directory holds them, and
|
||||||
|
`scripts/scriberr-rebuild` applies them to a pinned upstream sha, builds a
|
||||||
|
distinctly tagged image, and proves it before anyone deploys it.
|
||||||
|
|
||||||
|
| patch | against | status |
|
||||||
|
|---|---|---|
|
||||||
|
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | carried; upstream PR **prepared, not opened** (`upstream-pr/`) |
|
||||||
|
|
||||||
|
Ruling: Prime, 2026-09-30, "build the slicer". Opening the upstream PR is
|
||||||
|
outward-facing and needs Prime's explicit yes.
|
||||||
|
|
||||||
|
## 0001 — pause-aware Parakeet slicer
|
||||||
|
|
||||||
|
### What it changes
|
||||||
|
|
||||||
|
One file of product code, `internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py`,
|
||||||
|
plus one new test file beside it (`tests/test_parakeet_slicing.py`). No Go change.
|
||||||
|
|
||||||
|
Upstream cuts long audio at fixed `--chunk-len` marks with no overlap, so a
|
||||||
|
word that straddles a mark is chopped in two, lost, or transcribed twice. The
|
||||||
|
patch:
|
||||||
|
|
||||||
|
1. **Overlaps adjacent chunks** by `--overlap` seconds (default **4**), half on
|
||||||
|
each side of the cut, counted *inside* `--chunk-len`.
|
||||||
|
2. **Hands over at an agreed word.** In each overlap, the chunks switch at the
|
||||||
|
word nearest the cut that both transcribed alike: the same text after
|
||||||
|
lowercasing and stripping punctuation (punctuation alone never counts), with
|
||||||
|
start times within 0.5 s. The left chunk keeps the words before it and the
|
||||||
|
right chunk keeps the rest, the anchor taken from whichever chunk keeps the
|
||||||
|
words in time order. With no agreed word, both split at the cut by start time.
|
||||||
|
Segments are trimmed to the words their chunk keeps, and `transcription` is
|
||||||
|
the stitched words joined by spaces (upstream's text already equals that).
|
||||||
|
3. **Optional pause-aware cuts** (`--pause-search N`, default **off**): each cut
|
||||||
|
moves back to the middle of the quietest 0.3 s within the last N seconds
|
||||||
|
before the limit. It measured neutral once the stitch was right, so it is not
|
||||||
|
the default; Go never passes the flag.
|
||||||
|
4. **Imports NeMo inside `transcribe_buffered()`** so the pure helpers
|
||||||
|
(`plan_slices`, `stitch_slices`) import and test without a GPU or NeMo.
|
||||||
|
|
||||||
|
`--overlap 0` (with pause search off, the default) reproduces upstream's output
|
||||||
|
exactly: words, segments and text were byte-identical on all four test recordings.
|
||||||
|
|
||||||
|
Why the handover is by agreed word and not simply "each word goes to the chunk
|
||||||
|
its start time falls in" (the first design): at a quarter of the stitches the
|
||||||
|
two chunks put the *same* word on opposite sides of the cut, one frame apart,
|
||||||
|
so it was kept twice. Parakeet timestamps a word that follows a pause anywhere
|
||||||
|
inside the pause. Details in the bench doc.
|
||||||
|
|
||||||
|
### The seam it must keep (Go ↔ Python)
|
||||||
|
|
||||||
|
`parakeet_adapter.go` is not patched, so the script's CLI and JSON are frozen:
|
||||||
|
|
||||||
|
- **Invocation** (Go, `buildBufferedArgs`):
|
||||||
|
`uv run --native-tls --project <env> python parakeet_transcribe_buffered.py <audio> --output <json> --chunk-len <PARAKEET_CHUNK_THRESHOLD_SECS>`.
|
||||||
|
Go never passes `--overlap` or `--pause-search`, so **their defaults are the
|
||||||
|
shipped behaviour**. New flags must stay optional.
|
||||||
|
- **JSON** (Go, `parseResult`): `transcription` (str), `language` (str),
|
||||||
|
`word_timestamps` [{`word` str, `start_offset` **int**, `end_offset` **int**,
|
||||||
|
`start` float, `end` float}], `segment_timestamps` (same, with `segment`),
|
||||||
|
`audio_file`, `model`, `buffered`, `chunk_duration_secs`, `num_chunks`. An int
|
||||||
|
field that becomes a float fails the Go unmarshal. Extra keys are fine; the
|
||||||
|
patch adds `overlap_secs`, `pause_search_secs` and `cut_times`.
|
||||||
|
- `start_offset`/`end_offset` stay chunk-relative frame indices, as upstream
|
||||||
|
leaves them; Go does not read them.
|
||||||
|
- `scripts/scriberr-seam-check.py` asserts all of the above plus the stitch
|
||||||
|
invariants (word starts never go backwards; segments tile the words once).
|
||||||
|
|
||||||
|
### The memory bound
|
||||||
|
|
||||||
|
Scriberr shares fv-ml1 GPU 1 with intern-decision. **Parakeet's per-process
|
||||||
|
peak must stay ≤ 5,496 MiB** (nvidia-smi used_memory, sampled every 0.2 s), the
|
||||||
|
measured peak of upstream's 120 s slicer with `expandable_segments`. The patch
|
||||||
|
keeps it because the overlap counts **inside** `--chunk-len`: consecutive cuts
|
||||||
|
are at most `chunk-len − overlap` apart, so no chunk ever exceeds `--chunk-len`
|
||||||
|
(120 s in our compose). The pause search runs on the CPU copy of the waveform.
|
||||||
|
|
||||||
|
### Measured (2026-09-30, `docs/pfi/scriberr-slicer-bench-2026-09-30.md`)
|
||||||
|
|
||||||
|
Four recordings, 118 min in total: Prime's two uploads (private, metrics only),
|
||||||
|
the first 30 min of a Supreme Court oral argument, and a LibriVox dramatic
|
||||||
|
reading. Each was scored against a no-cut whole-file transcript, with every
|
||||||
|
variant at three cut placements.
|
||||||
|
|
||||||
|
| | upstream (fixed 120 s) | **patch (overlap 4 s, agreed-word handover)** |
|
||||||
|
|---|---|---|
|
||||||
|
| cuts with an error within ±3 s | 52 % (93/179) | **22 % (41/184)** |
|
||||||
|
| background: same test midway between cuts | 19 % | 18 % |
|
||||||
|
| near-cut error events | 102 | 47 |
|
||||||
|
| words duplicated at cuts | 17 | 2 |
|
||||||
|
| GPU peak, 35-min file, n=3 | 5,496 MiB | **5,496 MiB** (budget 5,496) |
|
||||||
|
|
||||||
|
Floor: ±0.08 on a pooled damaged-cut rate (2 SE). Pause-only, overlap-only with
|
||||||
|
the handover, and pause + overlap at 4 or 8 s all land inside that floor of
|
||||||
|
each other; the default is the simplest of them.
|
||||||
|
|
||||||
|
⚠ **Separate finding, not fixed by this patch:** Parakeet sometimes skips a run
|
||||||
|
of ≥10 consecutive words mid-chunk (12–17 runs and 500–720 words per 12
|
||||||
|
transcripts, for upstream's slicer too). See the bench doc.
|
||||||
|
|
||||||
|
### Upgrading upstream
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# on nh3-dev, from this repo
|
||||||
|
scripts/scriberr-rebuild --sha <full 40-char upstream sha> --suffix slicer1
|
||||||
|
```
|
||||||
|
|
||||||
|
It clones that sha into a new `/opt/docker/src/scriberr-<sha7>-<suffix>`,
|
||||||
|
`git apply --check`s each patch (a conflict stops the run and names the patch),
|
||||||
|
builds `scriberr:local-blackwell-<sha7>-<suffix>` without touching older tags,
|
||||||
|
then runs the embed, unit, seam and memory stages and prints a PASS/FAIL table.
|
||||||
|
Memory runs on GPU 3 and refuses a GPU that is not idle. A conflict means the
|
||||||
|
patch needs rebasing: in a checkout of the new sha, `git am -3` the old patch,
|
||||||
|
resolve, run the unit tests, then `git format-patch -1 --stdout >
|
||||||
|
0001-parakeet-pause-aware-slicer.patch` and re-run the rebuild.
|
||||||
|
|
||||||
|
Also re-measure after an upgrade that touches NeMo, torch or the slicer
|
||||||
|
(`docs/pfi/scriberr-slicer-bench-2026-09-30.md` has the harness).
|
||||||
|
|
||||||
|
**Disk.** Everything a rebuild writes lands on fv-ml1's root pool (zroot), not
|
||||||
|
`/tank`: the build dir under `/opt/docker/src` (~125 MB) and the image, which
|
||||||
|
shares ~6.1 GB of layers with the other Scriberr images and adds ~120 MB of its
|
||||||
|
own (measured 2026-09-30). The script refuses to build with less than 20 GB free
|
||||||
|
under Docker's root dir. Once a deploy has soaked, remove superseded builds by
|
||||||
|
their literal names: `docker rmi scriberr:local-blackwell-<sha7>-<suffix>` and
|
||||||
|
`sudo -n rm -rf /opt/docker/src/scriberr-<sha7>-<suffix>`, keeping the running
|
||||||
|
tag and the one before it for rollback.
|
||||||
|
|
||||||
|
⚠ `Dockerfile.cuda.12.9` installs the **latest** `uv`, `yt-dlp` and `deno` at
|
||||||
|
build time, so a rebuild changes those too, not just our patch. The seam stage
|
||||||
|
is what catches a `uv run` behaviour change.
|
||||||
|
|
||||||
|
### Deploy (manual; the rebuild script never does this)
|
||||||
|
|
||||||
|
`SCRIBERR_IMAGE` in `/opt/docker/compose/scriberr/.env` on fv-ml1 selects the
|
||||||
|
image (`.env.example` documents it). That `.env` is lkraven-owned mode 600,
|
||||||
|
kept out of git.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ssh infra-ops@10.251.50.54
|
||||||
|
cd /opt/docker/compose/scriberr
|
||||||
|
sudo -n cp -p .env .env.bak-$(date +%Y%m%d-%H%M) | cat # backup first
|
||||||
|
sudo -n grep -n '^SCRIBERR_IMAGE=' .env | cat # current value
|
||||||
|
# edit SCRIBERR_IMAGE=scriberr:local-blackwell-<sha7>-<suffix> with sudo -n
|
||||||
|
docker logs --since 2m scriberr 2>&1 | grep -c "Processing single-track job" # must be 0
|
||||||
|
sudo -n docker compose config >/dev/null | cat && sudo -n docker compose up -d scriberr | cat
|
||||||
|
```
|
||||||
|
|
||||||
|
Then prove the embed path live: the env's rewritten copy must match the patch.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker exec scriberr sha256sum /app/whisperx-env/parakeet/parakeet_transcribe_buffered.py
|
||||||
|
sha256sum /opt/docker/src/scriberr-<sha7>-<suffix>/internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py
|
||||||
|
```
|
||||||
|
|
||||||
|
Record the deploy with `scripts/ops-log record`.
|
||||||
|
|
||||||
|
**Rollback:** set `SCRIBERR_IMAGE` back to the previous tag (or restore the
|
||||||
|
`.env` backup) and `docker compose up -d scriberr`. The old image is never
|
||||||
|
deleted by the rebuild.
|
||||||
@@ -0,0 +1,113 @@
|
|||||||
|
# Upstream PR — prepared, NOT opened
|
||||||
|
|
||||||
|
**Status:** ready to submit to `rishikanthc/Scriberr`, **held for Prime's explicit
|
||||||
|
yes** (opening a PR is outward-facing). Nothing has been pushed to GitHub.
|
||||||
|
|
||||||
|
- **Diff:** `../0001-parakeet-pause-aware-slicer.patch`, a `git format-patch` of one
|
||||||
|
commit authored by Vuong Hoang against upstream `a353078` (current HEAD,
|
||||||
|
2026-09-20). It is the same file we carry, so the PR and our build cannot drift.
|
||||||
|
- **To open it (after the yes):** fork on GitHub, then
|
||||||
|
`git clone <fork> && cd Scriberr && git checkout -b parakeet-overlap-stitch a353078 &&
|
||||||
|
git am <path>/0001-parakeet-pause-aware-slicer.patch && git push -u origin parakeet-overlap-stitch`,
|
||||||
|
and open the PR with the title and body below.
|
||||||
|
- **Before sending, decide:** keep the opt-in `--pause-search` in the upstream
|
||||||
|
version, or drop it for a smaller diff (about 40 lines with its tests). It measured
|
||||||
|
neutral; see the body's last paragraph.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Title
|
||||||
|
|
||||||
|
Parakeet buffered transcription: overlap chunks and stitch at an agreed word
|
||||||
|
|
||||||
|
## Body
|
||||||
|
|
||||||
|
### Problem
|
||||||
|
|
||||||
|
For audio longer than `PARAKEET_CHUNK_THRESHOLD_SECS`, `parakeet_transcribe_buffered.py`
|
||||||
|
cuts the file at fixed `--chunk-len` marks with no overlap. A word that straddles a
|
||||||
|
mark is chopped in two, lost, or transcribed twice. We measured it on four
|
||||||
|
recordings (118 minutes, see below): **52 % of the cuts had a transcription error
|
||||||
|
within ±3 s of them, against 19 % at points midway between cuts.**
|
||||||
|
|
||||||
|
Lowering `PARAKEET_CHUNK_THRESHOLD_SECS` to save GPU memory, which is what the
|
||||||
|
knob is for on smaller cards, makes this worse, because there are more cuts.
|
||||||
|
|
||||||
|
### Change
|
||||||
|
|
||||||
|
One Python file, plus a new test file. No Go changes; the CLI and JSON that
|
||||||
|
`parakeet_adapter.go` reads are unchanged.
|
||||||
|
|
||||||
|
1. **Overlap.** Adjacent chunks share `--overlap` seconds (default 4), half on
|
||||||
|
each side of the cut. The overlap counts *inside* `--chunk-len`, so no chunk
|
||||||
|
gets longer and peak GPU memory is unchanged.
|
||||||
|
2. **Stitch at an agreed word.** In each overlap, the two chunks hand over at the
|
||||||
|
word nearest the cut that both transcribed alike: same text after lowercasing
|
||||||
|
and stripping punctuation, start times within 0.5 s. The left chunk keeps the
|
||||||
|
words before it and the right chunk the rest, with the anchor taken from
|
||||||
|
whichever chunk keeps word starts in time order. With no agreed word, both
|
||||||
|
split at the cut. Segments are trimmed to match, and `transcription` is the
|
||||||
|
stitched words joined by spaces (which is what it already equals for Parakeet).
|
||||||
|
|
||||||
|
The obvious rule, "keep each word from the chunk whose half its start time falls
|
||||||
|
in", is not enough. A word that follows a pause can be timestamped anywhere in
|
||||||
|
the pause, so the two chunks often put the *same* word on opposite sides of the
|
||||||
|
cut, one frame apart, and it comes out twice (or not at all).
|
||||||
|
3. `--pause-search N` (opt-in, off by default): move each cut back to the
|
||||||
|
quietest 0.3 s within the last N seconds before the limit.
|
||||||
|
4. NeMo is imported inside `transcribe_buffered()` so the pure helpers
|
||||||
|
(`plan_slices`, `stitch_slices`) can be unit-tested without a GPU.
|
||||||
|
|
||||||
|
New flags are optional with defaults, because the Go side does not pass them. The
|
||||||
|
JSON gains `overlap_secs`, `pause_search_secs` and `cut_times`. `--overlap 0`
|
||||||
|
reproduces the previous output exactly (words, segments and text were
|
||||||
|
byte-identical on all four test files).
|
||||||
|
|
||||||
|
### Measurements
|
||||||
|
|
||||||
|
Setup: RTX PRO 6000 Blackwell, the `Dockerfile.cuda.12.9` image, parakeet-tdt-0.6b-v3,
|
||||||
|
16 kHz mono input. Each run was scored against a **no-cut reference**: the same
|
||||||
|
model over the whole file in one pass with local attention (`parakeet_transcribe.py
|
||||||
|
--context-left 255 --context-right 255`). Words were aligned after lowercasing and
|
||||||
|
stripping punctuation. A cut counts as damaged if any error lies within ±3 s of it,
|
||||||
|
and the same test at points midway between cuts gives the background rate. Each
|
||||||
|
variant ran at three chunk lengths (120, 110 and 100 s) so the cuts land in
|
||||||
|
different places. Decoding is deterministic, so repeats are identical.
|
||||||
|
|
||||||
|
Recordings: the first 30 minutes of a U.S. Supreme Court oral argument (No. 22-451,
|
||||||
|
public domain), section 1 of the LibriVox dramatic reading *The Trial of Oscar
|
||||||
|
Wilde* (public domain), and two private conversational recordings (22 and 35 min;
|
||||||
|
numbers only).
|
||||||
|
|
||||||
|
| slicer | damaged cuts | background | error events near cuts | words duplicated at cuts |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| current (fixed, no overlap) | 93/179 = **52 %** | 19 % | 102 | 17 |
|
||||||
|
| overlap, start-time stitch | 52/184 = 28 % | 18 % | 60 | 18 |
|
||||||
|
| **overlap, agreed-word stitch (this PR)** | 41/184 = **22 %** | 18 % | 47 | 2 |
|
||||||
|
| pause-aware cut, no overlap | 53/201 = 26 % | 18 % | 59 | 0 |
|
||||||
|
| pause-aware + overlap, agreed-word stitch | 52/207 = 25 % | 18 % | 54 | 2 |
|
||||||
|
|
||||||
|
The 2-standard-error band on a difference of these rates is about ±0.08, so the
|
||||||
|
last three rows are statistically tied and all beat the current slicer by a wide
|
||||||
|
margin. The public files alone tell the same story: current 52/86 damaged cuts;
|
||||||
|
this PR 22/89. Peak GPU memory for a 35-minute file with `--chunk-len 120` is
|
||||||
|
unchanged (5,496 MiB before and after, n=3, `expandable_segments:True`).
|
||||||
|
|
||||||
|
Pause-aware cutting is included but off by default: once the stitch was right, it
|
||||||
|
did not measurably help, and without an overlap it drops words just before its
|
||||||
|
cuts. Happy to drop it from this PR if you prefer the smaller diff.
|
||||||
|
|
||||||
|
### Tests
|
||||||
|
|
||||||
|
```
|
||||||
|
cd internal/transcription/adapters/py/nvidia
|
||||||
|
python -m pytest tests/test_parakeet_slicing.py # needs numpy, librosa, soundfile; no GPU
|
||||||
|
```
|
||||||
|
|
||||||
|
The tests cover synthetic waveforms with known pauses, audio shorter than one chunk,
|
||||||
|
a pause at the very start or end, audio with no pause, the chunk limit with the
|
||||||
|
overlap included, and stitching with known word lists, including a word the two
|
||||||
|
chunks timestamp on either side of the cut, a word one chunk missed, a longer text
|
||||||
|
match at a different time, an anchor that would reverse time order, and
|
||||||
|
punctuation-only tokens. The existing `test_parakeet_transcribe_buffered.py`
|
||||||
|
still passes, since its `--chunk-len 10` run now also exercises the overlap.
|
||||||
Reference in New Issue
Block a user