# Scriberr local patches — contract We build Scriberr from source (no upstream sm_120 image; see `../README.md`), so we can carry patches on that build. This directory holds them, and `scripts/scriberr-rebuild` applies them to a pinned upstream sha, builds a distinctly tagged image, and proves it before anyone deploys it. | patch | against | status | |---|---|---| | `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | carried; upstream PR **prepared, not opened** (`upstream-pr/`) | Ruling: Prime, 2026-09-30, "build the slicer". Opening the upstream PR is outward-facing and needs Prime's explicit yes. ## 0001 — pause-aware Parakeet slicer ### What it changes One file of product code, `internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py`, plus one new test file beside it (`tests/test_parakeet_slicing.py`). No Go change. Upstream cuts long audio at fixed `--chunk-len` marks with no overlap, so a word that straddles a mark is chopped in two, lost, or transcribed twice. The patch: 1. **Overlaps adjacent chunks** by `--overlap` seconds (default **4**), half on each side of the cut, counted *inside* `--chunk-len`. 2. **Hands over at an agreed word.** In each overlap, the chunks switch at the word nearest the cut that both transcribed alike: the same text after lowercasing and stripping punctuation (punctuation alone never counts), with start times within 0.5 s. The left chunk keeps the words before it and the right chunk keeps the rest, the anchor taken from whichever chunk keeps the words in time order. With no agreed word, both split at the cut by start time. Segments are trimmed to the words their chunk keeps, and `transcription` is the stitched words joined by spaces (upstream's text already equals that). 3. **Optional pause-aware cuts** (`--pause-search N`, default **off**): each cut moves back to the middle of the quietest 0.3 s within the last N seconds before the limit. It measured neutral once the stitch was right, so it is not the default; Go never passes the flag. 4. **Imports NeMo inside `transcribe_buffered()`** so the pure helpers (`plan_slices`, `stitch_slices`) import and test without a GPU or NeMo. `--overlap 0` (with pause search off, the default) reproduces upstream's output exactly: words, segments and text were byte-identical on all four test recordings. Why the handover is by agreed word and not simply "each word goes to the chunk its start time falls in" (the first design): at a quarter of the stitches the two chunks put the *same* word on opposite sides of the cut, one frame apart, so it was kept twice. Parakeet timestamps a word that follows a pause anywhere inside the pause. Details in the bench doc. ### The seam it must keep (Go ↔ Python) `parakeet_adapter.go` is not patched, so the script's CLI and JSON are frozen: - **Invocation** (Go, `buildBufferedArgs`): `uv run --native-tls --project python parakeet_transcribe_buffered.py