scriberr: carry patch 0002 (gap retry + PARAKEET_MODEL_PATH), live as dropout2

Prime: Scriberr gets the basic fix, v3 stays (no NeMo 3.0.0 surgery). 0002 moves
from proposed/ into the carried set; scriberr-rebuild now applies 0001+0002 by
default (suffix dropout2) and its memory budget becomes a 5,600 MiB regression
guard (Scriberr is on GPU 3). Live on fv-ml1 1602: scripts rewritten from the
patched embed, a 20-min file at 5,502 MiB with retried_gaps reported.
This commit is contained in:
vh
2026-09-30 16:03:07 -07:00
parent f65e27b08f
commit 8c68bacf2e
3 changed files with 12 additions and 8 deletions
+1
View File
@@ -184,6 +184,7 @@ _As of 2026-09-30 ~0120 PT._
- **Dropout INVESTIGATED 2026-09-30 (Prime via coordinator; investigation only, nothing deployed):** `docs/pfi/parakeet-dropout-investigation-2026-09-30.md`. Real losses against ground truth (SCOTUS official transcript, Gutenberg #38916): v3 loses ~140 / 66 / 50 / 51 clean words per transcript (audiobook / argument / p1 / p2). Cause = v2/v3 0.6B weights collapse deep in long full-attention windows (encoder-side; 1.1B TDT/RNNT/CTC never do). No decoding, context, loudness or resampling fix. **Fix = re-transcribe ≥3 s speech gaps: −80–90 % everywhere** → `stacks/scriberr/patches/proposed/0002` (+ `PARAKEET_MODEL_PATH`), built as `scriberr:local-blackwell-a353078-dropout2`, NOT deployed; peak 5,506 MiB. Prime's calls: ship 0002?; v2 (0 on his files, collapses on read speech) vs keep v3; parakeet-unified-en-0.6b (needs NeMo 3.0.0 + NVIDIA Open Model License). Weights pulled pinned into `/tank/aimodels/huggingface` (~30 GB); throwaway env `/tank/spikes/scriberr-slicer/envs/nemo300`. - **Dropout INVESTIGATED 2026-09-30 (Prime via coordinator; investigation only, nothing deployed):** `docs/pfi/parakeet-dropout-investigation-2026-09-30.md`. Real losses against ground truth (SCOTUS official transcript, Gutenberg #38916): v3 loses ~140 / 66 / 50 / 51 clean words per transcript (audiobook / argument / p1 / p2). Cause = v2/v3 0.6B weights collapse deep in long full-attention windows (encoder-side; 1.1B TDT/RNNT/CTC never do). No decoding, context, loudness or resampling fix. **Fix = re-transcribe ≥3 s speech gaps: −80–90 % everywhere** → `stacks/scriberr/patches/proposed/0002` (+ `PARAKEET_MODEL_PATH`), built as `scriberr:local-blackwell-a353078-dropout2`, NOT deployed; peak 5,506 MiB. Prime's calls: ship 0002?; v2 (0 on his files, collapses on read speech) vs keep v3; parakeet-unified-en-0.6b (needs NeMo 3.0.0 + NVIDIA Open Model License). Weights pulled pinned into `/tank/aimodels/huggingface` (~30 GB); throwaway env `/tank/spikes/scriberr-slicer/envs/nemo300`.
- Scriberr moved to **fv-ml1 GPU 3** (coordinator, 2026-09-30); `scriberr-rebuild` memory stage now counts only its own PIDs and needs ≥20 GB free. Its default budget is still the retired GPU 1 5,496 MiB (0002 peaks 5,506 → pass `--budget`). - Scriberr moved to **fv-ml1 GPU 3** (coordinator, 2026-09-30); `scriberr-rebuild` memory stage now counts only its own PIDs and needs ≥20 GB free. Its default budget is still the retired GPU 1 5,496 MiB (0002 peaks 5,506 → pass `--budget`).
- Private bench data (copies of Prime's two uploads + transcripts) sits in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700), kept pending Prime; the public audio and metrics are beside it. - Private bench data (copies of Prime's two uploads + transcripts) sits in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700), kept pending Prime; the public audio and metrics are beside it.
- **Scriberr gap-retry fix (patch 0002) LIVE 1602 2026-09-30** as `scriberr:local-blackwell-a353078-dropout2` (Prime: "basic fix, no surgery for the new toolkit"; v3 stays). The investigation (`docs/pfi/parakeet-dropout-investigation-2026-09-30.md`) found the v3 drops are real against ground truth (50–140 words per transcript); the retry cuts them 80–90%. Live check: 5,502 MiB, `retried_gaps` reported. Rollback: `.env.bak-20260930-pre-dropout2` / slicer1. Leftovers kept pending Prime: 26 GB of candidate weights, `envs/nemo300`, and the private bench data in `/tank/spikes/scriberr-slicer/`.
- **2026-09-30 1322–1335, Prime: "Go GPU 3 now and extend the jev endpoint to hit 32k tokens".** DONE. - **2026-09-30 1322–1335, Prime: "Go GPU 3 now and extend the jev endpoint to hit 32k tokens".** DONE.
- **Scriberr is on fv-ml1 GPU 3** (`SCRIBERR_GPU_ID=3`; a 20-min file verified at 5,496 MiB). It is an on-demand tenant of the reserve, like Blender: it STEPS ASIDE when a full-size seat claims GPU 3, and it goes to **irv-ml1's A6000**, NOT back to GPU 1. - **Scriberr is on fv-ml1 GPU 3** (`SCRIBERR_GPU_ID=3`; a 20-min file verified at 5,496 MiB). It is an on-demand tenant of the reserve, like Blender: it STEPS ASIDE when a full-size seat claims GPU 3, and it goes to **irv-ml1's A6000**, NOT back to GPU 1.
- **intern-decision: `VRAM_CAP_GIB=14.4`, `MAX_TOKENS=32768`** (Jev's 32k). The measured card peak at the limit is 15,220 MiB (1 and 16 questions, n=3) against a 15,437 budget; 32,769 tokens → 422; latency 2.1 s at 32k. JevBench is still 202/231 with 0 diffs. - **intern-decision: `VRAM_CAP_GIB=14.4`, `MAX_TOKENS=32768`** (Jev's 32k). The measured card peak at the limit is 15,220 MiB (1 and 16 questions, n=3) against a 15,437 budget; 32,769 tokens → 422; latency 2.1 s at 32k. JevBench is still 202/231 with 0 diffs.
+5 -3
View File
@@ -38,10 +38,12 @@
# proposed patch on top of the carried ones, under its own --suffix: # proposed patch on top of the carried ones, under its own --suffix:
# --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed # --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed
# #
# Defaults: --sha PINNED_SHA below, --suffix slicer1, --gpu 3, --budget 5496, # Defaults: --sha PINNED_SHA below, --suffix dropout2, --gpu 3, --budget 5600,
# --memory-audio the public 30-min SCOTUS fixture. The GPU must have >= 20 GB # --memory-audio the public 30-min SCOTUS fixture. The GPU must have >= 20 GB
# free, which in practice means GPU 3 (Scriberr's own card since 2026-09-30, # free, which in practice means GPU 3 (Scriberr's own card since 2026-09-30,
# idle at 0 MiB, ~5.5 GB during a job); GPUs 0-2 are full of vLLM seats. # idle at 0 MiB, ~5.5 GB during a job); GPUs 0-2 are full of vLLM seats.
# The budget is a REGRESSION GUARD, not a card limit: on GPU 3 Scriberr has room, but a peak
# above 5,600 MiB means the patches changed memory behaviour (0001+0002 measured 5,506).
set -euo pipefail set -euo pipefail
PINNED_SHA=a353078fd96b8aca4002681813524b7397c90df1 # upstream HEAD 2026-09-20 PINNED_SHA=a353078fd96b8aca4002681813524b7397c90df1 # upstream HEAD 2026-09-20
@@ -54,7 +56,7 @@ STD_REL=internal/transcription/adapters/py/nvidia/parakeet_transcribe.py
TEST_REL=internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py TEST_REL=internal/transcription/adapters/py/nvidia/tests/test_parakeet_slicing.py
SEAM_AUDIO_REL=tests/data/AMI-Corpus-IB4002.Mix-Headset-clip.wav SEAM_AUDIO_REL=tests/data/AMI-Corpus-IB4002.Mix-Headset-clip.wav
SHA=$PINNED_SHA SUFFIX=slicer1 GPU=3 BUDGET=5496 REUSE_IMAGE=0 PATCH_DIR_ARG="" SHA=$PINNED_SHA SUFFIX=dropout2 GPU=3 BUDGET=5600 REUSE_IMAGE=0 PATCH_DIR_ARG=""
MEM_AUDIO=$TOOLS/fixtures/scotus-22-451-first30m.wav MEM_AUDIO=$TOOLS/fixtures/scotus-22-451-first30m.wav
while [ $# -gt 0 ]; do while [ $# -gt 0 ]; do
case $1 in case $1 in
@@ -278,7 +280,7 @@ read -r pids peak < <("${SSH[@]}" \
if [ "$peak" -le "$BUDGET" ]; then if [ "$peak" -le "$BUDGET" ]; then
pass memory "peak $peak MiB <= budget $BUDGET MiB (0.2 s samples, GPU $GPU); $(tail -1 <<<"$out")" pass memory "peak $peak MiB <= budget $BUDGET MiB (0.2 s samples, GPU $GPU); $(tail -1 <<<"$out")"
else else
fail memory "peak $peak MiB > budget $BUDGET MiB — do NOT deploy beside intern-decision" fail memory "peak $peak MiB > budget $BUDGET MiB — a regression vs the measured 5,506 (0001+0002); find out why before deploying"
fi fi
summary summary
+6 -5
View File
@@ -8,7 +8,7 @@ distinctly tagged image, and proves it before anyone deploys it.
| patch | against | status | | patch | against | status |
|---|---|---| |---|---|---|
| `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; upstream PR **prepared, not opened** (`upstream-pr/`) | | `0001-parakeet-pause-aware-slicer.patch` | upstream `a353078` (HEAD 2026-09-20) | **LIVE on fv-ml1 since 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1`; upstream PR **prepared, not opened** (`upstream-pr/`) |
| `proposed/0002-parakeet-model-path-and-gap-retry.patch` | 0001 | **proposed, not applied** (the rebuild script reads only this directory, not `proposed/`); built and tested as `scriberr:local-blackwell-a353078-dropout2`, not deployed. Why and how: `docs/pfi/parakeet-dropout-investigation-2026-09-30.md` | | `0002-parakeet-model-path-and-gap-retry.patch` | 0001 | **LIVE on fv-ml1 since 2026-09-30 1602 PT** as `scriberr:local-blackwell-a353078-dropout2` (Prime: "Scriberr gets the basic fix, no surgery for the new toolkit"; v3 stays). Live check: a 20-min file peaked at 5,502 MiB on GPU 3, `retried_gaps: 1`, +29 words against the slicer1 image. Why: `docs/pfi/parakeet-dropout-investigation-2026-09-30.md` |
Rollback for the live deploy: `SCRIBERR_IMAGE=scriberr:local-blackwell` (the Rollback for the live deploy: `SCRIBERR_IMAGE=scriberr:local-blackwell` (the
@@ -106,16 +106,17 @@ each other; the default is the simplest of them.
of ≥10 consecutive words mid-chunk (12–17 runs and 500–720 words per 12 of ≥10 consecutive words mid-chunk (12–17 runs and 500–720 words per 12
transcripts, for upstream's slicer too). See the bench doc. transcripts, for upstream's slicer too). See the bench doc.
### 0002 (proposed) — gap retry and an explicit model path ### 0002 (live since 2026-09-30 1602) — gap retry and an explicit model path
Parakeet v2/v3 sometimes stop producing words for tens of seconds inside a long Parakeet v2/v3 sometimes stop producing words for tens of seconds inside a long
chunk while someone is talking. 0002 re-transcribes any ≥ 3 s stretch where the chunk while someone is talking. 0002 re-transcribes any ≥ 3 s stretch where the
audio holds speech but no word came out (`--retry-gaps`, default 3; 0 off), which audio holds speech but no word came out (`--retry-gaps`, default 3; 0 off), which
cut those losses 80–90 % on four recordings. It also adds `PARAKEET_MODEL_PATH` (the cut those losses 80–90 % on four recordings. It also adds `PARAKEET_MODEL_PATH` (the
`.nemo` to load; default unchanged), reports the loaded model in the JSON, and makes `.nemo` to load; default unchanged), reports the loaded model in the JSON, and makes
the Go adapter record it as `ModelUsed`. To adopt: move it up into this directory and the Go adapter record it as `ModelUsed`. It is carried (in this directory) since 2026-09-30, and
rebuild. To test-build it on top of the carried set under its own suffix: `scripts/scriberr-rebuild` applies 0001+0002 by default (suffix `dropout2`, memory budget 5,600 MiB
`scripts/scriberr-rebuild --suffix <name> --patches stacks/scriberr/patches:stacks/scriberr/patches/proposed`. as a regression guard). Rollback: `SCRIBERR_IMAGE=scriberr:local-blackwell-a353078-slicer1`
(`.env.bak-20260930-pre-dropout2`). To test-build a future proposed patch: `--patches stacks/scriberr/patches:stacks/scriberr/patches/proposed`.
### Upgrading upstream ### Upgrading upstream