docs(scriberr): Parakeet dropout investigation; proposed 0002 (gap retry + model path)

Prime's ask (via the coordinator): investigate the "Parakeet skips
stretches of speech" finding, including other Parakeet weights.
Investigation only; nothing deployed.

Against ground truth (official SCOTUS transcript, Gutenberg #38916) the
drops are real: production v3 loses 140 / 66 clean words per transcript on
the two public files and ~50 on each private one (Whisper-referenced,
Canary-confirmed; adjudicator 129/129 correct on the calibration). Cause:
the v2/v3 0.6B weights collapse deep inside long full-attention windows;
the encoder output is degraded, the audio alone transcribes fine, and
1.1B TDT/RNNT/CTC and CTC-0.6B never do it. Decoding (CUDA graphs, greedy
variants, max_symbols, beam), slice length, local attention, loudness,
resampling and a noise floor do not fix it. Controls: A-vs-A, silence
positive control (>=15 words 36/36), null control, bootstrap floor.

Proposed patch 0002 re-transcribes >=3 s stretches where the audio holds
speech but no word came out (-80 to -90 % lost words on all four
recordings, lower WER, no invented text, +10 MiB) and adds an explicit
PARAKEET_MODEL_PATH with the loaded model recorded in JSON and ModelUsed.
Reviewed at high effort, all findings fixed; built and tested as
scriberr:local-blackwell-a353078-dropout2, not deployed.

scriberr-rebuild: --patches takes DIR[:DIR...]; embeds and seam-checks
both Parakeet scripts (seam-check --standard for the short-audio one).
This commit is contained in:
vh
2026-09-30 15:48:02 -07:00
parent 1189adbf18
commit 38015a1977
7 changed files with 1162 additions and 17 deletions
+13 -8
View File
@@ -7,7 +7,8 @@ job. This checks the shape Go reads plus the stitching invariants the slicer
patch promises. Stdlib only, so it runs under any python3. Prints counts, never
transcript text.
usage: scriberr-seam-check.py RESULT.json [--min-chunks N]
usage: scriberr-seam-check.py RESULT.json [--min-chunks N] [--standard]
--standard the short-audio script's result (no buffered/num_chunks keys)
"""
import argparse
import json
@@ -42,6 +43,8 @@ def main():
parser.add_argument("result", help="result JSON written by parakeet_transcribe_buffered.py")
parser.add_argument("--min-chunks", type=int, default=1,
help="fail unless the run used at least this many chunks")
parser.add_argument("--standard", action="store_true",
help="the short-audio script's result: no buffered/num_chunks keys")
args = parser.parse_args()
min_chunks = args.min_chunks
try:
@@ -56,12 +59,13 @@ def main():
for key, kind in required.items():
if not isinstance(data.get(key), kind):
fail(f"'{key}' missing or not {kind.__name__}")
if data.get("buffered") is not True:
fail("'buffered' is not true")
if not isinstance(data.get("chunk_duration_secs"), NUMBER):
fail("'chunk_duration_secs' is not a number")
if type(data.get("num_chunks")) is not int or data["num_chunks"] < min_chunks:
fail(f"'num_chunks' is not an integer >= {min_chunks}")
if not args.standard:
if data.get("buffered") is not True:
fail("'buffered' is not true")
if not isinstance(data.get("chunk_duration_secs"), NUMBER):
fail("'chunk_duration_secs' is not a number")
if type(data.get("num_chunks")) is not int or data["num_chunks"] < min_chunks:
fail(f"'num_chunks' is not an integer >= {min_chunks}")
words, segments = data["word_timestamps"], data["segment_timestamps"]
if not words or not data["transcription"].strip():
@@ -79,7 +83,8 @@ def main():
fail("segments do not cover the stitched words exactly once, in order")
print(f"SEAM OK: {len(words)} words, {len(segments)} segments, "
f"{data['num_chunks']} chunks, cuts at {len(data.get('cut_times', []))} points")
f"{data.get('num_chunks', 1)} chunks, cuts at {len(data.get('cut_times', []))} points"
f"{', model ' + data['model'] if data.get('model') else ''}")
if __name__ == "__main__":