docs(scriberr): Parakeet memory table, local-attention OOM, scripts are rewritten from the embedded copy

This commit is contained in:
vh
2026-09-30 09:21:51 -07:00
parent f21369e4ac
commit 618390c5fa
+30
View File
@@ -133,3 +133,33 @@ paid vendor. In the UI under the AI provider settings:
⚠ That key also reaches **paid** passthrough models (GLM, Kimi) on a shared ⚠ That key also reaches **paid** passthrough models (GLM, Kimi) on a shared
tab. Keep the configured model on a free local seat. tab. Keep the configured model on a free local seat.
## Parakeet memory and slicing (measured 2026-09-30)
Scriberr shares fv-ml1 GPU 1 with intern-decision (~10.3 GB cap). The budget left
for Scriberr is about 5.8 GB, so the compose file sets
`PARAKEET_CHUNK_THRESHOLD_SECS=120` and
`PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`. The numbers behind those
settings are in the compose comments.
Peak GPU memory on a 35-minute file:
| setting | peak |
|---|---|
| 300 s slices | 9,384 MiB |
| 300 s slices + expandable_segments | 6,962 MiB |
| 120 s slices + expandable_segments | 5,496 MiB (live setting) |
- **Whole file in one pass with local attention** (`--context-left/right 255`,
standard script): **CUDA OOM at >16 GB**. It asked for another 6.46 GiB at
9.8 GiB in use. It is not viable on this card budget.
- ⚠ **Hand-editing the Parakeet scripts does not persist.** Scriberr's
`PrepareEnvironment` rewrites `parakeet_transcribe.py` and
`parakeet_transcribe_buffered.py` into `whisperx-env/parakeet/` from the copies
embedded in the Go binary every time it prepares the environment. A change to the
slicer has to go into the source checkout
(`internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py`,
embedded at build) and be rebuilt. Or it goes upstream (MIT).
- The two env knobs are read by upstream's Go code (`parakeet_adapter.go`), so
they survive image upgrades for as long as upstream keeps them. Re-measure the
peak after any upgrade.