scriberr: correct the GPU 1 budget — nvidia-smi Free is 15,442 MiB, not total−used; 70 MiB spare beside intern-decision at 9.0 GiB
This commit is contained in:
@@ -179,7 +179,7 @@ _As of 2026-09-30 ~0120 PT._
|
|||||||
### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime)
|
### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime)
|
||||||
|
|
||||||
- **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now; we'll optimize scriberr later").** Stopped with `docker compose stop`, not removed, to give scriberr back its GPU 1 room. Scriberr's Parakeet path hardcodes `--chunk-len 300`, and the attention memory grows with the square of the slice, so it needs over 6 GB; it hit CUDA OOM at 0124 on a 35-min file with ~6.7 GB free. Stopping SemIf moved GPU 1 from 91,052 to 81,806 MiB used. The same job re-run at 0137 finished clean: 35m17s of audio in 44 s. That is n=1, and the peak memory was not captured. **Deferred fix (Prime: later):** shorten scriberr's slice to ~120 s in our local build, then SemIf can come back. Embedding cards were ruled out: esh-ml1 has ~4.4 GB free and nh3-ml1 ~5.1 GB. A replacement bench (brokkr's Jev candidates) is running on GPU 3 under a separate harness.
|
- **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now; we'll optimize scriberr later").** Stopped with `docker compose stop`, not removed, to give scriberr back its GPU 1 room. Scriberr's Parakeet path hardcodes `--chunk-len 300`, and the attention memory grows with the square of the slice, so it needs over 6 GB; it hit CUDA OOM at 0124 on a 35-min file with ~6.7 GB free. Stopping SemIf moved GPU 1 from 91,052 to 81,806 MiB used. The same job re-run at 0137 finished clean: 35m17s of audio in 44 s. That is n=1, and the peak memory was not captured. **Deferred fix (Prime: later):** shorten scriberr's slice to ~120 s in our local build, then SemIf can come back. Embedding cards were ruled out: esh-ml1 has ~4.4 GB free and nh3-ml1 ~5.1 GB. A replacement bench (brokkr's Jev candidates) is running on GPU 3 under a separate harness.
|
||||||
- **⚠ Prime ~0840 2026-09-30: "replace semif with intern-decision now", on GPU 1, with Scriberr fixed alongside (his pick).** A background agent is building `intern-decision-serve` (semif-compatible API, :8033, `stacks/intern-decision`, token `intern-decision/api-token`) under a hard cap: total footprint ≤ 10,300 MiB. **Scriberr fix LIVE 0900** (commit 0176ec0): `PARAKEET_CHUNK_THRESHOLD_SECS=120` + `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`, which drops the Parakeet peak from 9,384 to 5,496 MiB (n=3, deterministic). ⚠ My earlier claim that shorter slices cut memory ~6× was WRONG: a ~5.6 GB fixed floor dominates, and it is expandable_segments that cuts the fragmentation. GPU 1 budget: 16,081 free = 10,300 (intern-decision) + 5,496 (Scriberr) + 285 spare. The semif stack stays stopped as the rollback.
|
- **⚠ Prime ~0840 2026-09-30: "replace semif with intern-decision now", on GPU 1, with Scriberr fixed alongside (his pick).** A background agent is building `intern-decision-serve` (semif-compatible API, :8033, `stacks/intern-decision`, token `intern-decision/api-token`) under a hard cap: total footprint ≤ 10,300 MiB. **Scriberr fix LIVE 0900** (commit 0176ec0): `PARAKEET_CHUNK_THRESHOLD_SECS=120` + `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`, which drops the Parakeet peak from 9,384 to 5,496 MiB (n=3, deterministic). ⚠ My earlier claim that shorter slices cut memory ~6× was WRONG: a ~5.6 GB fixed floor dominates, and it is expandable_segments that cuts the fragmentation. GPU 1 budget: **15,442 MiB nvidia-smi Free** (my 16,081 was total minus used; the driver reserves ~640 MiB, which the build agent caught) = intern-decision at a 9.0 GiB cap (9,876 card peak; calls over ~7k tokens refused) + Scriberr 5,496 + 70 spare. The semif stack stays stopped as the rollback.
|
||||||
- **Jev replacement bench DONE 2026-09-30 0149–0456** (Prime's ask via brokkr, GPU 3, transient; the card is back to 2 MiB). **If SemIf is displaced, take Intern-Decision-4B on its own runtime.** It fits (9.7/10.3 GB) and is 1.5-2.3× faster (21 criteria in 88 vs 131 ms). It matches SemIf-with-rotations at ONE ordering (pooled +1.5, inside the ~4-pt floor) and is better on Wyrd. It is not a drop-in (new service + contract) and its card has no contamination statement. **JevBench rank does NOT transfer** to our sets: Plumb, the leader, is worse on Wyrd. The positive control reproduced exactly (SemIf 187/231, hard 0.613), and SemIf changed 0 labels across 4 restarts. About 33 GB of candidate weights stay on fv-ml1 `/tank/aimodels/huggingface/hub` pending Prime. The doc is `docs/pfi/jev-candidates-bench-2026-09-30.md` (475d6d6); the deliverable went to brokkr on thread `01M3RPS5MW5CXMAHPFFS0DF39Y`.
|
- **Jev replacement bench DONE 2026-09-30 0149–0456** (Prime's ask via brokkr, GPU 3, transient; the card is back to 2 MiB). **If SemIf is displaced, take Intern-Decision-4B on its own runtime.** It fits (9.7/10.3 GB) and is 1.5-2.3× faster (21 criteria in 88 vs 131 ms). It matches SemIf-with-rotations at ONE ordering (pooled +1.5, inside the ~4-pt floor) and is better on Wyrd. It is not a drop-in (new service + contract) and its card has no contamination statement. **JevBench rank does NOT transfer** to our sets: Plumb, the leader, is worse on Wyrd. The positive control reproduced exactly (SemIf 187/231, hard 0.613), and SemIf changed 0 labels across 4 restarts. About 33 GB of candidate weights stay on fv-ml1 `/tank/aimodels/huggingface/hub` pending Prime. The doc is `docs/pfi/jev-candidates-bench-2026-09-30.md` (475d6d6); the deliverable went to brokkr on thread `01M3RPS5MW5CXMAHPFFS0DF39Y`.
|
||||||
- **Was LIVE: `semif-serve` 0.1.4** (was 0.1.3 until 1014 on 2026-09-27) at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging
|
- **Was LIVE: `semif-serve` 0.1.4** (was 0.1.3 until 1014 on 2026-09-27) at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging
|
||||||
and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code +
|
and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code +
|
||||||
|
|||||||
@@ -136,8 +136,9 @@ tab. Keep the configured model on a free local seat.
|
|||||||
|
|
||||||
## Parakeet memory and slicing (measured 2026-09-30)
|
## Parakeet memory and slicing (measured 2026-09-30)
|
||||||
|
|
||||||
Scriberr shares fv-ml1 GPU 1 with intern-decision (~10.3 GB cap). The budget left
|
Scriberr shares fv-ml1 GPU 1 with intern-decision (9.0 GiB cap, 9,876 MiB card peak).
|
||||||
for Scriberr is about 5.8 GB, so the compose file sets
|
GPU 1's nvidia-smi Free is 15,442 MiB, so the budget left for Scriberr is about 5.5 GB
|
||||||
|
(70 MiB spare at both peaks), so the compose file sets
|
||||||
`PARAKEET_CHUNK_THRESHOLD_SECS=120` and
|
`PARAKEET_CHUNK_THRESHOLD_SECS=120` and
|
||||||
`PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`. The numbers behind those
|
`PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`. The numbers behind those
|
||||||
settings are in the compose comments.
|
settings are in the compose comments.
|
||||||
|
|||||||
@@ -80,7 +80,9 @@ services:
|
|||||||
# 120 s + expandable_segments 5,496 · 60 s + expandable_segments 5,502
|
# 120 s + expandable_segments 5,496 · 60 s + expandable_segments 5,502
|
||||||
# The floor, not the slice, dominates below ~120 s; expandable_segments is
|
# The floor, not the slice, dominates below ~120 s; expandable_segments is
|
||||||
# what removes the fragmentation on top of it. 120 s + expandable fits
|
# what removes the fragmentation on top of it. 120 s + expandable fits
|
||||||
# beside intern-decision even at both peaks (295 MiB spare). Transcripts
|
# beside intern-decision (9.0 GiB cap, 9,876 MiB card peak) even at both
|
||||||
|
# peaks: 9,876 + 5,496 = 15,372 of GPU 1's 15,442 MiB nvidia-smi Free
|
||||||
|
# (70 MiB spare; Free is NOT total−used, the driver reserves ~640 MiB). Transcripts
|
||||||
# change slightly: 95.8% word-sequence similarity vs 300 s (7,599 vs
|
# change slightly: 95.8% word-sequence similarity vs 300 s (7,599 vs
|
||||||
# 7,645 words); the diffs are mostly casing/punctuation spread through
|
# 7,645 words); the diffs are mostly casing/punctuation spread through
|
||||||
# the file, ~40 words at the 17 cuts. A-vs-A at 300 s: identical.
|
# the file, ~40 words at the 17 cuts. A-vs-A at 300 s: identical.
|
||||||
|
|||||||
Reference in New Issue
Block a user