docs(parakeet-nemo): steady-state 3,582 MiB and the boot-check arithmetic rule (audit findings)

This commit is contained in:
vh
2026-10-01 01:37:47 -07:00
parent 41d2014df9
commit 392660bd5f
+13 -5
View File
@@ -43,11 +43,19 @@ required; keep this section as the license note. Weights pinned at HF revision
## GPU 0 room
The seat rests ~2.5 GB, serves to ~2.8 GB, loads under ~3.0 GB (measured on GPU 0 at cut-over;
see the ops log). Room was taken from `vllm-gen-small`: `--gpu-memory-utilization` 0.48 → 0.46
(its `.env`), KV cache and concurrency re-read from its boot log at each change. ⚠ util does NOT
predict resident VRAM — after any gen-small restart, measure `nvidia-smi` Free on GPU 0 before
believing the fraction. GPU 1 is NOT an option: its free memory is intern-decision's 32k headroom.
The seat rests ~2.1 GB, serves to ~2.8 GB, loads under ~3.0 GB (see the ops log). But its
**steady state after ANY long (windowed) request is 3,582 MiB** — torch caches the window peak
and does not return it (measured flat across repeated 12-min requests, audit 2026-10-01). Plan
GPU 0 against 3,582 MiB, not the at-rest figure: in steady state the card sits at ~385 MiB Free.
Room was taken from `vllm-gen-small`: `--gpu-memory-utilization` 0.48 → 0.36 at cut-over → 0.33
after the audit (config-only; applies at its NEXT restart, and gives that restart ~3 GiB of
boot-check margin against the steady-state figure). Its KV is byte-pinned
(`--kv-cache-memory`), so the util number costs it nothing — boot log identical at 670,142
tokens / 2.56×. ⚠ Before restarting ANY vLLM seat on this card, do the boot-check arithmetic
against measured `nvidia-smi` Free: required = util × total, available = Free + that seat's own
resident memory. The cut-over iteration (0.46, 0.40 both refusing boot before 0.36 booted) took
gen-small down ~34 minutes for want of that one line of arithmetic. util does NOT predict
resident VRAM. GPU 1 is NOT an option: its free memory is intern-decision's 32k headroom.
## Rollback