From 392660bd5f85fd253a1a14e782baf45214127fe9 Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Thu, 1 Oct 2026 01:37:47 -0700 Subject: [PATCH] docs(parakeet-nemo): steady-state 3,582 MiB and the boot-check arithmetic rule (audit findings) --- stacks/parakeet-nemo/README.md | 18 +++++++++++++----- 1 file changed, 13 insertions(+), 5 deletions(-) diff --git a/stacks/parakeet-nemo/README.md b/stacks/parakeet-nemo/README.md index b0c19e9..405cf30 100644 --- a/stacks/parakeet-nemo/README.md +++ b/stacks/parakeet-nemo/README.md @@ -43,11 +43,19 @@ required; keep this section as the license note. Weights pinned at HF revision ## GPU 0 room -The seat rests ~2.5 GB, serves to ~2.8 GB, loads under ~3.0 GB (measured on GPU 0 at cut-over; -see the ops log). Room was taken from `vllm-gen-small`: `--gpu-memory-utilization` 0.48 → 0.46 -(its `.env`), KV cache and concurrency re-read from its boot log at each change. ⚠ util does NOT -predict resident VRAM — after any gen-small restart, measure `nvidia-smi` Free on GPU 0 before -believing the fraction. GPU 1 is NOT an option: its free memory is intern-decision's 32k headroom. +The seat rests ~2.1 GB, serves to ~2.8 GB, loads under ~3.0 GB (see the ops log). But its +**steady state after ANY long (windowed) request is 3,582 MiB** — torch caches the window peak +and does not return it (measured flat across repeated 12-min requests, audit 2026-10-01). Plan +GPU 0 against 3,582 MiB, not the at-rest figure: in steady state the card sits at ~385 MiB Free. +Room was taken from `vllm-gen-small`: `--gpu-memory-utilization` 0.48 → 0.36 at cut-over → 0.33 +after the audit (config-only; applies at its NEXT restart, and gives that restart ~3 GiB of +boot-check margin against the steady-state figure). Its KV is byte-pinned +(`--kv-cache-memory`), so the util number costs it nothing — boot log identical at 670,142 +tokens / 2.56×. ⚠ Before restarting ANY vLLM seat on this card, do the boot-check arithmetic +against measured `nvidia-smi` Free: required = util × total, available = Free + that seat's own +resident memory. The cut-over iteration (0.46, 0.40 both refusing boot before 0.36 booted) took +gen-small down ~34 minutes for want of that one line of arithmetic. util does NOT predict +resident VRAM. GPU 1 is NOT an option: its free memory is intern-decision's 32k headroom. ## Rollback