fix(parakeet-nemo): return the window cache and cap the process (nemo-0.1.1)
The gen-small EngineCore OOM (04:21 PT): our parked 3,582 MiB window cache left no room for vLLM's runtime workspace. Seat-side fix, three controls: - windowed path wraps every window in torch.cuda.empty_cache(), so the seat returns to ~2,108 MiB rest after a 12-min file instead of parking at the peak (measured: peak 3,028 MiB during, rest after, restarts=0); - MEM_CAP_MIB=3840 hard set_per_process_memory_fraction: over-cap requests answer 503 with the seat alive (proved at cap=2000), so the failure lands on us, never on a neighbour; - CUDA_GRAPHS=0: the graph decoder pins cache blocks that empty_cache must free (illegal-memory-access wedge when both were on first try). Cost: 12-min file 3.0 s vs 1.2 s, short bins 35-62 ms vs 33-42 ms -- still 4-15x under the sherpa seat. Measured, not computed: gen-small moved ZERO from 36,116 MiB across three realistic requests (1,351 in / ~180 out) -- its workspace lands at engine init; the growth window is restart-relative, matching infra-ops's observation. WINDOW_S is now a real compose tunable. README memory section rewritten.
This commit is contained in:
@@ -43,19 +43,29 @@ required; keep this section as the license note. Weights pinned at HF revision
|
||||
|
||||
## GPU 0 room
|
||||
|
||||
The seat rests ~2.1 GB, serves to ~2.8 GB, loads under ~3.0 GB (see the ops log). But its
|
||||
**steady state after ANY long (windowed) request is 3,582 MiB** — torch caches the window peak
|
||||
and does not return it (measured flat across repeated 12-min requests, audit 2026-10-01). Plan
|
||||
GPU 0 against 3,582 MiB, not the at-rest figure: in steady state the card sits at ~385 MiB Free.
|
||||
Room was taken from `vllm-gen-small`: `--gpu-memory-utilization` 0.48 → 0.36 at cut-over → 0.33
|
||||
after the audit (config-only; applies at its NEXT restart, and gives that restart ~3 GiB of
|
||||
boot-check margin against the steady-state figure). Its KV is byte-pinned
|
||||
(`--kv-cache-memory`), so the util number costs it nothing — boot log identical at 670,142
|
||||
tokens / 2.56×. ⚠ Before restarting ANY vLLM seat on this card, do the boot-check arithmetic
|
||||
against measured `nvidia-smi` Free: required = util × total, available = Free + that seat's own
|
||||
resident memory. The cut-over iteration (0.46, 0.40 both refusing boot before 0.36 booted) took
|
||||
gen-small down ~34 minutes for want of that one line of arithmetic. util does NOT predict
|
||||
resident VRAM. GPU 1 is NOT an option: its free memory is intern-decision's 32k headroom.
|
||||
The seat rests ~2.1 GB. **Since nemo-0.1.1 it returns to rest after long requests**: the
|
||||
windowed path calls `torch.cuda.empty_cache()` around each window, so a 12-min file peaks at
|
||||
~3,028 MiB during the request and falls back to ~2,108 after (measured 2026-10-01, restarts=0).
|
||||
Before 0.1.1 the seat PARKED at the window peak (3,582 MiB steady), and that cached peak left
|
||||
gen-small no room for its runtime workspace — an EngineCore CUDA-OOM incident at 04:21 PT.
|
||||
Three seat-side controls, all load-bearing:
|
||||
- **`MEM_CAP_MIB=3840`** (env, default): a hard `set_per_process_memory_fraction` ceiling. An
|
||||
over-cap request answers **503** with the seat still alive (proved at cap=2000: two 503s, then
|
||||
short requests fine) — the failure lands on us, never on a neighbour's allocation.
|
||||
- **`CUDA_GRAPHS=0`** (default): the CUDA-graph greedy decoder pins memory in torch's cache and
|
||||
died with an illegal-memory-access the first time `empty_cache` freed a graph-pool block.
|
||||
Graphs off costs latency (12-min file 3.0 s vs 1.2 s; short bins 35-62 ms vs 33-42 ms — still
|
||||
4-15x faster than the sherpa seat) and buys a lower, honestly-returned footprint.
|
||||
- **`WINDOW_S=360`** (default): see the windowing note above; mask is T×T even under local attention.
|
||||
Room came from `vllm-gen-small`: `--gpu-memory-utilization` 0.48 → 0.33 (config-only, applies at
|
||||
its next restart; KV byte-pinned, boot log identical at 670,142 tokens / 2.56×). gen-small's
|
||||
runtime growth was measured nvidia-smi-per-process, not computed: 3 realistic requests
|
||||
(1,351 prompt / ~180 completion tokens) moved it ZERO from its 36,116 MiB — the workspace
|
||||
allocation lands at engine init, right after restart (infra-ops saw the growth window at
|
||||
restart+3-requests). ⚠ Before restarting ANY vLLM seat on this card, do the boot-check
|
||||
arithmetic against measured `nvidia-smi` Free: required = util × total, available = Free + that
|
||||
seat's own resident memory. GPU 1 is NOT an option: its free memory is intern-decision's 32k
|
||||
headroom.
|
||||
|
||||
## Rollback
|
||||
|
||||
|
||||
Reference in New Issue
Block a user