fix(parakeet-nemo): return the window cache and cap the process (nemo-0.1.1)
The gen-small EngineCore OOM (04:21 PT): our parked 3,582 MiB window cache left no room for vLLM's runtime workspace. Seat-side fix, three controls: - windowed path wraps every window in torch.cuda.empty_cache(), so the seat returns to ~2,108 MiB rest after a 12-min file instead of parking at the peak (measured: peak 3,028 MiB during, rest after, restarts=0); - MEM_CAP_MIB=3840 hard set_per_process_memory_fraction: over-cap requests answer 503 with the seat alive (proved at cap=2000), so the failure lands on us, never on a neighbour; - CUDA_GRAPHS=0: the graph decoder pins cache blocks that empty_cache must free (illegal-memory-access wedge when both were on first try). Cost: 12-min file 3.0 s vs 1.2 s, short bins 35-62 ms vs 33-42 ms -- still 4-15x under the sherpa seat. Measured, not computed: gen-small moved ZERO from 36,116 MiB across three realistic requests (1,351 in / ~180 out) -- its workspace lands at engine init; the growth window is restart-relative, matching infra-ops's observation. WINDOW_S is now a real compose tunable. README memory section rewritten.
This commit is contained in:
@@ -32,6 +32,11 @@ services:
|
||||
environment:
|
||||
- MODEL_PATH=/hf/hub/models--nvidia--parakeet-unified-en-0.6b/snapshots/${PARAKEET_NEMO_REV}/parakeet-unified-en-0.6b.nemo
|
||||
- WARMUP_SECONDS=${PARAKEET_NEMO_WARMUP:-1,8,60}
|
||||
# CUDA graphs pin memory in torch's cache, which fights the empty_cache the windowed path
|
||||
# needs to return memory to GPU 0's neighbours (illegal-memory-access wedge, 2026-10-01).
|
||||
- CUDA_GRAPHS=${PARAKEET_NEMO_CUDA_GRAPHS:-0}
|
||||
- MEM_CAP_MIB=${PARAKEET_NEMO_MEM_CAP_MIB:-3840}
|
||||
- WINDOW_S=${PARAKEET_NEMO_WINDOW_S:-360}
|
||||
- LOG_LEVEL=${PARAKEET_NEMO_LOG_LEVEL:-INFO}
|
||||
volumes:
|
||||
- /tank/aimodels/huggingface:/hf:ro
|
||||
|
||||
Reference in New Issue
Block a user