fix(parakeet-nemo): return the window cache and cap the process (nemo-0.1.1)
The gen-small EngineCore OOM (04:21 PT): our parked 3,582 MiB window cache left no room for vLLM's runtime workspace. Seat-side fix, three controls: - windowed path wraps every window in torch.cuda.empty_cache(), so the seat returns to ~2,108 MiB rest after a 12-min file instead of parking at the peak (measured: peak 3,028 MiB during, rest after, restarts=0); - MEM_CAP_MIB=3840 hard set_per_process_memory_fraction: over-cap requests answer 503 with the seat alive (proved at cap=2000), so the failure lands on us, never on a neighbour; - CUDA_GRAPHS=0: the graph decoder pins cache blocks that empty_cache must free (illegal-memory-access wedge when both were on first try). Cost: 12-min file 3.0 s vs 1.2 s, short bins 35-62 ms vs 33-42 ms -- still 4-15x under the sherpa seat. Measured, not computed: gen-small moved ZERO from 36,116 MiB across three realistic requests (1,351 in / ~180 out) -- its workspace lands at engine init; the growth window is restart-relative, matching infra-ops's observation. WINDOW_S is now a real compose tunable. README memory section rewritten.
This commit is contained in:
@@ -1,5 +1,5 @@
|
||||
# Copy to .env next to compose.yaml on the host.
|
||||
PARAKEET_NEMO_TAG=nemo-0.1.0
|
||||
PARAKEET_NEMO_TAG=nemo-0.1.1
|
||||
# Port the seat listens on. 8300 is the seat port LiteLLM's ext-stt/whisper-1 point at;
|
||||
# run acceptance on a temporary port first, then cut over by changing this line.
|
||||
PARAKEET_NEMO_PORT=8300
|
||||
@@ -9,3 +9,10 @@ PARAKEET_NEMO_PORT=8300
|
||||
PARAKEET_NEMO_REV=fe53cd885760c96b6a5f51a0bfd362cb4584a98b
|
||||
# Ascending silent warm-up clips in seconds (CUDA-graph capture + longest-shape kernel warm).
|
||||
# PARAKEET_NEMO_WARMUP=1,8,60
|
||||
# Hard per-process VRAM ceiling, MiB (set_per_process_memory_fraction). Over-cap requests get
|
||||
# 503 and the seat stays alive; do not raise past ~3,840 — GPU 0's vLLM neighbours need the rest.
|
||||
# PARAKEET_NEMO_MEM_CAP_MIB=3840
|
||||
# 0 (default): the CUDA-graph decoder pins torch cache and wedges when the windowed path frees it.
|
||||
# PARAKEET_NEMO_CUDA_GRAPHS=0
|
||||
# Long-form window size in seconds (files over this are transcribed in windows of this size).
|
||||
# PARAKEET_NEMO_WINDOW_S=360
|
||||
|
||||
Reference in New Issue
Block a user