fix(parakeet-nemo): return the window cache and cap the process (nemo-0.1.1)
The gen-small EngineCore OOM (04:21 PT): our parked 3,582 MiB window cache left no room for vLLM's runtime workspace. Seat-side fix, three controls: - windowed path wraps every window in torch.cuda.empty_cache(), so the seat returns to ~2,108 MiB rest after a 12-min file instead of parking at the peak (measured: peak 3,028 MiB during, rest after, restarts=0); - MEM_CAP_MIB=3840 hard set_per_process_memory_fraction: over-cap requests answer 503 with the seat alive (proved at cap=2000), so the failure lands on us, never on a neighbour; - CUDA_GRAPHS=0: the graph decoder pins cache blocks that empty_cache must free (illegal-memory-access wedge when both were on first try). Cost: 12-min file 3.0 s vs 1.2 s, short bins 35-62 ms vs 33-42 ms -- still 4-15x under the sherpa seat. Measured, not computed: gen-small moved ZERO from 36,116 MiB across three realistic requests (1,351 in / ~180 out) -- its workspace lands at engine init; the growth window is restart-relative, matching infra-ops's observation. WINDOW_S is now a real compose tunable. README memory section rewritten.
This commit is contained in:
@@ -31,6 +31,11 @@ from fastapi.responses import JSONResponse
|
||||
|
||||
MODEL_PATH = os.environ["MODEL_PATH"]
|
||||
WARMUP_SECONDS = [int(x) for x in os.environ.get("WARMUP_SECONDS", "1,8,60").split(",")]
|
||||
# Hard ceiling for the whole process, MiB. The measured window peak is 3,582 (audit 2026-10-01);
|
||||
# 3,840 gives a little headroom and NO more. GPU 0 is shared with vLLM seats that grow at RUNTIME
|
||||
# (~0.8 GB for gen-small), and an unbounded window cache OOMed gen-small's EngineCore on its first
|
||||
# request after the switch — a cached peak is an unpaid debt to the neighbours.
|
||||
MEM_CAP_MIB = int(os.environ.get("MEM_CAP_MIB", "3840"))
|
||||
SR = 16000
|
||||
|
||||
logger = logging.getLogger("parakeet-nemo")
|
||||
@@ -50,7 +55,7 @@ def _load():
|
||||
d = m.cfg.decoding
|
||||
with open_dict(d):
|
||||
d.strategy = "greedy_batch"
|
||||
d.greedy["use_cuda_graph_decoder"] = True
|
||||
d.greedy["use_cuda_graph_decoder"] = os.environ.get("CUDA_GRAPHS", "1") == "1"
|
||||
m.change_decoding_strategy(d, verbose=False)
|
||||
# transcribe() sets these on entry; the direct path must match, and must not dither (dither is
|
||||
# a training-time augmentation and makes the same file decode differently on each call).
|
||||
@@ -71,6 +76,10 @@ def _load():
|
||||
|
||||
|
||||
model = _load()
|
||||
# Hard per-process ceiling on torch allocations (see MEM_CAP_MIB): an over-size request must
|
||||
# fail HERE, at us, instead of stealing runtime room from a neighbour's process. torch counts
|
||||
# RESERVED bytes against this, which is exactly the ledger we want capped.
|
||||
torch.cuda.set_per_process_memory_fraction(MEM_CAP_MIB / (torch.cuda.get_device_properties(0).total_memory / 2**20))
|
||||
|
||||
|
||||
def _hyp_text(h) -> str:
|
||||
@@ -127,7 +136,18 @@ def _decode(raw: bytes) -> str:
|
||||
win = int(os.environ.get("WINDOW_S", "360")) * SR
|
||||
if len(samples) <= win:
|
||||
return _infer(samples)
|
||||
parts = [_infer(samples[i:i + win]) for i in range(0, len(samples), win)]
|
||||
# Windowed path: empty BEFORE each window too — torch's cap counts RESERVED bytes, and cached
|
||||
# blocks from the previous window would otherwise count against it and bite spuriously.
|
||||
torch.cuda.empty_cache()
|
||||
try:
|
||||
parts = []
|
||||
for i in range(0, len(samples), win):
|
||||
parts.append(_infer(samples[i:i + win]))
|
||||
torch.cuda.empty_cache()
|
||||
except torch.OutOfMemoryError as exc:
|
||||
# Our cap (or a neighbour's pressure) bit: answer 503, give the cache back either way.
|
||||
torch.cuda.empty_cache()
|
||||
raise HTTPException(503, f"GPU memory limit reached for this file: {exc}") from exc
|
||||
return " ".join(p for p in parts if p)
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user