Commit Graph
3 Commits
Author SHA1 Message Date
vh 963f9ed8c0 fix(parakeet-nemo): return the window cache and cap the process (nemo-0.1.1)
The gen-small EngineCore OOM (04:21 PT): our parked 3,582 MiB window cache left
no room for vLLM's runtime workspace. Seat-side fix, three controls:

- windowed path wraps every window in torch.cuda.empty_cache(), so the seat
  returns to ~2,108 MiB rest after a 12-min file instead of parking at the
  peak (measured: peak 3,028 MiB during, rest after, restarts=0);
- MEM_CAP_MIB=3840 hard set_per_process_memory_fraction: over-cap requests
  answer 503 with the seat alive (proved at cap=2000), so the failure lands
  on us, never on a neighbour;
- CUDA_GRAPHS=0: the graph decoder pins cache blocks that empty_cache must
  free (illegal-memory-access wedge when both were on first try). Cost:
  12-min file 3.0 s vs 1.2 s, short bins 35-62 ms vs 33-42 ms -- still 4-15x
  under the sherpa seat.

Measured, not computed: gen-small moved ZERO from 36,116 MiB across three
realistic requests (1,351 in / ~180 out) -- its workspace lands at engine
init; the growth window is restart-relative, matching infra-ops's observation.
WINDOW_S is now a real compose tunable. README memory section rewritten.
2026-10-01 04:39:39 -07:00
vh 03826d2029 docs: gen-small .env.example records util 0.33 + boot-check rule; parakeet-nemo compose note corrected (slack, not KV; steady state 3,582 MiB) 2026-10-01 01:38:42 -07:00
vh de6ea32f34 feat(parakeet-nemo): speech seat moves to parakeet-unified-en under NeMo (bf16 weights)
Prime-approved switch of the fleet STT seat (fv-ml1 :8300, LiteLLM ext-stt/
whisper-1, caller talk) from the sherpa-onnx int8 seat to arm B-bf16w of the
2026-09-30 A/B (docs/pfi/parakeet-seat-ab-2026-09-30.md): p50 33/36/42/71 ms
vs the old seat's 187/308/626 measured on the same card today, WER 1.965/3.026
vs the A/B floor 1.97/3.09. All three seat defects fixed: 12-min file 200s
(windowed at 360 s after a GPU 0 OOM on one whole request; the A/B's own
long-form method), no pause truncation, no long-form dropout.

GPU 0 room: gen-small --gpu-memory-utilization 0.48 -> 0.36 (0.46 and 0.40
refuse their boot check; cyberprev+voices hold the card). Its KV is byte-
pinned, so the boot log is token-identical: 670,142 tokens / 2.56x before
and after. Seat rests 2,088 MiB; GPU 0 keeps ~1.9 GB free.

Two runtime landmines documented in the README: NeMo's attention mask is
materialised T x T even under local attention (hence the window), and
httptools 0.8.0 writes a NUL into the HTTP status line that httpx — i.e.
LiteLLM — rejects, so the image ships plain uvicorn with --http h11.

Old seat stopped, not removed: docker stop parakeet-nemo && docker start
parakeet is the rollback.

License: NVIDIA Open Model License (accepted by Prime 2026-09-30); note in
stacks/parakeet-nemo/README.md.
2026-10-01 01:32:48 -07:00