fix(intern-decision-serve): load and score on one dedicated inference thread

torch keeps CUDA state per host thread (cuBLAS handles and workspaces), partly outside the
per-process VRAM cap. Scoring on anyio's threadpool let 40 threads each create it: measured on
fv-ml1 GPU 3, +252 MiB outside the cap and +326 MiB inside, which pushed the process past the
10,300 MiB GPU 1 budget. Load, warm-up and every call now run on the same single thread.
This commit is contained in:
vh
2026-09-30 09:18:46 -07:00
parent f7415db5c9
commit f21369e4ac
5 changed files with 87 additions and 11 deletions
@@ -88,9 +88,15 @@ request. A violation is a 422.
`calibration` and `input_tokens` is what `predict()` returned. The wrapper only re-keys it.
If an answer lacks one of the decision's option ids, that is a 500 `scoring_failed`, never a
guess.
- **INV-2 one model, one inference at a time.** The model loads at startup, and a process-wide
lock serialises every request's calls (all of a request's chunks run inside one hold). The app
runs one worker. Calls run off the event loop, so `/health` answers during one.
- **INV-2 one model, one inference thread.** The model loads at startup on a dedicated
single-thread executor. The warm-ups and every later call run on that **same host thread**,
never on the event loop's threadpool. A lock also serialises each request's calls, so all of a
request's chunks run inside one hold. The app runs one worker, and `/health` answers during a
call.
- **Why one thread:** torch keeps CUDA state per host thread (cuBLAS handles and workspaces),
and part of it sits outside the VRAM cap.
- **Measured on 2026-09-30, fv-ml1 GPU 3:** anyio's 40 worker threads added 252 MiB outside the
cap and 326 MiB inside it. That pushed the footprint past the GPU 1 budget.
- **INV-3 fail-closed startup.** Before the service serves, all of these must hold:
- `inference.py` in the checkpoint hashes to the pinned sha256 (it is executed code, loaded
from a data mount);
@@ -237,7 +243,7 @@ on the host, the cap is the single knob `VRAM_CAP_GIB`.
- **Engine failures.** An engine `OutOfMemory` is 503, and any other failure is 500.
- **Concurrency.** Requests are serialised: two never overlap inside the engine, and the
chunks of one request are not interleaved with another's. `/health` answers while a call is
blocked.
blocked. Every call runs on one thread, which is the thread the engine was loaded on.
- **Engine against a fake torch.**
- An OOM is re-raised unchained, and `empty_cache` runs only after the failed call's tensors
are freed.