fix(intern-decision-serve): load and score on one dedicated inference thread
torch keeps CUDA state per host thread (cuBLAS handles and workspaces), partly outside the per-process VRAM cap. Scoring on anyio's threadpool let 40 threads each create it: measured on fv-ml1 GPU 3, +252 MiB outside the cap and +326 MiB inside, which pushed the process past the 10,300 MiB GPU 1 budget. Load, warm-up and every call now run on the same single thread.
This commit is contained in:
@@ -88,9 +88,15 @@ request. A violation is a 422.
|
||||
`calibration` and `input_tokens` is what `predict()` returned. The wrapper only re-keys it.
|
||||
If an answer lacks one of the decision's option ids, that is a 500 `scoring_failed`, never a
|
||||
guess.
|
||||
- **INV-2 one model, one inference at a time.** The model loads at startup, and a process-wide
|
||||
lock serialises every request's calls (all of a request's chunks run inside one hold). The app
|
||||
runs one worker. Calls run off the event loop, so `/health` answers during one.
|
||||
- **INV-2 one model, one inference thread.** The model loads at startup on a dedicated
|
||||
single-thread executor. The warm-ups and every later call run on that **same host thread**,
|
||||
never on the event loop's threadpool. A lock also serialises each request's calls, so all of a
|
||||
request's chunks run inside one hold. The app runs one worker, and `/health` answers during a
|
||||
call.
|
||||
- **Why one thread:** torch keeps CUDA state per host thread (cuBLAS handles and workspaces),
|
||||
and part of it sits outside the VRAM cap.
|
||||
- **Measured on 2026-09-30, fv-ml1 GPU 3:** anyio's 40 worker threads added 252 MiB outside the
|
||||
cap and 326 MiB inside it. That pushed the footprint past the GPU 1 budget.
|
||||
- **INV-3 fail-closed startup.** Before the service serves, all of these must hold:
|
||||
- `inference.py` in the checkpoint hashes to the pinned sha256 (it is executed code, loaded
|
||||
from a data mount);
|
||||
@@ -237,7 +243,7 @@ on the host, the cap is the single knob `VRAM_CAP_GIB`.
|
||||
- **Engine failures.** An engine `OutOfMemory` is 503, and any other failure is 500.
|
||||
- **Concurrency.** Requests are serialised: two never overlap inside the engine, and the
|
||||
chunks of one request are not interleaved with another's. `/health` answers while a call is
|
||||
blocked.
|
||||
blocked. Every call runs on one thread, which is the thread the engine was loaded on.
|
||||
- **Engine against a fake torch.**
|
||||
- An OOM is re-raised unchained, and `empty_cache` runs only after the failed call's tensors
|
||||
are freed.
|
||||
|
||||
Reference in New Issue
Block a user