fix(intern-decision-serve): load and score on one dedicated inference thread

torch keeps CUDA state per host thread (cuBLAS handles and workspaces), partly outside the
per-process VRAM cap. Scoring on anyio's threadpool let 40 threads each create it: measured on
fv-ml1 GPU 3, +252 MiB outside the cap and +326 MiB inside, which pushed the process past the
10,300 MiB GPU 1 budget. Load, warm-up and every call now run on the same single thread.
This commit is contained in:
vh
2026-09-30 09:18:46 -07:00
parent f7415db5c9
commit f21369e4ac
5 changed files with 87 additions and 11 deletions
@@ -5,7 +5,7 @@ import os
from fastapi import FastAPI
from .app import create_app
from .app import create_app, inference_thread
from .config import Settings
@@ -17,4 +17,7 @@ def app_from_env() -> FastAPI:
os.environ["TRANSFORMERS_OFFLINE"] = "1"
from .engine import TorchEngine # torch loads only inside load(), never in the unit tests
return create_app(settings, TorchEngine.load(settings))
# INV-2: load, warm-up and every later call run on this one thread (per-thread CUDA state).
executor = inference_thread()
engine = executor.submit(TorchEngine.load, settings).result()
return create_app(settings, engine, executor)