fix(intern-decision-serve): code-review fixes

- a failure while building the response (a non-finite number included) is a 500 inside the
  envelope, never a 422 or a render crash outside it
- an engine ValueError keeps its message but is released and raised unchained, like an OOM
- the prompt is built (and the model's own validation run) before the forward
- the row cap is counted before any ordering is built
- /health reads a device name cached at load, so it makes no driver call off the inference thread
This commit is contained in:
vh
2026-09-30 09:30:40 -07:00
parent 618390c5fa
commit 21d16d7ad8
5 changed files with 110 additions and 21 deletions
@@ -43,7 +43,12 @@ request. A violation is a 422.
- **One call** is one `predict()` request:
`{"state": state, "questions": {<field>: {"type": "choice", "instructions": question, "criteria": {option.id: option.description, ...}}}}`.
The criteria keep the caller's option order. This is the format the bench measured.
The criteria keep the caller's option order.
- **A one-question call is exactly the bench's native request.** Acceptance showed it
bit-identical, row for row.
- **A multi-question call differs from the bench's multifield run only in the field names.**
Those are positional here and were the decision names in the bench. Measured: 1 of 84 Wyrd
rows differs (78/84 against 77/84).
- **Field names are positional:** `q` when a call carries one question, and `q1`..`qN` in
request order when it carries several. Decision ids never reach the prompt. **Option ids do:**
the model prints `A = <id>: <description>`.
@@ -87,12 +92,15 @@ request. A violation is a 422.
- **INV-1 pass-through.** Every number in `native`, `probabilities`, `top`, `confidence`,
`calibration` and `input_tokens` is what `predict()` returned. The wrapper only re-keys it.
If an answer lacks one of the decision's option ids, that is a 500 `scoring_failed`, never a
guess.
guess. So is any failure while building the response, a non-finite number included. That stays
inside the error envelope and is never a 422 or a bare 500.
- **INV-2 one model, one inference thread.** The model loads at startup on a dedicated
single-thread executor. The warm-ups and every later call run on that **same host thread**,
never on the event loop's threadpool. A lock also serialises each request's calls, so all of a
request's chunks run inside one hold. The app runs one worker, and `/health` answers during a
call.
- `/health` reads only allocator counters and a device name cached at load, so it adds no
per-thread CUDA state.
- **Why one thread:** torch keeps CUDA state per host thread (cuBLAS handles and workspaces),
and part of it sits outside the VRAM cap.
- **Measured on 2026-09-30, fv-ml1 GPU 3:** anyio's 40 worker threads added 252 MiB outside the
@@ -110,10 +118,12 @@ request. A violation is a 422.
- An OOM in any call makes the whole request a 503 `out_of_memory`. So does a RuntimeError
whose first line says "out of memory". The engine then frees the failed call's frames,
runs `gc.collect()` and `empty_cache()`, and raises unchained. The process stays up.
- A `ValueError` (the token limit, or one raised inside a forward) keeps its message and maps
to 422. It is released and raised unchained the same way.
- After every call, if reserved memory exceeds the post-warm-up baseline by more than
`RELEASE_SLACK_MIB` (default 512), the engine runs `empty_cache()`.
- Any other failure except `ValueError` is logged with its traceback and released the same
way, then raised unchained as `ScoringFailed`.
- Any other failure is logged with its traceback and released the same way, then raised
unchained as `ScoringFailed`.
- **INV-5 no network.** The entry point sets `HF_HUB_OFFLINE=1` and `TRANSFORMERS_OFFLINE=1`
before torch or transformers load. The weights are read from the mounted, read-only HF cache.
- **INV-6 constant-time auth.** The token is compared with `hmac.compare_digest`. It must be at
@@ -249,7 +259,8 @@ on the host, the cap is the single knob `VRAM_CAP_GIB`.
are freed.
- A RuntimeError saying "out of memory" becomes `OutOfMemory`.
- Another failure becomes `ScoringFailed`, unchained.
- `ValueError` passes through.
- A `ValueError` keeps its message but is released and raised unchained.
- A request the prompt builder rejects never reaches the model.
- A burst over baseline plus slack is released, and one at or under it is left alone.
- **Engine load (fake).**
- A wrong `inference.py` hash, or a checkpoint path that is not the pinned snapshot, refuses