fix(intern-decision-serve): code-review fixes
- a failure while building the response (a non-finite number included) is a 500 inside the envelope, never a 422 or a render crash outside it - an engine ValueError keeps its message but is released and raised unchained, like an OOM - the prompt is built (and the model's own validation run) before the forward - the row cap is counted before any ordering is built - /health reads a device name cached at load, so it makes no driver call off the inference thread
This commit is contained in:
@@ -43,7 +43,12 @@ request. A violation is a 422.
|
||||
|
||||
- **One call** is one `predict()` request:
|
||||
`{"state": state, "questions": {<field>: {"type": "choice", "instructions": question, "criteria": {option.id: option.description, ...}}}}`.
|
||||
The criteria keep the caller's option order. This is the format the bench measured.
|
||||
The criteria keep the caller's option order.
|
||||
- **A one-question call is exactly the bench's native request.** Acceptance showed it
|
||||
bit-identical, row for row.
|
||||
- **A multi-question call differs from the bench's multifield run only in the field names.**
|
||||
Those are positional here and were the decision names in the bench. Measured: 1 of 84 Wyrd
|
||||
rows differs (78/84 against 77/84).
|
||||
- **Field names are positional:** `q` when a call carries one question, and `q1`..`qN` in
|
||||
request order when it carries several. Decision ids never reach the prompt. **Option ids do:**
|
||||
the model prints `A = <id>: <description>`.
|
||||
@@ -87,12 +92,15 @@ request. A violation is a 422.
|
||||
- **INV-1 pass-through.** Every number in `native`, `probabilities`, `top`, `confidence`,
|
||||
`calibration` and `input_tokens` is what `predict()` returned. The wrapper only re-keys it.
|
||||
If an answer lacks one of the decision's option ids, that is a 500 `scoring_failed`, never a
|
||||
guess.
|
||||
guess. So is any failure while building the response, a non-finite number included. That stays
|
||||
inside the error envelope and is never a 422 or a bare 500.
|
||||
- **INV-2 one model, one inference thread.** The model loads at startup on a dedicated
|
||||
single-thread executor. The warm-ups and every later call run on that **same host thread**,
|
||||
never on the event loop's threadpool. A lock also serialises each request's calls, so all of a
|
||||
request's chunks run inside one hold. The app runs one worker, and `/health` answers during a
|
||||
call.
|
||||
- `/health` reads only allocator counters and a device name cached at load, so it adds no
|
||||
per-thread CUDA state.
|
||||
- **Why one thread:** torch keeps CUDA state per host thread (cuBLAS handles and workspaces),
|
||||
and part of it sits outside the VRAM cap.
|
||||
- **Measured on 2026-09-30, fv-ml1 GPU 3:** anyio's 40 worker threads added 252 MiB outside the
|
||||
@@ -110,10 +118,12 @@ request. A violation is a 422.
|
||||
- An OOM in any call makes the whole request a 503 `out_of_memory`. So does a RuntimeError
|
||||
whose first line says "out of memory". The engine then frees the failed call's frames,
|
||||
runs `gc.collect()` and `empty_cache()`, and raises unchained. The process stays up.
|
||||
- A `ValueError` (the token limit, or one raised inside a forward) keeps its message and maps
|
||||
to 422. It is released and raised unchained the same way.
|
||||
- After every call, if reserved memory exceeds the post-warm-up baseline by more than
|
||||
`RELEASE_SLACK_MIB` (default 512), the engine runs `empty_cache()`.
|
||||
- Any other failure except `ValueError` is logged with its traceback and released the same
|
||||
way, then raised unchained as `ScoringFailed`.
|
||||
- Any other failure is logged with its traceback and released the same way, then raised
|
||||
unchained as `ScoringFailed`.
|
||||
- **INV-5 no network.** The entry point sets `HF_HUB_OFFLINE=1` and `TRANSFORMERS_OFFLINE=1`
|
||||
before torch or transformers load. The weights are read from the mounted, read-only HF cache.
|
||||
- **INV-6 constant-time auth.** The token is compared with `hmac.compare_digest`. It must be at
|
||||
@@ -249,7 +259,8 @@ on the host, the cap is the single knob `VRAM_CAP_GIB`.
|
||||
are freed.
|
||||
- A RuntimeError saying "out of memory" becomes `OutOfMemory`.
|
||||
- Another failure becomes `ScoringFailed`, unchained.
|
||||
- `ValueError` passes through.
|
||||
- A `ValueError` keeps its message but is released and raised unchained.
|
||||
- A request the prompt builder rejects never reaches the model.
|
||||
- A burst over baseline plus slack is released, and one at or under it is left alone.
|
||||
- **Engine load (fake).**
|
||||
- A wrong `inference.py` hash, or a checkpoint path that is not the pinned snapshot, refuses
|
||||
|
||||
Reference in New Issue
Block a user