feat(intern-decision): cap 9.0 GiB with MAX_TOKENS 7168, the largest call measured to fit

Both are required in compose because they are coupled: MAX_TOKENS is checked before the forward
pass, so an oversized call is a clear 422 instead of reaching the cap as a 503. Pre-deploy floor
is nvidia-smi Free >= 15,400 MiB on GPU 1 (card peak 9,876 + scriberr 5,496).
This commit is contained in:
vh
2026-09-30 09:40:08 -07:00
parent 66034cc69e
commit 750675e391
4 changed files with 38 additions and 13 deletions
+7 -2
View File
@@ -190,9 +190,14 @@ cd /opt/docker/compose/intern-decision && new=$(sed 's/^IMAGE=.*/IMAGE=intern-de
&& printf '%s\n' "$new" > .env && docker compose config -q && docker compose up -d
```
**Before any deploy onto GPU 1:** GPU 1 must have at least 15,800 MiB free
(`nvidia-smi -i 1 --query-gpu=memory.free --format=csv`). If it has less, stop. Do not squeeze
**Before any deploy onto GPU 1**, check two things. If either fails, stop; do not squeeze
scriberr.
- nvidia-smi's own `Free` on GPU 1 must be at least **15,400 MiB**:
`nvidia-smi -i 1 --query-gpu=memory.free --format=csv`. That is our card peak of 9,876 MiB plus
scriberr's 5,496, rounded up. Do not use total − used, which misses the driver's 640 MiB
reserve.
- Scriberr must not be running a job. This command must print 0:
`docker logs --since 2m scriberr | grep -c "Processing single-track job"`.
Startup fails closed. A container that never reaches healthy did not pass its own checks: the
`inference.py` hash, the pinned snapshot, the warm-up, the text-only swap and the prompt hash.