scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)
Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens. Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1 freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front. JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
This commit is contained in:
@@ -1,5 +1,6 @@
|
||||
# intern-decision: Intern-Decision-4B (internlm, Apache-2.0) behind intern-decision-serve, on
|
||||
# fv-ml1 GPU 1 (the utility card, beside vllm-coder, the erp/meromero seats and scriberr).
|
||||
# fv-ml1 GPU 1 (the utility card, beside vllm-coder and the erp/meromero seats; scriberr moved to
|
||||
# GPU 3 on 2026-09-30 1322, Prime, to free this card's headroom for 32k-token calls).
|
||||
# Replaces semif (Prime, 2026-09-30: "replace semif with intern-decision now").
|
||||
#
|
||||
# One forward pass per call, scored by the checkpoint's OWN inference.py (sha256-pinned); the
|
||||
@@ -8,8 +9,8 @@
|
||||
# fv-ml1 from that dir.
|
||||
#
|
||||
# ⚠ VRAM_CAP_GIB is a HARD cap on torch's allocator (per-process memory fraction), set so the
|
||||
# container's WHOLE nvidia-smi footprint, CUDA context included, fits beside scriberr's peak
|
||||
# (infra-ops budget, 2026-09-30); MAX_TOKENS keeps every accepted call under the cap. A request that needs more
|
||||
# container's WHOLE nvidia-smi footprint, CUDA context included, fits GPU 1's free memory beside
|
||||
# the static vLLM seats (infra-ops budget, 2026-09-30); MAX_TOKENS keeps every accepted call under the cap. A request that needs more
|
||||
# gets 503 out_of_memory and the service stays up. See the README before changing it.
|
||||
#
|
||||
# .env (tunables): IMAGE, PORT, GPU_ID, VRAM_CAP_GIB, MAX_TOKENS, HOST_IP, INTERN_DECISION_API_TOKEN
|
||||
|
||||
Reference in New Issue
Block a user