memory: intern-decision warm-up cache tasked to infra-hermes

This commit is contained in:
vh
2026-09-30 15:44:45 -07:00
parent af450f4ef7
commit 1189adbf18
+1 -1
View File
@@ -187,7 +187,7 @@ _As of 2026-09-30 ~0120 PT._
- **2026-09-30 1322–1335, Prime: "Go GPU 3 now and extend the jev endpoint to hit 32k tokens".** DONE. - **2026-09-30 1322–1335, Prime: "Go GPU 3 now and extend the jev endpoint to hit 32k tokens".** DONE.
- **Scriberr is on fv-ml1 GPU 3** (`SCRIBERR_GPU_ID=3`; a 20-min file verified at 5,496 MiB). It is an on-demand tenant of the reserve, like Blender: it STEPS ASIDE when a full-size seat claims GPU 3, and it goes to **irv-ml1's A6000**, NOT back to GPU 1. - **Scriberr is on fv-ml1 GPU 3** (`SCRIBERR_GPU_ID=3`; a 20-min file verified at 5,496 MiB). It is an on-demand tenant of the reserve, like Blender: it STEPS ASIDE when a full-size seat claims GPU 3, and it goes to **irv-ml1's A6000**, NOT back to GPU 1.
- **intern-decision: `VRAM_CAP_GIB=14.4`, `MAX_TOKENS=32768`** (Jev's 32k). The measured card peak at the limit is 15,220 MiB (1 and 16 questions, n=3) against a 15,437 budget; 32,769 tokens → 422; latency 2.1 s at 32k. JevBench is still 202/231 with 0 diffs. - **intern-decision: `VRAM_CAP_GIB=14.4`, `MAX_TOKENS=32768`** (Jev's 32k). The measured card peak at the limit is 15,220 MiB (1 and 16 questions, n=3) against a 15,437 budget; 32,769 tokens → 422; latency 2.1 s at 32k. JevBench is still 202/231 with 0 diffs.
- ⚠ The first call in a new length bucket costs ~6.5 s: Triton/fla autotune, apparently in 2,048-token buckets, ~16 of them up to 32k. The cache (`/tmp/triton-cache`, writable layer) SURVIVES `docker restart` (measured) but is LOST on recreate. Fix, not done: a named volume for the cache plus a ~2 min warm-up after an image change. - ⚠ The first call in a new length bucket costs ~6.5 s: Triton/fla autotune, apparently in 2,048-token buckets, ~16 of them up to 32k. The cache (`/tmp/triton-cache`, writable layer) SURVIVES `docker restart` (measured) but is LOST on recreate. **Prime 1545: build it. TASKED to infra-hermes (a named volume for the cache plus `scripts/intern-decision-warmup`, proving the bucket width with random sizes); infra-ops AUDITS when it lands.**
- **intern-decision LIVE on fv-ml1 GPU 1 since 0941 2026-09-30, REPLACING SemIf (Prime: "replace semif with intern-decision now", with Scriberr fixed alongside).** - **intern-decision LIVE on fv-ml1 GPU 1 since 0941 2026-09-30, REPLACING SemIf (Prime: "replace semif with intern-decision now", with Scriberr fixed alongside).**
- Where: `http://intern-decision.fv.internal:8033`, image `intern-decision-serve:0.1.0`, token `intern-decision/api-token`. Code and contract are in `services/intern-decision-serve/`, the stack in `stacks/intern-decision`. - Where: `http://intern-decision.fv.internal:8033`, image `intern-decision-serve:0.1.0`, token `intern-decision/api-token`. Code and contract are in `services/intern-decision-serve/`, the stack in `stacks/intern-decision`.
- Surface: semif-compatible `/decide`, `/decide/shared`, `/health`. It has 12 documented deltas; the main one is that the questions in one call share a prompt, in calls of at most 16. - Surface: semif-compatible `/decide`, `/decide/shared`, `/health`. It has 12 documented deltas; the main one is that the questions in one call share a prompt, in calls of at most 16.