memory: snapshot — U11a off + U11b gate; SemIf→intern-decision (Jev, 32k); Scriberr GPU 3 + slicer + gap retry; Parakeet seat switch approved for next session; 26 entries archived

This commit is contained in:
vh
2026-09-30 23:52:27 -07:00
parent 212b736836
commit 1ae324d576
27 changed files with 1608 additions and 1467 deletions
@@ -0,0 +1,24 @@
# SemIf replaced by intern-decision (Intern-Decision-4B) on fv-ml1 GPU 1 (2026-09-30)
**Jev candidate bench** (Prime's ask, relayed by brokkr; run on GPU 3, 0149–0456; `docs/pfi/jev-candidates-bench-2026-09-30.md`, 475d6d6):
- Intern-Decision-4B on its own runtime matches SemIf-with-rotations at ONE ordering (pooled +1.5, inside the ~4-pt floor), is better on Wyrd, and is 1.5–2.3× faster.
- JevBench rank does NOT transfer to our sets; Plumb, the board leader, is worse on Wyrd.
- The positive control reproduced exactly: SemIf 187/231, hard 0.613.
- The losing weights (plumb, JevK5, imajev; ~24 GB) were deleted on Prime's word at 1234.
**Prime: "replace semif with intern-decision now."**
- `intern-decision-serve` (`services/intern-decision-serve/`, `stacks/intern-decision`, :8033, token `intern-decision/api-token`) went live at 0941.
- It keeps semif's `/decide` and `/decide/shared`. 12 deltas are documented; the main one is that questions in one call share a prompt, at most 16 per call.
- The SemIf container was REMOVED at 0949. Its image, files and token are kept; the rollback is in `stacks/semif/README.md`.
**Jev compatibility:**
- infra-hermes coded `POST /v1/systemone` as 0.1.1 and 0.1.2; my audit passed twice.
- The pass line: JevBench v1.2.16's `typesafe` adapter gives 202/231 (hard 83) with 0 row diffs against the bench.
- More than 16 questions → 422 (never chunked). Non-empty images → 422. True Jev is TEXT ONLY per docs.typesafe.ai, so that matches.
**32k context (Prime: scriberr to GPU 3, then extend the Jev endpoint to 32k):**
- `VRAM_CAP_GIB=14.4` and `MAX_TOKENS=32768`. The card peak at the limit is 15,220 MiB against a 15,437 budget, n=3, for 1 and 16 questions. 32,769 tokens → 422. 2.1 s at 32k.
- Budget from nvidia-smi `Free`, never total − used: the driver reserves ~640 MiB.
- Kernel warm-up: Triton/fla autotune runs per 2,048-token bucket, 16 buckets. The cache lives on the named volume `intern-decision_triton-cache` (0.1.3, infra-hermes, audited), so it survives a recreate. Run `scripts/intern-decision-warmup` after an IMAGE change: 109 s cold, 17 s warm.
Open: label ~50 real Wyrd/Cicada turns before trusting it in production; its card makes no contamination claim.