Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md
T

25 lines
2.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# SemIf replaced by intern-decision (Intern-Decision-4B) on fv-ml1 GPU 1 (2026-09-30)
**Jev candidate bench** (Prime's ask, relayed by brokkr; run on GPU 3, 0149–0456; `docs/pfi/jev-candidates-bench-2026-09-30.md`, 475d6d6):
- Intern-Decision-4B on its own runtime matches SemIf-with-rotations at ONE ordering (pooled +1.5, inside the ~4-pt floor), is better on Wyrd, and is 1.5–2.3× faster.
- JevBench rank does NOT transfer to our sets; Plumb, the board leader, is worse on Wyrd.
- The positive control reproduced exactly: SemIf 187/231, hard 0.613.
- The losing weights (plumb, JevK5, imajev; ~24 GB) were deleted on Prime's word at 1234.
**Prime: "replace semif with intern-decision now."**
- `intern-decision-serve` (`services/intern-decision-serve/`, `stacks/intern-decision`, :8033, token `intern-decision/api-token`) went live at 0941.
- It keeps semif's `/decide` and `/decide/shared`. 12 deltas are documented; the main one is that questions in one call share a prompt, at most 16 per call.
- The SemIf container was REMOVED at 0949. Its image, files and token are kept; the rollback is in `stacks/semif/README.md`.
**Jev compatibility:**
- infra-hermes coded `POST /v1/systemone` as 0.1.1 and 0.1.2; my audit passed twice.
- The pass line: JevBench v1.2.16's `typesafe` adapter gives 202/231 (hard 83) with 0 row diffs against the bench.
- More than 16 questions → 422 (never chunked). Non-empty images → 422. True Jev is TEXT ONLY per docs.typesafe.ai, so that matches.
**32k context (Prime: scriberr to GPU 3, then extend the Jev endpoint to 32k):**
- `VRAM_CAP_GIB=14.4` and `MAX_TOKENS=32768`. The card peak at the limit is 15,220 MiB against a 15,437 budget, n=3, for 1 and 16 questions. 32,769 tokens → 422. 2.1 s at 32k.
- Budget from nvidia-smi `Free`, never total − used: the driver reserves ~640 MiB.
- Kernel warm-up: Triton/fla autotune runs per 2,048-token bucket, 16 buckets. The cache lives on the named volume `intern-decision_triton-cache` (0.1.3, infra-hermes, audited), so it survives a recreate. Run `scripts/intern-decision-warmup` after an IMAGE change: 109 s cold, 17 s warm.
Open: label ~50 real Wyrd/Cicada turns before trusting it in production; its card makes no contamination claim.