Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md
T

2.2 KiB
Raw Blame History

SemIf replaced by intern-decision (Intern-Decision-4B) on fv-ml1 GPU 1 (2026-09-30)

Jev candidate bench (Prime's ask, relayed by brokkr; run on GPU 3, 0149–0456; docs/pfi/jev-candidates-bench-2026-09-30.md, 475d6d6):

  • Intern-Decision-4B on its own runtime matches SemIf-with-rotations at ONE ordering (pooled +1.5, inside the ~4-pt floor), is better on Wyrd, and is 1.5–2.3× faster.
  • JevBench rank does NOT transfer to our sets; Plumb, the board leader, is worse on Wyrd.
  • The positive control reproduced exactly: SemIf 187/231, hard 0.613.
  • The losing weights (plumb, JevK5, imajev; ~24 GB) were deleted on Prime's word at 1234.

Prime: "replace semif with intern-decision now."

  • intern-decision-serve (services/intern-decision-serve/, stacks/intern-decision, :8033, token intern-decision/api-token) went live at 0941.
  • It keeps semif's /decide and /decide/shared. 12 deltas are documented; the main one is that questions in one call share a prompt, at most 16 per call.
  • The SemIf container was REMOVED at 0949. Its image, files and token are kept; the rollback is in stacks/semif/README.md.

Jev compatibility:

  • infra-hermes coded POST /v1/systemone as 0.1.1 and 0.1.2; my audit passed twice.
  • The pass line: JevBench v1.2.16's typesafe adapter gives 202/231 (hard 83) with 0 row diffs against the bench.
  • More than 16 questions → 422 (never chunked). Non-empty images → 422. True Jev is TEXT ONLY per docs.typesafe.ai, so that matches.

32k context (Prime: scriberr to GPU 3, then extend the Jev endpoint to 32k):

  • VRAM_CAP_GIB=14.4 and MAX_TOKENS=32768. The card peak at the limit is 15,220 MiB against a 15,437 budget, n=3, for 1 and 16 questions. 32,769 tokens → 422. 2.1 s at 32k.
  • Budget from nvidia-smi Free, never total − used: the driver reserves ~640 MiB.
  • Kernel warm-up: Triton/fla autotune runs per 2,048-token bucket, 16 buckets. The cache lives on the named volume intern-decision_triton-cache (0.1.3, infra-hermes, audited), so it survives a recreate. Run scripts/intern-decision-warmup after an IMAGE change: 109 s cold, 17 s warm.

Open: label ~50 real Wyrd/Cicada turns before trusting it in production; its card makes no contamination claim.