Files
esh-pfi-infrastructure/docs/pfi/jev-candidates-bench-2026-09-30.md
T
vh 475d6d6bcb docs: Jev candidate bench vs SemIf (fv-ml1 GPU 3) — Intern-Decision-4B is the replacement if SemIf is displaced
Operator ask relayed by brokkr-smithy-dev. Positive control (SemIf 187/231,
hard 0.613) reproduced exactly; negative control and a 4-restart noise floor
(0 flips) measured. On our replaced-baseline sets no candidate beats SemIf-with-
rotations beyond the ~4-pt floor; Intern-Decision-4B native matches it at one
ordering, is better on Wyrd, fits 9.7/10.3 GB and is 1.5-2.3x faster. JevBench
rank does not transfer. Raw per-item data kept out of git.
2026-09-30 05:00:45 -07:00

52 KiB
Raw Blame History

Jev candidates bench: a local replacement for SemIf on the utility GPU (2026-09-30)

Ask (Prime, relayed by brokkr-smithy-dev, 2026-09-30): "Find a local Jev that will fit on our utility GPU at good latency if SemIf is displaced." This is a bench, not a switch. semif-serve and its stack were not touched. The live semif container was stopped by the operator at 0135 for scriberr's sake; this bench did not use it or restart it. Where: fv-ml1 GPU 3 only, transiently, one model at a time; GPU 3 read 2 MiB before every repeat and after cleanup. Runs 0149–0456 PDT. Raw data: services/semif-serve/bench-jev-2026-09-30/ (code/ = every script, raw/out/<system>/r<k>/ = every run, summary.json, tables.md).

Verdict

None of the candidates is resolvably better than SemIf when SemIf is run the way its README says to run it (with rotations). If SemIf is displaced, take Intern-Decision-4B, served by its own runtime. It is the only candidate that is at least as accurate as SemIf on our sets, fits the 12 GiB cap with room to spare, and is faster than SemIf at our shape.

  • Intern-Decision-4B (native) fits: 9.7 GB at rest and a 10.3 GB peak, against SemIf's 9.2 / 12.9 GB. It is faster at our shape: 21 binary criteria take 88 ms against 131 ms, and 16 criteria over a ~3,900-token state take 215 ms against 505 ms, because it scores every question in one forward pass.
    • Against SemIf as deployed (one ordering): better. The pooled 259 rows go from 217 to 240, +8.9 pts (95% CI +4.7..+13.2). On Wyrd, the private consumer set, it scores 79 against 71 of 84 (+9.5, CI +3.6..+16.7).
    • Against SemIf with rotations: the same on the pooled set (+1.5, CI −2.0..+5.0). It stays better on Wyrd (77 vs 71, +7.1, CI +1.2..+13.1, McNemar p = 0.07).
    • At one ordering it matches SemIf-with-rotations on the pooled set (+3.1, CI −0.4..+6.7) and beats it on Wyrd (+9.5, CI +2.4..+16.7, p = 0.02). So it gets the rotation-level accuracy without paying for rotations. Its answers barely depend on option order: reversing the options changes 8/144 labels, against SemIf's 30/144.
    • Caveats.
      • Its card has no contamination statement, and it is not on the JevBench board, so no held-out or sealed number exists for it.
      • It is not a drop-in. Its request format is the Jev schema, it scores all questions in one prompt, and it takes at most 16 questions per request. Pushed through semif-serve unchanged, it keeps only part of the gain: pooled +5.8 at one ordering and +0.4 with rotations, and on Wyrd +3.6, which is not significant.
      • Its card pins transformers 5.14.1; we ran 5.17.0. Its hard tier came out at 83/111 against the vendor's 82.
  • The JevBench ranking does not carry over to our workload. Every candidate beats SemIf on JevBench's public items: 198–207/231 against 187. On our sets, the gains all sit on SemIf's own public NLI-style rows (authored144 and perturbations108, both evidence → supported / insufficient / contradicted, the family these models train on). On the private Cicada and Wyrd sets they vanish or reverse:
    • Plumb, top of JevBench at 207/231, scores 65 against 71 on Wyrd.
    • JevK5 v0.3 scores 26 against 30 on Cicada w1.
    • Imajev's +15 pts on authored144 becomes −3.6 on Wyrd.

One line per candidate (against the measured floor: SemIf moved 0 labels in 4 restarts, so every difference below is item sampling; the pooled set resolves about ±4 pts)

candidate fits the 12 GiB cap? vs SemIf, replaced-baseline sets latency at our shape licence at the raw file
Intern-Decision-4B native yes, 9.7 / 10.3 GB better than SemIf-single (pooled +8.9, Wyrd +9.5); same as SemIf-rotations pooled (+1.5), better on Wyrd (+7.1) faster (88 vs 131 ms; 215 vs 505 ms) Apache-2.0 (Shanghai AI Lab) + Qwen Apache-2.0. No contamination statement
Intern-Decision-4B drop-in yes, same as SemIf better than SemIf-single (+5.8); same as SemIf-rotations (+0.4); Wyrd same same as SemIf as above
Imajev-4B native (no drop-in possible) yes, 10.4 / 11.1 GB better than SemIf-single (+6.9) but same as SemIf-rotations (+1.5); the gain is authored/perturbations only; Wyrd same (−3.6) 9–10× slower (1,219 ms; 5,125 ms); 8 questions per request Apache-2.0 on GitHub; no LICENSE file in the HF repo
Plumb-4B native yes, but pins 12.06 GB at rest (CUDA graphs) same (+2.7 / −1.9); worse on Wyrd (−7.1, CI −13.1..−1.2) 2.2× (294 ms) to 6.7× (3.4 s) slower Apache-2.0 (raw LICENSE + NOTICE)
Plumb-4B drop-in yes, same as SemIf same (+1.5 / −3.1); Cicada w1 worse single (26 vs 30) same as SemIf as above
JevK5 v0.2 native yes, pins 12.06 GB same (+4.6 / −1.5); Wyrd same (−4.8) like Plumb Apache-2.0 on GitHub; no LICENSE file in the HF repo; trained on 940 MMLU-Pro test items (disclosed)
JevK5 v0.2 drop-in yes, same as SemIf better than SemIf-single (+5.8, CI +1.1..+10.7); same as SemIf-rotations (+1.9) same as SemIf as above
JevK5 v0.3 (current HF main; extra) yes / same same pooled (native +3.9 / −2.3, drop-in +3.9 / −1.9); worse on Cicada w1 in both formats (26 vs 30); native worse on Wyrd (−6.0 / −7.1) like v0.2 as above; its NOTICE says 14k training questions were written by GPT-6 Luna "under OpenAI's terms"

Key table (median over repeats, min–max where they differ; loopback on fv-ml1 GPU 3)

system fits 12 GiB? rest / peak MiB JevBench all /231 · hard /111 pooled /259 single · rot Wyrd /84 single · rot Δ pooled vs SemIf single (95% CI) Δ pooled vs SemIf rot (95% CI) 21 criteria, ms 16 × ~3,900 tok, ms
SemIf (Qwen3.5-4B), baseline yes: 9242 (9242–9250) / 12918 (12886–12918) 187 · 68 217 · 232 71 · 71 baseline baseline 131 (131–132) 505 (502–519)
Plumb-4B, native yes: 12064 / 12064 207 (206–207) · 89 (88–89) 224 (224–226) · 227 (227–228) 65 · 67 +2.7 (-2.7..+7.9) -1.9 (-5.4..+1.5) 294 (292–295) 3378 (3374–3384)
Plumb-4B, drop-in yes: 9242 / 12918 206 · 88 221 · 224 66 · 67 +1.5 (-3.7..+6.6) -3.1 (-6.8..+0.4) 131 (131–132) 506 (506–507)
Imajev-4B, native yes: 10404 / 11056 199 · 80 235 · 236 68 · 69 +6.9 (+1.6..+12.2) +1.5 (-2.5..+5.6) 1219 (1212–1230) 5125 (5054–5140)
Intern-Decision-4B, native yes: 9736 / 10290 202 · 83 240 · 236 79 · 77 +8.9 (+4.7..+13.2) +1.5 (-2.0..+5.0) 88 (88–89) 215 (213–216)
Intern-Decision-4B, drop-in yes: 9242 / 12918 202 · 85 232 · 233 74 · 73 +5.8 (+2.7..+9.0) +0.4 (-2.3..+3.1) 132 508 (500–516)
JevK5 v0.2, native yes: 12064 / 12064 198 (198–199) · 81 (81–82) 229 (228–229) · 228 67 (66–67) · 67 +4.6 (-0.8..+9.9) -1.5 (-4.9..+1.9) 291 (291–292) 3360 (3358–3368)
JevK5 v0.2, drop-in yes: 9242 / 12918 200 · 82 232 · 237 72 · 74 +5.8 (+1.1..+10.7) +1.9 (-1.1..+5.1) 132 518 (505–519)
JevK5 v0.3, native yes: 12064 / 12064 202 · 86 227 · 226 66 · 65 +3.9 (-0.9..+8.5) -2.3 (-6.0..+1.4) 291 (291–292) 3363 (3355–3370)
JevK5 v0.3, drop-in yes: 9242 / 12918 201 · 85 227 · 227 69 · 66 +3.9 (-0.8..+8.2) -1.9 (-5.5..+1.8) 133 (131–133) 518 (506–521)

Pooled = authored144 + Cicada w1 + Wyrd (all 4 decisions) = 259 labelled rows. "single" = one ordering, "rot" = n rotations averaged. Δ is paired, per row, on each row's majority answer over the repeats, with a group-bootstrap 95% CI.

Recommendation (each claim with its own strength)

  • "If SemIf is displaced, the replacement is Intern-Decision-4B run on its own runtime": recommend [measured: best on the pooled set and on Wyrd, fastest at our shape, fits the cap; reversible, since it is a service swap and the weights are already on /tank].
  • "Displace SemIf today on accuracy alone": lean against [measured: against SemIf-with-rotations the pooled gain (+1.5 to +3.1) is inside the ~±4-pt floor; reversible]. What would justify a switch is compute: rotation-level accuracy at one ordering, and 1.5–2.3× faster multi-criteria requests. That is a build: a small new service, contract and TDD. It is not a config flip.
  • "Pick Plumb or Imajev because of their JevBench rank": strongly recommend against [measured: no gain on our sets; Plumb is worse on Wyrd and Imajev is 9× slower; reversible].
  • "Swap weights inside semif-serve (drop-in) for a free gain": lean against [measured: the best drop-in, JevK5 v0.2, is +1.9 against SemIf-rotations, CI −1.1..+5.1, which is the same; reversible].
  • Before trusting Intern-Decision in production: label ~50 real Wyrd/Cicada turns as a held-out set. Its card makes no contamination claim, and Wyrd and Cicada are the only sets here that no model could have seen. recommend [reasoned; reversible].

Plain-language decision tree

  • Do we only need to free the utility GPU, or do we want a better decider?
    • Only free the GPU / keep what works → keep SemIf with rotations. Nothing here beats it by enough to justify a new service. (Every candidate and format is within ±3.1 pts of SemIf-with-rotations on the pooled 259 rows; the floor is ~±4.)
    • Want the same answers cheaper, or better answers on the Wyrd-style multi-question turns → Intern-Decision-4B on its own runtime.
      • It needs a small new service, because its request format differs (Jev schema; one forward over up to 16 questions; inference.py ships in the model repo).
      • First, check it on ~50 real labelled turns (its card makes no contamination claim, and it is not on the JevBench board).
    • Want a zero-code swap through semif-serve → not worth it. (Best drop-in: JevK5 v0.2, +1.9 pts, CI spans 0; same latency.)

Harness (stated once; every number below was measured under it)

Card fv-ml1 GPU 3 (RTX PRO 6000 Blackwell Max-Q, 96 GB, sm_120), empty at start (2 MiB) and nothing foreign on it during the bench (guard log per repeat). Nothing on GPU 1.
Stack, every system torch 2.10.0+cu128, transformers 5.17.0, flash-linear-attention / fla-core 0.5.2, causal-conv1d 1.7.0, BF16, SDPA. This is semif-serve:0.1.4's image (sha256:36a3e1d2…); the native runtimes run in jevbench-native:2026-09-30, the same image plus jevk5 0.2.0 (85238d7b), peft 0.21.1, pillow 12.3.0, torchvision 0.25.0+cu128.
VRAM cap 12 GiB for every served system: semif-serve's own SEMIF_VRAM_CAP_GIB=12; the native servers under capped.py, which applies the same torch.cuda.set_per_process_memory_fraction before any weights load.
Clients fv-ml1 host python 3.13, stdlib urllib, loopback (127.0.0.1), one request at a time. No network in the latency numbers (the README's nh3-dev figures carry ~27–33 ms of network on top).
Repeat one repeat = one fresh process (container) lifetime. N = 3 per cell; 4 for SemIf, Imajev and Intern-Decision native (one repeat added each: Imajev's first repeat ran its latency shapes without its 8-questions-per-request cap and was discarded for latency only, r1/shape-maxq-unset-invalid.json). The Wyrd one-request-per-turn condition was added after the first repeats, so it has N = 3 except Plumb native (N = 2; that runtime loops over questions, so it equals single by construction).
Prompts native = each candidate's own runtime and prompt as its card documents it, reached over its TypeSafe-style /v1/systemone server; a row becomes one choice question with criteria = {option id: description} in the row's order. drop-in = the candidate's weights served by semif-serve 0.1.4 unchanged (SEMIF_MODEL/SEMIF_REVISION), i.e. SemIf's own prompt, letter readout and shared-prefix path: what a switch through the semif-serve contract would get.
Orderings single = the caller's order. rotations = the n cyclic rotations averaged by log-mean (semif-serve's own "orderings": "rotations" for semif-serve; the same arithmetic client-side, n separate requests, for native runtimes).
JevBench fstandhartinger/jevbench tag v1.2.16 (5e95f23cbb7be098a9061fea924c4421620ab1a5), the 231 public items (easy 48, standard/original 72, hard 111), its own jevbench.cli run, one ordering, v1.2 scoring (argmax over the exact label set). SemIf and every drop-in through JevBench's own semif_direct adapter (in-process, SemIf's code path); native candidates through its typesafe adapter to the candidate's loopback server, which is how the board measured them.

Controls, and what each one showed

Positive control 1 — the JevBench harness (step 1). SemIf on our box, through JevBench v1.2.16's own semif_direct adapter: easy 48/48, standard 71/72, hard 68/111 (0.613), all 187/231 (0.810), identical in all 4 repeats. That is the cards' 0.810 / 0.613 exactly, and the same outcome JevK5's card reports for the untrained base. The harness is right. Each candidate's public number was also reproduced, which makes every candidate row its own positive control: Plumb 89/111 hard (its card: 0.802 = 89), Imajev 80/111 (board hard_public 0.7207 = 80), JevK5 v0.2 81–82/111 native, 82 drop-in (card 0.739 = 82), Intern-Decision 83/111 against a vendor 82/111 (one item over, on transformers 5.17.0 where the card pins 5.14.1).

Positive control 2 — the replaced-baseline harness. The SemIf repeats reproduce semif-serve 0.1.4's acceptance and the 09-27 spikes row for row: parity with SemIf's committed torch predictions 144/144 top choice, 144/144 identical prompt SHA-256, max prob gap 0.060 (acceptance: 144/144, 0.060); negative control 14/144 same-top (acceptance: 14/144); Cicada w1 30/31; Wyrd with rotations place 16/21, place2 21/21, exit 18/21, exit2 16/21 (spike: identical); SemIf's labelled sets 78.6% → 88.1% with rotations in the averaging acceptance, here 199/252 = 79.0% → 224/252 = 88.9%. Capacity under 12 GiB: 16 rows at ~3,900 prefix tokens, 18 is a 503 (README: 16). Latency: 21 binary criteria in 131 ms on loopback; the README's 159 ms was measured from nh3-dev, whose round trip is 22–28 ms (ping averages today); a bridge run from nh3-dev against the GPU 3 instance gave 151.5 ms (below, under what could not be measured).

Negative control (step 6). authored144 with each row's option descriptions rotated one place and the ids kept. A model that reads the descriptions must move its top to the id that now carries the right description. Reported three ways: same top as unrotated (should be LOW), follows the description (should be HIGH), right against the original gold (should be LOW).

Null control. Every Cicada and Wyrd row is also asked over a content-free state ("(No evidence is available for this turn.)"). Whatever that gets right is the prior carried by the question and options alone; the evidence condition is read against it.

Noise floors. A-vs-A inside a process: authored144 scored twice in the same process. Across restarts: every repeat is a fresh container, so each pair of repeats is an A-vs-A across restarts over all 560 labelled+blind rows (single) and 560 (rotations). The paired comparisons use each row's majority top over the repeats, so a row that flips across restarts cannot manufacture a difference by itself.

Sensitivity floor. Item sampling dominates, not run noise (see the floors table: SemIf flips 0 rows across restarts). Paired group-bootstrap 95% CIs on the pooled 259 rows have a half-width of about ±3 to ±5 points, so this bench cannot resolve a pooled difference smaller than ~4 points; on Cicada (31 rows) nothing under ~10–15 points is resolvable, on Wyrd (84 rows, 21 turns) nothing under ~6–7.

Decision rule, stated before the final repeats were in (after the first repeat of each system had been seen): a candidate is better on a set when the paired 95% CI against SemIf excludes zero in its favour AND the gap exceeds the number of rows either side flips across restarts; worse the mirror image; otherwise same. It is judged against SemIf both as deployed today (single ordering, the default) and with rotations (what semif-serve's README says to use "for anything real").

Candidates: identity, pins, licence as read at the raw file

Every HF repo ID was verified with an authenticated API call (/api/models/<id>, HF token from the vault bundle nh3-dev/.config/secrets/env.sh) before any pull: all returned 200, none gated. Weights were pulled to /tank/aimodels/huggingface on fv-ml1 as llmuser and every weight file's sha256 was checked against the Hub's LFS oid (all match; raw/weights-sha256.txt).

candidate HF repo @ pinned revision what it is runtime used (native) licence, read at the raw file contamination statement JevBench board (v1.4.2.2)
SemIf (baseline) Qwen/Qwen3.5-4B @ 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a frozen base; SemIf 23cf1f39 reads the option letters semif-serve 0.1.4 Qwen LICENSE: Apache-2.0 (Copyright 2026 Alibaba Cloud). SemIf: MIT n/a (frozen) #13, 47.7; hard(220) 0.595, sealed 0.263
Plumb-4B crh225/plumb-4b @ 55de037801a8a9b9de3db5c0e16cef86210c2186 (the board's pin; HF main 24f7bf77 changes only the card) JevK5 v0.2 + LoRA, merged, 8.4 GB jevk5 runtime 0.2.0 (85238d7b), the board's serving path LICENSE: Apache-2.0 (unfilled template). NOTICE: JevK5 v0.2 (Apache-2.0), Qwen (Apache-2.0), SemIf prompt/readout (MIT) yes: no JevBench item trained/tuned/selected; 8-gram scan + exact-text audit #2, 65.8; public 0.896, sealed 0.380
Imajev-4B mohit67890/imajev-4b @ c9e5f132465da85d31735ec502d5557982671a7d (the board's pin) PEFT LoRA (r64) on Qwen3.5-4B + its own 256-code decision readout head; 0.5 GB adapter imajev server a0134749 (the board's commit), torch backend, --rotations 1 --calibration calibration.json HF repo has no LICENSE file (card tag only). GitHub mohit67890/imajev LICENSE: Apache-2.0; card: "Code and adapters: Apache-2.0" yes: "No JevBench items (8-gram lint)" #1, 67.4; public 0.861, sealed 0.370
Intern-Decision-4B internlm/Intern-Decision-4B @ 0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd full Qwen3.5-4B finetune (incl. vision tower), 9.1 GB, adds a <decision> token its own inference.py (DecisionEngine, shipped in the repo) behind a 60-line /v1/systemone wrapper LICENSE: Apache-2.0 (Copyright 2026-2027 Shanghai AI Laboratory); LICENSE-QWEN: Qwen's Apache-2.0 none on the card (flagged) not on the board (released 09-26); vendor-only claims
JevK5 v0.2 alibiserikbay/JevK5 @ 27d2d6b8d4714807f6293b0623bd7370b27e42f8 (= Hub tag v0.2; tag object ea4804e9; weights identical to the v0.2 release 844e4d0a, LFS 0fba3bba…) Qwen3.5-4B + distilled LoRA, merged, 8.4 GB. allebee/jevk5 (GitHub) and alibiserikbay/JevK5 (HF) are the same project: the NOTICE names both handles jevk5 runtime 0.2.0 (85238d7b) HF repo has no LICENSE file. GitHub allebee/jevk5 LICENSE: Apache-2.0 yes (v0.2 card); discloses 940 MMLU-Pro test items in its training replay #5, 62.0; public 0.853, sealed 0.331
JevK5 v0.3 (extra) alibiserikbay/JevK5 @ c4f7fdb3aeab5582336406e78d3bef11bf98833d (= Hub tag v0.3 and current main) as above, retrained; 8.4 GB jevk5 runtime 0.3.3 (f944fe37) as above; NOTICE: 14,138 training questions written by GPT-6 Luna "under OpenAI's terms" yes not on the board

Not run, with the reason:

  • Hopper (HopitAI/hopper, last in brokkr's order): the adapter's card sets license: other / research-and-demo and says "Do not use this adapter commercially" (RACE training data, non-commercial terms). There is no LICENSE file. That is the same class as openjev/openjev (CC-BY-NC), which the brief excluded, so it was excluded by the same rule. It is one queue line if Prime wants it measured anyway.
  • Imajev drop-in: not applicable. Imajev is a LoRA plus its own readout head, not a full checkpoint whose LM-head letters carry the answer, so semif-serve cannot load it without new code.
  • Intern-Decision-2B/0.8B, AlexWortega/openjev: dropped per brokkr's revised order. akhilaaa3/Jev-Omni, openjev/openjev, Cygnet: excluded as briefed (size / licence).

Full tables

Generated by code/tables.py from summary.json (code/analyze.py over raw/out/), all under services/semif-serve/bench-jev-2026-09-30/.

JevBench v1.2.16 public items (positive control + candidate score)

system repeats easy /48 standard /72 hard /111 all /231 all p50 ms
SemIf (Qwen3.5-4B), baseline 4 48 71 68 187 0.810 36
Plumb-4B, native 3 48 70 89 (88–89) 207 (206–207) 0.896 (0.892–0.896) 20
Plumb-4B, drop-in 3 48 70 88 206 0.892 36
Imajev-4B, native 4 48 71 80 199 0.861 62 (62–63)
Intern-Decision-4B, native 4 48 71 83 202 0.874 40
Intern-Decision-4B, drop-in 3 48 69 85 202 0.874 36
JevK5 v0.2, native 3 48 69 81 (81–82) 198 (198–199) 0.857 (0.857–0.861) 20 (20–21)
JevK5 v0.2, drop-in 3 48 70 82 200 0.866 36 (36–37)
JevK5 v0.3, native 3 48 68 86 202 0.874 20 (20–21)
JevK5 v0.3, drop-in 3 48 68 85 201 0.870 36

Replaced-baseline sets, single (caller order)

system pooled /259 authored144 perturb108 Cicada w1 /31 Cicada w2 /31 Wyrd /84 Wyrd place2 /21 Wyrd exit /21
SemIf (Qwen3.5-4B), baseline 217 116 83 30 20 71 21 16
Plumb-4B, native 224 (224–226) 130 (130–132) 84 29 19 65 21 10
Plumb-4B, drop-in 221 129 88 26 20 66 21 12
Imajev-4B, native 235 138 102 29 22 68 20 13
Intern-Decision-4B, native 240 132 100 29 22 79 21 19
Intern-Decision-4B, drop-in 232 128 102 30 19 74 21 17
JevK5 v0.2, native 229 (228–229) 132 85 (85–86) 30 16 67 (66–67) 21 12 (11–12)
JevK5 v0.2, drop-in 232 130 92 30 17 72 21 17
JevK5 v0.3, native 227 135 101 (101–102) 26 17 66 21 12
JevK5 v0.3, drop-in 227 132 99 26 21 69 21 15

Replaced-baseline sets, rotations (n rotations averaged)

system pooled /259 authored144 perturb108 Cicada w1 /31 Cicada w2 /31 Wyrd /84 Wyrd place2 /21 Wyrd exit /21
SemIf (Qwen3.5-4B), baseline 232 131 93 30 19 71 21 18
Plumb-4B, native 227 (227–228) 130 (130–131) 87 30 18 (17–18) 67 21 11
Plumb-4B, drop-in 224 129 92 28 18 67 21 13
Imajev-4B, native 236 138 101 29 23 69 20 13
Intern-Decision-4B, native 236 130 99 29 22 77 21 19
Intern-Decision-4B, drop-in 233 130 102 30 19 73 21 17
JevK5 v0.2, native 228 131 91 (90–91) 30 16 67 21 12
JevK5 v0.2, drop-in 237 133 93 30 16 74 21 20
JevK5 v0.3, native 226 135 101 26 16 65 20 12
JevK5 v0.3, drop-in 227 135 101 26 19 66 20 13

Wyrd as one request per turn (4 decisions over one state: SemIf /decide/shared, native multi-question request)

system repeats Wyrd /84 place /21 place2 /21 exit /21 exit2 /21
SemIf (Qwen3.5-4B), baseline 3 71 17 21 16 17
Plumb-4B, native 2 65 17 21 10 17
Plumb-4B, drop-in 3 66 17 21 12 16
Imajev-4B, native 3 68 19 20 13 16
Intern-Decision-4B, native 3 77 16 21 20 20
Intern-Decision-4B, drop-in 3 74 18 21 17 18
JevK5 v0.2, native 3 67 (66–67) 17 21 12 (11–12) 17
JevK5 v0.2, drop-in 3 73 18 21 18 16
JevK5 v0.3, native 3 66 18 21 12 15
JevK5 v0.3, drop-in 3 69 18 21 15 15

Null control (content-free state) and positive controls inside the spike sets, single ordering

system Cicada w1 blind /31 Wyrd blind /84 Cicada w1 controls /5 Wyrd controls /28
SemIf (Qwen3.5-4B), baseline 16 60 5 21
Plumb-4B, native 15 50 5 20
Plumb-4B, drop-in 15 60 4 21
Imajev-4B, native 16 48 5 24
Intern-Decision-4B, native 16 60 5 27
Intern-Decision-4B, drop-in 16 60 5 23
JevK5 v0.2, native 15 53 5 22 (21–22)
JevK5 v0.2, drop-in 15 60 5 23
JevK5 v0.3, native 16 60 5 20
JevK5 v0.3, drop-in 16 60 5 22

Paired against SemIf (each row's majority top over the repeats; group bootstrap 95% CI; exact McNemar)

system set cond SemIf cand fixed broken Δ pts 95% CI p
Plumb-4B, native pooled single 217/259 224/259 23 16 +2.7 -2.7..+7.9 0.3368
Plumb-4B, native pooled rotations 232/259 227/259 10 15 -1.9 -5.4..+1.5 0.4244
Plumb-4B, native authored144 single 116/144 130/144 22 8 +9.7 +1.4..+18.1 0.0161
Plumb-4B, native authored144 rotations 131/144 130/144 7 8 -0.7 -5.6..+3.5 1.0
Plumb-4B, native cicada-w1 single 30/31 29/31 0 1 -3.2 -9.7..+0.0 1.0
Plumb-4B, native cicada-w1 rotations 30/31 30/31 0 0 +0.0 +0.0..+0.0 1.0
Plumb-4B, native wyrd single 71/84 65/84 1 7 -7.1 -13.1..-1.2 0.0703
Plumb-4B, native wyrd rotations 71/84 67/84 3 7 -4.8 -11.9..+2.4 0.3438
Plumb-4B, native perturbations108 single 83/108 84/108 10 9 +0.9 -9.3..+11.1 1.0
Plumb-4B, native perturbations108 rotations 93/108 87/108 4 10 -5.6 -15.7..+3.7 0.1796
Plumb-4B, native cicada-w2 single 20/31 19/31 1 2 -3.2 -12.9..+6.5 1.0
Plumb-4B, native cicada-w2 rotations 19/31 18/31 1 2 -3.2 -16.1..+6.5 1.0
Plumb-4B, native pooled cand-single-vs-semif-rotations 232/259 224/259 10 18 -3.1 -7.3..+0.8 0.1849
Plumb-4B, native wyrd cand-single-vs-semif-rotations 71/84 65/84 3 9 -7.1 -15.5..+1.2 0.146
Plumb-4B, drop-in pooled single 217/259 221/259 21 17 +1.5 -3.7..+6.6 0.6271
Plumb-4B, drop-in pooled rotations 232/259 224/259 7 15 -3.1 -6.8..+0.4 0.1338
Plumb-4B, drop-in authored144 single 116/144 129/144 20 7 +9.0 +1.4..+16.7 0.0192
Plumb-4B, drop-in authored144 rotations 131/144 129/144 5 7 -1.4 -6.2..+3.5 0.7744
Plumb-4B, drop-in cicada-w1 single 30/31 26/31 0 4 -12.9 -25.8..-3.2 0.125
Plumb-4B, drop-in cicada-w1 rotations 30/31 28/31 0 2 -6.5 -16.1..+0.0 0.5
Plumb-4B, drop-in wyrd single 71/84 66/84 1 6 -6.0 -11.9..+0.0 0.125
Plumb-4B, drop-in wyrd rotations 71/84 67/84 2 6 -4.8 -11.9..+2.4 0.2891
Plumb-4B, drop-in perturbations108 single 83/108 88/108 12 7 +4.6 -5.6..+14.8 0.3593
Plumb-4B, drop-in perturbations108 rotations 93/108 92/108 6 7 -0.9 -10.2..+7.4 1.0
Plumb-4B, drop-in cicada-w2 single 20/31 20/31 1 1 +0.0 -9.7..+9.7 1.0
Plumb-4B, drop-in cicada-w2 rotations 19/31 18/31 0 1 -3.2 -9.7..+0.0 1.0
Plumb-4B, drop-in pooled cand-single-vs-semif-rotations 232/259 221/259 8 19 -4.2 -8.3..-0.4 0.0522
Plumb-4B, drop-in wyrd cand-single-vs-semif-rotations 71/84 66/84 2 7 -6.0 -13.1..+1.2 0.1797
Imajev-4B, native pooled single 217/259 235/259 30 12 +6.9 +1.6..+12.2 0.0079
Imajev-4B, native pooled rotations 232/259 236/259 14 10 +1.5 -2.5..+5.6 0.5413
Imajev-4B, native authored144 single 116/144 138/144 25 3 +15.3 +9.0..+21.5 0.0
Imajev-4B, native authored144 rotations 131/144 138/144 8 1 +4.9 +0.7..+9.0 0.0391
Imajev-4B, native cicada-w1 single 30/31 29/31 0 1 -3.2 -9.7..+0.0 1.0
Imajev-4B, native cicada-w1 rotations 30/31 29/31 0 1 -3.2 -9.7..+0.0 1.0
Imajev-4B, native wyrd single 71/84 68/84 5 8 -3.6 -13.1..+6.0 0.5811
Imajev-4B, native wyrd rotations 71/84 69/84 6 8 -2.4 -11.9..+7.1 0.7905
Imajev-4B, native perturbations108 single 83/108 102/108 20 1 +17.6 +8.3..+27.8 0.0
Imajev-4B, native perturbations108 rotations 93/108 101/108 9 1 +7.4 +1.9..+13.9 0.0215
Imajev-4B, native cicada-w2 single 20/31 22/31 3 1 +6.5 -6.5..+19.4 0.625
Imajev-4B, native cicada-w2 rotations 19/31 23/31 4 0 +12.9 +3.2..+25.8 0.125
Imajev-4B, native pooled cand-single-vs-semif-rotations 232/259 235/259 15 12 +1.2 -3.2..+5.5 0.7011
Imajev-4B, native wyrd cand-single-vs-semif-rotations 71/84 68/84 6 9 -3.6 -14.3..+6.0 0.6072
Intern-Decision-4B, native pooled single 217/259 240/259 28 5 +8.9 +4.7..+13.2 0.0001
Intern-Decision-4B, native pooled rotations 232/259 236/259 12 8 +1.5 -2.0..+5.0 0.5034
Intern-Decision-4B, native authored144 single 116/144 132/144 20 4 +11.1 +4.9..+18.1 0.0015
Intern-Decision-4B, native authored144 rotations 131/144 130/144 5 6 -0.7 -5.6..+4.2 1.0
Intern-Decision-4B, native cicada-w1 single 30/31 29/31 0 1 -3.2 -9.7..+0.0 1.0
Intern-Decision-4B, native cicada-w1 rotations 30/31 29/31 0 1 -3.2 -9.7..+0.0 1.0
Intern-Decision-4B, native wyrd single 71/84 79/84 8 0 +9.5 +3.6..+16.7 0.0078
Intern-Decision-4B, native wyrd rotations 71/84 77/84 7 1 +7.1 +1.2..+13.1 0.0703
Intern-Decision-4B, native perturbations108 single 83/108 100/108 20 3 +15.7 +4.6..+26.9 0.0005
Intern-Decision-4B, native perturbations108 rotations 93/108 99/108 9 3 +5.6 -2.8..+13.0 0.146
Intern-Decision-4B, native cicada-w2 single 20/31 22/31 2 0 +6.5 +0.0..+16.1 0.5
Intern-Decision-4B, native cicada-w2 rotations 19/31 22/31 4 1 +9.7 -3.2..+22.6 0.375
Intern-Decision-4B, native pooled cand-single-vs-semif-rotations 232/259 240/259 14 6 +3.1 -0.4..+6.7 0.1153
Intern-Decision-4B, native wyrd cand-single-vs-semif-rotations 71/84 79/84 9 1 +9.5 +2.4..+16.7 0.0215
Intern-Decision-4B, drop-in pooled single 217/259 232/259 18 3 +5.8 +2.7..+9.0 0.0015
Intern-Decision-4B, drop-in pooled rotations 232/259 233/259 8 7 +0.4 -2.3..+3.1 1.0
Intern-Decision-4B, drop-in authored144 single 116/144 128/144 15 3 +8.3 +3.5..+13.9 0.0075
Intern-Decision-4B, drop-in authored144 rotations 131/144 130/144 5 6 -0.7 -4.9..+3.5 1.0
Intern-Decision-4B, drop-in cicada-w1 single 30/31 30/31 0 0 +0.0 +0.0..+0.0 1.0
Intern-Decision-4B, drop-in cicada-w1 rotations 30/31 30/31 0 0 +0.0 +0.0..+0.0 1.0
Intern-Decision-4B, drop-in wyrd single 71/84 74/84 3 0 +3.6 +0.0..+8.3 0.25
Intern-Decision-4B, drop-in wyrd rotations 71/84 73/84 3 1 +2.4 -2.4..+7.1 0.625
Intern-Decision-4B, drop-in perturbations108 single 83/108 102/108 20 1 +17.6 +7.4..+28.7 0.0
Intern-Decision-4B, drop-in perturbations108 rotations 93/108 102/108 9 0 +8.3 +2.8..+15.7 0.0039
Intern-Decision-4B, drop-in cicada-w2 single 20/31 19/31 0 1 -3.2 -9.7..+0.0 1.0
Intern-Decision-4B, drop-in cicada-w2 rotations 19/31 19/31 0 0 +0.0 +0.0..+0.0 1.0
Intern-Decision-4B, drop-in pooled cand-single-vs-semif-rotations 232/259 232/259 10 10 +0.0 -3.0..+3.1 1.0
Intern-Decision-4B, drop-in wyrd cand-single-vs-semif-rotations 71/84 74/84 4 1 +3.6 -1.2..+8.3 0.375
JevK5 v0.2, native pooled single 217/259 229/259 25 13 +4.6 -0.8..+9.9 0.073
JevK5 v0.2, native pooled rotations 232/259 228/259 9 13 -1.5 -4.9..+1.9 0.5235
JevK5 v0.2, native authored144 single 116/144 132/144 23 7 +11.1 +2.8..+18.8 0.0052
JevK5 v0.2, native authored144 rotations 131/144 131/144 7 7 +0.0 -4.9..+4.9 1.0
JevK5 v0.2, native cicada-w1 single 30/31 30/31 0 0 +0.0 +0.0..+0.0 1.0
JevK5 v0.2, native cicada-w1 rotations 30/31 30/31 0 0 +0.0 +0.0..+0.0 1.0
JevK5 v0.2, native wyrd single 71/84 67/84 2 6 -4.8 -10.7..+1.2 0.2891
JevK5 v0.2, native wyrd rotations 71/84 67/84 2 6 -4.8 -10.7..+1.2 0.2891
JevK5 v0.2, native perturbations108 single 83/108 85/108 10 8 +1.9 -7.4..+11.1 0.8145
JevK5 v0.2, native perturbations108 rotations 93/108 91/108 4 6 -1.9 -10.2..+5.6 0.7539
JevK5 v0.2, native cicada-w2 single 20/31 16/31 0 4 -12.9 -25.8..-3.2 0.125
JevK5 v0.2, native cicada-w2 rotations 19/31 16/31 0 3 -9.7 -19.4..+0.0 0.25
JevK5 v0.2, native pooled cand-single-vs-semif-rotations 232/259 229/259 11 14 -1.2 -5.0..+2.6 0.69
JevK5 v0.2, native wyrd cand-single-vs-semif-rotations 71/84 67/84 3 7 -4.8 -13.1..+2.4 0.3438
JevK5 v0.2, drop-in pooled single 217/259 232/259 24 9 +5.8 +1.1..+10.7 0.0135
JevK5 v0.2, drop-in pooled rotations 232/259 237/259 11 6 +1.9 -1.1..+5.1 0.3323
JevK5 v0.2, drop-in authored144 single 116/144 130/144 21 7 +9.7 +2.1..+17.4 0.0125
JevK5 v0.2, drop-in authored144 rotations 131/144 133/144 6 4 +1.4 -2.8..+5.6 0.7539
JevK5 v0.2, drop-in cicada-w1 single 30/31 30/31 0 0 +0.0 +0.0..+0.0 1.0
JevK5 v0.2, drop-in cicada-w1 rotations 30/31 30/31 0 0 +0.0 +0.0..+0.0 1.0
JevK5 v0.2, drop-in wyrd single 71/84 72/84 3 2 +1.2 -3.6..+6.0 1.0
JevK5 v0.2, drop-in wyrd rotations 71/84 74/84 5 2 +3.6 -2.4..+10.7 0.4531
JevK5 v0.2, drop-in perturbations108 single 83/108 92/108 13 4 +8.3 -0.9..+17.6 0.049
JevK5 v0.2, drop-in perturbations108 rotations 93/108 93/108 4 4 +0.0 -7.4..+6.5 1.0
JevK5 v0.2, drop-in cicada-w2 single 20/31 17/31 0 3 -9.7 -22.6..+0.0 0.25
JevK5 v0.2, drop-in cicada-w2 rotations 19/31 16/31 0 3 -9.7 -19.4..+0.0 0.25
JevK5 v0.2, drop-in pooled cand-single-vs-semif-rotations 232/259 232/259 10 10 +0.0 -3.2..+3.2 1.0
JevK5 v0.2, drop-in wyrd cand-single-vs-semif-rotations 71/84 72/84 3 2 +1.2 -3.6..+6.0 1.0
JevK5 v0.3, native pooled single 217/259 227/259 25 15 +3.9 -0.9..+8.5 0.1539
JevK5 v0.3, native pooled rotations 232/259 226/259 9 15 -2.3 -6.0..+1.4 0.3075
JevK5 v0.3, native authored144 single 116/144 135/144 24 5 +13.2 +6.9..+19.4 0.0005
JevK5 v0.3, native authored144 rotations 131/144 135/144 7 3 +2.8 -1.4..+7.6 0.3438
JevK5 v0.3, native cicada-w1 single 30/31 26/31 0 4 -12.9 -25.8..-3.2 0.125
JevK5 v0.3, native cicada-w1 rotations 30/31 26/31 0 4 -12.9 -25.8..-3.2 0.125
JevK5 v0.3, native wyrd single 71/84 66/84 1 6 -6.0 -10.7..-1.2 0.125
JevK5 v0.3, native wyrd rotations 71/84 65/84 2 8 -7.1 -13.1..-1.2 0.1094
JevK5 v0.3, native perturbations108 single 83/108 101/108 18 0 +16.7 +8.3..+25.9 0.0
JevK5 v0.3, native perturbations108 rotations 93/108 101/108 8 0 +7.4 +1.9..+13.9 0.0078
JevK5 v0.3, native cicada-w2 single 20/31 17/31 0 3 -9.7 -19.4..+0.0 0.25
JevK5 v0.3, native cicada-w2 rotations 19/31 16/31 0 3 -9.7 -19.4..+0.0 0.25
JevK5 v0.3, native pooled cand-single-vs-semif-rotations 232/259 227/259 11 16 -1.9 -5.5..+1.5 0.4421
JevK5 v0.3, native wyrd cand-single-vs-semif-rotations 71/84 66/84 2 7 -6.0 -11.9..+0.0 0.1797
JevK5 v0.3, drop-in pooled single 217/259 227/259 24 14 +3.9 -0.8..+8.2 0.1433
JevK5 v0.3, drop-in pooled rotations 232/259 227/259 10 15 -1.9 -5.5..+1.8 0.4244
JevK5 v0.3, drop-in authored144 single 116/144 132/144 22 6 +11.1 +4.9..+17.4 0.0037
JevK5 v0.3, drop-in authored144 rotations 131/144 135/144 8 4 +2.8 -1.4..+7.6 0.3877
JevK5 v0.3, drop-in cicada-w1 single 30/31 26/31 0 4 -12.9 -25.8..-3.2 0.125
JevK5 v0.3, drop-in cicada-w1 rotations 30/31 26/31 0 4 -12.9 -25.8..-3.2 0.125
JevK5 v0.3, drop-in wyrd single 71/84 69/84 2 4 -2.4 -7.1..+2.4 0.6875
JevK5 v0.3, drop-in wyrd rotations 71/84 66/84 2 7 -6.0 -11.9..+0.0 0.1797
JevK5 v0.3, drop-in perturbations108 single 83/108 99/108 17 1 +14.8 +6.5..+24.1 0.0001
JevK5 v0.3, drop-in perturbations108 rotations 93/108 101/108 8 0 +7.4 +1.9..+13.9 0.0078
JevK5 v0.3, drop-in cicada-w2 single 20/31 21/31 1 0 +3.2 +0.0..+9.7 1.0
JevK5 v0.3, drop-in cicada-w2 rotations 19/31 19/31 0 0 +0.0 +0.0..+0.0 1.0
JevK5 v0.3, drop-in pooled cand-single-vs-semif-rotations 232/259 227/259 10 15 -1.9 -5.2..+1.2 0.4244
JevK5 v0.3, drop-in wyrd cand-single-vs-semif-rotations 71/84 69/84 3 5 -2.4 -8.3..+3.6 0.7266

Noise floors

system A-vs-A in process (authored144): flips, max Δp across restarts, single (560 rows/pair): flips per pair, max Δp across restarts, rotations: flips per pair labelled rows whose top moved in ANY restart pair, single: pooled /259 · perturb /108 · Cicada w2 /31
SemIf (Qwen3.5-4B), baseline 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) 0, 0, 0, 0, 0, 0 0 · 0 · 0 (4 repeats)
Plumb-4B, native 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 0 (Δp≤0.000), 2 (Δp≤0.043), 2 (Δp≤0.043) 0, 2, 2 2 · 0 · 0 (3 repeats)
Plumb-4B, drop-in 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) 0, 0, 0 0 · 0 · 0 (3 repeats)
Imajev-4B, native 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) 0, 0, 0, 0, 0, 0 0 · 0 · 0 (4 repeats)
Intern-Decision-4B, native 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) 0, 0, 0, 0, 0, 0 0 · 0 · 0 (4 repeats)
Intern-Decision-4B, drop-in 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 0 (Δp≤0.128), 0 (Δp≤0.128), 0 (Δp≤0.000) 0, 0, 0 0 · 0 · 0 (3 repeats)
JevK5 v0.2, native 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 0 (Δp≤0.000), 3 (Δp≤0.036), 3 (Δp≤0.036) 0, 2, 2 1 · 1 · 0 (3 repeats)
JevK5 v0.2, drop-in 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) 0, 0, 0 0 · 0 · 0 (3 repeats)
JevK5 v0.3, native 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 1 (Δp≤0.051), 1 (Δp≤0.051), 0 (Δp≤0.000) 1, 1, 0 0 · 1 · 0 (3 repeats)
JevK5 v0.3, drop-in 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) 0, 0, 0 0 · 0 · 0 (3 repeats)

Order sensitivity (authored144, single ordering vs the same request reordered) and the negative control

system reversed: label changes /144 reversed: max Δp shuffled: label changes shuffled: max Δp NEG: same top as unrotated /144 NEG: follows the description NEG: right vs original gold
SemIf (Qwen3.5-4B), baseline 30 0.89 29 0.89 14 119 17
Plumb-4B, native 15 0.54 (0.54–0.56) 10 0.44 (0.44–0.45) 24 (24–25) 106 (105–106) 27 (27–28)
Plumb-4B, drop-in 12 0.73 8 0.67 3 126 12
Imajev-4B, native 3 0.72 4 0.68 4 135 5
Intern-Decision-4B, native 8 0.60 5 0.46 10 122 14
Intern-Decision-4B, drop-in 10 0.98 12 (11–12) 0.94 (0.93–0.94) 5 127 10
JevK5 v0.2, native 13 (12–13) 0.72 10 0.64 (0.63–0.64) 20 114 21
JevK5 v0.2, drop-in 15 0.70 12 0.70 4 128 9
JevK5 v0.3, native 8 (7–8) 0.61 11 0.49 20 118 18
JevK5 v0.3, drop-in 4 0.53 5 0.51 4 133 5

Latency at our shape (loopback on fv-ml1, GPU 3; ms; median of all requests, run-median range)

system short1 e2e crit21 e2e crit21 server long1 e2e long16 e2e long16 server tokens crit21 / long16
SemIf (Qwen3.5-4B), baseline 37 (37–38) 131 (131–132) 119 (119–120) 186 (186–187) 505 (502–519) 486 (482–500) 62 / 3900
Plumb-4B, native 16 294 (292–295) 292 (290–293) 215 (212–215) 3378 (3374–3384) 3376 (3372–3382) 2489 / 63302
Plumb-4B, drop-in 37 131 (131–132) 120 (119–120) 186 506 (506–507) 487 (487–488) 62 / 3900
Imajev-4B, native 60 (60–61) 1219 (1212–1230) 1212 (1204–1223) 315 (310–315) 5125 (5054–5140) 5121 (5049–5135) 374 / 7925
Intern-Decision-4B, native 39 88 (88–89) 84 (84–85) 190 (189–191) 215 (213–216) 213 (211–214) 1083 / 4579
Intern-Decision-4B, drop-in 38 (37–38) 132 120 186 (186–188) 508 (500–516) 488 (481–497) 62 / 3900
JevK5 v0.2, native 16 291 (291–292) 289 (289–290) 212 3360 (3358–3368) 3358 (3356–3366) 2489 / 63302
JevK5 v0.2, drop-in 38 132 120 (119–120) 187 (185–188) 518 (505–519) 499 (486–500) 62 / 3900
JevK5 v0.3, native 16 291 (291–292) 289 (289–290) 212 3363 (3355–3370) 3361 (3353–3368) 2489 / 63302
JevK5 v0.3, drop-in 38 (37–39) 133 (131–133) 120 (119–121) 187 (186–188) 518 (506–521) 498 (487–501) 62 / 3900

VRAM (nvidia-smi, whole GPU 3, only our process on it) and capacity under the 12 GiB cap

system rest after warm-up MiB peak in sets+shapes MiB peak in capacity sweep MiB max rows @ short state max rows @ ~3,900 tok
SemIf (Qwen3.5-4B), baseline 9242 (9242–9250) 12918 (12886–12918) 12920 64 (next 96: 422) 16 (next 18: 503)
Plumb-4B, native 12064 12064 12064 128 (no failure up to the last step) 64 (no failure up to the last step)
Plumb-4B, drop-in 9242 12918 12920 64 (next 96: 422) 16 (next 18: 503)
Imajev-4B, native 10404 11056 11056 128 (no failure up to the last step) 64 (no failure up to the last step)
Intern-Decision-4B, native 9736 10290 10290 128 (no failure up to the last step) 64 (no failure up to the last step)
Intern-Decision-4B, drop-in 9242 12918 12920 64 (next 96: 422) 16 (next 18: 503)
JevK5 v0.2, native 12064 12064 12064 128 (no failure up to the last step) 64 (no failure up to the last step)
JevK5 v0.2, drop-in 9242 12918 – – –
JevK5 v0.3, native 12064 12064 – – –
JevK5 v0.3, drop-in 9242 12918 – – –

Reading the tables

  • Fit. For every served system, "peak" is nvidia-smi for the whole card while only our process was on it, so it includes the CUDA context (~0.6 GB) that sits outside torch's 12 GiB fraction. SemIf's own peak of 12.9 GB is the same shape the live service shows. The JevK5 runtime (Plumb, JevK5 native) captures its CUDA graphs at startup and holds 12.06 GB from the first second. It fits the cap, but unlike semif-serve it never hands a burst back to the neighbours on GPU 1.
  • Capacity. The native runtimes chunk: Intern-Decision takes 16 questions per request and Imajev 8 (jev_api.MAX_QUESTIONS). The JevK5 runtime runs questions one by one. "No failure up to the last step" means 64 criteria over the ~3,900-token state and 128 over the short state were answered under the cap, in several requests where the runtime chunks. semif-serve's limit is SemIf's replicated prefix cache: 16 rows at ~3,900 tokens, exactly the README's figure.
  • Order sensitivity and the agreement signal. Every candidate is far less position-biased than SemIf: reversing the options changes 3–15 of 144 labels against SemIf's 30. So rotations buy them little, and in several cases nothing. With rotations, SemIf's rows split between orderings 90 times out of 252 (unanimous rows 94.4% right, split rows 78.9%); Intern-Decision's split only 16 times.
  • Negative control. Every system fails it as it must: on 106–135 of 144 rows the top moves to the option that now carries the right description. The native formats show the option id next to its description (id: description, or A = id: description), which gives a model a second, now contradictory, cue. That is why Plumb and JevK5 native keep their unrotated answer on 20–24 rows against SemIf's 14. Through SemIf's prompt, which shows only descriptions, the same weights keep it on 3–5.
  • Null control. Over a content-free state, Cicada falls to the base rate (15–16 of 31 = always "ordinary") for every system, and Wyrd to 48–60 of 84. Every evidence score above is read against that.
  • Wyrd as one request per turn. Asking the 4 decisions of a turn together changes nothing for SemIf (shared prefix, 71) or for the runtimes that loop over questions. Intern-Decision, which puts all 4 in one prompt, drops from 79 to 77, still above SemIf's 71.

What could not be measured, and why

  • Held-out and sealed JevBench. 109 hard items plus 308 sealed items stay with the evaluator, so our JevBench numbers are public-item numbers. For the board's rows the sealed scores exist (SemIf 0.263; Plumb 0.380, Imajev 0.370, JevK5 v0.2 0.331). They point the same way as our private sets: every candidate is ahead on public items by far more than on unseen ones. Intern-Decision has no sealed number.
  • Contamination of authored144 / perturbations108. Both have been public on GitHub (SemIf) since mid-September. No candidate's repo mentions them as training data (searched at the pinned commits), but that cannot be verified. Cicada and Wyrd were written in this repo on 09-27 and are the only sets no model could have seen. They are also small (31 and 84 rows, one labeller), and their winning wordings were tuned on SemIf, which favours SemIf.
  • Hopper: not run (research-and-demo licence; see Candidates). Imajev drop-in: not applicable (LoRA plus its own readout head). Intern-Decision-2B/0.8B, AlexWortega/openjev: dropped by brokkr's revised order.
  • Contention on GPU 1. Everything ran alone on the empty GPU 3. Next to vllm-coder, the erp/meromero seats and scriberr on GPU 1, latency would be worse for all systems alike. The 09-27 Cicada spike measured SemIf going 136 → 196 ms under a concurrent decode.
  • Latency from a caller. The loopback numbers carry no network. One bridge run from nh3-dev to SemIf on GPU 3 gave 21 criteria in 151.5 ms end to end (runs 150.1–151.8; server 119.3 ms; ping 22.5 ms avg), against the README's 159 ms. The latency instrument agrees with the reference (raw/shape-semif-bridge-from-nh3dev.json).
  • The README's "a bf16 near-tie can flip across restarts". Not reproduced: SemIf flipped 0 of 560 rows across 4 restarts, bit-identical probabilities. Only the CUDA-graph runtime (Plumb, JevK5 native: 1–3 rows per pair, Δp ≤ 0.05) and the Intern-Decision drop-in (0 flips, Δp up to 0.128, one bf16 step) moved at all.

Exact commands

Everything ran from /tmp/jevbench-2026-09-30 on fv-ml1 (removed afterwards); code/ here is the same tree.

# weights, as llmuser, public repos (no token on fv-ml1); pins in code/pull.py, code/pull2.py
docker run -d --name jevbench-pull --user 1001:1001 -e HF_HOME=/hf -e HF_HUB_OFFLINE=0 \
  -v /tank/aimodels/huggingface:/hf -v $W/code:/code:ro --entrypoint /app/.venv/bin/python semif-serve:0.1.4 /code/pull.py
# native image (removed afterwards)
docker build -f code/Dockerfile.native -t jevbench-native:2026-09-30 code/
# the ~3,900-token state
docker run --rm -e HF_HOME=/hf -v /tank/aimodels/huggingface:/hf:ro -v $W:/w --entrypoint /app/.venv/bin/python \
  semif-serve:0.1.4 /w/code/make_long_state.py /w/src/jevbench-v1.2.16/datasets/public/hard.jsonl /w/out/long_state.txt
# one repeat of one system = one line (code/run.sh); the full matrix was code/queue.sh over these lines:
code/run.sh semif-qwen35-4b            semif  Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a            <r> [capacity]
code/run.sh plumb-4b-native            native jevk5:crh225/plumb-4b@55de037801a8a9b9de3db5c0e16cef86210c2186      <r>
code/run.sh plumb-4b-dropin            semif  crh225/plumb-4b@55de037801a8a9b9de3db5c0e16cef86210c2186            <r>
code/run.sh imajev-4b-native           native imajev:mohit67890/imajev-4b@c9e5f132465da85d31735ec502d5557982671a7d <r>
code/run.sh intern-decision-4b-native  native intern:internlm/Intern-Decision-4B@0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd <r>
code/run.sh intern-decision-4b-dropin  semif  internlm/Intern-Decision-4B@0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd <r>
code/run.sh jevk5-v02-native           native jevk5:alibiserikbay/JevK5@27d2d6b8d4714807f6293b0623bd7370b27e42f8  <r>
code/run.sh jevk5-v02-dropin           semif  alibiserikbay/JevK5@27d2d6b8d4714807f6293b0623bd7370b27e42f8        <r>
code/run.sh jevk5-v03-native           native jevk5v03:alibiserikbay/JevK5@c4f7fdb3aeab5582336406e78d3bef11bf98833d <r>
code/run.sh jevk5-v03-dropin           semif  alibiserikbay/JevK5@c4f7fdb3aeab5582336406e78d3bef11bf98833d        <r>
# inside run.sh, per repeat: JevBench (semif_direct in-process, or typesafe to the loopback server) ->
#   bench_sets.py (single, rotations, repeat, reversed, shuffled, negative, multifield) -> bench_shape.py
# analysis (on nh3-dev):
python3 code/analyze.py raw/out raw/jevbench-public-v1.2.16 > summary.json && python3 code/tables.py summary.json > tables.md

Sources used, pinned: JevBench 5e95f23c (tag v1.2.16), jevk5 runtime 85238d7b (v0.2.0) and f944fe37 (v0.3.3), imajev server a0134749, SemIf 23cf1f39 (inside semif-serve:0.1.4). Intern-Decision's inference.py is the one inside its pinned HF snapshot.

Host changes (fv-ml1; all in scripts/ops-log)

when (PDT) change state now
0141 pulled plumb-4b, JevK5 v0.2, Intern-Decision-4B, imajev-4b to /tank/aimodels/huggingface as llmuser; sha256 checked against the Hub kept (7.9 + 8.4 + 8.5 + 0.5 GB)
0209 pulled JevK5 v0.3 (c4f7fdb3) kept (8.4 GB; the JevK5 dir totals 16 GB)
0147 built jevbench-native:2026-09-30 removed 0456
0149–0455 transient containers jevbench-pull, jevbench-jb, jevbench-serve (0.0.0.0:18032, bearer-protected), jevbench-native (127.0.0.1:18090), GPU 3 only, one at a time all removed; GPU 3 at 2 MiB, no compute apps (0456)
0141–0456 scratch /tmp/jevbench-2026-09-30 removed (sudo -n rm -r, because the containers wrote some files as uid 10001)

Nothing touched GPU 1, the semif stack, its config, any alias, LiteLLM, or any seat. Nothing was committed and nothing was sent over althing.