Operator ask relayed by brokkr-smithy-dev. Positive control (SemIf 187/231, hard 0.613) reproduced exactly; negative control and a 4-restart noise floor (0 flips) measured. On our replaced-baseline sets no candidate beats SemIf-with- rotations beyond the ~4-pt floor; Intern-Decision-4B native matches it at one ordering, is better on Wyrd, fits 9.7/10.3 GB and is 1.5-2.3x faster. JevBench rank does not transfer. Raw per-item data kept out of git.
52 KiB
Jev candidates bench: a local replacement for SemIf on the utility GPU (2026-09-30)
Ask (Prime, relayed by brokkr-smithy-dev, 2026-09-30): "Find a local Jev that will fit on our utility GPU at good latency if SemIf is displaced." This is a bench, not a switch. semif-serve and its stack were not touched. The live semif container was stopped by the operator at 0135 for scriberr's sake; this bench did not use it or restart it.
Where: fv-ml1 GPU 3 only, transiently, one model at a time; GPU 3 read 2 MiB before every repeat and after cleanup. Runs 0149–0456 PDT.
Raw data: services/semif-serve/bench-jev-2026-09-30/ (code/ = every script, raw/out/<system>/r<k>/ = every run, summary.json, tables.md).
Verdict
None of the candidates is resolvably better than SemIf when SemIf is run the way its README says to run it (with rotations). If SemIf is displaced, take Intern-Decision-4B, served by its own runtime. It is the only candidate that is at least as accurate as SemIf on our sets, fits the 12 GiB cap with room to spare, and is faster than SemIf at our shape.
- Intern-Decision-4B (native) fits: 9.7 GB at rest and a 10.3 GB peak, against SemIf's 9.2 / 12.9 GB. It is faster at our shape: 21 binary criteria take 88 ms against 131 ms, and 16 criteria over a ~3,900-token state take 215 ms against 505 ms, because it scores every question in one forward pass.
- Against SemIf as deployed (one ordering): better. The pooled 259 rows go from 217 to 240, +8.9 pts (95% CI +4.7..+13.2). On Wyrd, the private consumer set, it scores 79 against 71 of 84 (+9.5, CI +3.6..+16.7).
- Against SemIf with rotations: the same on the pooled set (+1.5, CI −2.0..+5.0). It stays better on Wyrd (77 vs 71, +7.1, CI +1.2..+13.1, McNemar p = 0.07).
- At one ordering it matches SemIf-with-rotations on the pooled set (+3.1, CI −0.4..+6.7) and beats it on Wyrd (+9.5, CI +2.4..+16.7, p = 0.02). So it gets the rotation-level accuracy without paying for rotations. Its answers barely depend on option order: reversing the options changes 8/144 labels, against SemIf's 30/144.
- Caveats.
- Its card has no contamination statement, and it is not on the JevBench board, so no held-out or sealed number exists for it.
- It is not a drop-in. Its request format is the Jev schema, it scores all questions in one prompt, and it takes at most 16 questions per request. Pushed through semif-serve unchanged, it keeps only part of the gain: pooled +5.8 at one ordering and +0.4 with rotations, and on Wyrd +3.6, which is not significant.
- Its card pins transformers 5.14.1; we ran 5.17.0. Its hard tier came out at 83/111 against the vendor's 82.
- The JevBench ranking does not carry over to our workload. Every candidate beats SemIf on JevBench's public items: 198–207/231 against 187. On our sets, the gains all sit on SemIf's own public NLI-style rows (authored144 and perturbations108, both evidence → supported / insufficient / contradicted, the family these models train on). On the private Cicada and Wyrd sets they vanish or reverse:
- Plumb, top of JevBench at 207/231, scores 65 against 71 on Wyrd.
- JevK5 v0.3 scores 26 against 30 on Cicada w1.
- Imajev's +15 pts on authored144 becomes −3.6 on Wyrd.
One line per candidate (against the measured floor: SemIf moved 0 labels in 4 restarts, so every difference below is item sampling; the pooled set resolves about ±4 pts)
| candidate | fits the 12 GiB cap? | vs SemIf, replaced-baseline sets | latency at our shape | licence at the raw file |
|---|---|---|---|---|
| Intern-Decision-4B native | yes, 9.7 / 10.3 GB | better than SemIf-single (pooled +8.9, Wyrd +9.5); same as SemIf-rotations pooled (+1.5), better on Wyrd (+7.1) | faster (88 vs 131 ms; 215 vs 505 ms) | Apache-2.0 (Shanghai AI Lab) + Qwen Apache-2.0. No contamination statement |
| Intern-Decision-4B drop-in | yes, same as SemIf | better than SemIf-single (+5.8); same as SemIf-rotations (+0.4); Wyrd same | same as SemIf | as above |
| Imajev-4B native (no drop-in possible) | yes, 10.4 / 11.1 GB | better than SemIf-single (+6.9) but same as SemIf-rotations (+1.5); the gain is authored/perturbations only; Wyrd same (−3.6) | 9–10× slower (1,219 ms; 5,125 ms); 8 questions per request | Apache-2.0 on GitHub; no LICENSE file in the HF repo |
| Plumb-4B native | yes, but pins 12.06 GB at rest (CUDA graphs) | same (+2.7 / −1.9); worse on Wyrd (−7.1, CI −13.1..−1.2) | 2.2× (294 ms) to 6.7× (3.4 s) slower | Apache-2.0 (raw LICENSE + NOTICE) |
| Plumb-4B drop-in | yes, same as SemIf | same (+1.5 / −3.1); Cicada w1 worse single (26 vs 30) | same as SemIf | as above |
| JevK5 v0.2 native | yes, pins 12.06 GB | same (+4.6 / −1.5); Wyrd same (−4.8) | like Plumb | Apache-2.0 on GitHub; no LICENSE file in the HF repo; trained on 940 MMLU-Pro test items (disclosed) |
| JevK5 v0.2 drop-in | yes, same as SemIf | better than SemIf-single (+5.8, CI +1.1..+10.7); same as SemIf-rotations (+1.9) | same as SemIf | as above |
| JevK5 v0.3 (current HF main; extra) | yes / same | same pooled (native +3.9 / −2.3, drop-in +3.9 / −1.9); worse on Cicada w1 in both formats (26 vs 30); native worse on Wyrd (−6.0 / −7.1) | like v0.2 | as above; its NOTICE says 14k training questions were written by GPT-6 Luna "under OpenAI's terms" |
Key table (median over repeats, min–max where they differ; loopback on fv-ml1 GPU 3)
| system | fits 12 GiB? rest / peak MiB | JevBench all /231 · hard /111 | pooled /259 single · rot | Wyrd /84 single · rot | Δ pooled vs SemIf single (95% CI) | Δ pooled vs SemIf rot (95% CI) | 21 criteria, ms | 16 × ~3,900 tok, ms |
|---|---|---|---|---|---|---|---|---|
| SemIf (Qwen3.5-4B), baseline | yes: 9242 (9242–9250) / 12918 (12886–12918) | 187 · 68 | 217 · 232 | 71 · 71 | baseline | baseline | 131 (131–132) | 505 (502–519) |
| Plumb-4B, native | yes: 12064 / 12064 | 207 (206–207) · 89 (88–89) | 224 (224–226) · 227 (227–228) | 65 · 67 | +2.7 (-2.7..+7.9) | -1.9 (-5.4..+1.5) | 294 (292–295) | 3378 (3374–3384) |
| Plumb-4B, drop-in | yes: 9242 / 12918 | 206 · 88 | 221 · 224 | 66 · 67 | +1.5 (-3.7..+6.6) | -3.1 (-6.8..+0.4) | 131 (131–132) | 506 (506–507) |
| Imajev-4B, native | yes: 10404 / 11056 | 199 · 80 | 235 · 236 | 68 · 69 | +6.9 (+1.6..+12.2) | +1.5 (-2.5..+5.6) | 1219 (1212–1230) | 5125 (5054–5140) |
| Intern-Decision-4B, native | yes: 9736 / 10290 | 202 · 83 | 240 · 236 | 79 · 77 | +8.9 (+4.7..+13.2) | +1.5 (-2.0..+5.0) | 88 (88–89) | 215 (213–216) |
| Intern-Decision-4B, drop-in | yes: 9242 / 12918 | 202 · 85 | 232 · 233 | 74 · 73 | +5.8 (+2.7..+9.0) | +0.4 (-2.3..+3.1) | 132 | 508 (500–516) |
| JevK5 v0.2, native | yes: 12064 / 12064 | 198 (198–199) · 81 (81–82) | 229 (228–229) · 228 | 67 (66–67) · 67 | +4.6 (-0.8..+9.9) | -1.5 (-4.9..+1.9) | 291 (291–292) | 3360 (3358–3368) |
| JevK5 v0.2, drop-in | yes: 9242 / 12918 | 200 · 82 | 232 · 237 | 72 · 74 | +5.8 (+1.1..+10.7) | +1.9 (-1.1..+5.1) | 132 | 518 (505–519) |
| JevK5 v0.3, native | yes: 12064 / 12064 | 202 · 86 | 227 · 226 | 66 · 65 | +3.9 (-0.9..+8.5) | -2.3 (-6.0..+1.4) | 291 (291–292) | 3363 (3355–3370) |
| JevK5 v0.3, drop-in | yes: 9242 / 12918 | 201 · 85 | 227 · 227 | 69 · 66 | +3.9 (-0.8..+8.2) | -1.9 (-5.5..+1.8) | 133 (131–133) | 518 (506–521) |
Pooled = authored144 + Cicada w1 + Wyrd (all 4 decisions) = 259 labelled rows. "single" = one ordering, "rot" = n rotations averaged. Δ is paired, per row, on each row's majority answer over the repeats, with a group-bootstrap 95% CI.
Recommendation (each claim with its own strength)
- "If SemIf is displaced, the replacement is Intern-Decision-4B run on its own runtime": recommend [measured: best on the pooled set and on Wyrd, fastest at our shape, fits the cap; reversible, since it is a service swap and the weights are already on
/tank]. - "Displace SemIf today on accuracy alone": lean against [measured: against SemIf-with-rotations the pooled gain (+1.5 to +3.1) is inside the ~±4-pt floor; reversible]. What would justify a switch is compute: rotation-level accuracy at one ordering, and 1.5–2.3× faster multi-criteria requests. That is a build: a small new service, contract and TDD. It is not a config flip.
- "Pick Plumb or Imajev because of their JevBench rank": strongly recommend against [measured: no gain on our sets; Plumb is worse on Wyrd and Imajev is 9× slower; reversible].
- "Swap weights inside semif-serve (drop-in) for a free gain": lean against [measured: the best drop-in, JevK5 v0.2, is +1.9 against SemIf-rotations, CI −1.1..+5.1, which is the same; reversible].
- Before trusting Intern-Decision in production: label ~50 real Wyrd/Cicada turns as a held-out set. Its card makes no contamination claim, and Wyrd and Cicada are the only sets here that no model could have seen. recommend [reasoned; reversible].
Plain-language decision tree
- Do we only need to free the utility GPU, or do we want a better decider?
- Only free the GPU / keep what works → keep SemIf with rotations. Nothing here beats it by enough to justify a new service. (Every candidate and format is within ±3.1 pts of SemIf-with-rotations on the pooled 259 rows; the floor is ~±4.)
- Want the same answers cheaper, or better answers on the Wyrd-style multi-question turns → Intern-Decision-4B on its own runtime.
- It needs a small new service, because its request format differs (Jev schema; one forward over up to 16 questions;
inference.pyships in the model repo). - First, check it on ~50 real labelled turns (its card makes no contamination claim, and it is not on the JevBench board).
- It needs a small new service, because its request format differs (Jev schema; one forward over up to 16 questions;
- Want a zero-code swap through semif-serve → not worth it. (Best drop-in: JevK5 v0.2, +1.9 pts, CI spans 0; same latency.)
Harness (stated once; every number below was measured under it)
| Card | fv-ml1 GPU 3 (RTX PRO 6000 Blackwell Max-Q, 96 GB, sm_120), empty at start (2 MiB) and nothing foreign on it during the bench (guard log per repeat). Nothing on GPU 1. |
| Stack, every system | torch 2.10.0+cu128, transformers 5.17.0, flash-linear-attention / fla-core 0.5.2, causal-conv1d 1.7.0, BF16, SDPA. This is semif-serve:0.1.4's image (sha256:36a3e1d2…); the native runtimes run in jevbench-native:2026-09-30, the same image plus jevk5 0.2.0 (85238d7b), peft 0.21.1, pillow 12.3.0, torchvision 0.25.0+cu128. |
| VRAM cap | 12 GiB for every served system: semif-serve's own SEMIF_VRAM_CAP_GIB=12; the native servers under capped.py, which applies the same torch.cuda.set_per_process_memory_fraction before any weights load. |
| Clients | fv-ml1 host python 3.13, stdlib urllib, loopback (127.0.0.1), one request at a time. No network in the latency numbers (the README's nh3-dev figures carry ~27–33 ms of network on top). |
| Repeat | one repeat = one fresh process (container) lifetime. N = 3 per cell; 4 for SemIf, Imajev and Intern-Decision native (one repeat added each: Imajev's first repeat ran its latency shapes without its 8-questions-per-request cap and was discarded for latency only, r1/shape-maxq-unset-invalid.json). The Wyrd one-request-per-turn condition was added after the first repeats, so it has N = 3 except Plumb native (N = 2; that runtime loops over questions, so it equals single by construction). |
| Prompts | native = each candidate's own runtime and prompt as its card documents it, reached over its TypeSafe-style /v1/systemone server; a row becomes one choice question with criteria = {option id: description} in the row's order. drop-in = the candidate's weights served by semif-serve 0.1.4 unchanged (SEMIF_MODEL/SEMIF_REVISION), i.e. SemIf's own prompt, letter readout and shared-prefix path: what a switch through the semif-serve contract would get. |
| Orderings | single = the caller's order. rotations = the n cyclic rotations averaged by log-mean (semif-serve's own "orderings": "rotations" for semif-serve; the same arithmetic client-side, n separate requests, for native runtimes). |
| JevBench | fstandhartinger/jevbench tag v1.2.16 (5e95f23cbb7be098a9061fea924c4421620ab1a5), the 231 public items (easy 48, standard/original 72, hard 111), its own jevbench.cli run, one ordering, v1.2 scoring (argmax over the exact label set). SemIf and every drop-in through JevBench's own semif_direct adapter (in-process, SemIf's code path); native candidates through its typesafe adapter to the candidate's loopback server, which is how the board measured them. |
Controls, and what each one showed
Positive control 1 — the JevBench harness (step 1). SemIf on our box, through JevBench v1.2.16's own semif_direct adapter: easy 48/48, standard 71/72, hard 68/111 (0.613), all 187/231 (0.810), identical in all 4 repeats. That is the cards' 0.810 / 0.613 exactly, and the same outcome JevK5's card reports for the untrained base. The harness is right. Each candidate's public number was also reproduced, which makes every candidate row its own positive control: Plumb 89/111 hard (its card: 0.802 = 89), Imajev 80/111 (board hard_public 0.7207 = 80), JevK5 v0.2 81–82/111 native, 82 drop-in (card 0.739 = 82), Intern-Decision 83/111 against a vendor 82/111 (one item over, on transformers 5.17.0 where the card pins 5.14.1).
Positive control 2 — the replaced-baseline harness. The SemIf repeats reproduce semif-serve 0.1.4's acceptance and the 09-27 spikes row for row: parity with SemIf's committed torch predictions 144/144 top choice, 144/144 identical prompt SHA-256, max prob gap 0.060 (acceptance: 144/144, 0.060); negative control 14/144 same-top (acceptance: 14/144); Cicada w1 30/31; Wyrd with rotations place 16/21, place2 21/21, exit 18/21, exit2 16/21 (spike: identical); SemIf's labelled sets 78.6% → 88.1% with rotations in the averaging acceptance, here 199/252 = 79.0% → 224/252 = 88.9%. Capacity under 12 GiB: 16 rows at ~3,900 prefix tokens, 18 is a 503 (README: 16). Latency: 21 binary criteria in 131 ms on loopback; the README's 159 ms was measured from nh3-dev, whose round trip is 22–28 ms (ping averages today); a bridge run from nh3-dev against the GPU 3 instance gave 151.5 ms (below, under what could not be measured).
Negative control (step 6). authored144 with each row's option descriptions rotated one place and the ids kept. A model that reads the descriptions must move its top to the id that now carries the right description. Reported three ways: same top as unrotated (should be LOW), follows the description (should be HIGH), right against the original gold (should be LOW).
Null control. Every Cicada and Wyrd row is also asked over a content-free state ("(No evidence is available for this turn.)"). Whatever that gets right is the prior carried by the question and options alone; the evidence condition is read against it.
Noise floors. A-vs-A inside a process: authored144 scored twice in the same process. Across restarts: every repeat is a fresh container, so each pair of repeats is an A-vs-A across restarts over all 560 labelled+blind rows (single) and 560 (rotations). The paired comparisons use each row's majority top over the repeats, so a row that flips across restarts cannot manufacture a difference by itself.
Sensitivity floor. Item sampling dominates, not run noise (see the floors table: SemIf flips 0 rows across restarts). Paired group-bootstrap 95% CIs on the pooled 259 rows have a half-width of about ±3 to ±5 points, so this bench cannot resolve a pooled difference smaller than ~4 points; on Cicada (31 rows) nothing under ~10–15 points is resolvable, on Wyrd (84 rows, 21 turns) nothing under ~6–7.
Decision rule, stated before the final repeats were in (after the first repeat of each system had been seen): a candidate is better on a set when the paired 95% CI against SemIf excludes zero in its favour AND the gap exceeds the number of rows either side flips across restarts; worse the mirror image; otherwise same. It is judged against SemIf both as deployed today (single ordering, the default) and with rotations (what semif-serve's README says to use "for anything real").
Candidates: identity, pins, licence as read at the raw file
Every HF repo ID was verified with an authenticated API call (/api/models/<id>, HF token from the vault bundle nh3-dev/.config/secrets/env.sh) before any pull: all returned 200, none gated. Weights were pulled to /tank/aimodels/huggingface on fv-ml1 as llmuser and every weight file's sha256 was checked against the Hub's LFS oid (all match; raw/weights-sha256.txt).
| candidate | HF repo @ pinned revision | what it is | runtime used (native) | licence, read at the raw file | contamination statement | JevBench board (v1.4.2.2) |
|---|---|---|---|---|---|---|
| SemIf (baseline) | Qwen/Qwen3.5-4B @ 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a |
frozen base; SemIf 23cf1f39 reads the option letters |
semif-serve 0.1.4 | Qwen LICENSE: Apache-2.0 (Copyright 2026 Alibaba Cloud). SemIf: MIT |
n/a (frozen) | #13, 47.7; hard(220) 0.595, sealed 0.263 |
| Plumb-4B | crh225/plumb-4b @ 55de037801a8a9b9de3db5c0e16cef86210c2186 (the board's pin; HF main 24f7bf77 changes only the card) |
JevK5 v0.2 + LoRA, merged, 8.4 GB | jevk5 runtime 0.2.0 (85238d7b), the board's serving path |
LICENSE: Apache-2.0 (unfilled template). NOTICE: JevK5 v0.2 (Apache-2.0), Qwen (Apache-2.0), SemIf prompt/readout (MIT) |
yes: no JevBench item trained/tuned/selected; 8-gram scan + exact-text audit | #2, 65.8; public 0.896, sealed 0.380 |
| Imajev-4B | mohit67890/imajev-4b @ c9e5f132465da85d31735ec502d5557982671a7d (the board's pin) |
PEFT LoRA (r64) on Qwen3.5-4B + its own 256-code decision readout head; 0.5 GB adapter | imajev server a0134749 (the board's commit), torch backend, --rotations 1 --calibration calibration.json |
HF repo has no LICENSE file (card tag only). GitHub mohit67890/imajev LICENSE: Apache-2.0; card: "Code and adapters: Apache-2.0" |
yes: "No JevBench items (8-gram lint)" | #1, 67.4; public 0.861, sealed 0.370 |
| Intern-Decision-4B | internlm/Intern-Decision-4B @ 0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd |
full Qwen3.5-4B finetune (incl. vision tower), 9.1 GB, adds a <decision> token |
its own inference.py (DecisionEngine, shipped in the repo) behind a 60-line /v1/systemone wrapper |
LICENSE: Apache-2.0 (Copyright 2026-2027 Shanghai AI Laboratory); LICENSE-QWEN: Qwen's Apache-2.0 |
none on the card (flagged) | not on the board (released 09-26); vendor-only claims |
| JevK5 v0.2 | alibiserikbay/JevK5 @ 27d2d6b8d4714807f6293b0623bd7370b27e42f8 (= Hub tag v0.2; tag object ea4804e9; weights identical to the v0.2 release 844e4d0a, LFS 0fba3bba…) |
Qwen3.5-4B + distilled LoRA, merged, 8.4 GB. allebee/jevk5 (GitHub) and alibiserikbay/JevK5 (HF) are the same project: the NOTICE names both handles |
jevk5 runtime 0.2.0 (85238d7b) |
HF repo has no LICENSE file. GitHub allebee/jevk5 LICENSE: Apache-2.0 |
yes (v0.2 card); discloses 940 MMLU-Pro test items in its training replay | #5, 62.0; public 0.853, sealed 0.331 |
| JevK5 v0.3 (extra) | alibiserikbay/JevK5 @ c4f7fdb3aeab5582336406e78d3bef11bf98833d (= Hub tag v0.3 and current main) |
as above, retrained; 8.4 GB | jevk5 runtime 0.3.3 (f944fe37) |
as above; NOTICE: 14,138 training questions written by GPT-6 Luna "under OpenAI's terms" | yes | not on the board |
Not run, with the reason:
- Hopper (
HopitAI/hopper, last in brokkr's order): the adapter's card setslicense: other / research-and-demoand says "Do not use this adapter commercially" (RACE training data, non-commercial terms). There is no LICENSE file. That is the same class asopenjev/openjev(CC-BY-NC), which the brief excluded, so it was excluded by the same rule. It is one queue line if Prime wants it measured anyway. - Imajev drop-in: not applicable. Imajev is a LoRA plus its own readout head, not a full checkpoint whose LM-head letters carry the answer, so semif-serve cannot load it without new code.
Intern-Decision-2B/0.8B,AlexWortega/openjev: dropped per brokkr's revised order.akhilaaa3/Jev-Omni,openjev/openjev, Cygnet: excluded as briefed (size / licence).
Full tables
Generated by code/tables.py from summary.json (code/analyze.py over raw/out/), all under services/semif-serve/bench-jev-2026-09-30/.
JevBench v1.2.16 public items (positive control + candidate score)
| system | repeats | easy /48 | standard /72 | hard /111 | all /231 | all | p50 ms |
|---|---|---|---|---|---|---|---|
| SemIf (Qwen3.5-4B), baseline | 4 | 48 | 71 | 68 | 187 | 0.810 | 36 |
| Plumb-4B, native | 3 | 48 | 70 | 89 (88–89) | 207 (206–207) | 0.896 (0.892–0.896) | 20 |
| Plumb-4B, drop-in | 3 | 48 | 70 | 88 | 206 | 0.892 | 36 |
| Imajev-4B, native | 4 | 48 | 71 | 80 | 199 | 0.861 | 62 (62–63) |
| Intern-Decision-4B, native | 4 | 48 | 71 | 83 | 202 | 0.874 | 40 |
| Intern-Decision-4B, drop-in | 3 | 48 | 69 | 85 | 202 | 0.874 | 36 |
| JevK5 v0.2, native | 3 | 48 | 69 | 81 (81–82) | 198 (198–199) | 0.857 (0.857–0.861) | 20 (20–21) |
| JevK5 v0.2, drop-in | 3 | 48 | 70 | 82 | 200 | 0.866 | 36 (36–37) |
| JevK5 v0.3, native | 3 | 48 | 68 | 86 | 202 | 0.874 | 20 (20–21) |
| JevK5 v0.3, drop-in | 3 | 48 | 68 | 85 | 201 | 0.870 | 36 |
Replaced-baseline sets, single (caller order)
| system | pooled /259 | authored144 | perturb108 | Cicada w1 /31 | Cicada w2 /31 | Wyrd /84 | Wyrd place2 /21 | Wyrd exit /21 |
|---|---|---|---|---|---|---|---|---|
| SemIf (Qwen3.5-4B), baseline | 217 | 116 | 83 | 30 | 20 | 71 | 21 | 16 |
| Plumb-4B, native | 224 (224–226) | 130 (130–132) | 84 | 29 | 19 | 65 | 21 | 10 |
| Plumb-4B, drop-in | 221 | 129 | 88 | 26 | 20 | 66 | 21 | 12 |
| Imajev-4B, native | 235 | 138 | 102 | 29 | 22 | 68 | 20 | 13 |
| Intern-Decision-4B, native | 240 | 132 | 100 | 29 | 22 | 79 | 21 | 19 |
| Intern-Decision-4B, drop-in | 232 | 128 | 102 | 30 | 19 | 74 | 21 | 17 |
| JevK5 v0.2, native | 229 (228–229) | 132 | 85 (85–86) | 30 | 16 | 67 (66–67) | 21 | 12 (11–12) |
| JevK5 v0.2, drop-in | 232 | 130 | 92 | 30 | 17 | 72 | 21 | 17 |
| JevK5 v0.3, native | 227 | 135 | 101 (101–102) | 26 | 17 | 66 | 21 | 12 |
| JevK5 v0.3, drop-in | 227 | 132 | 99 | 26 | 21 | 69 | 21 | 15 |
Replaced-baseline sets, rotations (n rotations averaged)
| system | pooled /259 | authored144 | perturb108 | Cicada w1 /31 | Cicada w2 /31 | Wyrd /84 | Wyrd place2 /21 | Wyrd exit /21 |
|---|---|---|---|---|---|---|---|---|
| SemIf (Qwen3.5-4B), baseline | 232 | 131 | 93 | 30 | 19 | 71 | 21 | 18 |
| Plumb-4B, native | 227 (227–228) | 130 (130–131) | 87 | 30 | 18 (17–18) | 67 | 21 | 11 |
| Plumb-4B, drop-in | 224 | 129 | 92 | 28 | 18 | 67 | 21 | 13 |
| Imajev-4B, native | 236 | 138 | 101 | 29 | 23 | 69 | 20 | 13 |
| Intern-Decision-4B, native | 236 | 130 | 99 | 29 | 22 | 77 | 21 | 19 |
| Intern-Decision-4B, drop-in | 233 | 130 | 102 | 30 | 19 | 73 | 21 | 17 |
| JevK5 v0.2, native | 228 | 131 | 91 (90–91) | 30 | 16 | 67 | 21 | 12 |
| JevK5 v0.2, drop-in | 237 | 133 | 93 | 30 | 16 | 74 | 21 | 20 |
| JevK5 v0.3, native | 226 | 135 | 101 | 26 | 16 | 65 | 20 | 12 |
| JevK5 v0.3, drop-in | 227 | 135 | 101 | 26 | 19 | 66 | 20 | 13 |
Wyrd as one request per turn (4 decisions over one state: SemIf /decide/shared, native multi-question request)
| system | repeats | Wyrd /84 | place /21 | place2 /21 | exit /21 | exit2 /21 |
|---|---|---|---|---|---|---|
| SemIf (Qwen3.5-4B), baseline | 3 | 71 | 17 | 21 | 16 | 17 |
| Plumb-4B, native | 2 | 65 | 17 | 21 | 10 | 17 |
| Plumb-4B, drop-in | 3 | 66 | 17 | 21 | 12 | 16 |
| Imajev-4B, native | 3 | 68 | 19 | 20 | 13 | 16 |
| Intern-Decision-4B, native | 3 | 77 | 16 | 21 | 20 | 20 |
| Intern-Decision-4B, drop-in | 3 | 74 | 18 | 21 | 17 | 18 |
| JevK5 v0.2, native | 3 | 67 (66–67) | 17 | 21 | 12 (11–12) | 17 |
| JevK5 v0.2, drop-in | 3 | 73 | 18 | 21 | 18 | 16 |
| JevK5 v0.3, native | 3 | 66 | 18 | 21 | 12 | 15 |
| JevK5 v0.3, drop-in | 3 | 69 | 18 | 21 | 15 | 15 |
Null control (content-free state) and positive controls inside the spike sets, single ordering
| system | Cicada w1 blind /31 | Wyrd blind /84 | Cicada w1 controls /5 | Wyrd controls /28 |
|---|---|---|---|---|
| SemIf (Qwen3.5-4B), baseline | 16 | 60 | 5 | 21 |
| Plumb-4B, native | 15 | 50 | 5 | 20 |
| Plumb-4B, drop-in | 15 | 60 | 4 | 21 |
| Imajev-4B, native | 16 | 48 | 5 | 24 |
| Intern-Decision-4B, native | 16 | 60 | 5 | 27 |
| Intern-Decision-4B, drop-in | 16 | 60 | 5 | 23 |
| JevK5 v0.2, native | 15 | 53 | 5 | 22 (21–22) |
| JevK5 v0.2, drop-in | 15 | 60 | 5 | 23 |
| JevK5 v0.3, native | 16 | 60 | 5 | 20 |
| JevK5 v0.3, drop-in | 16 | 60 | 5 | 22 |
Paired against SemIf (each row's majority top over the repeats; group bootstrap 95% CI; exact McNemar)
| system | set | cond | SemIf | cand | fixed | broken | Δ pts | 95% CI | p |
|---|---|---|---|---|---|---|---|---|---|
| Plumb-4B, native | pooled | single | 217/259 | 224/259 | 23 | 16 | +2.7 | -2.7..+7.9 | 0.3368 |
| Plumb-4B, native | pooled | rotations | 232/259 | 227/259 | 10 | 15 | -1.9 | -5.4..+1.5 | 0.4244 |
| Plumb-4B, native | authored144 | single | 116/144 | 130/144 | 22 | 8 | +9.7 | +1.4..+18.1 | 0.0161 |
| Plumb-4B, native | authored144 | rotations | 131/144 | 130/144 | 7 | 8 | -0.7 | -5.6..+3.5 | 1.0 |
| Plumb-4B, native | cicada-w1 | single | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
| Plumb-4B, native | cicada-w1 | rotations | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
| Plumb-4B, native | wyrd | single | 71/84 | 65/84 | 1 | 7 | -7.1 | -13.1..-1.2 | 0.0703 |
| Plumb-4B, native | wyrd | rotations | 71/84 | 67/84 | 3 | 7 | -4.8 | -11.9..+2.4 | 0.3438 |
| Plumb-4B, native | perturbations108 | single | 83/108 | 84/108 | 10 | 9 | +0.9 | -9.3..+11.1 | 1.0 |
| Plumb-4B, native | perturbations108 | rotations | 93/108 | 87/108 | 4 | 10 | -5.6 | -15.7..+3.7 | 0.1796 |
| Plumb-4B, native | cicada-w2 | single | 20/31 | 19/31 | 1 | 2 | -3.2 | -12.9..+6.5 | 1.0 |
| Plumb-4B, native | cicada-w2 | rotations | 19/31 | 18/31 | 1 | 2 | -3.2 | -16.1..+6.5 | 1.0 |
| Plumb-4B, native | pooled | cand-single-vs-semif-rotations | 232/259 | 224/259 | 10 | 18 | -3.1 | -7.3..+0.8 | 0.1849 |
| Plumb-4B, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 65/84 | 3 | 9 | -7.1 | -15.5..+1.2 | 0.146 |
| Plumb-4B, drop-in | pooled | single | 217/259 | 221/259 | 21 | 17 | +1.5 | -3.7..+6.6 | 0.6271 |
| Plumb-4B, drop-in | pooled | rotations | 232/259 | 224/259 | 7 | 15 | -3.1 | -6.8..+0.4 | 0.1338 |
| Plumb-4B, drop-in | authored144 | single | 116/144 | 129/144 | 20 | 7 | +9.0 | +1.4..+16.7 | 0.0192 |
| Plumb-4B, drop-in | authored144 | rotations | 131/144 | 129/144 | 5 | 7 | -1.4 | -6.2..+3.5 | 0.7744 |
| Plumb-4B, drop-in | cicada-w1 | single | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
| Plumb-4B, drop-in | cicada-w1 | rotations | 30/31 | 28/31 | 0 | 2 | -6.5 | -16.1..+0.0 | 0.5 |
| Plumb-4B, drop-in | wyrd | single | 71/84 | 66/84 | 1 | 6 | -6.0 | -11.9..+0.0 | 0.125 |
| Plumb-4B, drop-in | wyrd | rotations | 71/84 | 67/84 | 2 | 6 | -4.8 | -11.9..+2.4 | 0.2891 |
| Plumb-4B, drop-in | perturbations108 | single | 83/108 | 88/108 | 12 | 7 | +4.6 | -5.6..+14.8 | 0.3593 |
| Plumb-4B, drop-in | perturbations108 | rotations | 93/108 | 92/108 | 6 | 7 | -0.9 | -10.2..+7.4 | 1.0 |
| Plumb-4B, drop-in | cicada-w2 | single | 20/31 | 20/31 | 1 | 1 | +0.0 | -9.7..+9.7 | 1.0 |
| Plumb-4B, drop-in | cicada-w2 | rotations | 19/31 | 18/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
| Plumb-4B, drop-in | pooled | cand-single-vs-semif-rotations | 232/259 | 221/259 | 8 | 19 | -4.2 | -8.3..-0.4 | 0.0522 |
| Plumb-4B, drop-in | wyrd | cand-single-vs-semif-rotations | 71/84 | 66/84 | 2 | 7 | -6.0 | -13.1..+1.2 | 0.1797 |
| Imajev-4B, native | pooled | single | 217/259 | 235/259 | 30 | 12 | +6.9 | +1.6..+12.2 | 0.0079 |
| Imajev-4B, native | pooled | rotations | 232/259 | 236/259 | 14 | 10 | +1.5 | -2.5..+5.6 | 0.5413 |
| Imajev-4B, native | authored144 | single | 116/144 | 138/144 | 25 | 3 | +15.3 | +9.0..+21.5 | 0.0 |
| Imajev-4B, native | authored144 | rotations | 131/144 | 138/144 | 8 | 1 | +4.9 | +0.7..+9.0 | 0.0391 |
| Imajev-4B, native | cicada-w1 | single | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
| Imajev-4B, native | cicada-w1 | rotations | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
| Imajev-4B, native | wyrd | single | 71/84 | 68/84 | 5 | 8 | -3.6 | -13.1..+6.0 | 0.5811 |
| Imajev-4B, native | wyrd | rotations | 71/84 | 69/84 | 6 | 8 | -2.4 | -11.9..+7.1 | 0.7905 |
| Imajev-4B, native | perturbations108 | single | 83/108 | 102/108 | 20 | 1 | +17.6 | +8.3..+27.8 | 0.0 |
| Imajev-4B, native | perturbations108 | rotations | 93/108 | 101/108 | 9 | 1 | +7.4 | +1.9..+13.9 | 0.0215 |
| Imajev-4B, native | cicada-w2 | single | 20/31 | 22/31 | 3 | 1 | +6.5 | -6.5..+19.4 | 0.625 |
| Imajev-4B, native | cicada-w2 | rotations | 19/31 | 23/31 | 4 | 0 | +12.9 | +3.2..+25.8 | 0.125 |
| Imajev-4B, native | pooled | cand-single-vs-semif-rotations | 232/259 | 235/259 | 15 | 12 | +1.2 | -3.2..+5.5 | 0.7011 |
| Imajev-4B, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 68/84 | 6 | 9 | -3.6 | -14.3..+6.0 | 0.6072 |
| Intern-Decision-4B, native | pooled | single | 217/259 | 240/259 | 28 | 5 | +8.9 | +4.7..+13.2 | 0.0001 |
| Intern-Decision-4B, native | pooled | rotations | 232/259 | 236/259 | 12 | 8 | +1.5 | -2.0..+5.0 | 0.5034 |
| Intern-Decision-4B, native | authored144 | single | 116/144 | 132/144 | 20 | 4 | +11.1 | +4.9..+18.1 | 0.0015 |
| Intern-Decision-4B, native | authored144 | rotations | 131/144 | 130/144 | 5 | 6 | -0.7 | -5.6..+4.2 | 1.0 |
| Intern-Decision-4B, native | cicada-w1 | single | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
| Intern-Decision-4B, native | cicada-w1 | rotations | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
| Intern-Decision-4B, native | wyrd | single | 71/84 | 79/84 | 8 | 0 | +9.5 | +3.6..+16.7 | 0.0078 |
| Intern-Decision-4B, native | wyrd | rotations | 71/84 | 77/84 | 7 | 1 | +7.1 | +1.2..+13.1 | 0.0703 |
| Intern-Decision-4B, native | perturbations108 | single | 83/108 | 100/108 | 20 | 3 | +15.7 | +4.6..+26.9 | 0.0005 |
| Intern-Decision-4B, native | perturbations108 | rotations | 93/108 | 99/108 | 9 | 3 | +5.6 | -2.8..+13.0 | 0.146 |
| Intern-Decision-4B, native | cicada-w2 | single | 20/31 | 22/31 | 2 | 0 | +6.5 | +0.0..+16.1 | 0.5 |
| Intern-Decision-4B, native | cicada-w2 | rotations | 19/31 | 22/31 | 4 | 1 | +9.7 | -3.2..+22.6 | 0.375 |
| Intern-Decision-4B, native | pooled | cand-single-vs-semif-rotations | 232/259 | 240/259 | 14 | 6 | +3.1 | -0.4..+6.7 | 0.1153 |
| Intern-Decision-4B, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 79/84 | 9 | 1 | +9.5 | +2.4..+16.7 | 0.0215 |
| Intern-Decision-4B, drop-in | pooled | single | 217/259 | 232/259 | 18 | 3 | +5.8 | +2.7..+9.0 | 0.0015 |
| Intern-Decision-4B, drop-in | pooled | rotations | 232/259 | 233/259 | 8 | 7 | +0.4 | -2.3..+3.1 | 1.0 |
| Intern-Decision-4B, drop-in | authored144 | single | 116/144 | 128/144 | 15 | 3 | +8.3 | +3.5..+13.9 | 0.0075 |
| Intern-Decision-4B, drop-in | authored144 | rotations | 131/144 | 130/144 | 5 | 6 | -0.7 | -4.9..+3.5 | 1.0 |
| Intern-Decision-4B, drop-in | cicada-w1 | single | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
| Intern-Decision-4B, drop-in | cicada-w1 | rotations | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
| Intern-Decision-4B, drop-in | wyrd | single | 71/84 | 74/84 | 3 | 0 | +3.6 | +0.0..+8.3 | 0.25 |
| Intern-Decision-4B, drop-in | wyrd | rotations | 71/84 | 73/84 | 3 | 1 | +2.4 | -2.4..+7.1 | 0.625 |
| Intern-Decision-4B, drop-in | perturbations108 | single | 83/108 | 102/108 | 20 | 1 | +17.6 | +7.4..+28.7 | 0.0 |
| Intern-Decision-4B, drop-in | perturbations108 | rotations | 93/108 | 102/108 | 9 | 0 | +8.3 | +2.8..+15.7 | 0.0039 |
| Intern-Decision-4B, drop-in | cicada-w2 | single | 20/31 | 19/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
| Intern-Decision-4B, drop-in | cicada-w2 | rotations | 19/31 | 19/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
| Intern-Decision-4B, drop-in | pooled | cand-single-vs-semif-rotations | 232/259 | 232/259 | 10 | 10 | +0.0 | -3.0..+3.1 | 1.0 |
| Intern-Decision-4B, drop-in | wyrd | cand-single-vs-semif-rotations | 71/84 | 74/84 | 4 | 1 | +3.6 | -1.2..+8.3 | 0.375 |
| JevK5 v0.2, native | pooled | single | 217/259 | 229/259 | 25 | 13 | +4.6 | -0.8..+9.9 | 0.073 |
| JevK5 v0.2, native | pooled | rotations | 232/259 | 228/259 | 9 | 13 | -1.5 | -4.9..+1.9 | 0.5235 |
| JevK5 v0.2, native | authored144 | single | 116/144 | 132/144 | 23 | 7 | +11.1 | +2.8..+18.8 | 0.0052 |
| JevK5 v0.2, native | authored144 | rotations | 131/144 | 131/144 | 7 | 7 | +0.0 | -4.9..+4.9 | 1.0 |
| JevK5 v0.2, native | cicada-w1 | single | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
| JevK5 v0.2, native | cicada-w1 | rotations | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
| JevK5 v0.2, native | wyrd | single | 71/84 | 67/84 | 2 | 6 | -4.8 | -10.7..+1.2 | 0.2891 |
| JevK5 v0.2, native | wyrd | rotations | 71/84 | 67/84 | 2 | 6 | -4.8 | -10.7..+1.2 | 0.2891 |
| JevK5 v0.2, native | perturbations108 | single | 83/108 | 85/108 | 10 | 8 | +1.9 | -7.4..+11.1 | 0.8145 |
| JevK5 v0.2, native | perturbations108 | rotations | 93/108 | 91/108 | 4 | 6 | -1.9 | -10.2..+5.6 | 0.7539 |
| JevK5 v0.2, native | cicada-w2 | single | 20/31 | 16/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
| JevK5 v0.2, native | cicada-w2 | rotations | 19/31 | 16/31 | 0 | 3 | -9.7 | -19.4..+0.0 | 0.25 |
| JevK5 v0.2, native | pooled | cand-single-vs-semif-rotations | 232/259 | 229/259 | 11 | 14 | -1.2 | -5.0..+2.6 | 0.69 |
| JevK5 v0.2, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 67/84 | 3 | 7 | -4.8 | -13.1..+2.4 | 0.3438 |
| JevK5 v0.2, drop-in | pooled | single | 217/259 | 232/259 | 24 | 9 | +5.8 | +1.1..+10.7 | 0.0135 |
| JevK5 v0.2, drop-in | pooled | rotations | 232/259 | 237/259 | 11 | 6 | +1.9 | -1.1..+5.1 | 0.3323 |
| JevK5 v0.2, drop-in | authored144 | single | 116/144 | 130/144 | 21 | 7 | +9.7 | +2.1..+17.4 | 0.0125 |
| JevK5 v0.2, drop-in | authored144 | rotations | 131/144 | 133/144 | 6 | 4 | +1.4 | -2.8..+5.6 | 0.7539 |
| JevK5 v0.2, drop-in | cicada-w1 | single | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
| JevK5 v0.2, drop-in | cicada-w1 | rotations | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
| JevK5 v0.2, drop-in | wyrd | single | 71/84 | 72/84 | 3 | 2 | +1.2 | -3.6..+6.0 | 1.0 |
| JevK5 v0.2, drop-in | wyrd | rotations | 71/84 | 74/84 | 5 | 2 | +3.6 | -2.4..+10.7 | 0.4531 |
| JevK5 v0.2, drop-in | perturbations108 | single | 83/108 | 92/108 | 13 | 4 | +8.3 | -0.9..+17.6 | 0.049 |
| JevK5 v0.2, drop-in | perturbations108 | rotations | 93/108 | 93/108 | 4 | 4 | +0.0 | -7.4..+6.5 | 1.0 |
| JevK5 v0.2, drop-in | cicada-w2 | single | 20/31 | 17/31 | 0 | 3 | -9.7 | -22.6..+0.0 | 0.25 |
| JevK5 v0.2, drop-in | cicada-w2 | rotations | 19/31 | 16/31 | 0 | 3 | -9.7 | -19.4..+0.0 | 0.25 |
| JevK5 v0.2, drop-in | pooled | cand-single-vs-semif-rotations | 232/259 | 232/259 | 10 | 10 | +0.0 | -3.2..+3.2 | 1.0 |
| JevK5 v0.2, drop-in | wyrd | cand-single-vs-semif-rotations | 71/84 | 72/84 | 3 | 2 | +1.2 | -3.6..+6.0 | 1.0 |
| JevK5 v0.3, native | pooled | single | 217/259 | 227/259 | 25 | 15 | +3.9 | -0.9..+8.5 | 0.1539 |
| JevK5 v0.3, native | pooled | rotations | 232/259 | 226/259 | 9 | 15 | -2.3 | -6.0..+1.4 | 0.3075 |
| JevK5 v0.3, native | authored144 | single | 116/144 | 135/144 | 24 | 5 | +13.2 | +6.9..+19.4 | 0.0005 |
| JevK5 v0.3, native | authored144 | rotations | 131/144 | 135/144 | 7 | 3 | +2.8 | -1.4..+7.6 | 0.3438 |
| JevK5 v0.3, native | cicada-w1 | single | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
| JevK5 v0.3, native | cicada-w1 | rotations | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
| JevK5 v0.3, native | wyrd | single | 71/84 | 66/84 | 1 | 6 | -6.0 | -10.7..-1.2 | 0.125 |
| JevK5 v0.3, native | wyrd | rotations | 71/84 | 65/84 | 2 | 8 | -7.1 | -13.1..-1.2 | 0.1094 |
| JevK5 v0.3, native | perturbations108 | single | 83/108 | 101/108 | 18 | 0 | +16.7 | +8.3..+25.9 | 0.0 |
| JevK5 v0.3, native | perturbations108 | rotations | 93/108 | 101/108 | 8 | 0 | +7.4 | +1.9..+13.9 | 0.0078 |
| JevK5 v0.3, native | cicada-w2 | single | 20/31 | 17/31 | 0 | 3 | -9.7 | -19.4..+0.0 | 0.25 |
| JevK5 v0.3, native | cicada-w2 | rotations | 19/31 | 16/31 | 0 | 3 | -9.7 | -19.4..+0.0 | 0.25 |
| JevK5 v0.3, native | pooled | cand-single-vs-semif-rotations | 232/259 | 227/259 | 11 | 16 | -1.9 | -5.5..+1.5 | 0.4421 |
| JevK5 v0.3, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 66/84 | 2 | 7 | -6.0 | -11.9..+0.0 | 0.1797 |
| JevK5 v0.3, drop-in | pooled | single | 217/259 | 227/259 | 24 | 14 | +3.9 | -0.8..+8.2 | 0.1433 |
| JevK5 v0.3, drop-in | pooled | rotations | 232/259 | 227/259 | 10 | 15 | -1.9 | -5.5..+1.8 | 0.4244 |
| JevK5 v0.3, drop-in | authored144 | single | 116/144 | 132/144 | 22 | 6 | +11.1 | +4.9..+17.4 | 0.0037 |
| JevK5 v0.3, drop-in | authored144 | rotations | 131/144 | 135/144 | 8 | 4 | +2.8 | -1.4..+7.6 | 0.3877 |
| JevK5 v0.3, drop-in | cicada-w1 | single | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
| JevK5 v0.3, drop-in | cicada-w1 | rotations | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
| JevK5 v0.3, drop-in | wyrd | single | 71/84 | 69/84 | 2 | 4 | -2.4 | -7.1..+2.4 | 0.6875 |
| JevK5 v0.3, drop-in | wyrd | rotations | 71/84 | 66/84 | 2 | 7 | -6.0 | -11.9..+0.0 | 0.1797 |
| JevK5 v0.3, drop-in | perturbations108 | single | 83/108 | 99/108 | 17 | 1 | +14.8 | +6.5..+24.1 | 0.0001 |
| JevK5 v0.3, drop-in | perturbations108 | rotations | 93/108 | 101/108 | 8 | 0 | +7.4 | +1.9..+13.9 | 0.0078 |
| JevK5 v0.3, drop-in | cicada-w2 | single | 20/31 | 21/31 | 1 | 0 | +3.2 | +0.0..+9.7 | 1.0 |
| JevK5 v0.3, drop-in | cicada-w2 | rotations | 19/31 | 19/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
| JevK5 v0.3, drop-in | pooled | cand-single-vs-semif-rotations | 232/259 | 227/259 | 10 | 15 | -1.9 | -5.2..+1.2 | 0.4244 |
| JevK5 v0.3, drop-in | wyrd | cand-single-vs-semif-rotations | 71/84 | 69/84 | 3 | 5 | -2.4 | -8.3..+3.6 | 0.7266 |
Noise floors
| system | A-vs-A in process (authored144): flips, max Δp | across restarts, single (560 rows/pair): flips per pair, max Δp | across restarts, rotations: flips per pair | labelled rows whose top moved in ANY restart pair, single: pooled /259 · perturb /108 · Cicada w2 /31 |
|---|---|---|---|---|
| SemIf (Qwen3.5-4B), baseline | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0, 0, 0, 0 | 0 · 0 · 0 (4 repeats) |
| Plumb-4B, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 2 (Δp≤0.043), 2 (Δp≤0.043) | 0, 2, 2 | 2 · 0 · 0 (3 repeats) |
| Plumb-4B, drop-in | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0 | 0 · 0 · 0 (3 repeats) |
| Imajev-4B, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0, 0, 0, 0 | 0 · 0 · 0 (4 repeats) |
| Intern-Decision-4B, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0, 0, 0, 0 | 0 · 0 · 0 (4 repeats) |
| Intern-Decision-4B, drop-in | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.128), 0 (Δp≤0.128), 0 (Δp≤0.000) | 0, 0, 0 | 0 · 0 · 0 (3 repeats) |
| JevK5 v0.2, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 3 (Δp≤0.036), 3 (Δp≤0.036) | 0, 2, 2 | 1 · 1 · 0 (3 repeats) |
| JevK5 v0.2, drop-in | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0 | 0 · 0 · 0 (3 repeats) |
| JevK5 v0.3, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 1 (Δp≤0.051), 1 (Δp≤0.051), 0 (Δp≤0.000) | 1, 1, 0 | 0 · 1 · 0 (3 repeats) |
| JevK5 v0.3, drop-in | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0 | 0 · 0 · 0 (3 repeats) |
Order sensitivity (authored144, single ordering vs the same request reordered) and the negative control
| system | reversed: label changes /144 | reversed: max Δp | shuffled: label changes | shuffled: max Δp | NEG: same top as unrotated /144 | NEG: follows the description | NEG: right vs original gold |
|---|---|---|---|---|---|---|---|
| SemIf (Qwen3.5-4B), baseline | 30 | 0.89 | 29 | 0.89 | 14 | 119 | 17 |
| Plumb-4B, native | 15 | 0.54 (0.54–0.56) | 10 | 0.44 (0.44–0.45) | 24 (24–25) | 106 (105–106) | 27 (27–28) |
| Plumb-4B, drop-in | 12 | 0.73 | 8 | 0.67 | 3 | 126 | 12 |
| Imajev-4B, native | 3 | 0.72 | 4 | 0.68 | 4 | 135 | 5 |
| Intern-Decision-4B, native | 8 | 0.60 | 5 | 0.46 | 10 | 122 | 14 |
| Intern-Decision-4B, drop-in | 10 | 0.98 | 12 (11–12) | 0.94 (0.93–0.94) | 5 | 127 | 10 |
| JevK5 v0.2, native | 13 (12–13) | 0.72 | 10 | 0.64 (0.63–0.64) | 20 | 114 | 21 |
| JevK5 v0.2, drop-in | 15 | 0.70 | 12 | 0.70 | 4 | 128 | 9 |
| JevK5 v0.3, native | 8 (7–8) | 0.61 | 11 | 0.49 | 20 | 118 | 18 |
| JevK5 v0.3, drop-in | 4 | 0.53 | 5 | 0.51 | 4 | 133 | 5 |
Latency at our shape (loopback on fv-ml1, GPU 3; ms; median of all requests, run-median range)
| system | short1 e2e | crit21 e2e | crit21 server | long1 e2e | long16 e2e | long16 server | tokens crit21 / long16 |
|---|---|---|---|---|---|---|---|
| SemIf (Qwen3.5-4B), baseline | 37 (37–38) | 131 (131–132) | 119 (119–120) | 186 (186–187) | 505 (502–519) | 486 (482–500) | 62 / 3900 |
| Plumb-4B, native | 16 | 294 (292–295) | 292 (290–293) | 215 (212–215) | 3378 (3374–3384) | 3376 (3372–3382) | 2489 / 63302 |
| Plumb-4B, drop-in | 37 | 131 (131–132) | 120 (119–120) | 186 | 506 (506–507) | 487 (487–488) | 62 / 3900 |
| Imajev-4B, native | 60 (60–61) | 1219 (1212–1230) | 1212 (1204–1223) | 315 (310–315) | 5125 (5054–5140) | 5121 (5049–5135) | 374 / 7925 |
| Intern-Decision-4B, native | 39 | 88 (88–89) | 84 (84–85) | 190 (189–191) | 215 (213–216) | 213 (211–214) | 1083 / 4579 |
| Intern-Decision-4B, drop-in | 38 (37–38) | 132 | 120 | 186 (186–188) | 508 (500–516) | 488 (481–497) | 62 / 3900 |
| JevK5 v0.2, native | 16 | 291 (291–292) | 289 (289–290) | 212 | 3360 (3358–3368) | 3358 (3356–3366) | 2489 / 63302 |
| JevK5 v0.2, drop-in | 38 | 132 | 120 (119–120) | 187 (185–188) | 518 (505–519) | 499 (486–500) | 62 / 3900 |
| JevK5 v0.3, native | 16 | 291 (291–292) | 289 (289–290) | 212 | 3363 (3355–3370) | 3361 (3353–3368) | 2489 / 63302 |
| JevK5 v0.3, drop-in | 38 (37–39) | 133 (131–133) | 120 (119–121) | 187 (186–188) | 518 (506–521) | 498 (487–501) | 62 / 3900 |
VRAM (nvidia-smi, whole GPU 3, only our process on it) and capacity under the 12 GiB cap
| system | rest after warm-up MiB | peak in sets+shapes MiB | peak in capacity sweep MiB | max rows @ short state | max rows @ ~3,900 tok |
|---|---|---|---|---|---|
| SemIf (Qwen3.5-4B), baseline | 9242 (9242–9250) | 12918 (12886–12918) | 12920 | 64 (next 96: 422) | 16 (next 18: 503) |
| Plumb-4B, native | 12064 | 12064 | 12064 | 128 (no failure up to the last step) | 64 (no failure up to the last step) |
| Plumb-4B, drop-in | 9242 | 12918 | 12920 | 64 (next 96: 422) | 16 (next 18: 503) |
| Imajev-4B, native | 10404 | 11056 | 11056 | 128 (no failure up to the last step) | 64 (no failure up to the last step) |
| Intern-Decision-4B, native | 9736 | 10290 | 10290 | 128 (no failure up to the last step) | 64 (no failure up to the last step) |
| Intern-Decision-4B, drop-in | 9242 | 12918 | 12920 | 64 (next 96: 422) | 16 (next 18: 503) |
| JevK5 v0.2, native | 12064 | 12064 | 12064 | 128 (no failure up to the last step) | 64 (no failure up to the last step) |
| JevK5 v0.2, drop-in | 9242 | 12918 | – | – | – |
| JevK5 v0.3, native | 12064 | 12064 | – | – | – |
| JevK5 v0.3, drop-in | 9242 | 12918 | – | – | – |
Reading the tables
- Fit. For every served system, "peak" is nvidia-smi for the whole card while only our process was on it, so it includes the CUDA context (~0.6 GB) that sits outside torch's 12 GiB fraction. SemIf's own peak of 12.9 GB is the same shape the live service shows. The JevK5 runtime (Plumb, JevK5 native) captures its CUDA graphs at startup and holds 12.06 GB from the first second. It fits the cap, but unlike semif-serve it never hands a burst back to the neighbours on GPU 1.
- Capacity. The native runtimes chunk: Intern-Decision takes 16 questions per request and Imajev 8 (
jev_api.MAX_QUESTIONS). The JevK5 runtime runs questions one by one. "No failure up to the last step" means 64 criteria over the ~3,900-token state and 128 over the short state were answered under the cap, in several requests where the runtime chunks. semif-serve's limit is SemIf's replicated prefix cache: 16 rows at ~3,900 tokens, exactly the README's figure. - Order sensitivity and the agreement signal. Every candidate is far less position-biased than SemIf: reversing the options changes 3–15 of 144 labels against SemIf's 30. So rotations buy them little, and in several cases nothing. With rotations, SemIf's rows split between orderings 90 times out of 252 (unanimous rows 94.4% right, split rows 78.9%); Intern-Decision's split only 16 times.
- Negative control. Every system fails it as it must: on 106–135 of 144 rows the top moves to the option that now carries the right description. The native formats show the option id next to its description (
id: description, orA = id: description), which gives a model a second, now contradictory, cue. That is why Plumb and JevK5 native keep their unrotated answer on 20–24 rows against SemIf's 14. Through SemIf's prompt, which shows only descriptions, the same weights keep it on 3–5. - Null control. Over a content-free state, Cicada falls to the base rate (15–16 of 31 = always "ordinary") for every system, and Wyrd to 48–60 of 84. Every evidence score above is read against that.
- Wyrd as one request per turn. Asking the 4 decisions of a turn together changes nothing for SemIf (shared prefix, 71) or for the runtimes that loop over questions. Intern-Decision, which puts all 4 in one prompt, drops from 79 to 77, still above SemIf's 71.
What could not be measured, and why
- Held-out and sealed JevBench. 109 hard items plus 308 sealed items stay with the evaluator, so our JevBench numbers are public-item numbers. For the board's rows the sealed scores exist (SemIf 0.263; Plumb 0.380, Imajev 0.370, JevK5 v0.2 0.331). They point the same way as our private sets: every candidate is ahead on public items by far more than on unseen ones. Intern-Decision has no sealed number.
- Contamination of authored144 / perturbations108. Both have been public on GitHub (SemIf) since mid-September. No candidate's repo mentions them as training data (searched at the pinned commits), but that cannot be verified. Cicada and Wyrd were written in this repo on 09-27 and are the only sets no model could have seen. They are also small (31 and 84 rows, one labeller), and their winning wordings were tuned on SemIf, which favours SemIf.
- Hopper: not run (research-and-demo licence; see Candidates). Imajev drop-in: not applicable (LoRA plus its own readout head). Intern-Decision-2B/0.8B, AlexWortega/openjev: dropped by brokkr's revised order.
- Contention on GPU 1. Everything ran alone on the empty GPU 3. Next to vllm-coder, the erp/meromero seats and scriberr on GPU 1, latency would be worse for all systems alike. The 09-27 Cicada spike measured SemIf going 136 → 196 ms under a concurrent decode.
- Latency from a caller. The loopback numbers carry no network. One bridge run from nh3-dev to SemIf on GPU 3 gave 21 criteria in 151.5 ms end to end (runs 150.1–151.8; server 119.3 ms; ping 22.5 ms avg), against the README's 159 ms. The latency instrument agrees with the reference (
raw/shape-semif-bridge-from-nh3dev.json). - The README's "a bf16 near-tie can flip across restarts". Not reproduced: SemIf flipped 0 of 560 rows across 4 restarts, bit-identical probabilities. Only the CUDA-graph runtime (Plumb, JevK5 native: 1–3 rows per pair, Δp ≤ 0.05) and the Intern-Decision drop-in (0 flips, Δp up to 0.128, one bf16 step) moved at all.
Exact commands
Everything ran from /tmp/jevbench-2026-09-30 on fv-ml1 (removed afterwards); code/ here is the same tree.
# weights, as llmuser, public repos (no token on fv-ml1); pins in code/pull.py, code/pull2.py
docker run -d --name jevbench-pull --user 1001:1001 -e HF_HOME=/hf -e HF_HUB_OFFLINE=0 \
-v /tank/aimodels/huggingface:/hf -v $W/code:/code:ro --entrypoint /app/.venv/bin/python semif-serve:0.1.4 /code/pull.py
# native image (removed afterwards)
docker build -f code/Dockerfile.native -t jevbench-native:2026-09-30 code/
# the ~3,900-token state
docker run --rm -e HF_HOME=/hf -v /tank/aimodels/huggingface:/hf:ro -v $W:/w --entrypoint /app/.venv/bin/python \
semif-serve:0.1.4 /w/code/make_long_state.py /w/src/jevbench-v1.2.16/datasets/public/hard.jsonl /w/out/long_state.txt
# one repeat of one system = one line (code/run.sh); the full matrix was code/queue.sh over these lines:
code/run.sh semif-qwen35-4b semif Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a <r> [capacity]
code/run.sh plumb-4b-native native jevk5:crh225/plumb-4b@55de037801a8a9b9de3db5c0e16cef86210c2186 <r>
code/run.sh plumb-4b-dropin semif crh225/plumb-4b@55de037801a8a9b9de3db5c0e16cef86210c2186 <r>
code/run.sh imajev-4b-native native imajev:mohit67890/imajev-4b@c9e5f132465da85d31735ec502d5557982671a7d <r>
code/run.sh intern-decision-4b-native native intern:internlm/Intern-Decision-4B@0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd <r>
code/run.sh intern-decision-4b-dropin semif internlm/Intern-Decision-4B@0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd <r>
code/run.sh jevk5-v02-native native jevk5:alibiserikbay/JevK5@27d2d6b8d4714807f6293b0623bd7370b27e42f8 <r>
code/run.sh jevk5-v02-dropin semif alibiserikbay/JevK5@27d2d6b8d4714807f6293b0623bd7370b27e42f8 <r>
code/run.sh jevk5-v03-native native jevk5v03:alibiserikbay/JevK5@c4f7fdb3aeab5582336406e78d3bef11bf98833d <r>
code/run.sh jevk5-v03-dropin semif alibiserikbay/JevK5@c4f7fdb3aeab5582336406e78d3bef11bf98833d <r>
# inside run.sh, per repeat: JevBench (semif_direct in-process, or typesafe to the loopback server) ->
# bench_sets.py (single, rotations, repeat, reversed, shuffled, negative, multifield) -> bench_shape.py
# analysis (on nh3-dev):
python3 code/analyze.py raw/out raw/jevbench-public-v1.2.16 > summary.json && python3 code/tables.py summary.json > tables.md
Sources used, pinned: JevBench 5e95f23c (tag v1.2.16), jevk5 runtime 85238d7b (v0.2.0) and f944fe37 (v0.3.3), imajev server a0134749, SemIf 23cf1f39 (inside semif-serve:0.1.4). Intern-Decision's inference.py is the one inside its pinned HF snapshot.
Host changes (fv-ml1; all in scripts/ops-log)
| when (PDT) | change | state now |
|---|---|---|
| 0141 | pulled plumb-4b, JevK5 v0.2, Intern-Decision-4B, imajev-4b to /tank/aimodels/huggingface as llmuser; sha256 checked against the Hub |
kept (7.9 + 8.4 + 8.5 + 0.5 GB) |
| 0209 | pulled JevK5 v0.3 (c4f7fdb3) |
kept (8.4 GB; the JevK5 dir totals 16 GB) |
| 0147 | built jevbench-native:2026-09-30 |
removed 0456 |
| 0149–0455 | transient containers jevbench-pull, jevbench-jb, jevbench-serve (0.0.0.0:18032, bearer-protected), jevbench-native (127.0.0.1:18090), GPU 3 only, one at a time |
all removed; GPU 3 at 2 MiB, no compute apps (0456) |
| 0141–0456 | scratch /tmp/jevbench-2026-09-30 |
removed (sudo -n rm -r, because the containers wrote some files as uid 10001) |
Nothing touched GPU 1, the semif stack, its config, any alias, LiteLLM, or any seat. Nothing was committed and nothing was sent over althing.