docs: Jev candidate bench vs SemIf (fv-ml1 GPU 3) — Intern-Decision-4B is the replacement if SemIf is displaced
Operator ask relayed by brokkr-smithy-dev. Positive control (SemIf 187/231, hard 0.613) reproduced exactly; negative control and a 4-restart noise floor (0 flips) measured. On our replaced-baseline sets no candidate beats SemIf-with- rotations beyond the ~4-pt floor; Intern-Decision-4B native matches it at one ordering, is better on Wyrd, fits 9.7/10.3 GB and is 1.5-2.3x faster. JevBench rank does not transfer. Raw per-item data kept out of git.
This commit is contained in:
@@ -47,3 +47,6 @@ stacks/lobe-chat/.env
|
||||
# and every commit here is attributed to Vuong Hoang anyway, so git could
|
||||
# not carry the attribution this file exists to provide.
|
||||
.ops-log/
|
||||
|
||||
# Jev bench raw per-item responses (83 MB): kept on disk (restic via /home) and in the Booth, not in git
|
||||
services/semif-serve/bench-jev-2026-09-30/raw/
|
||||
|
||||
@@ -0,0 +1,448 @@
|
||||
# Jev candidates bench: a local replacement for SemIf on the utility GPU (2026-09-30)
|
||||
|
||||
**Ask (Prime, relayed by brokkr-smithy-dev, 2026-09-30):** "Find a local Jev that will fit on our utility GPU at good latency if SemIf is displaced." This is a bench, not a switch. semif-serve and its stack were not touched. The live `semif` container was stopped by the operator at 0135 for scriberr's sake; this bench did not use it or restart it.
|
||||
**Where:** fv-ml1 **GPU 3** only, transiently, one model at a time; GPU 3 read 2 MiB before every repeat and after cleanup. Runs 0149–0456 PDT.
|
||||
**Raw data:** `services/semif-serve/bench-jev-2026-09-30/` (`code/` = every script, `raw/out/<system>/r<k>/` = every run, `summary.json`, `tables.md`).
|
||||
|
||||
## Verdict
|
||||
|
||||
**None of the candidates is resolvably better than SemIf when SemIf is run the way its README says to run it (with rotations). If SemIf is displaced, take Intern-Decision-4B, served by its own runtime.** It is the only candidate that is at least as accurate as SemIf on our sets, fits the 12 GiB cap with room to spare, and is faster than SemIf at our shape.
|
||||
|
||||
- **Intern-Decision-4B (native)** fits: 9.7 GB at rest and a 10.3 GB peak, against SemIf's 9.2 / 12.9 GB. It is faster at our shape: 21 binary criteria take 88 ms against 131 ms, and 16 criteria over a ~3,900-token state take 215 ms against 505 ms, because it scores every question in one forward pass.
|
||||
- **Against SemIf as deployed (one ordering):** better. The pooled 259 rows go from 217 to 240, **+8.9 pts (95% CI +4.7..+13.2)**. On Wyrd, the private consumer set, it scores 79 against 71 of 84 (**+9.5, CI +3.6..+16.7**).
|
||||
- **Against SemIf with rotations:** the same on the pooled set (+1.5, CI −2.0..+5.0). It stays better on Wyrd (77 vs 71, +7.1, CI +1.2..+13.1, McNemar p = 0.07).
|
||||
- **At one ordering it matches SemIf-with-rotations** on the pooled set (+3.1, CI −0.4..+6.7) and beats it on Wyrd (+9.5, CI +2.4..+16.7, p = 0.02). So it gets the rotation-level accuracy without paying for rotations. Its answers barely depend on option order: reversing the options changes 8/144 labels, against SemIf's 30/144.
|
||||
- **Caveats.**
|
||||
- Its card has **no contamination statement**, and it is not on the JevBench board, so no held-out or sealed number exists for it.
|
||||
- It is **not a drop-in.** Its request format is the Jev schema, it scores all questions in one prompt, and it takes at most 16 questions per request. Pushed through semif-serve unchanged, it keeps only part of the gain: pooled +5.8 at one ordering and +0.4 with rotations, and on Wyrd +3.6, which is not significant.
|
||||
- Its card pins transformers 5.14.1; we ran 5.17.0. Its hard tier came out at 83/111 against the vendor's 82.
|
||||
- **The JevBench ranking does not carry over to our workload.** Every candidate beats SemIf on JevBench's public items: 198–207/231 against 187. On our sets, the gains all sit on SemIf's own public NLI-style rows (authored144 and perturbations108, both evidence → supported / insufficient / contradicted, the family these models train on). On the private Cicada and Wyrd sets they vanish or reverse:
|
||||
- Plumb, top of JevBench at 207/231, scores 65 against 71 on Wyrd.
|
||||
- JevK5 v0.3 scores 26 against 30 on Cicada w1.
|
||||
- Imajev's +15 pts on authored144 becomes −3.6 on Wyrd.
|
||||
|
||||
### One line per candidate (against the measured floor: SemIf moved **0 labels in 4 restarts**, so every difference below is item sampling; the pooled set resolves about ±4 pts)
|
||||
|
||||
| candidate | fits the 12 GiB cap? | vs SemIf, replaced-baseline sets | latency at our shape | licence at the raw file |
|
||||
|---|---|---|---|---|
|
||||
| **Intern-Decision-4B** native | **yes**, 9.7 / 10.3 GB | **better** than SemIf-single (pooled +8.9, Wyrd +9.5); **same** as SemIf-rotations pooled (+1.5), **better** on Wyrd (+7.1) | **faster** (88 vs 131 ms; 215 vs 505 ms) | Apache-2.0 (Shanghai AI Lab) + Qwen Apache-2.0. **No contamination statement** |
|
||||
| Intern-Decision-4B drop-in | yes, same as SemIf | better than SemIf-single (+5.8); same as SemIf-rotations (+0.4); Wyrd same | same as SemIf | as above |
|
||||
| Imajev-4B native (no drop-in possible) | yes, 10.4 / 11.1 GB | better than SemIf-single (+6.9) but **same** as SemIf-rotations (+1.5); the gain is authored/perturbations only; Wyrd same (−3.6) | **9–10× slower** (1,219 ms; 5,125 ms); 8 questions per request | Apache-2.0 on GitHub; **no LICENSE file in the HF repo** |
|
||||
| Plumb-4B native | yes, but pins 12.06 GB at rest (CUDA graphs) | **same** (+2.7 / −1.9); **worse on Wyrd** (−7.1, CI −13.1..−1.2) | 2.2× (294 ms) to 6.7× (3.4 s) slower | Apache-2.0 (raw LICENSE + NOTICE) |
|
||||
| Plumb-4B drop-in | yes, same as SemIf | same (+1.5 / −3.1); Cicada w1 worse single (26 vs 30) | same as SemIf | as above |
|
||||
| JevK5 v0.2 native | yes, pins 12.06 GB | same (+4.6 / −1.5); Wyrd same (−4.8) | like Plumb | Apache-2.0 on GitHub; **no LICENSE file in the HF repo**; trained on 940 MMLU-Pro *test* items (disclosed) |
|
||||
| JevK5 v0.2 drop-in | yes, same as SemIf | better than SemIf-single (+5.8, CI +1.1..+10.7); same as SemIf-rotations (+1.9) | same as SemIf | as above |
|
||||
| JevK5 v0.3 (current HF main; extra) | yes / same | same pooled (native +3.9 / −2.3, drop-in +3.9 / −1.9); **worse on Cicada w1 in both formats (26 vs 30)**; native worse on Wyrd (−6.0 / −7.1) | like v0.2 | as above; its NOTICE says 14k training questions were written by GPT-6 Luna "under OpenAI's terms" |
|
||||
|
||||
### Key table (median over repeats, min–max where they differ; loopback on fv-ml1 GPU 3)
|
||||
|
||||
| system | fits 12 GiB? rest / peak MiB | JevBench all /231 · hard /111 | pooled /259 single · rot | Wyrd /84 single · rot | Δ pooled vs SemIf single (95% CI) | Δ pooled vs SemIf rot (95% CI) | 21 criteria, ms | 16 × ~3,900 tok, ms |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | yes: 9242 (9242–9250) / 12918 (12886–12918) | 187 · 68 | 217 · 232 | 71 · 71 | baseline | baseline | 131 (131–132) | 505 (502–519) |
|
||||
| Plumb-4B, native | yes: 12064 / 12064 | 207 (206–207) · 89 (88–89) | 224 (224–226) · 227 (227–228) | 65 · 67 | +2.7 (-2.7..+7.9) | -1.9 (-5.4..+1.5) | 294 (292–295) | 3378 (3374–3384) |
|
||||
| Plumb-4B, drop-in | yes: 9242 / 12918 | 206 · 88 | 221 · 224 | 66 · 67 | +1.5 (-3.7..+6.6) | -3.1 (-6.8..+0.4) | 131 (131–132) | 506 (506–507) |
|
||||
| Imajev-4B, native | yes: 10404 / 11056 | 199 · 80 | 235 · 236 | 68 · 69 | +6.9 (+1.6..+12.2) | +1.5 (-2.5..+5.6) | 1219 (1212–1230) | 5125 (5054–5140) |
|
||||
| Intern-Decision-4B, native | yes: 9736 / 10290 | 202 · 83 | 240 · 236 | 79 · 77 | +8.9 (+4.7..+13.2) | +1.5 (-2.0..+5.0) | 88 (88–89) | 215 (213–216) |
|
||||
| Intern-Decision-4B, drop-in | yes: 9242 / 12918 | 202 · 85 | 232 · 233 | 74 · 73 | +5.8 (+2.7..+9.0) | +0.4 (-2.3..+3.1) | 132 | 508 (500–516) |
|
||||
| JevK5 v0.2, native | yes: 12064 / 12064 | 198 (198–199) · 81 (81–82) | 229 (228–229) · 228 | 67 (66–67) · 67 | +4.6 (-0.8..+9.9) | -1.5 (-4.9..+1.9) | 291 (291–292) | 3360 (3358–3368) |
|
||||
| JevK5 v0.2, drop-in | yes: 9242 / 12918 | 200 · 82 | 232 · 237 | 72 · 74 | +5.8 (+1.1..+10.7) | +1.9 (-1.1..+5.1) | 132 | 518 (505–519) |
|
||||
| JevK5 v0.3, native | yes: 12064 / 12064 | 202 · 86 | 227 · 226 | 66 · 65 | +3.9 (-0.9..+8.5) | -2.3 (-6.0..+1.4) | 291 (291–292) | 3363 (3355–3370) |
|
||||
| JevK5 v0.3, drop-in | yes: 9242 / 12918 | 201 · 85 | 227 · 227 | 69 · 66 | +3.9 (-0.8..+8.2) | -1.9 (-5.5..+1.8) | 133 (131–133) | 518 (506–521) |
|
||||
|
||||
Pooled = authored144 + Cicada w1 + Wyrd (all 4 decisions) = 259 labelled rows. "single" = one ordering, "rot" = n rotations averaged. Δ is paired, per row, on each row's majority answer over the repeats, with a group-bootstrap 95% CI.
|
||||
|
||||
### Recommendation (each claim with its own strength)
|
||||
|
||||
- **"If SemIf is displaced, the replacement is Intern-Decision-4B run on its own runtime"**: **recommend** [measured: best on the pooled set and on Wyrd, fastest at our shape, fits the cap; reversible, since it is a service swap and the weights are already on `/tank`].
|
||||
- **"Displace SemIf today on accuracy alone"**: **lean against** [measured: against SemIf-with-rotations the pooled gain (+1.5 to +3.1) is inside the ~±4-pt floor; reversible]. What would justify a switch is compute: rotation-level accuracy at one ordering, and 1.5–2.3× faster multi-criteria requests. That is a build: a small new service, contract and TDD. It is not a config flip.
|
||||
- **"Pick Plumb or Imajev because of their JevBench rank"**: **strongly recommend against** [measured: no gain on our sets; Plumb is worse on Wyrd and Imajev is 9× slower; reversible].
|
||||
- **"Swap weights inside semif-serve (drop-in) for a free gain"**: **lean against** [measured: the best drop-in, JevK5 v0.2, is +1.9 against SemIf-rotations, CI −1.1..+5.1, which is the same; reversible].
|
||||
- **Before trusting Intern-Decision in production:** label ~50 real Wyrd/Cicada turns as a held-out set. Its card makes no contamination claim, and Wyrd and Cicada are the only sets here that no model could have seen. **recommend** [reasoned; reversible].
|
||||
|
||||
### Plain-language decision tree
|
||||
|
||||
- **Do we only need to free the utility GPU, or do we want a better decider?**
|
||||
- **Only free the GPU / keep what works** → keep SemIf with rotations. Nothing here beats it by enough to justify a new service. *(Every candidate and format is within ±3.1 pts of SemIf-with-rotations on the pooled 259 rows; the floor is ~±4.)*
|
||||
- **Want the same answers cheaper, or better answers on the Wyrd-style multi-question turns** → Intern-Decision-4B on its own runtime.
|
||||
- It needs a small new service, because its request format differs *(Jev schema; one forward over up to 16 questions; `inference.py` ships in the model repo)*.
|
||||
- First, check it on ~50 real labelled turns *(its card makes no contamination claim, and it is not on the JevBench board)*.
|
||||
- **Want a zero-code swap through semif-serve** → not worth it. *(Best drop-in: JevK5 v0.2, +1.9 pts, CI spans 0; same latency.)*
|
||||
|
||||
## Harness (stated once; every number below was measured under it)
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Card | fv-ml1 **GPU 3** (RTX PRO 6000 Blackwell Max-Q, 96 GB, sm_120), empty at start (2 MiB) and nothing foreign on it during the bench (guard log per repeat). Nothing on GPU 1. |
|
||||
| Stack, every system | torch 2.10.0+cu128, transformers 5.17.0, flash-linear-attention / fla-core 0.5.2, causal-conv1d 1.7.0, BF16, SDPA. This is semif-serve:0.1.4's image (`sha256:36a3e1d2…`); the native runtimes run in `jevbench-native:2026-09-30`, the same image plus jevk5 0.2.0 (`85238d7b`), peft 0.21.1, pillow 12.3.0, torchvision 0.25.0+cu128. |
|
||||
| VRAM cap | 12 GiB for every served system: semif-serve's own `SEMIF_VRAM_CAP_GIB=12`; the native servers under `capped.py`, which applies the same `torch.cuda.set_per_process_memory_fraction` before any weights load. |
|
||||
| Clients | fv-ml1 host python 3.13, stdlib urllib, loopback (`127.0.0.1`), one request at a time. No network in the latency numbers (the README's nh3-dev figures carry ~27–33 ms of network on top). |
|
||||
| Repeat | one repeat = one fresh process (container) lifetime. N = 3 per cell; 4 for SemIf, Imajev and Intern-Decision native (one repeat added each: Imajev's first repeat ran its latency shapes without its 8-questions-per-request cap and was discarded for latency only, `r1/shape-maxq-unset-invalid.json`). The Wyrd one-request-per-turn condition was added after the first repeats, so it has N = 3 except Plumb native (N = 2; that runtime loops over questions, so it equals single by construction). |
|
||||
| Prompts | **native** = each candidate's own runtime and prompt as its card documents it, reached over its TypeSafe-style `/v1/systemone` server; a row becomes one `choice` question with `criteria = {option id: description}` in the row's order. **drop-in** = the candidate's weights served by semif-serve 0.1.4 unchanged (`SEMIF_MODEL`/`SEMIF_REVISION`), i.e. SemIf's own prompt, letter readout and shared-prefix path: what a switch through the semif-serve contract would get. |
|
||||
| Orderings | **single** = the caller's order. **rotations** = the n cyclic rotations averaged by log-mean (semif-serve's own `"orderings": "rotations"` for semif-serve; the same arithmetic client-side, n separate requests, for native runtimes). |
|
||||
| JevBench | fstandhartinger/jevbench tag **v1.2.16** (`5e95f23cbb7be098a9061fea924c4421620ab1a5`), the 231 public items (easy 48, standard/original 72, hard 111), its own `jevbench.cli run`, one ordering, v1.2 scoring (argmax over the exact label set). SemIf and every drop-in through JevBench's own `semif_direct` adapter (in-process, SemIf's code path); native candidates through its `typesafe` adapter to the candidate's loopback server, which is how the board measured them. |
|
||||
|
||||
## Controls, and what each one showed
|
||||
|
||||
**Positive control 1 — the JevBench harness (step 1).** SemIf on our box, through JevBench v1.2.16's own `semif_direct` adapter: easy 48/48, standard 71/72, hard 68/111 (0.613), all 187/231 (**0.810**), identical in all 4 repeats. That is the cards' 0.810 / 0.613 exactly, and the same outcome JevK5's card reports for the untrained base. The harness is right. Each candidate's public number was also reproduced, which makes every candidate row its own positive control: Plumb 89/111 hard (its card: 0.802 = 89), Imajev 80/111 (board `hard_public` 0.7207 = 80), JevK5 v0.2 81–82/111 native, 82 drop-in (card 0.739 = 82), Intern-Decision 83/111 against a vendor 82/111 (one item over, on transformers 5.17.0 where the card pins 5.14.1).
|
||||
|
||||
**Positive control 2 — the replaced-baseline harness.** The SemIf repeats reproduce semif-serve 0.1.4's acceptance and the 09-27 spikes row for row: parity with SemIf's committed torch predictions 144/144 top choice, 144/144 identical prompt SHA-256, max prob gap 0.060 (acceptance: 144/144, 0.060); negative control 14/144 same-top (acceptance: 14/144); Cicada w1 30/31; Wyrd with rotations place 16/21, place2 21/21, exit 18/21, exit2 16/21 (spike: identical); SemIf's labelled sets 78.6% → 88.1% with rotations in the averaging acceptance, here 199/252 = 79.0% → 224/252 = 88.9%. Capacity under 12 GiB: 16 rows at ~3,900 prefix tokens, 18 is a 503 (README: 16). Latency: 21 binary criteria in 131 ms on loopback; the README's 159 ms was measured from nh3-dev, whose round trip is 22–28 ms (ping averages today); a bridge run from nh3-dev against the GPU 3 instance gave 151.5 ms (below, under what could not be measured).
|
||||
|
||||
**Negative control (step 6).** authored144 with each row's option descriptions rotated one place and the ids kept. A model that reads the descriptions must move its top to the id that now carries the right description. Reported three ways: same top as unrotated (should be LOW), follows the description (should be HIGH), right against the original gold (should be LOW).
|
||||
|
||||
**Null control.** Every Cicada and Wyrd row is also asked over a content-free state ("(No evidence is available for this turn.)"). Whatever that gets right is the prior carried by the question and options alone; the evidence condition is read against it.
|
||||
|
||||
**Noise floors.** A-vs-A inside a process: authored144 scored twice in the same process. Across restarts: every repeat is a fresh container, so each pair of repeats is an A-vs-A across restarts over all 560 labelled+blind rows (single) and 560 (rotations). The paired comparisons use each row's majority top over the repeats, so a row that flips across restarts cannot manufacture a difference by itself.
|
||||
|
||||
**Sensitivity floor.** Item sampling dominates, not run noise (see the floors table: SemIf flips 0 rows across restarts). Paired group-bootstrap 95% CIs on the pooled 259 rows have a half-width of about ±3 to ±5 points, so **this bench cannot resolve a pooled difference smaller than ~4 points**; on Cicada (31 rows) nothing under ~10–15 points is resolvable, on Wyrd (84 rows, 21 turns) nothing under ~6–7.
|
||||
|
||||
**Decision rule, stated before the final repeats were in (after the first repeat of each system had been seen):** a candidate is *better* on a set when the paired 95% CI against SemIf excludes zero in its favour AND the gap exceeds the number of rows either side flips across restarts; *worse* the mirror image; otherwise *same*. It is judged against SemIf **both** as deployed today (single ordering, the default) **and** with rotations (what semif-serve's README says to use "for anything real").
|
||||
|
||||
## Candidates: identity, pins, licence as read at the raw file
|
||||
|
||||
Every HF repo ID was verified with an authenticated API call (`/api/models/<id>`, HF token from the vault bundle `nh3-dev/.config/secrets/env.sh`) before any pull: all returned 200, none gated. Weights were pulled to `/tank/aimodels/huggingface` on fv-ml1 as `llmuser` and every weight file's sha256 was checked against the Hub's LFS oid (all match; `raw/weights-sha256.txt`).
|
||||
|
||||
| candidate | HF repo @ pinned revision | what it is | runtime used (native) | licence, read at the raw file | contamination statement | JevBench board (v1.4.2.2) |
|
||||
|---|---|---|---|---|---|---|
|
||||
| **SemIf** (baseline) | `Qwen/Qwen3.5-4B` @ `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` | frozen base; SemIf `23cf1f39` reads the option letters | semif-serve 0.1.4 | Qwen `LICENSE`: Apache-2.0 (Copyright 2026 Alibaba Cloud). SemIf: MIT | n/a (frozen) | #13, 47.7; hard(220) 0.595, sealed 0.263 |
|
||||
| **Plumb-4B** | `crh225/plumb-4b` @ `55de037801a8a9b9de3db5c0e16cef86210c2186` (the board's pin; HF main `24f7bf77` changes only the card) | JevK5 v0.2 + LoRA, merged, 8.4 GB | jevk5 runtime 0.2.0 (`85238d7b`), the board's serving path | `LICENSE`: Apache-2.0 (unfilled template). `NOTICE`: JevK5 v0.2 (Apache-2.0), Qwen (Apache-2.0), SemIf prompt/readout (MIT) | yes: no JevBench item trained/tuned/selected; 8-gram scan + exact-text audit | #2, 65.8; public 0.896, sealed 0.380 |
|
||||
| **Imajev-4B** | `mohit67890/imajev-4b` @ `c9e5f132465da85d31735ec502d5557982671a7d` (the board's pin) | PEFT LoRA (r64) on Qwen3.5-4B + its own 256-code decision readout head; 0.5 GB adapter | imajev server `a0134749` (the board's commit), torch backend, `--rotations 1 --calibration calibration.json` | **HF repo has no LICENSE file** (card tag only). GitHub `mohit67890/imajev` `LICENSE`: Apache-2.0; card: "Code and adapters: Apache-2.0" | yes: "No JevBench items (8-gram lint)" | #1, 67.4; public 0.861, sealed 0.370 |
|
||||
| **Intern-Decision-4B** | `internlm/Intern-Decision-4B` @ `0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd` | full Qwen3.5-4B finetune (incl. vision tower), 9.1 GB, adds a `<decision>` token | its own `inference.py` (`DecisionEngine`, shipped in the repo) behind a 60-line `/v1/systemone` wrapper | `LICENSE`: Apache-2.0 (Copyright 2026-2027 Shanghai AI Laboratory); `LICENSE-QWEN`: Qwen's Apache-2.0 | **none on the card** (flagged) | **not on the board** (released 09-26); vendor-only claims |
|
||||
| **JevK5 v0.2** | `alibiserikbay/JevK5` @ `27d2d6b8d4714807f6293b0623bd7370b27e42f8` (= Hub tag `v0.2`; tag object `ea4804e9`; weights identical to the v0.2 release `844e4d0a`, LFS `0fba3bba…`) | Qwen3.5-4B + distilled LoRA, merged, 8.4 GB. `allebee/jevk5` (GitHub) and `alibiserikbay/JevK5` (HF) are the same project: the NOTICE names both handles | jevk5 runtime 0.2.0 (`85238d7b`) | **HF repo has no LICENSE file**. GitHub `allebee/jevk5` `LICENSE`: Apache-2.0 | yes (v0.2 card); discloses 940 MMLU-Pro **test** items in its training replay | #5, 62.0; public 0.853, sealed 0.331 |
|
||||
| JevK5 v0.3 (extra) | `alibiserikbay/JevK5` @ `c4f7fdb3aeab5582336406e78d3bef11bf98833d` (= Hub tag `v0.3` and current main) | as above, retrained; 8.4 GB | jevk5 runtime 0.3.3 (`f944fe37`) | as above; NOTICE: 14,138 training questions written by GPT-6 Luna "under OpenAI's terms" | yes | not on the board |
|
||||
|
||||
Not run, with the reason:
|
||||
- **Hopper** (`HopitAI/hopper`, last in brokkr's order): the adapter's card sets `license: other / research-and-demo` and says "Do not use this adapter commercially" (RACE training data, non-commercial terms). There is no LICENSE file. That is the same class as `openjev/openjev` (CC-BY-NC), which the brief excluded, so it was excluded by the same rule. It is one queue line if Prime wants it measured anyway.
|
||||
- **Imajev drop-in**: not applicable. Imajev is a LoRA plus its own readout head, not a full checkpoint whose LM-head letters carry the answer, so semif-serve cannot load it without new code.
|
||||
- `Intern-Decision-2B/0.8B`, `AlexWortega/openjev`: dropped per brokkr's revised order. `akhilaaa3/Jev-Omni`, `openjev/openjev`, Cygnet: excluded as briefed (size / licence).
|
||||
|
||||
## Full tables
|
||||
|
||||
Generated by `code/tables.py` from `summary.json` (`code/analyze.py` over `raw/out/`), all under `services/semif-serve/bench-jev-2026-09-30/`.
|
||||
|
||||
#### JevBench v1.2.16 public items (positive control + candidate score)
|
||||
|
||||
| system | repeats | easy /48 | standard /72 | hard /111 | all /231 | all | p50 ms |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 4 | 48 | 71 | 68 | 187 | 0.810 | 36 |
|
||||
| Plumb-4B, native | 3 | 48 | 70 | 89 (88–89) | 207 (206–207) | 0.896 (0.892–0.896) | 20 |
|
||||
| Plumb-4B, drop-in | 3 | 48 | 70 | 88 | 206 | 0.892 | 36 |
|
||||
| Imajev-4B, native | 4 | 48 | 71 | 80 | 199 | 0.861 | 62 (62–63) |
|
||||
| Intern-Decision-4B, native | 4 | 48 | 71 | 83 | 202 | 0.874 | 40 |
|
||||
| Intern-Decision-4B, drop-in | 3 | 48 | 69 | 85 | 202 | 0.874 | 36 |
|
||||
| JevK5 v0.2, native | 3 | 48 | 69 | 81 (81–82) | 198 (198–199) | 0.857 (0.857–0.861) | 20 (20–21) |
|
||||
| JevK5 v0.2, drop-in | 3 | 48 | 70 | 82 | 200 | 0.866 | 36 (36–37) |
|
||||
| JevK5 v0.3, native | 3 | 48 | 68 | 86 | 202 | 0.874 | 20 (20–21) |
|
||||
| JevK5 v0.3, drop-in | 3 | 48 | 68 | 85 | 201 | 0.870 | 36 |
|
||||
|
||||
#### Replaced-baseline sets, single (caller order)
|
||||
|
||||
| system | pooled /259 | authored144 | perturb108 | Cicada w1 /31 | Cicada w2 /31 | Wyrd /84 | Wyrd place2 /21 | Wyrd exit /21 |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 217 | 116 | 83 | 30 | 20 | 71 | 21 | 16 |
|
||||
| Plumb-4B, native | 224 (224–226) | 130 (130–132) | 84 | 29 | 19 | 65 | 21 | 10 |
|
||||
| Plumb-4B, drop-in | 221 | 129 | 88 | 26 | 20 | 66 | 21 | 12 |
|
||||
| Imajev-4B, native | 235 | 138 | 102 | 29 | 22 | 68 | 20 | 13 |
|
||||
| Intern-Decision-4B, native | 240 | 132 | 100 | 29 | 22 | 79 | 21 | 19 |
|
||||
| Intern-Decision-4B, drop-in | 232 | 128 | 102 | 30 | 19 | 74 | 21 | 17 |
|
||||
| JevK5 v0.2, native | 229 (228–229) | 132 | 85 (85–86) | 30 | 16 | 67 (66–67) | 21 | 12 (11–12) |
|
||||
| JevK5 v0.2, drop-in | 232 | 130 | 92 | 30 | 17 | 72 | 21 | 17 |
|
||||
| JevK5 v0.3, native | 227 | 135 | 101 (101–102) | 26 | 17 | 66 | 21 | 12 |
|
||||
| JevK5 v0.3, drop-in | 227 | 132 | 99 | 26 | 21 | 69 | 21 | 15 |
|
||||
|
||||
#### Replaced-baseline sets, rotations (n rotations averaged)
|
||||
|
||||
| system | pooled /259 | authored144 | perturb108 | Cicada w1 /31 | Cicada w2 /31 | Wyrd /84 | Wyrd place2 /21 | Wyrd exit /21 |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 232 | 131 | 93 | 30 | 19 | 71 | 21 | 18 |
|
||||
| Plumb-4B, native | 227 (227–228) | 130 (130–131) | 87 | 30 | 18 (17–18) | 67 | 21 | 11 |
|
||||
| Plumb-4B, drop-in | 224 | 129 | 92 | 28 | 18 | 67 | 21 | 13 |
|
||||
| Imajev-4B, native | 236 | 138 | 101 | 29 | 23 | 69 | 20 | 13 |
|
||||
| Intern-Decision-4B, native | 236 | 130 | 99 | 29 | 22 | 77 | 21 | 19 |
|
||||
| Intern-Decision-4B, drop-in | 233 | 130 | 102 | 30 | 19 | 73 | 21 | 17 |
|
||||
| JevK5 v0.2, native | 228 | 131 | 91 (90–91) | 30 | 16 | 67 | 21 | 12 |
|
||||
| JevK5 v0.2, drop-in | 237 | 133 | 93 | 30 | 16 | 74 | 21 | 20 |
|
||||
| JevK5 v0.3, native | 226 | 135 | 101 | 26 | 16 | 65 | 20 | 12 |
|
||||
| JevK5 v0.3, drop-in | 227 | 135 | 101 | 26 | 19 | 66 | 20 | 13 |
|
||||
|
||||
#### Wyrd as one request per turn (4 decisions over one state: SemIf /decide/shared, native multi-question request)
|
||||
|
||||
| system | repeats | Wyrd /84 | place /21 | place2 /21 | exit /21 | exit2 /21 |
|
||||
|---|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 3 | 71 | 17 | 21 | 16 | 17 |
|
||||
| Plumb-4B, native | 2 | 65 | 17 | 21 | 10 | 17 |
|
||||
| Plumb-4B, drop-in | 3 | 66 | 17 | 21 | 12 | 16 |
|
||||
| Imajev-4B, native | 3 | 68 | 19 | 20 | 13 | 16 |
|
||||
| Intern-Decision-4B, native | 3 | 77 | 16 | 21 | 20 | 20 |
|
||||
| Intern-Decision-4B, drop-in | 3 | 74 | 18 | 21 | 17 | 18 |
|
||||
| JevK5 v0.2, native | 3 | 67 (66–67) | 17 | 21 | 12 (11–12) | 17 |
|
||||
| JevK5 v0.2, drop-in | 3 | 73 | 18 | 21 | 18 | 16 |
|
||||
| JevK5 v0.3, native | 3 | 66 | 18 | 21 | 12 | 15 |
|
||||
| JevK5 v0.3, drop-in | 3 | 69 | 18 | 21 | 15 | 15 |
|
||||
|
||||
#### Null control (content-free state) and positive controls inside the spike sets, single ordering
|
||||
|
||||
| system | Cicada w1 blind /31 | Wyrd blind /84 | Cicada w1 controls /5 | Wyrd controls /28 |
|
||||
|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 16 | 60 | 5 | 21 |
|
||||
| Plumb-4B, native | 15 | 50 | 5 | 20 |
|
||||
| Plumb-4B, drop-in | 15 | 60 | 4 | 21 |
|
||||
| Imajev-4B, native | 16 | 48 | 5 | 24 |
|
||||
| Intern-Decision-4B, native | 16 | 60 | 5 | 27 |
|
||||
| Intern-Decision-4B, drop-in | 16 | 60 | 5 | 23 |
|
||||
| JevK5 v0.2, native | 15 | 53 | 5 | 22 (21–22) |
|
||||
| JevK5 v0.2, drop-in | 15 | 60 | 5 | 23 |
|
||||
| JevK5 v0.3, native | 16 | 60 | 5 | 20 |
|
||||
| JevK5 v0.3, drop-in | 16 | 60 | 5 | 22 |
|
||||
|
||||
#### Paired against SemIf (each row's majority top over the repeats; group bootstrap 95% CI; exact McNemar)
|
||||
|
||||
| system | set | cond | SemIf | cand | fixed | broken | Δ pts | 95% CI | p |
|
||||
|---|---|---|---|---|---|---|---|---|---|
|
||||
| Plumb-4B, native | pooled | single | 217/259 | 224/259 | 23 | 16 | +2.7 | -2.7..+7.9 | 0.3368 |
|
||||
| Plumb-4B, native | pooled | rotations | 232/259 | 227/259 | 10 | 15 | -1.9 | -5.4..+1.5 | 0.4244 |
|
||||
| Plumb-4B, native | authored144 | single | 116/144 | 130/144 | 22 | 8 | +9.7 | +1.4..+18.1 | 0.0161 |
|
||||
| Plumb-4B, native | authored144 | rotations | 131/144 | 130/144 | 7 | 8 | -0.7 | -5.6..+3.5 | 1.0 |
|
||||
| Plumb-4B, native | cicada-w1 | single | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
|
||||
| Plumb-4B, native | cicada-w1 | rotations | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| Plumb-4B, native | wyrd | single | 71/84 | 65/84 | 1 | 7 | -7.1 | -13.1..-1.2 | 0.0703 |
|
||||
| Plumb-4B, native | wyrd | rotations | 71/84 | 67/84 | 3 | 7 | -4.8 | -11.9..+2.4 | 0.3438 |
|
||||
| Plumb-4B, native | perturbations108 | single | 83/108 | 84/108 | 10 | 9 | +0.9 | -9.3..+11.1 | 1.0 |
|
||||
| Plumb-4B, native | perturbations108 | rotations | 93/108 | 87/108 | 4 | 10 | -5.6 | -15.7..+3.7 | 0.1796 |
|
||||
| Plumb-4B, native | cicada-w2 | single | 20/31 | 19/31 | 1 | 2 | -3.2 | -12.9..+6.5 | 1.0 |
|
||||
| Plumb-4B, native | cicada-w2 | rotations | 19/31 | 18/31 | 1 | 2 | -3.2 | -16.1..+6.5 | 1.0 |
|
||||
| Plumb-4B, native | pooled | cand-single-vs-semif-rotations | 232/259 | 224/259 | 10 | 18 | -3.1 | -7.3..+0.8 | 0.1849 |
|
||||
| Plumb-4B, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 65/84 | 3 | 9 | -7.1 | -15.5..+1.2 | 0.146 |
|
||||
| Plumb-4B, drop-in | pooled | single | 217/259 | 221/259 | 21 | 17 | +1.5 | -3.7..+6.6 | 0.6271 |
|
||||
| Plumb-4B, drop-in | pooled | rotations | 232/259 | 224/259 | 7 | 15 | -3.1 | -6.8..+0.4 | 0.1338 |
|
||||
| Plumb-4B, drop-in | authored144 | single | 116/144 | 129/144 | 20 | 7 | +9.0 | +1.4..+16.7 | 0.0192 |
|
||||
| Plumb-4B, drop-in | authored144 | rotations | 131/144 | 129/144 | 5 | 7 | -1.4 | -6.2..+3.5 | 0.7744 |
|
||||
| Plumb-4B, drop-in | cicada-w1 | single | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
|
||||
| Plumb-4B, drop-in | cicada-w1 | rotations | 30/31 | 28/31 | 0 | 2 | -6.5 | -16.1..+0.0 | 0.5 |
|
||||
| Plumb-4B, drop-in | wyrd | single | 71/84 | 66/84 | 1 | 6 | -6.0 | -11.9..+0.0 | 0.125 |
|
||||
| Plumb-4B, drop-in | wyrd | rotations | 71/84 | 67/84 | 2 | 6 | -4.8 | -11.9..+2.4 | 0.2891 |
|
||||
| Plumb-4B, drop-in | perturbations108 | single | 83/108 | 88/108 | 12 | 7 | +4.6 | -5.6..+14.8 | 0.3593 |
|
||||
| Plumb-4B, drop-in | perturbations108 | rotations | 93/108 | 92/108 | 6 | 7 | -0.9 | -10.2..+7.4 | 1.0 |
|
||||
| Plumb-4B, drop-in | cicada-w2 | single | 20/31 | 20/31 | 1 | 1 | +0.0 | -9.7..+9.7 | 1.0 |
|
||||
| Plumb-4B, drop-in | cicada-w2 | rotations | 19/31 | 18/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
|
||||
| Plumb-4B, drop-in | pooled | cand-single-vs-semif-rotations | 232/259 | 221/259 | 8 | 19 | -4.2 | -8.3..-0.4 | 0.0522 |
|
||||
| Plumb-4B, drop-in | wyrd | cand-single-vs-semif-rotations | 71/84 | 66/84 | 2 | 7 | -6.0 | -13.1..+1.2 | 0.1797 |
|
||||
| Imajev-4B, native | pooled | single | 217/259 | 235/259 | 30 | 12 | +6.9 | +1.6..+12.2 | 0.0079 |
|
||||
| Imajev-4B, native | pooled | rotations | 232/259 | 236/259 | 14 | 10 | +1.5 | -2.5..+5.6 | 0.5413 |
|
||||
| Imajev-4B, native | authored144 | single | 116/144 | 138/144 | 25 | 3 | +15.3 | +9.0..+21.5 | 0.0 |
|
||||
| Imajev-4B, native | authored144 | rotations | 131/144 | 138/144 | 8 | 1 | +4.9 | +0.7..+9.0 | 0.0391 |
|
||||
| Imajev-4B, native | cicada-w1 | single | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
|
||||
| Imajev-4B, native | cicada-w1 | rotations | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
|
||||
| Imajev-4B, native | wyrd | single | 71/84 | 68/84 | 5 | 8 | -3.6 | -13.1..+6.0 | 0.5811 |
|
||||
| Imajev-4B, native | wyrd | rotations | 71/84 | 69/84 | 6 | 8 | -2.4 | -11.9..+7.1 | 0.7905 |
|
||||
| Imajev-4B, native | perturbations108 | single | 83/108 | 102/108 | 20 | 1 | +17.6 | +8.3..+27.8 | 0.0 |
|
||||
| Imajev-4B, native | perturbations108 | rotations | 93/108 | 101/108 | 9 | 1 | +7.4 | +1.9..+13.9 | 0.0215 |
|
||||
| Imajev-4B, native | cicada-w2 | single | 20/31 | 22/31 | 3 | 1 | +6.5 | -6.5..+19.4 | 0.625 |
|
||||
| Imajev-4B, native | cicada-w2 | rotations | 19/31 | 23/31 | 4 | 0 | +12.9 | +3.2..+25.8 | 0.125 |
|
||||
| Imajev-4B, native | pooled | cand-single-vs-semif-rotations | 232/259 | 235/259 | 15 | 12 | +1.2 | -3.2..+5.5 | 0.7011 |
|
||||
| Imajev-4B, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 68/84 | 6 | 9 | -3.6 | -14.3..+6.0 | 0.6072 |
|
||||
| Intern-Decision-4B, native | pooled | single | 217/259 | 240/259 | 28 | 5 | +8.9 | +4.7..+13.2 | 0.0001 |
|
||||
| Intern-Decision-4B, native | pooled | rotations | 232/259 | 236/259 | 12 | 8 | +1.5 | -2.0..+5.0 | 0.5034 |
|
||||
| Intern-Decision-4B, native | authored144 | single | 116/144 | 132/144 | 20 | 4 | +11.1 | +4.9..+18.1 | 0.0015 |
|
||||
| Intern-Decision-4B, native | authored144 | rotations | 131/144 | 130/144 | 5 | 6 | -0.7 | -5.6..+4.2 | 1.0 |
|
||||
| Intern-Decision-4B, native | cicada-w1 | single | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
|
||||
| Intern-Decision-4B, native | cicada-w1 | rotations | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
|
||||
| Intern-Decision-4B, native | wyrd | single | 71/84 | 79/84 | 8 | 0 | +9.5 | +3.6..+16.7 | 0.0078 |
|
||||
| Intern-Decision-4B, native | wyrd | rotations | 71/84 | 77/84 | 7 | 1 | +7.1 | +1.2..+13.1 | 0.0703 |
|
||||
| Intern-Decision-4B, native | perturbations108 | single | 83/108 | 100/108 | 20 | 3 | +15.7 | +4.6..+26.9 | 0.0005 |
|
||||
| Intern-Decision-4B, native | perturbations108 | rotations | 93/108 | 99/108 | 9 | 3 | +5.6 | -2.8..+13.0 | 0.146 |
|
||||
| Intern-Decision-4B, native | cicada-w2 | single | 20/31 | 22/31 | 2 | 0 | +6.5 | +0.0..+16.1 | 0.5 |
|
||||
| Intern-Decision-4B, native | cicada-w2 | rotations | 19/31 | 22/31 | 4 | 1 | +9.7 | -3.2..+22.6 | 0.375 |
|
||||
| Intern-Decision-4B, native | pooled | cand-single-vs-semif-rotations | 232/259 | 240/259 | 14 | 6 | +3.1 | -0.4..+6.7 | 0.1153 |
|
||||
| Intern-Decision-4B, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 79/84 | 9 | 1 | +9.5 | +2.4..+16.7 | 0.0215 |
|
||||
| Intern-Decision-4B, drop-in | pooled | single | 217/259 | 232/259 | 18 | 3 | +5.8 | +2.7..+9.0 | 0.0015 |
|
||||
| Intern-Decision-4B, drop-in | pooled | rotations | 232/259 | 233/259 | 8 | 7 | +0.4 | -2.3..+3.1 | 1.0 |
|
||||
| Intern-Decision-4B, drop-in | authored144 | single | 116/144 | 128/144 | 15 | 3 | +8.3 | +3.5..+13.9 | 0.0075 |
|
||||
| Intern-Decision-4B, drop-in | authored144 | rotations | 131/144 | 130/144 | 5 | 6 | -0.7 | -4.9..+3.5 | 1.0 |
|
||||
| Intern-Decision-4B, drop-in | cicada-w1 | single | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| Intern-Decision-4B, drop-in | cicada-w1 | rotations | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| Intern-Decision-4B, drop-in | wyrd | single | 71/84 | 74/84 | 3 | 0 | +3.6 | +0.0..+8.3 | 0.25 |
|
||||
| Intern-Decision-4B, drop-in | wyrd | rotations | 71/84 | 73/84 | 3 | 1 | +2.4 | -2.4..+7.1 | 0.625 |
|
||||
| Intern-Decision-4B, drop-in | perturbations108 | single | 83/108 | 102/108 | 20 | 1 | +17.6 | +7.4..+28.7 | 0.0 |
|
||||
| Intern-Decision-4B, drop-in | perturbations108 | rotations | 93/108 | 102/108 | 9 | 0 | +8.3 | +2.8..+15.7 | 0.0039 |
|
||||
| Intern-Decision-4B, drop-in | cicada-w2 | single | 20/31 | 19/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
|
||||
| Intern-Decision-4B, drop-in | cicada-w2 | rotations | 19/31 | 19/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| Intern-Decision-4B, drop-in | pooled | cand-single-vs-semif-rotations | 232/259 | 232/259 | 10 | 10 | +0.0 | -3.0..+3.1 | 1.0 |
|
||||
| Intern-Decision-4B, drop-in | wyrd | cand-single-vs-semif-rotations | 71/84 | 74/84 | 4 | 1 | +3.6 | -1.2..+8.3 | 0.375 |
|
||||
| JevK5 v0.2, native | pooled | single | 217/259 | 229/259 | 25 | 13 | +4.6 | -0.8..+9.9 | 0.073 |
|
||||
| JevK5 v0.2, native | pooled | rotations | 232/259 | 228/259 | 9 | 13 | -1.5 | -4.9..+1.9 | 0.5235 |
|
||||
| JevK5 v0.2, native | authored144 | single | 116/144 | 132/144 | 23 | 7 | +11.1 | +2.8..+18.8 | 0.0052 |
|
||||
| JevK5 v0.2, native | authored144 | rotations | 131/144 | 131/144 | 7 | 7 | +0.0 | -4.9..+4.9 | 1.0 |
|
||||
| JevK5 v0.2, native | cicada-w1 | single | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| JevK5 v0.2, native | cicada-w1 | rotations | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| JevK5 v0.2, native | wyrd | single | 71/84 | 67/84 | 2 | 6 | -4.8 | -10.7..+1.2 | 0.2891 |
|
||||
| JevK5 v0.2, native | wyrd | rotations | 71/84 | 67/84 | 2 | 6 | -4.8 | -10.7..+1.2 | 0.2891 |
|
||||
| JevK5 v0.2, native | perturbations108 | single | 83/108 | 85/108 | 10 | 8 | +1.9 | -7.4..+11.1 | 0.8145 |
|
||||
| JevK5 v0.2, native | perturbations108 | rotations | 93/108 | 91/108 | 4 | 6 | -1.9 | -10.2..+5.6 | 0.7539 |
|
||||
| JevK5 v0.2, native | cicada-w2 | single | 20/31 | 16/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
|
||||
| JevK5 v0.2, native | cicada-w2 | rotations | 19/31 | 16/31 | 0 | 3 | -9.7 | -19.4..+0.0 | 0.25 |
|
||||
| JevK5 v0.2, native | pooled | cand-single-vs-semif-rotations | 232/259 | 229/259 | 11 | 14 | -1.2 | -5.0..+2.6 | 0.69 |
|
||||
| JevK5 v0.2, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 67/84 | 3 | 7 | -4.8 | -13.1..+2.4 | 0.3438 |
|
||||
| JevK5 v0.2, drop-in | pooled | single | 217/259 | 232/259 | 24 | 9 | +5.8 | +1.1..+10.7 | 0.0135 |
|
||||
| JevK5 v0.2, drop-in | pooled | rotations | 232/259 | 237/259 | 11 | 6 | +1.9 | -1.1..+5.1 | 0.3323 |
|
||||
| JevK5 v0.2, drop-in | authored144 | single | 116/144 | 130/144 | 21 | 7 | +9.7 | +2.1..+17.4 | 0.0125 |
|
||||
| JevK5 v0.2, drop-in | authored144 | rotations | 131/144 | 133/144 | 6 | 4 | +1.4 | -2.8..+5.6 | 0.7539 |
|
||||
| JevK5 v0.2, drop-in | cicada-w1 | single | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| JevK5 v0.2, drop-in | cicada-w1 | rotations | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| JevK5 v0.2, drop-in | wyrd | single | 71/84 | 72/84 | 3 | 2 | +1.2 | -3.6..+6.0 | 1.0 |
|
||||
| JevK5 v0.2, drop-in | wyrd | rotations | 71/84 | 74/84 | 5 | 2 | +3.6 | -2.4..+10.7 | 0.4531 |
|
||||
| JevK5 v0.2, drop-in | perturbations108 | single | 83/108 | 92/108 | 13 | 4 | +8.3 | -0.9..+17.6 | 0.049 |
|
||||
| JevK5 v0.2, drop-in | perturbations108 | rotations | 93/108 | 93/108 | 4 | 4 | +0.0 | -7.4..+6.5 | 1.0 |
|
||||
| JevK5 v0.2, drop-in | cicada-w2 | single | 20/31 | 17/31 | 0 | 3 | -9.7 | -22.6..+0.0 | 0.25 |
|
||||
| JevK5 v0.2, drop-in | cicada-w2 | rotations | 19/31 | 16/31 | 0 | 3 | -9.7 | -19.4..+0.0 | 0.25 |
|
||||
| JevK5 v0.2, drop-in | pooled | cand-single-vs-semif-rotations | 232/259 | 232/259 | 10 | 10 | +0.0 | -3.2..+3.2 | 1.0 |
|
||||
| JevK5 v0.2, drop-in | wyrd | cand-single-vs-semif-rotations | 71/84 | 72/84 | 3 | 2 | +1.2 | -3.6..+6.0 | 1.0 |
|
||||
| JevK5 v0.3, native | pooled | single | 217/259 | 227/259 | 25 | 15 | +3.9 | -0.9..+8.5 | 0.1539 |
|
||||
| JevK5 v0.3, native | pooled | rotations | 232/259 | 226/259 | 9 | 15 | -2.3 | -6.0..+1.4 | 0.3075 |
|
||||
| JevK5 v0.3, native | authored144 | single | 116/144 | 135/144 | 24 | 5 | +13.2 | +6.9..+19.4 | 0.0005 |
|
||||
| JevK5 v0.3, native | authored144 | rotations | 131/144 | 135/144 | 7 | 3 | +2.8 | -1.4..+7.6 | 0.3438 |
|
||||
| JevK5 v0.3, native | cicada-w1 | single | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
|
||||
| JevK5 v0.3, native | cicada-w1 | rotations | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
|
||||
| JevK5 v0.3, native | wyrd | single | 71/84 | 66/84 | 1 | 6 | -6.0 | -10.7..-1.2 | 0.125 |
|
||||
| JevK5 v0.3, native | wyrd | rotations | 71/84 | 65/84 | 2 | 8 | -7.1 | -13.1..-1.2 | 0.1094 |
|
||||
| JevK5 v0.3, native | perturbations108 | single | 83/108 | 101/108 | 18 | 0 | +16.7 | +8.3..+25.9 | 0.0 |
|
||||
| JevK5 v0.3, native | perturbations108 | rotations | 93/108 | 101/108 | 8 | 0 | +7.4 | +1.9..+13.9 | 0.0078 |
|
||||
| JevK5 v0.3, native | cicada-w2 | single | 20/31 | 17/31 | 0 | 3 | -9.7 | -19.4..+0.0 | 0.25 |
|
||||
| JevK5 v0.3, native | cicada-w2 | rotations | 19/31 | 16/31 | 0 | 3 | -9.7 | -19.4..+0.0 | 0.25 |
|
||||
| JevK5 v0.3, native | pooled | cand-single-vs-semif-rotations | 232/259 | 227/259 | 11 | 16 | -1.9 | -5.5..+1.5 | 0.4421 |
|
||||
| JevK5 v0.3, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 66/84 | 2 | 7 | -6.0 | -11.9..+0.0 | 0.1797 |
|
||||
| JevK5 v0.3, drop-in | pooled | single | 217/259 | 227/259 | 24 | 14 | +3.9 | -0.8..+8.2 | 0.1433 |
|
||||
| JevK5 v0.3, drop-in | pooled | rotations | 232/259 | 227/259 | 10 | 15 | -1.9 | -5.5..+1.8 | 0.4244 |
|
||||
| JevK5 v0.3, drop-in | authored144 | single | 116/144 | 132/144 | 22 | 6 | +11.1 | +4.9..+17.4 | 0.0037 |
|
||||
| JevK5 v0.3, drop-in | authored144 | rotations | 131/144 | 135/144 | 8 | 4 | +2.8 | -1.4..+7.6 | 0.3877 |
|
||||
| JevK5 v0.3, drop-in | cicada-w1 | single | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
|
||||
| JevK5 v0.3, drop-in | cicada-w1 | rotations | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
|
||||
| JevK5 v0.3, drop-in | wyrd | single | 71/84 | 69/84 | 2 | 4 | -2.4 | -7.1..+2.4 | 0.6875 |
|
||||
| JevK5 v0.3, drop-in | wyrd | rotations | 71/84 | 66/84 | 2 | 7 | -6.0 | -11.9..+0.0 | 0.1797 |
|
||||
| JevK5 v0.3, drop-in | perturbations108 | single | 83/108 | 99/108 | 17 | 1 | +14.8 | +6.5..+24.1 | 0.0001 |
|
||||
| JevK5 v0.3, drop-in | perturbations108 | rotations | 93/108 | 101/108 | 8 | 0 | +7.4 | +1.9..+13.9 | 0.0078 |
|
||||
| JevK5 v0.3, drop-in | cicada-w2 | single | 20/31 | 21/31 | 1 | 0 | +3.2 | +0.0..+9.7 | 1.0 |
|
||||
| JevK5 v0.3, drop-in | cicada-w2 | rotations | 19/31 | 19/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| JevK5 v0.3, drop-in | pooled | cand-single-vs-semif-rotations | 232/259 | 227/259 | 10 | 15 | -1.9 | -5.2..+1.2 | 0.4244 |
|
||||
| JevK5 v0.3, drop-in | wyrd | cand-single-vs-semif-rotations | 71/84 | 69/84 | 3 | 5 | -2.4 | -8.3..+3.6 | 0.7266 |
|
||||
|
||||
#### Noise floors
|
||||
|
||||
| system | A-vs-A in process (authored144): flips, max Δp | across restarts, single (560 rows/pair): flips per pair, max Δp | across restarts, rotations: flips per pair | labelled rows whose top moved in ANY restart pair, single: pooled /259 · perturb /108 · Cicada w2 /31 |
|
||||
|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0, 0, 0, 0 | 0 · 0 · 0 (4 repeats) |
|
||||
| Plumb-4B, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 2 (Δp≤0.043), 2 (Δp≤0.043) | 0, 2, 2 | 2 · 0 · 0 (3 repeats) |
|
||||
| Plumb-4B, drop-in | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0 | 0 · 0 · 0 (3 repeats) |
|
||||
| Imajev-4B, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0, 0, 0, 0 | 0 · 0 · 0 (4 repeats) |
|
||||
| Intern-Decision-4B, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0, 0, 0, 0 | 0 · 0 · 0 (4 repeats) |
|
||||
| Intern-Decision-4B, drop-in | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.128), 0 (Δp≤0.128), 0 (Δp≤0.000) | 0, 0, 0 | 0 · 0 · 0 (3 repeats) |
|
||||
| JevK5 v0.2, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 3 (Δp≤0.036), 3 (Δp≤0.036) | 0, 2, 2 | 1 · 1 · 0 (3 repeats) |
|
||||
| JevK5 v0.2, drop-in | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0 | 0 · 0 · 0 (3 repeats) |
|
||||
| JevK5 v0.3, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 1 (Δp≤0.051), 1 (Δp≤0.051), 0 (Δp≤0.000) | 1, 1, 0 | 0 · 1 · 0 (3 repeats) |
|
||||
| JevK5 v0.3, drop-in | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0 | 0 · 0 · 0 (3 repeats) |
|
||||
|
||||
#### Order sensitivity (authored144, single ordering vs the same request reordered) and the negative control
|
||||
|
||||
| system | reversed: label changes /144 | reversed: max Δp | shuffled: label changes | shuffled: max Δp | NEG: same top as unrotated /144 | NEG: follows the description | NEG: right vs original gold |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 30 | 0.89 | 29 | 0.89 | 14 | 119 | 17 |
|
||||
| Plumb-4B, native | 15 | 0.54 (0.54–0.56) | 10 | 0.44 (0.44–0.45) | 24 (24–25) | 106 (105–106) | 27 (27–28) |
|
||||
| Plumb-4B, drop-in | 12 | 0.73 | 8 | 0.67 | 3 | 126 | 12 |
|
||||
| Imajev-4B, native | 3 | 0.72 | 4 | 0.68 | 4 | 135 | 5 |
|
||||
| Intern-Decision-4B, native | 8 | 0.60 | 5 | 0.46 | 10 | 122 | 14 |
|
||||
| Intern-Decision-4B, drop-in | 10 | 0.98 | 12 (11–12) | 0.94 (0.93–0.94) | 5 | 127 | 10 |
|
||||
| JevK5 v0.2, native | 13 (12–13) | 0.72 | 10 | 0.64 (0.63–0.64) | 20 | 114 | 21 |
|
||||
| JevK5 v0.2, drop-in | 15 | 0.70 | 12 | 0.70 | 4 | 128 | 9 |
|
||||
| JevK5 v0.3, native | 8 (7–8) | 0.61 | 11 | 0.49 | 20 | 118 | 18 |
|
||||
| JevK5 v0.3, drop-in | 4 | 0.53 | 5 | 0.51 | 4 | 133 | 5 |
|
||||
|
||||
#### Latency at our shape (loopback on fv-ml1, GPU 3; ms; median of all requests, run-median range)
|
||||
|
||||
| system | short1 e2e | crit21 e2e | crit21 server | long1 e2e | long16 e2e | long16 server | tokens crit21 / long16 |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 37 (37–38) | 131 (131–132) | 119 (119–120) | 186 (186–187) | 505 (502–519) | 486 (482–500) | 62 / 3900 |
|
||||
| Plumb-4B, native | 16 | 294 (292–295) | 292 (290–293) | 215 (212–215) | 3378 (3374–3384) | 3376 (3372–3382) | 2489 / 63302 |
|
||||
| Plumb-4B, drop-in | 37 | 131 (131–132) | 120 (119–120) | 186 | 506 (506–507) | 487 (487–488) | 62 / 3900 |
|
||||
| Imajev-4B, native | 60 (60–61) | 1219 (1212–1230) | 1212 (1204–1223) | 315 (310–315) | 5125 (5054–5140) | 5121 (5049–5135) | 374 / 7925 |
|
||||
| Intern-Decision-4B, native | 39 | 88 (88–89) | 84 (84–85) | 190 (189–191) | 215 (213–216) | 213 (211–214) | 1083 / 4579 |
|
||||
| Intern-Decision-4B, drop-in | 38 (37–38) | 132 | 120 | 186 (186–188) | 508 (500–516) | 488 (481–497) | 62 / 3900 |
|
||||
| JevK5 v0.2, native | 16 | 291 (291–292) | 289 (289–290) | 212 | 3360 (3358–3368) | 3358 (3356–3366) | 2489 / 63302 |
|
||||
| JevK5 v0.2, drop-in | 38 | 132 | 120 (119–120) | 187 (185–188) | 518 (505–519) | 499 (486–500) | 62 / 3900 |
|
||||
| JevK5 v0.3, native | 16 | 291 (291–292) | 289 (289–290) | 212 | 3363 (3355–3370) | 3361 (3353–3368) | 2489 / 63302 |
|
||||
| JevK5 v0.3, drop-in | 38 (37–39) | 133 (131–133) | 120 (119–121) | 187 (186–188) | 518 (506–521) | 498 (487–501) | 62 / 3900 |
|
||||
|
||||
#### VRAM (nvidia-smi, whole GPU 3, only our process on it) and capacity under the 12 GiB cap
|
||||
|
||||
| system | rest after warm-up MiB | peak in sets+shapes MiB | peak in capacity sweep MiB | max rows @ short state | max rows @ ~3,900 tok |
|
||||
|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 9242 (9242–9250) | 12918 (12886–12918) | 12920 | 64 (next 96: 422) | 16 (next 18: 503) |
|
||||
| Plumb-4B, native | 12064 | 12064 | 12064 | 128 (no failure up to the last step) | 64 (no failure up to the last step) |
|
||||
| Plumb-4B, drop-in | 9242 | 12918 | 12920 | 64 (next 96: 422) | 16 (next 18: 503) |
|
||||
| Imajev-4B, native | 10404 | 11056 | 11056 | 128 (no failure up to the last step) | 64 (no failure up to the last step) |
|
||||
| Intern-Decision-4B, native | 9736 | 10290 | 10290 | 128 (no failure up to the last step) | 64 (no failure up to the last step) |
|
||||
| Intern-Decision-4B, drop-in | 9242 | 12918 | 12920 | 64 (next 96: 422) | 16 (next 18: 503) |
|
||||
| JevK5 v0.2, native | 12064 | 12064 | 12064 | 128 (no failure up to the last step) | 64 (no failure up to the last step) |
|
||||
| JevK5 v0.2, drop-in | 9242 | 12918 | – | – | – |
|
||||
| JevK5 v0.3, native | 12064 | 12064 | – | – | – |
|
||||
| JevK5 v0.3, drop-in | 9242 | 12918 | – | – | – |
|
||||
|
||||
### Reading the tables
|
||||
|
||||
- **Fit.** For every served system, "peak" is nvidia-smi for the whole card while only our process was on it, so it includes the CUDA context (~0.6 GB) that sits outside torch's 12 GiB fraction. SemIf's own peak of 12.9 GB is the same shape the live service shows. The JevK5 runtime (Plumb, JevK5 native) captures its CUDA graphs at startup and holds **12.06 GB from the first second**. It fits the cap, but unlike semif-serve it never hands a burst back to the neighbours on GPU 1.
|
||||
- **Capacity.** The native runtimes chunk: Intern-Decision takes 16 questions per request and Imajev 8 (`jev_api.MAX_QUESTIONS`). The JevK5 runtime runs questions one by one. "No failure up to the last step" means 64 criteria over the ~3,900-token state and 128 over the short state were answered under the cap, in several requests where the runtime chunks. semif-serve's limit is SemIf's replicated prefix cache: 16 rows at ~3,900 tokens, exactly the README's figure.
|
||||
- **Order sensitivity and the agreement signal.** Every candidate is far less position-biased than SemIf: reversing the options changes 3–15 of 144 labels against SemIf's 30. So rotations buy them little, and in several cases nothing. With rotations, SemIf's rows split between orderings 90 times out of 252 (unanimous rows 94.4% right, split rows 78.9%); Intern-Decision's split only 16 times.
|
||||
- **Negative control.** Every system fails it as it must: on 106–135 of 144 rows the top moves to the option that now carries the right description. The native formats show the option **id** next to its description (`id: description`, or `A = id: description`), which gives a model a second, now contradictory, cue. That is why Plumb and JevK5 native keep their unrotated answer on 20–24 rows against SemIf's 14. Through SemIf's prompt, which shows only descriptions, the same weights keep it on 3–5.
|
||||
- **Null control.** Over a content-free state, Cicada falls to the base rate (15–16 of 31 = always "ordinary") for every system, and Wyrd to 48–60 of 84. Every evidence score above is read against that.
|
||||
- **Wyrd as one request per turn.** Asking the 4 decisions of a turn together changes nothing for SemIf (shared prefix, 71) or for the runtimes that loop over questions. Intern-Decision, which puts all 4 in one prompt, drops from 79 to 77, still above SemIf's 71.
|
||||
|
||||
## What could not be measured, and why
|
||||
|
||||
- **Held-out and sealed JevBench.** 109 hard items plus 308 sealed items stay with the evaluator, so our JevBench numbers are public-item numbers. For the board's rows the sealed scores exist (SemIf 0.263; Plumb 0.380, Imajev 0.370, JevK5 v0.2 0.331). They point the same way as our private sets: every candidate is ahead on public items by far more than on unseen ones. Intern-Decision has **no** sealed number.
|
||||
- **Contamination of authored144 / perturbations108.** Both have been public on GitHub (SemIf) since mid-September. No candidate's repo mentions them as training data (searched at the pinned commits), but that cannot be verified. Cicada and Wyrd were written in this repo on 09-27 and are the only sets no model could have seen. They are also small (31 and 84 rows, one labeller), and their winning wordings were tuned on SemIf, which favours SemIf.
|
||||
- **Hopper**: not run (research-and-demo licence; see Candidates). **Imajev drop-in**: not applicable (LoRA plus its own readout head). **Intern-Decision-2B/0.8B, AlexWortega/openjev**: dropped by brokkr's revised order.
|
||||
- **Contention on GPU 1.** Everything ran alone on the empty GPU 3. Next to vllm-coder, the erp/meromero seats and scriberr on GPU 1, latency would be worse for all systems alike. The 09-27 Cicada spike measured SemIf going 136 → 196 ms under a concurrent decode.
|
||||
- **Latency from a caller.** The loopback numbers carry no network. One bridge run from nh3-dev to SemIf on GPU 3 gave 21 criteria in **151.5 ms** end to end (runs 150.1–151.8; server 119.3 ms; ping 22.5 ms avg), against the README's 159 ms. The latency instrument agrees with the reference (`raw/shape-semif-bridge-from-nh3dev.json`).
|
||||
- **The README's "a bf16 near-tie can flip across restarts".** Not reproduced: SemIf flipped 0 of 560 rows across 4 restarts, bit-identical probabilities. Only the CUDA-graph runtime (Plumb, JevK5 native: 1–3 rows per pair, Δp ≤ 0.05) and the Intern-Decision drop-in (0 flips, Δp up to 0.128, one bf16 step) moved at all.
|
||||
|
||||
## Exact commands
|
||||
|
||||
Everything ran from `/tmp/jevbench-2026-09-30` on fv-ml1 (removed afterwards); `code/` here is the same tree.
|
||||
|
||||
```bash
|
||||
# weights, as llmuser, public repos (no token on fv-ml1); pins in code/pull.py, code/pull2.py
|
||||
docker run -d --name jevbench-pull --user 1001:1001 -e HF_HOME=/hf -e HF_HUB_OFFLINE=0 \
|
||||
-v /tank/aimodels/huggingface:/hf -v $W/code:/code:ro --entrypoint /app/.venv/bin/python semif-serve:0.1.4 /code/pull.py
|
||||
# native image (removed afterwards)
|
||||
docker build -f code/Dockerfile.native -t jevbench-native:2026-09-30 code/
|
||||
# the ~3,900-token state
|
||||
docker run --rm -e HF_HOME=/hf -v /tank/aimodels/huggingface:/hf:ro -v $W:/w --entrypoint /app/.venv/bin/python \
|
||||
semif-serve:0.1.4 /w/code/make_long_state.py /w/src/jevbench-v1.2.16/datasets/public/hard.jsonl /w/out/long_state.txt
|
||||
# one repeat of one system = one line (code/run.sh); the full matrix was code/queue.sh over these lines:
|
||||
code/run.sh semif-qwen35-4b semif Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a <r> [capacity]
|
||||
code/run.sh plumb-4b-native native jevk5:crh225/plumb-4b@55de037801a8a9b9de3db5c0e16cef86210c2186 <r>
|
||||
code/run.sh plumb-4b-dropin semif crh225/plumb-4b@55de037801a8a9b9de3db5c0e16cef86210c2186 <r>
|
||||
code/run.sh imajev-4b-native native imajev:mohit67890/imajev-4b@c9e5f132465da85d31735ec502d5557982671a7d <r>
|
||||
code/run.sh intern-decision-4b-native native intern:internlm/Intern-Decision-4B@0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd <r>
|
||||
code/run.sh intern-decision-4b-dropin semif internlm/Intern-Decision-4B@0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd <r>
|
||||
code/run.sh jevk5-v02-native native jevk5:alibiserikbay/JevK5@27d2d6b8d4714807f6293b0623bd7370b27e42f8 <r>
|
||||
code/run.sh jevk5-v02-dropin semif alibiserikbay/JevK5@27d2d6b8d4714807f6293b0623bd7370b27e42f8 <r>
|
||||
code/run.sh jevk5-v03-native native jevk5v03:alibiserikbay/JevK5@c4f7fdb3aeab5582336406e78d3bef11bf98833d <r>
|
||||
code/run.sh jevk5-v03-dropin semif alibiserikbay/JevK5@c4f7fdb3aeab5582336406e78d3bef11bf98833d <r>
|
||||
# inside run.sh, per repeat: JevBench (semif_direct in-process, or typesafe to the loopback server) ->
|
||||
# bench_sets.py (single, rotations, repeat, reversed, shuffled, negative, multifield) -> bench_shape.py
|
||||
# analysis (on nh3-dev):
|
||||
python3 code/analyze.py raw/out raw/jevbench-public-v1.2.16 > summary.json && python3 code/tables.py summary.json > tables.md
|
||||
```
|
||||
|
||||
Sources used, pinned: JevBench `5e95f23c` (tag v1.2.16), jevk5 runtime `85238d7b` (v0.2.0) and `f944fe37` (v0.3.3), imajev server `a0134749`, SemIf `23cf1f39` (inside semif-serve:0.1.4). Intern-Decision's `inference.py` is the one inside its pinned HF snapshot.
|
||||
|
||||
## Host changes (fv-ml1; all in `scripts/ops-log`)
|
||||
|
||||
| when (PDT) | change | state now |
|
||||
|---|---|---|
|
||||
| 0141 | pulled plumb-4b, JevK5 v0.2, Intern-Decision-4B, imajev-4b to `/tank/aimodels/huggingface` as `llmuser`; sha256 checked against the Hub | **kept** (7.9 + 8.4 + 8.5 + 0.5 GB) |
|
||||
| 0209 | pulled JevK5 v0.3 (`c4f7fdb3`) | **kept** (8.4 GB; the JevK5 dir totals 16 GB) |
|
||||
| 0147 | built `jevbench-native:2026-09-30` | removed 0456 |
|
||||
| 0149–0455 | transient containers `jevbench-pull`, `jevbench-jb`, `jevbench-serve` (0.0.0.0:18032, bearer-protected), `jevbench-native` (127.0.0.1:18090), GPU 3 only, one at a time | all removed; GPU 3 at 2 MiB, no compute apps (0456) |
|
||||
| 0141–0456 | scratch `/tmp/jevbench-2026-09-30` | removed (`sudo -n rm -r`, because the containers wrote some files as uid 10001) |
|
||||
|
||||
Nothing touched GPU 1, the `semif` stack, its config, any alias, LiteLLM, or any seat. Nothing was committed and nothing was sent over althing.
|
||||
@@ -0,0 +1,10 @@
|
||||
# Transient bench image (2026-09-30): semif-serve:0.1.4's exact torch 2.10.0+cu128 / transformers 5.17.0 /
|
||||
# fla 0.5.2 / causal-conv1d 1.7.0 stack, plus what the candidates' own runtimes need. Removed after the bench.
|
||||
FROM semif-serve:0.1.4
|
||||
USER root
|
||||
RUN uv pip install --python /app/.venv/bin/python --no-cache --no-deps \
|
||||
"jevk5 @ git+https://github.com/allebee/jevk5@85238d7be5527370c43206fe54cd752eb3134c1b" \
|
||||
&& uv pip install --python /app/.venv/bin/python --no-cache peft pillow python-multipart \
|
||||
&& uv pip install --python /app/.venv/bin/python --no-cache --no-deps \
|
||||
--index-url https://download.pytorch.org/whl/cu128 "torchvision==0.25.0"
|
||||
USER semif
|
||||
@@ -0,0 +1,360 @@
|
||||
"""Jev-candidate bench (2026-09-30): turn the raw runs into the tables of the results doc.
|
||||
|
||||
python3 analyze.py <out dir> > summary.json (and prints markdown tables to stderr)
|
||||
|
||||
Every cell is a median over the repeats (one repeat = one process lifetime) with min..max.
|
||||
The noise floors come from the same data:
|
||||
* A-vs-A inside a process: authored144 "single" vs "repeat"
|
||||
* across restarts: every pair of repeats' "single" rows (label flips, max |dp|)
|
||||
Paired candidate-vs-SemIf comparisons use each row's majority top over the repeats (ties
|
||||
never arise with 3 repeats and a strict majority rule; a row without a majority keeps r1).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import csv
|
||||
import itertools
|
||||
import json
|
||||
import math
|
||||
import random
|
||||
import statistics as st
|
||||
import sys
|
||||
from collections import Counter, defaultdict
|
||||
from pathlib import Path
|
||||
|
||||
OUT = Path(sys.argv[1])
|
||||
JB = OUT.parent / "src/jevbench-v1.2.16/datasets/public"
|
||||
if not JB.exists():
|
||||
JB = Path(sys.argv[2]) if len(sys.argv) > 2 else JB
|
||||
TIERS = {}
|
||||
for t in ("easy", "original", "hard"):
|
||||
for line in (JB / f"{t}.jsonl").read_text().splitlines():
|
||||
if line.strip():
|
||||
TIERS[json.loads(line)["id"]] = t
|
||||
|
||||
|
||||
def med(xs):
|
||||
xs = [x for x in xs if x is not None]
|
||||
if not xs:
|
||||
return None
|
||||
return {"median": st.median(xs), "min": min(xs), "max": max(xs), "n": len(xs)}
|
||||
|
||||
|
||||
def load_runs(label):
|
||||
runs = {}
|
||||
for d in sorted((OUT / label).glob("r*")):
|
||||
r = {"dir": d}
|
||||
for name in ("sets", "shape"):
|
||||
p = d / f"{name}.json"
|
||||
r[name] = json.loads(p.read_text()) if p.exists() else None
|
||||
p = d / "jevbench/results.jsonl"
|
||||
r["jb"] = [json.loads(l) for l in p.read_text().splitlines() if l.strip()] if p.exists() else None
|
||||
r["vram"] = []
|
||||
p = d / "vram.csv"
|
||||
if p.exists():
|
||||
for row in csv.reader(p.read_text().splitlines()):
|
||||
try:
|
||||
r["vram"].append((float(row[0]), float(row[1])))
|
||||
except (ValueError, IndexError):
|
||||
pass
|
||||
r["events"] = {}
|
||||
p = d / "events.txt"
|
||||
if p.exists():
|
||||
for line in p.read_text().splitlines():
|
||||
ts, name = line.split(" ", 1)
|
||||
r["events"][name] = float(ts)
|
||||
p = d / "up_at.txt"
|
||||
if p.exists():
|
||||
r["events"]["up"] = float(p.read_text().strip())
|
||||
runs[d.name] = r
|
||||
return runs
|
||||
|
||||
|
||||
def jevbench(runs):
|
||||
per = defaultdict(list)
|
||||
for r in runs.values():
|
||||
if not r["jb"]:
|
||||
continue
|
||||
c, n = Counter(), Counter()
|
||||
for x in r["jb"]:
|
||||
t = TIERS[x["task_id"]]
|
||||
n[t] += 1
|
||||
c[t] += bool(x["correct"])
|
||||
for t in ("easy", "original", "hard"):
|
||||
per[t].append(c[t] / n[t] if n[t] else None)
|
||||
per[t + "_n"].append(c[t])
|
||||
per["all"].append(sum(c.values()) / sum(n.values()))
|
||||
per["all_n"].append(sum(c.values()))
|
||||
per["failed"].append(sum(not x["ok"] for x in r["jb"]))
|
||||
lat = [x["latency_s"] for x in r["jb"] if x.get("latency_s") is not None]
|
||||
per["p50_ms"].append(st.median(lat) * 1000 if lat else None)
|
||||
return {k: med(v) for k, v in per.items()}
|
||||
|
||||
|
||||
def rows(run, cond):
|
||||
s = run["sets"]
|
||||
return {x["id"]: x for x in s["conditions"][cond]["rows"]} if s and cond in s["conditions"] else {}
|
||||
|
||||
|
||||
SETS = ("authored144", "perturbations108", "cicada-w1", "cicada-w2", "wyrd")
|
||||
POOLED = ("authored144", "cicada-w1", "wyrd") # the brief's three replaced-baseline sets, 259 rows
|
||||
|
||||
|
||||
def acc(rs, key=None):
|
||||
lab = [x for x in rs if x["ok"] and x["gold"]]
|
||||
if key:
|
||||
lab = [x for x in lab if key(x)]
|
||||
return (sum(x["top"] in x["gold"] for x in lab), len(lab))
|
||||
|
||||
|
||||
def set_table(runs):
|
||||
out = {}
|
||||
for cond in ("single", "rotations", "multifield"):
|
||||
for name in ("pooled",) + SETS + ("wyrd:place", "wyrd:place2", "wyrd:exit", "wyrd:exit2"):
|
||||
base, _, dec = name.partition(":")
|
||||
wanted = set(POOLED) if base == "pooled" else {base}
|
||||
vals, blind, ctrl, fails = [], [], [], []
|
||||
for r in runs.values():
|
||||
rr = list(rows(r, cond).values())
|
||||
if not rr:
|
||||
continue
|
||||
sub = [x for x in rr if x["set"] in wanted and (not dec or x["decision"] == dec)]
|
||||
c, n = acc([x for x in sub if x["cond"] == "evidence"])
|
||||
vals.append((c, n))
|
||||
b = acc([x for x in sub if x["cond"] == "blind"])
|
||||
blind.append(b)
|
||||
k = acc([x for x in sub if x["cond"] == "evidence" and x["tag"] == "control"])
|
||||
ctrl.append(k)
|
||||
fails.append(sum(not x["ok"] for x in sub))
|
||||
if vals and vals[0][1]:
|
||||
out[f"{cond}/{name}"] = {
|
||||
"correct": med([c for c, _ in vals]), "n": vals[0][1],
|
||||
"acc": med([c / n for c, n in vals if n]),
|
||||
"blind_correct": med([c for c, _ in blind]) if blind and blind[0][1] else None,
|
||||
"controls": med([c for c, _ in ctrl]) if ctrl and ctrl[0][1] else None,
|
||||
"controls_n": ctrl[0][1] if ctrl else None,
|
||||
"failures": med(fails)}
|
||||
# agreement signal (rotations): accuracy of unanimous vs split rows, authored144+perturbations108
|
||||
for r in runs.values():
|
||||
rr = [x for x in rows(r, "rotations").values() if x["set"] in ("authored144", "perturbations108") and x["ok"]]
|
||||
if rr:
|
||||
un = [x for x in rr if x.get("agreement") == 1.0]
|
||||
sp = [x for x in rr if x.get("agreement") is not None and x["agreement"] < 1.0]
|
||||
out.setdefault("agreement_signal", []).append({"unanimous": acc(un), "split": acc(sp)})
|
||||
return out
|
||||
|
||||
|
||||
def top_of(x):
|
||||
return x["top"] if x.get("ok") else None
|
||||
|
||||
|
||||
def pdiff(a, b):
|
||||
pa = dict(zip(a["option_ids"], a["probs"]))
|
||||
pb = dict(zip(b["option_ids"], b["probs"]))
|
||||
return max(abs(pa[i] - pb[i]) for i in pa)
|
||||
|
||||
|
||||
def floors(runs):
|
||||
out = {}
|
||||
# A-vs-A inside the process
|
||||
aa = []
|
||||
for r in runs.values():
|
||||
s, rep = rows(r, "single"), rows(r, "repeat")
|
||||
ids = [i for i in rep if i in s and s[i]["ok"] and rep[i]["ok"]]
|
||||
if ids:
|
||||
aa.append({"flips": sum(s[i]["top"] != rep[i]["top"] for i in ids),
|
||||
"max_dp": max(pdiff(s[i], rep[i]) for i in ids), "n": len(ids)})
|
||||
out["a_vs_a_in_process"] = aa
|
||||
# across restarts, every row of "single" and of "rotations"
|
||||
keys = sorted(runs)
|
||||
for cond in ("single", "rotations"):
|
||||
pairs = []
|
||||
for a, b in itertools.combinations(keys, 2):
|
||||
ra, rb = rows(runs[a], cond), rows(runs[b], cond)
|
||||
ids = [i for i in ra if i in rb and ra[i]["ok"] and rb[i]["ok"]]
|
||||
if ids:
|
||||
flips = [i for i in ids if ra[i]["top"] != rb[i]["top"]]
|
||||
pairs.append({"pair": f"{a}-{b}", "n": len(ids), "flips": len(flips),
|
||||
"flipped_ids": flips[:10], "max_dp": max(pdiff(ra[i], rb[i]) for i in ids)})
|
||||
out[f"cross_restart/{cond}"] = pairs
|
||||
# rows whose top differs between ANY two repeats, per set (labelled evidence rows only)
|
||||
tops = defaultdict(set)
|
||||
meta = {}
|
||||
for k in keys:
|
||||
for i, x in rows(runs[k], cond).items():
|
||||
if x["ok"]:
|
||||
tops[i].add(x["top"])
|
||||
meta[i] = x
|
||||
unstable = [i for i, t in tops.items() if len(t) > 1]
|
||||
out[f"unstable_rows/{cond}"] = {
|
||||
s: sum(1 for i in unstable if meta[i]["set"] == s and meta[i]["cond"] == "evidence" and meta[i]["gold"])
|
||||
for s in SETS} | {"pooled": sum(1 for i in unstable if meta[i]["set"] in POOLED
|
||||
and meta[i]["cond"] == "evidence" and meta[i]["gold"]),
|
||||
"repeats": len(keys)}
|
||||
return out
|
||||
|
||||
|
||||
def order_sensitivity(runs):
|
||||
out = defaultdict(list)
|
||||
for r in runs.values():
|
||||
s = rows(r, "single")
|
||||
for cond in ("reversed", "shuffled"):
|
||||
o = rows(r, cond)
|
||||
ids = [i for i in o if i in s and s[i]["ok"] and o[i]["ok"]]
|
||||
if ids:
|
||||
out[cond].append({"n": len(ids), "label_changes": sum(s[i]["top"] != o[i]["top"] for i in ids),
|
||||
"max_dp": max(pdiff(s[i], o[i]) for i in ids),
|
||||
"median_max_dp": st.median(pdiff(s[i], o[i]) for i in ids),
|
||||
"acc": acc([o[i] for i in ids])})
|
||||
return {k: {"label_changes": med([x["label_changes"] for x in v]), "n": v[0]["n"],
|
||||
"max_dp": med([x["max_dp"] for x in v]), "median_row_max_dp": med([x["median_max_dp"] for x in v]),
|
||||
"acc": med([x["acc"][0] for x in v])} for k, v in out.items()}
|
||||
|
||||
|
||||
def negative(runs):
|
||||
res = []
|
||||
for r in runs.values():
|
||||
s, neg = rows(r, "single"), rows(r, "negative")
|
||||
ids = [i for i in neg if i in s and s[i]["ok"] and neg[i]["ok"]]
|
||||
if ids:
|
||||
res.append({"n": len(ids),
|
||||
"same_top_as_unrotated": sum(neg[i]["top"] == s[i]["top"] for i in ids),
|
||||
"vs_original_gold": sum(neg[i]["top"] in neg[i]["gold"] for i in ids),
|
||||
"follows_description": sum(neg[i]["top"] == neg[i]["gold_desc_id"] for i in ids)})
|
||||
return {k: med([x[k] for x in res]) for k in ("same_top_as_unrotated", "vs_original_gold", "follows_description")} | (
|
||||
{"n": res[0]["n"]} if res else {})
|
||||
|
||||
|
||||
def majority_tops(runs, cond):
|
||||
votes = defaultdict(list)
|
||||
meta = {}
|
||||
for k in sorted(runs):
|
||||
for i, x in rows(runs[k], cond).items():
|
||||
votes[i].append(top_of(x))
|
||||
meta[i] = x
|
||||
out = {}
|
||||
for i, v in votes.items():
|
||||
c = Counter(t for t in v if t is not None).most_common()
|
||||
out[i] = c[0][0] if c and c[0][1] > len(v) / 2 else v[0]
|
||||
return out, meta
|
||||
|
||||
|
||||
def mcnemar_exact(b, c):
|
||||
n = b + c
|
||||
if n == 0:
|
||||
return 1.0
|
||||
k = min(b, c)
|
||||
p = sum(math.comb(n, i) for i in range(0, k + 1)) / 2 ** n
|
||||
return min(1.0, 2 * p)
|
||||
|
||||
|
||||
def paired(base_runs, cand_runs, cond, setname, cand_cond=None):
|
||||
bt, meta = majority_tops(base_runs, cond)
|
||||
ct, _ = majority_tops(cand_runs, cand_cond or cond)
|
||||
wanted = set(POOLED) if setname == "pooled" else {setname.split(":")[0]}
|
||||
ids = [i for i in bt if i in ct and meta[i]["set"] in wanted and meta[i]["cond"] == "evidence"
|
||||
and meta[i]["gold"] and (":" not in setname or meta[i]["decision"] == setname.split(":")[1])]
|
||||
if not ids:
|
||||
return None
|
||||
b_ok = {i: bt[i] in meta[i]["gold"] for i in ids}
|
||||
c_ok = {i: ct[i] in meta[i]["gold"] for i in ids}
|
||||
fixed = sum(c_ok[i] and not b_ok[i] for i in ids)
|
||||
broken = sum(b_ok[i] and not c_ok[i] for i in ids)
|
||||
groups = defaultdict(list)
|
||||
for i in ids:
|
||||
groups[meta[i]["group"]].append(i)
|
||||
rng, keys, deltas = random.Random(7), list(groups), []
|
||||
for _ in range(5000):
|
||||
sample = [i for g in (rng.choice(keys) for _ in keys) for i in groups[g]]
|
||||
deltas.append((sum(c_ok[i] for i in sample) - sum(b_ok[i] for i in sample)) / len(sample))
|
||||
deltas.sort()
|
||||
return {"n": len(ids), "semif": sum(b_ok.values()), "cand": sum(c_ok.values()), "fixed": fixed,
|
||||
"broken": broken, "mcnemar_p": round(mcnemar_exact(fixed, broken), 4),
|
||||
"delta_pts": round(100 * (sum(c_ok.values()) - sum(b_ok.values())) / len(ids), 1) if ids else None,
|
||||
"delta_95ci_pts": [round(100 * deltas[125], 1), round(100 * deltas[4875], 1)] if ids else None}
|
||||
|
||||
|
||||
def vram(runs):
|
||||
res = []
|
||||
for r in runs.values():
|
||||
ev, v = r["events"], r["vram"]
|
||||
if not v:
|
||||
continue
|
||||
rest = [m for t, m in v if "rest:start" in ev and ev["rest:start"] <= t <= ev.get("rest:end", 0)]
|
||||
if not rest and "up" in ev: # r1 of the SemIf baseline predates the rest window
|
||||
rest = [m for t, m in v if ev["up"] <= t <= ev["up"] + 1.5]
|
||||
serve0 = ev.get("serve:start", 0)
|
||||
sets0 = ev.get("sets:start", ev.get("up", serve0))
|
||||
shape = r["shape"]
|
||||
cap0 = None
|
||||
if shape:
|
||||
cap_ev = [t for name, t in shape["events"] if name.startswith("capacity:")]
|
||||
cap0 = min(cap_ev) if cap_ev else None
|
||||
work = [m for t, m in v if t >= sets0 and (cap0 is None or t < cap0) and t <= ev.get("serve:end", 1e18)]
|
||||
capw = [m for t, m in v if cap0 is not None and t >= cap0 and t <= ev.get("serve:end", 1e18)]
|
||||
jb = [m for t, m in v if ev.get("jevbench:start", 1e18) <= t <= ev.get("jevbench:end", 0)]
|
||||
res.append({"rest_mib": st.median(rest) if rest else None, "peak_work_mib": max(work) if work else None,
|
||||
"peak_capacity_mib": max(capw) if capw else None, "peak_jevbench_mib": max(jb) if jb else None})
|
||||
return {k: med([x[k] for x in res]) for k in ("rest_mib", "peak_work_mib", "peak_capacity_mib", "peak_jevbench_mib")}
|
||||
|
||||
|
||||
def latency(runs):
|
||||
out = defaultdict(lambda: defaultdict(list))
|
||||
cap = {}
|
||||
for k, r in sorted(runs.items()):
|
||||
sh = r["shape"]
|
||||
if not sh:
|
||||
continue
|
||||
for name, s in sh["shapes"].items():
|
||||
out[name]["e2e"].append(s["median_e2e_ms"])
|
||||
out[name]["server"].append(s["median_server_ms"])
|
||||
out[name]["run_medians"].extend(s["run_medians_e2e_ms"])
|
||||
out[name]["tokens"].append(s["tokens"])
|
||||
out[name]["non_200"].append(s["non_200"])
|
||||
if sh.get("capacity"):
|
||||
cap[k] = {lab: [(x["rows"], x["status"]) for x in series] for lab, series in sh["capacity"].items()}
|
||||
return {name: {"e2e_ms": med(v["e2e"]), "server_ms": med(v["server"]),
|
||||
"run_medians_ms": [min(v["run_medians"]), max(v["run_medians"])],
|
||||
"tokens": v["tokens"][0], "non_200": sum(v["non_200"])} for name, v in out.items()}, cap
|
||||
|
||||
|
||||
def parity(runs):
|
||||
"""semif-format runs only: authored144 'single' against SemIf's committed torch predictions."""
|
||||
p = Path(__file__).resolve().parent / "data/direct-authored144.jsonl"
|
||||
ref = {x["id"]: x for x in map(json.loads, p.read_text().splitlines()) if x}
|
||||
res = []
|
||||
for r in runs.values():
|
||||
s = rows(r, "single")
|
||||
ids = [i for i in ref if i in s and s[i].get("prompt_sha256")]
|
||||
if not ids:
|
||||
continue
|
||||
top_ref = {i: ref[i]["option_ids"][max(range(len(ref[i]["probabilities"])), key=ref[i]["probabilities"].__getitem__)]
|
||||
for i in ids}
|
||||
res.append({"n": len(ids), "top_agree": sum(s[i]["top"] == top_ref[i] for i in ids),
|
||||
"prompt_sha_equal": sum(s[i]["prompt_sha256"] == ref[i]["prompt_sha256"] for i in ids),
|
||||
"max_dp": max(max(abs(a - b) for a, b in zip(s[i]["probs"], ref[i]["probabilities"])) for i in ids)})
|
||||
return res
|
||||
|
||||
|
||||
def main():
|
||||
labels = [p.name for p in sorted(OUT.iterdir()) if p.is_dir() and any(p.glob("r*/"))]
|
||||
allruns = {l: load_runs(l) for l in labels}
|
||||
summary = {}
|
||||
base = allruns.get("semif-qwen35-4b")
|
||||
for l, runs in allruns.items():
|
||||
lat, cap = latency(runs)
|
||||
s = {"repeats": sorted(runs), "jevbench": jevbench(runs), "sets": set_table(runs), "floors": floors(runs),
|
||||
"order": order_sensitivity(runs), "negative": negative(runs), "vram": vram(runs),
|
||||
"latency": lat, "capacity": cap, "parity_vs_upstream_semif": parity(runs)}
|
||||
if base and l != "semif-qwen35-4b":
|
||||
s["paired_vs_semif"] = {f"{cond}/{name}": paired(base, runs, cond, name)
|
||||
for cond in ("single", "rotations")
|
||||
for name in ("pooled",) + SETS + ("wyrd:place2", "wyrd:exit")}
|
||||
# the cost question: the candidate at ONE ordering against SemIf WITH rotations
|
||||
s["paired_vs_semif"].update({f"cand-single-vs-semif-rotations/{name}": paired(base, runs, "rotations", name, "single")
|
||||
for name in ("pooled",) + SETS})
|
||||
summary[l] = s
|
||||
json.dump(summary, sys.stdout, indent=1, default=str)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,303 @@
|
||||
"""Jev-candidate bench (2026-09-30): the replaced-baseline sets through one backend.
|
||||
|
||||
Stdlib only (urllib), so it runs on fv-ml1's host python against a loopback server.
|
||||
|
||||
Backends
|
||||
semif semif-serve (SemIf's own scorer and prompt). /decide; order averaging uses the
|
||||
service's own `"orderings": "rotations"` (what a switch through the contract gets).
|
||||
systemone a TypeSafe-style POST /v1/systemone server (a candidate's NATIVE runtime and prompt).
|
||||
Each row is one `choice` question, criteria {option id: description} in the row's
|
||||
order. Order averaging is done here: the n cyclic rotations are n separate requests,
|
||||
combined exactly as semif-serve does (per-ordering log p, mean per option id,
|
||||
renormalised; agreement = share of orderings whose top equals the combined top).
|
||||
|
||||
Sets (all labelled rows; see ../README in the results doc)
|
||||
authored144, perturbations108 SemIf's own labelled sets (evidence interpretation, 3 options)
|
||||
cicada-w1, cicada-w2 the 2026-09-27 Cicada affect-gate spike, both wordings, 35 cases
|
||||
wyrd the 2026-09-27 Wyrd scene-change spike, 21 cases x 4 decisions
|
||||
Spike rows are also asked over a content-free state (the null control, cond=blind).
|
||||
|
||||
Conditions (one process lifetime = one repeat r)
|
||||
single every row, caller's order
|
||||
rotations every row, n cyclic rotations averaged
|
||||
repeat authored144 again, caller's order (A-vs-A inside the process)
|
||||
reversed authored144, options reversed (order sensitivity)
|
||||
shuffled authored144, options shuffled, fixed seed (order sensitivity)
|
||||
negative authored144, descriptions rotated one place, ids fixed (NEGATIVE CONTROL)
|
||||
|
||||
python3 bench_sets.py --backend semif --url http://127.0.0.1:18032 --token-file T --out out.json
|
||||
python3 bench_sets.py --backend systemone --url http://127.0.0.1:18090 --out out.json
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import math
|
||||
import random
|
||||
import sys
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from pathlib import Path
|
||||
|
||||
DATA = Path(__file__).resolve().parent / "data"
|
||||
BLIND = "(No evidence is available for this turn.)"
|
||||
|
||||
|
||||
def load_items() -> list[dict]:
|
||||
items = []
|
||||
for name in ("authored144", "perturbations108"):
|
||||
for line in (DATA / f"{name}.jsonl").read_text().splitlines():
|
||||
if not line.strip():
|
||||
continue
|
||||
r = json.loads(line)
|
||||
items.append({"set": name, "id": r["id"], "cond": "evidence", "state": r["state"],
|
||||
"question": r["question"],
|
||||
"options": [{"id": o["id"], "description": o["description"]} for o in r["options"]],
|
||||
"gold": [r["options"][r["label"]]["id"]], "group": r["group_id"], "tag": "case",
|
||||
"decision": "evidence", "forbid": []})
|
||||
for f in sorted((DATA / "spike").glob("*.json")):
|
||||
scen = json.loads(f.read_text())
|
||||
name = scen["name"]
|
||||
setname = ("cicada-w1" if "-w1-" in name else "cicada-w2") if name.startswith("cicada") else "wyrd"
|
||||
for case in scen["cases"]:
|
||||
for d in scen["decisions"]:
|
||||
lab = case.get("labels", {}).get(d["id"])
|
||||
gold = None if lab is None else ([lab] if isinstance(lab, str) else list(lab))
|
||||
base = {"set": setname, "decision": d["id"], "question": d["question"],
|
||||
"options": d["options"], "gold": gold, "group": f"{name}/{case['id']}",
|
||||
"tag": case.get("tag", "case"), "forbid": case.get("forbid", {}).get(d["id"], [])}
|
||||
items.append({**base, "id": f"{name}/{case['id']}/{d['id']}", "cond": "evidence",
|
||||
"state": case["state"]})
|
||||
items.append({**base, "id": f"{name}/{case['id']}/{d['id']}#blind", "cond": "blind",
|
||||
"state": BLIND})
|
||||
return items
|
||||
|
||||
|
||||
def post(url: str, body: dict, headers: dict, timeout: float = 300) -> tuple[int, dict, float]:
|
||||
data = json.dumps(body).encode()
|
||||
req = urllib.request.Request(url, data=data, headers={"Content-Type": "application/json", **headers})
|
||||
t = time.perf_counter()
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=timeout) as resp:
|
||||
status, raw = resp.status, resp.read()
|
||||
except urllib.error.HTTPError as e:
|
||||
status, raw = e.code, e.read()
|
||||
ms = (time.perf_counter() - t) * 1000
|
||||
try:
|
||||
parsed = json.loads(raw)
|
||||
except ValueError:
|
||||
parsed = {"raw": raw[:500].decode(errors="replace")}
|
||||
return status, parsed, ms
|
||||
|
||||
|
||||
def combine(per_ordering: list[dict[str, float]], ids: list[str]) -> dict:
|
||||
"""semif-serve's method: mean of log p per option id, renormalised."""
|
||||
logs = {i: sum(math.log(max(p[i], 1e-300)) for p in per_ordering) / len(per_ordering) for i in ids}
|
||||
m = max(logs.values())
|
||||
w = {i: math.exp(v - m) for i, v in logs.items()}
|
||||
z = sum(w.values())
|
||||
probs = {i: w[i] / z for i in ids}
|
||||
top = max(ids, key=lambda i: probs[i])
|
||||
tops = [max(ids, key=lambda i: p[i]) for p in per_ordering]
|
||||
return {"probabilities": [probs[i] for i in ids], "top": top,
|
||||
"agreement": sum(t == top for t in tops) / len(tops), "orderings": len(per_ordering)}
|
||||
|
||||
|
||||
class Semif:
|
||||
def __init__(self, url: str, token: str):
|
||||
self.url, self.h = url.rstrip("/"), {"Authorization": f"Bearer {token}"}
|
||||
|
||||
def score(self, item: dict, options: list[dict], averaged: bool) -> dict:
|
||||
body = {"id": item["id"], "state": item["state"], "question": item["question"], "options": options}
|
||||
if averaged:
|
||||
body["orderings"] = "rotations"
|
||||
status, j, ms = post(f"{self.url}/decide", body, self.h)
|
||||
if status != 200:
|
||||
return {"ok": False, "status": status, "error": str(j)[:300], "e2e_ms": ms}
|
||||
ids = [o["id"] for o in options]
|
||||
if averaged:
|
||||
c = j["combined"]
|
||||
probs = dict(zip(j["option_ids"], c["probabilities"]))
|
||||
return {"ok": True, "e2e_ms": ms, "probs": [probs[i] for i in ids], "top": c["top"],
|
||||
"agreement": c["agreement"],
|
||||
"input_tokens": j["orderings"][0]["input_tokens"]}
|
||||
probs = dict(zip(j["option_ids"], j["probabilities"]))
|
||||
return {"ok": True, "e2e_ms": ms, "probs": [probs[i] for i in ids],
|
||||
"top": max(ids, key=lambda i: probs[i]), "option_logits": j.get("option_logits"),
|
||||
"prompt_sha256": j.get("prompt_sha256"), "input_tokens": j.get("input_tokens"),
|
||||
"server_ms": round(j.get("total_seconds", 0) * 1000, 2)}
|
||||
|
||||
def score_many(self, items: list[dict]) -> list[dict]:
|
||||
"""All decisions over one state in ONE /decide/shared (SemIf's shared-prefix path)."""
|
||||
body = {"state": items[0]["state"], "decisions": [
|
||||
{"id": i["id"], "question": i["question"], "options": i["options"]} for i in items]}
|
||||
status, j, ms = post(f"{self.url}/decide/shared", body, self.h)
|
||||
if status != 200:
|
||||
return [{"ok": False, "status": status, "error": str(j)[:300], "e2e_ms": ms} for _ in items]
|
||||
out = []
|
||||
for item, res in zip(items, j["results"]):
|
||||
p = dict(zip(res["option_ids"], res["probabilities"]))
|
||||
ids = [o["id"] for o in item["options"]]
|
||||
out.append({"ok": True, "e2e_ms": ms, "probs": [p[i] for i in ids], "top": max(ids, key=lambda i: p[i])})
|
||||
return out
|
||||
|
||||
|
||||
class SystemOne:
|
||||
def __init__(self, url: str):
|
||||
self.url = url.rstrip("/")
|
||||
|
||||
def _one(self, item: dict, options: list[dict]) -> dict:
|
||||
body = {"state": item["state"], "questions": {"q": {
|
||||
"type": "choice", "instructions": item["question"],
|
||||
"criteria": {o["id"]: o["description"] for o in options}}}}
|
||||
status, j, ms = post(f"{self.url}/v1/systemone", body, {})
|
||||
if status != 200:
|
||||
return {"ok": False, "status": status, "error": str(j)[:300], "e2e_ms": ms}
|
||||
a = j["answers"]["q"]
|
||||
p = a["probabilities"]
|
||||
missing = [o["id"] for o in options if o["id"] not in p]
|
||||
if missing:
|
||||
return {"ok": False, "status": status, "error": f"missing ids {missing}", "e2e_ms": ms}
|
||||
out = {"ok": True, "e2e_ms": ms, "p": {o["id"]: float(p[o["id"]]) for o in options},
|
||||
"input_tokens": (j.get("usage") or {}).get("input_tokens")}
|
||||
for k in ("unknown_probability", "abstained"):
|
||||
if k in a:
|
||||
out[k] = a[k]
|
||||
return out
|
||||
|
||||
def score_many(self, items: list[dict]) -> list[dict]:
|
||||
"""All decisions over one state as ONE request with several questions (the runtime's own
|
||||
multi-question mode; Intern-Decision scores them in one forward over one prompt)."""
|
||||
body = {"state": items[0]["state"], "questions": {
|
||||
it["decision"]: {"type": "choice", "instructions": it["question"],
|
||||
"criteria": {o["id"]: o["description"] for o in it["options"]}} for it in items}}
|
||||
status, j, ms = post(f"{self.url}/v1/systemone", body, {})
|
||||
if status != 200:
|
||||
return [{"ok": False, "status": status, "error": str(j)[:300], "e2e_ms": ms} for _ in items]
|
||||
out = []
|
||||
for it in items:
|
||||
p = j["answers"][it["decision"]]["probabilities"]
|
||||
ids = [o["id"] for o in it["options"]]
|
||||
out.append({"ok": True, "e2e_ms": ms, "probs": [float(p[i]) for i in ids],
|
||||
"top": max(ids, key=lambda i: float(p[i]))})
|
||||
return out
|
||||
|
||||
def score(self, item: dict, options: list[dict], averaged: bool) -> dict:
|
||||
ids = [o["id"] for o in options]
|
||||
if not averaged:
|
||||
r = self._one(item, options)
|
||||
if not r["ok"]:
|
||||
return r
|
||||
p = r.pop("p")
|
||||
return {**r, "probs": [p[i] for i in ids], "top": max(ids, key=lambda i: p[i])}
|
||||
per, ms, extra = [], 0.0, {}
|
||||
for k in range(len(options)):
|
||||
r = self._one(item, options[k:] + options[:k])
|
||||
if not r["ok"]:
|
||||
return r
|
||||
per.append(r["p"])
|
||||
ms += r["e2e_ms"]
|
||||
if "abstained" in r:
|
||||
extra.setdefault("abstained_orderings", 0)
|
||||
extra["abstained_orderings"] += bool(r["abstained"])
|
||||
c = combine(per, ids)
|
||||
return {"ok": True, "e2e_ms": ms, "probs": c["probabilities"], "top": c["top"],
|
||||
"agreement": c["agreement"], **extra}
|
||||
|
||||
|
||||
def shuffled(options: list[dict], key: str) -> list[dict]:
|
||||
rng = random.Random(f"jev-bench-2026-09-30/{key}")
|
||||
out = list(options)
|
||||
for _ in range(10):
|
||||
rng.shuffle(out)
|
||||
if [o["id"] for o in out] != [o["id"] for o in options]:
|
||||
break
|
||||
return out
|
||||
|
||||
|
||||
def rotated_descriptions(options: list[dict]) -> list[dict]:
|
||||
descs = [o["description"] for o in options]
|
||||
descs = descs[1:] + descs[:1]
|
||||
return [{**o, "description": d} for o, d in zip(options, descs)]
|
||||
|
||||
|
||||
def gold_desc_id(item: dict) -> str | None:
|
||||
"""After rotated_descriptions, position j carries the description that was at j+1, so the gold
|
||||
description (position g) now sits at g-1."""
|
||||
if not item["gold"]:
|
||||
return None
|
||||
ids = [o["id"] for o in item["options"]]
|
||||
g = ids.index(item["gold"][0])
|
||||
return ids[(g - 1) % len(ids)]
|
||||
|
||||
|
||||
def main() -> int:
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--backend", choices=("semif", "systemone"), required=True)
|
||||
ap.add_argument("--url", required=True)
|
||||
ap.add_argument("--token-file")
|
||||
ap.add_argument("--label", required=True)
|
||||
ap.add_argument("--out", required=True)
|
||||
ap.add_argument("--conditions", default="single,rotations,repeat,reversed,shuffled,negative,multifield")
|
||||
args = ap.parse_args()
|
||||
backend = (Semif(args.url, Path(args.token_file).read_text().strip()) if args.backend == "semif"
|
||||
else SystemOne(args.url))
|
||||
items = load_items()
|
||||
a144 = [i for i in items if i["set"] == "authored144"]
|
||||
plan = {
|
||||
"single": [(i, i["options"], False) for i in items],
|
||||
"rotations": [(i, i["options"], True) for i in items],
|
||||
"repeat": [(i, i["options"], False) for i in a144],
|
||||
"reversed": [(i, list(reversed(i["options"])), False) for i in a144],
|
||||
"shuffled": [(i, shuffled(i["options"], i["id"]), False) for i in a144],
|
||||
"negative": [(i, rotated_descriptions(i["options"]), False) for i in a144],
|
||||
}
|
||||
report = {"label": args.label, "backend": args.backend, "url": args.url,
|
||||
"started_utc": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()), "conditions": {}}
|
||||
for cond in args.conditions.split(","):
|
||||
if cond == "multifield":
|
||||
# Wyrd's real shape: the 4 decisions of one turn over one state, in one request
|
||||
groups: dict[str, list[dict]] = {}
|
||||
for it in items:
|
||||
if it["set"] == "wyrd" and it["cond"] == "evidence":
|
||||
groups.setdefault(it["group"], []).append(it)
|
||||
rows, t0 = [], time.perf_counter()
|
||||
for g, its in groups.items():
|
||||
for it, r in zip(its, backend.score_many(its)):
|
||||
rows.append({"id": it["id"], "set": it["set"], "cond": it["cond"], "decision": it["decision"],
|
||||
"tag": it["tag"], "group": it["group"], "gold": it["gold"], "forbid": it["forbid"],
|
||||
"option_ids": [o["id"] for o in it["options"]], "gold_desc_id": None, **r})
|
||||
fails = sum(not r["ok"] for r in rows)
|
||||
report["conditions"][cond] = {"rows": rows, "seconds": round(time.perf_counter() - t0, 1), "failures": fails}
|
||||
ok = [x for x in rows if x["ok"] and x["gold"]]
|
||||
print(f"[{args.label}] {cond}: {len(rows)} rows in {len(groups)} requests, {fails} failed, "
|
||||
f"acc {sum(x['top'] in x['gold'] for x in ok) / max(1, len(ok)):.3f}", flush=True)
|
||||
continue
|
||||
rows, t0, fails = [], time.perf_counter(), 0
|
||||
for item, options, averaged in plan[cond]:
|
||||
r = backend.score(item, options, averaged)
|
||||
fails += not r["ok"]
|
||||
rows.append({"id": item["id"], "set": item["set"], "cond": item["cond"],
|
||||
"decision": item["decision"], "tag": item["tag"], "group": item["group"],
|
||||
"gold": item["gold"], "forbid": item["forbid"],
|
||||
"option_ids": [o["id"] for o in options],
|
||||
# negative control: the id that now carries the gold option's description
|
||||
"gold_desc_id": gold_desc_id(item) if cond == "negative" else None,
|
||||
**r})
|
||||
if fails >= 5 and fails == len(rows):
|
||||
print(f"[{args.label}] {cond}: first {fails} rows all failed: {r}", file=sys.stderr)
|
||||
return 2
|
||||
report["conditions"][cond] = {"rows": rows, "seconds": round(time.perf_counter() - t0, 1),
|
||||
"failures": fails}
|
||||
ok = [x for x in rows if x["ok"] and x["gold"] and x["cond"] == "evidence"]
|
||||
acc = sum(x["top"] in x["gold"] for x in ok) / len(ok) if ok else float("nan")
|
||||
print(f"[{args.label}] {cond}: {len(rows)} rows, {fails} failed, "
|
||||
f"{report['conditions'][cond]['seconds']}s, labelled-evidence acc {acc:.3f}", flush=True)
|
||||
report["finished_utc"] = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
|
||||
Path(args.out).write_text(json.dumps(report))
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,162 @@
|
||||
"""Jev-candidate bench (2026-09-30): latency and capacity at OUR shape, through one backend.
|
||||
|
||||
Shapes (binary yes/no criteria, single ordering, as in semif-serve's acceptance step 5):
|
||||
short1 one decision over authored144 row 0 (~130 input tokens)
|
||||
crit21 21 binary criteria over authored144 row 0's state in ONE request
|
||||
(semif: /decide/shared; systemone: one request with 21 questions, split into
|
||||
chunks of --max-questions when the runtime caps questions per request)
|
||||
long1 one binary criterion over the ~3,900-token state
|
||||
long16 16 binary criteria over the ~3,900-token state (SemIf's 12 GiB capacity at that size)
|
||||
Each shape: 2 warm-ups, then RUNS runs x PER requests; every request's end-to-end ms and the
|
||||
server-reported ms are kept.
|
||||
|
||||
Capacity (--capacity): rows = binary criteria in one request, stepped up at the short and the long
|
||||
state until the first non-200. The VRAM cap is whatever the server process was started under.
|
||||
|
||||
Phase markers (unix time) are written so the nvidia-smi poll can be split into rest / load.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import statistics as st
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
from bench_sets import DATA, post
|
||||
|
||||
YESNO = [{"id": "yes", "description": "Yes"}, {"id": "no", "description": "No"}]
|
||||
|
||||
|
||||
def criteria(n: int) -> list[dict]:
|
||||
return [{"id": f"c{i}", "question": f"Does the evidence mention item number {i}?", "options": YESNO}
|
||||
for i in range(n)]
|
||||
|
||||
|
||||
def server_ms(j: dict) -> float | None:
|
||||
if "timing" in j and isinstance(j["timing"], dict) and "total_seconds" in j["timing"]:
|
||||
return j["timing"]["total_seconds"] * 1000 # semif /decide/shared
|
||||
if "total_seconds" in j:
|
||||
return j["total_seconds"] * 1000 # semif /decide
|
||||
if "latency_ms" in j:
|
||||
return j["latency_ms"] # jevk5-serve
|
||||
u = j.get("usage") or {}
|
||||
if "total_ms" in u:
|
||||
return u["total_ms"] # imajev playground server
|
||||
t = j.get("timing") or {}
|
||||
if "server_ms" in t:
|
||||
return t["server_ms"] # our Intern-Decision wrapper
|
||||
return None
|
||||
|
||||
|
||||
class Semif:
|
||||
def __init__(self, url, token):
|
||||
self.url, self.h = url.rstrip("/"), {"Authorization": f"Bearer {token}"}
|
||||
|
||||
def request(self, state, crits):
|
||||
if len(crits) == 1:
|
||||
c = crits[0]
|
||||
s, j, ms = post(f"{self.url}/decide", {"id": c["id"], "state": state, "question": c["question"],
|
||||
"options": c["options"]}, self.h)
|
||||
return s, ms, server_ms(j) if s == 200 else None, j.get("input_tokens") if s == 200 else None, j
|
||||
s, j, ms = post(f"{self.url}/decide/shared", {"state": state, "decisions": crits}, self.h)
|
||||
tokens = j["timing"].get("prefix_tokens") if s == 200 else None
|
||||
return s, ms, server_ms(j) if s == 200 else None, tokens, j
|
||||
|
||||
|
||||
class SystemOne:
|
||||
def __init__(self, url, max_questions):
|
||||
self.url, self.maxq = url.rstrip("/"), max_questions
|
||||
|
||||
def request(self, state, crits):
|
||||
total_ms, total_srv, tokens, status, last = 0.0, 0.0, 0, 200, {}
|
||||
for k in range(0, len(crits), self.maxq):
|
||||
chunk = crits[k:k + self.maxq]
|
||||
body = {"state": state, "questions": {c["id"]: {"type": "choice", "instructions": c["question"],
|
||||
"criteria": {o["id"]: o["description"] for o in c["options"]}}
|
||||
for c in chunk}}
|
||||
s, j, ms = post(f"{self.url}/v1/systemone", body, {})
|
||||
total_ms += ms
|
||||
last = j
|
||||
if s != 200:
|
||||
return s, total_ms, None, None, j
|
||||
srv = server_ms(j)
|
||||
total_srv = None if srv is None or total_srv is None else total_srv + srv
|
||||
tokens += (j.get("usage") or {}).get("input_tokens") or 0
|
||||
return status, total_ms, total_srv, tokens, last
|
||||
|
||||
|
||||
def main() -> int:
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--backend", choices=("semif", "systemone"), required=True)
|
||||
ap.add_argument("--url", required=True)
|
||||
ap.add_argument("--token-file")
|
||||
ap.add_argument("--max-questions", type=int, default=10_000)
|
||||
ap.add_argument("--long-state-file", required=True)
|
||||
ap.add_argument("--label", required=True)
|
||||
ap.add_argument("--out", required=True)
|
||||
ap.add_argument("--runs", type=int, default=3)
|
||||
ap.add_argument("--per", type=int, default=20)
|
||||
ap.add_argument("--capacity", action="store_true")
|
||||
ap.add_argument("--short-steps", default="16,32,48,56,60,63,64,96,128")
|
||||
ap.add_argument("--long-steps", default="1,4,8,12,14,16,18,20,24,32,48,64")
|
||||
args = ap.parse_args()
|
||||
b = (Semif(args.url, Path(args.token_file).read_text().strip()) if args.backend == "semif"
|
||||
else SystemOne(args.url, args.max_questions))
|
||||
short_state = json.loads((DATA / "authored144.jsonl").read_text().splitlines()[0])["state"]
|
||||
long_state = Path(args.long_state_file).read_text()
|
||||
shapes = {"short1": (short_state, criteria(1)), "crit21": (short_state, criteria(21)),
|
||||
"long1": (long_state, criteria(1)), "long16": (long_state, criteria(16))}
|
||||
report = {"label": args.label, "backend": args.backend, "runs": args.runs, "per": args.per,
|
||||
"max_questions": args.max_questions, "events": [["start", time.time()]], "shapes": {}}
|
||||
for name, (state, crits) in shapes.items():
|
||||
for _ in range(2):
|
||||
b.request(state, crits)
|
||||
report["events"].append([f"{name}:measure", time.time()])
|
||||
runs = []
|
||||
for _ in range(args.runs):
|
||||
e2e, srv, statuses, tokens = [], [], [], None
|
||||
for _ in range(args.per):
|
||||
s, ms, sm, tk, _ = b.request(state, crits)
|
||||
statuses.append(s)
|
||||
e2e.append(ms)
|
||||
if sm is not None:
|
||||
srv.append(sm)
|
||||
tokens = tk if tk is not None else tokens
|
||||
runs.append({"e2e_ms": e2e, "server_ms": srv, "statuses": statuses, "tokens": tokens,
|
||||
"median_e2e": st.median(e2e), "median_server": st.median(srv) if srv else None})
|
||||
all_e2e = [x for r in runs for x in r["e2e_ms"]]
|
||||
report["shapes"][name] = {
|
||||
"rows": len(crits), "runs": runs, "tokens": runs[-1]["tokens"],
|
||||
"median_e2e_ms": round(st.median(all_e2e), 1),
|
||||
"run_medians_e2e_ms": [round(r["median_e2e"], 1) for r in runs],
|
||||
"median_server_ms": (round(st.median([x for r in runs for x in r["server_ms"]]), 1)
|
||||
if all(r["server_ms"] for r in runs) else None),
|
||||
"non_200": sum(s != 200 for r in runs for s in r["statuses"])}
|
||||
print(f"[{args.label}] {name}: rows={len(crits)} tokens={runs[-1]['tokens']} "
|
||||
f"e2e median {report['shapes'][name]['median_e2e_ms']} ms (runs {report['shapes'][name]['run_medians_e2e_ms']}) "
|
||||
f"server {report['shapes'][name]['median_server_ms']} non200={report['shapes'][name]['non_200']}", flush=True)
|
||||
report["events"].append(["shapes:done", time.time()])
|
||||
if args.capacity:
|
||||
report["capacity"] = {}
|
||||
for label, state, steps in (("short", short_state, args.short_steps), ("long", long_state, args.long_steps)):
|
||||
series = []
|
||||
for n in [int(x) for x in steps.split(",")]:
|
||||
report["events"].append([f"capacity:{label}:{n}", time.time()])
|
||||
s, ms, sm, tk, j = b.request(state, criteria(n))
|
||||
err = None if s == 200 else (j.get("error") if isinstance(j, dict) else str(j))
|
||||
series.append({"rows": n, "status": s, "e2e_ms": round(ms, 1), "tokens": tk,
|
||||
"error": str(err)[:300] if err else None})
|
||||
print(f"[{args.label}] capacity {label} rows={n}: {s} {round(ms)} ms tokens={tk} {str(err)[:120] if err else ''}",
|
||||
flush=True)
|
||||
if s != 200:
|
||||
break
|
||||
report["capacity"][label] = series
|
||||
report["events"].append(["capacity:done", time.time()])
|
||||
report["events"].append(["end", time.time()])
|
||||
Path(args.out).write_text(json.dumps(report))
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,22 @@
|
||||
"""Run a Python module or script under a hard per-process VRAM cap, the way semif-serve applies
|
||||
SEMIF_VRAM_CAP_GIB (torch.cuda.set_per_process_memory_fraction BEFORE any weights load).
|
||||
BENCH_VRAM_CAP_GIB=12 python capped.py <module-or-script.py> [args...]
|
||||
Unset or 0 = no cap. The cap covers torch's allocator only, as in semif-serve; the CUDA context
|
||||
(~0.5-0.9 GiB) sits outside it, so nvidia-smi reads higher than the cap would suggest."""
|
||||
import os
|
||||
import runpy
|
||||
import sys
|
||||
|
||||
import torch
|
||||
|
||||
cap = float(os.environ.get("BENCH_VRAM_CAP_GIB") or 0)
|
||||
if cap:
|
||||
total = torch.cuda.get_device_properties(0).total_memory
|
||||
torch.cuda.set_per_process_memory_fraction(cap * 2**30 / total, 0)
|
||||
print(f"[capped] per-process VRAM cap {cap} GiB of {total / 2**30:.1f} GiB", flush=True)
|
||||
target, sys.argv = sys.argv[1], sys.argv[1:]
|
||||
if target.endswith(".py"):
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(target)))
|
||||
runpy.run_path(target, run_name="__main__")
|
||||
else:
|
||||
runpy.run_module(target, run_name="__main__", alter_sys=True)
|
||||
@@ -0,0 +1,144 @@
|
||||
{"family": "evidence_interpretation", "group_id": "2fdaa8a61e6e989dd866", "id": "a3f18f3a63d45345942b", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The optician ordered replacement lenses. The workshop confirms they have not yet been fitted to the customer's glasses.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_intent_outcome", "partition": "familiar_mechanism_new_situation", "rationale": "Fitting is explicitly not complete.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e01", "variant": "original"}, "question": "Assess the claim: the replacement lenses have been fitted.", "split": "test", "state": "The optician ordered replacement lenses. The workshop confirms they have not yet been fitted to the customer's glasses.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "2fdaa8a61e6e989dd866", "id": "f40beba9088c8db8bbd6", "label": 0, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The optician ordered replacement lenses. The workshop confirms they have not yet been fitted to the customer's glasses.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_intent_outcome", "partition": "familiar_mechanism_new_situation", "rationale": "Ordering is explicitly complete.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e01", "variant": "criterion_reversal"}, "question": "Assess the claim: replacement lenses were ordered.", "split": "test", "state": "The optician ordered replacement lenses. The workshop confirms they have not yet been fitted to the customer's glasses.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "2fdaa8a61e6e989dd866", "id": "0b43ea8e24e74d621f1b", "label": 1, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The workshop confirms it fitted the replacement lenses to the customer's glasses.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_intent_outcome", "partition": "familiar_mechanism_new_situation", "rationale": "Fitting is now explicitly complete.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e01", "variant": "evidence_change"}, "question": "Assess the claim: the replacement lenses have been fitted.", "split": "test", "state": "The workshop confirms it fitted the replacement lenses to the customer's glasses.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "2fdaa8a61e6e989dd866", "id": "942ec65dac0bfb9874f9", "label": 0, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The optician records a glasses prescription but no order or fitting status.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_intent_outcome", "partition": "familiar_mechanism_new_situation", "rationale": "Prescription details settle no fitting status.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e01", "variant": "missing"}, "question": "Assess the claim: the replacement lenses have been fitted.", "split": "test", "state": "The optician records a glasses prescription but no order or fitting status.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "f4bea3d3bcc01056740d", "id": "e0c140e9222d1bce5647", "label": 0, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The guest writes, \"The listing says towels are supplied. I agree: fresh towels were provided during my stay.\"", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_stance", "partition": "familiar_mechanism_new_situation", "rationale": "The guest explicitly agrees.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e02", "variant": "original"}, "question": "Assess the claim: the guest endorses the listing's towel statement.", "split": "test", "state": "The guest writes, \"The listing says towels are supplied. I agree: fresh towels were provided during my stay.\"", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "f4bea3d3bcc01056740d", "id": "8e8c3804a3c15ebb31e3", "label": 2, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The guest writes, \"The listing says towels are supplied. I agree: fresh towels were provided during my stay.\"", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_stance", "partition": "familiar_mechanism_new_situation", "rationale": "Agreement contradicts disputing the statement.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e02", "variant": "criterion_reversal"}, "question": "Assess the claim: the guest disputes the towel statement.", "split": "test", "state": "The guest writes, \"The listing says towels are supplied. I agree: fresh towels were provided during my stay.\"", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "f4bea3d3bcc01056740d", "id": "56c419b12cd37cbe790c", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The guest writes, \"The listing says towels are supplied, but none were provided. I dispute that statement.\"", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_stance", "partition": "familiar_mechanism_new_situation", "rationale": "The guest now explicitly disputes it.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e02", "variant": "evidence_change"}, "question": "Assess the claim: the guest endorses the listing's towel statement.", "split": "test", "state": "The guest writes, \"The listing says towels are supplied, but none were provided. I dispute that statement.\"", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "f4bea3d3bcc01056740d", "id": "75b470586834df8c568d", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The guest forwards the listing but provides no comment about towels or the stay.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_stance", "partition": "familiar_mechanism_new_situation", "rationale": "Forwarding does not settle endorsement.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e02", "variant": "missing"}, "question": "Assess the claim: the guest endorses the listing's towel statement.", "split": "test", "state": "The guest forwards the listing but provides no comment about towels or the stay.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "b51b300da2936decc6b7", "id": "834f3b3e10cc61b33d25", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The technician says terminal Cedar is online. Terminal Maple is explicitly offline.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_identity_scope", "partition": "familiar_mechanism_new_situation", "rationale": "Maple is explicitly offline.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e03", "variant": "original"}, "question": "Assess the claim: terminal Maple is online.", "split": "test", "state": "The technician says terminal Cedar is online. Terminal Maple is explicitly offline.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "b51b300da2936decc6b7", "id": "76c7a79b1ff2363a60c5", "label": 0, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The technician says terminal Cedar is online. Terminal Maple is explicitly offline.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_identity_scope", "partition": "familiar_mechanism_new_situation", "rationale": "Cedar is explicitly online.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e03", "variant": "criterion_reversal"}, "question": "Assess the claim: terminal Cedar is online.", "split": "test", "state": "The technician says terminal Cedar is online. Terminal Maple is explicitly offline.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "b51b300da2936decc6b7", "id": "36c71bb649b7d5eccb1f", "label": 0, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The technician confirms terminal Maple is online and terminal Cedar is offline.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_identity_scope", "partition": "familiar_mechanism_new_situation", "rationale": "Maple is now explicitly online.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e03", "variant": "evidence_change"}, "question": "Assess the claim: terminal Maple is online.", "split": "test", "state": "The technician confirms terminal Maple is online and terminal Cedar is offline.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "b51b300da2936decc6b7", "id": "eac5121a4a1bf391a08c", "label": 1, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The technician lists both terminals but omits their connection status.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_identity_scope", "partition": "familiar_mechanism_new_situation", "rationale": "Connection status is absent.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e03", "variant": "missing"}, "question": "Assess the claim: terminal Maple is online.", "split": "test", "state": "The technician lists both terminals but omits their connection status.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "e4cfd694c80e08d3e6a5", "id": "44574c44e243ee1b2f52", "label": 2, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The catalogue originally dated the letter to June. Its author issued a correction stating July was correct and June was mistaken.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_correction", "partition": "familiar_mechanism_new_situation", "rationale": "The correction establishes July.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e04", "variant": "original"}, "question": "Assess the claim: the corrected date is in July.", "split": "test", "state": "The catalogue originally dated the letter to June. Its author issued a correction stating July was correct and June was mistaken.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "e4cfd694c80e08d3e6a5", "id": "2608370d01783735c3ea", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The catalogue originally dated the letter to June. Its author issued a correction stating July was correct and June was mistaken.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_correction", "partition": "familiar_mechanism_new_situation", "rationale": "The correction rejects June.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e04", "variant": "criterion_reversal"}, "question": "Assess the claim: the corrected date is in June.", "split": "test", "state": "The catalogue originally dated the letter to June. Its author issued a correction stating July was correct and June was mistaken.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "e4cfd694c80e08d3e6a5", "id": "ec94ef909f7918695ab3", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The catalogue author corrects the date to June and expressly rejects the earlier July date.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_correction", "partition": "familiar_mechanism_new_situation", "rationale": "The correction now rejects July.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e04", "variant": "evidence_change"}, "question": "Assess the claim: the corrected date is in July.", "split": "test", "state": "The catalogue author corrects the date to June and expressly rejects the earlier July date.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "e4cfd694c80e08d3e6a5", "id": "03d919fd0e9b3a7cd1e6", "label": 0, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The catalogue mentions a date correction but gives no corrected date.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_correction", "partition": "familiar_mechanism_new_situation", "rationale": "The corrected date is missing.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e04", "variant": "missing"}, "question": "Assess the claim: the corrected date is in July.", "split": "test", "state": "The catalogue mentions a date correction but gives no corrected date.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "20460d4ae0db46091e84", "id": "86a15618ba4cb90a307e", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The intern attached instructions for weighing the parcel and explicitly says no weighing has been performed.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_result_procedure", "partition": "familiar_mechanism_new_situation", "rationale": "The intern explicitly denies performed weighing.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e05", "variant": "original"}, "question": "Assess the claim: the parcel has been weighed.", "split": "test", "state": "The intern attached instructions for weighing the parcel and explicitly says no weighing has been performed.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "20460d4ae0db46091e84", "id": "3dd9c3ac98f8ddf7bdd3", "label": 1, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The intern attached instructions for weighing the parcel and explicitly says no weighing has been performed.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_result_procedure", "partition": "familiar_mechanism_new_situation", "rationale": "Instructions are explicitly attached.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e05", "variant": "criterion_reversal"}, "question": "Assess the claim: the document contains weighing instructions.", "split": "test", "state": "The intern attached instructions for weighing the parcel and explicitly says no weighing has been performed.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "20460d4ae0db46091e84", "id": "36273dd4777876a26338", "label": 1, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The intern performed the weighing and recorded the parcel's measured weight.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_result_procedure", "partition": "familiar_mechanism_new_situation", "rationale": "Performed weighing is now explicit.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e05", "variant": "evidence_change"}, "question": "Assess the claim: the parcel has been weighed.", "split": "test", "state": "The intern performed the weighing and recorded the parcel's measured weight.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "20460d4ae0db46091e84", "id": "07ee70cfe4b01330ce89", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The intern supplies a parcel identifier but no procedure or measurement information.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_result_procedure", "partition": "familiar_mechanism_new_situation", "rationale": "No measurement information is supplied.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e05", "variant": "missing"}, "question": "Assess the claim: the parcel has been weighed.", "split": "test", "state": "The intern supplies a parcel identifier but no procedure or measurement information.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "51b91ba31984c5939c1c", "id": "51ceab3ce622c7c759f1", "label": 2, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "During a puppet show, a character announces the discovery of a dragon egg. The production note states this is fictional and no real egg was discovered.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_fiction", "partition": "familiar_mechanism_new_situation", "rationale": "The production note explicitly identifies fiction.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e06", "variant": "original"}, "question": "Assess the claim: the egg discovery was fictional.", "split": "test", "state": "During a puppet show, a character announces the discovery of a dragon egg. The production note states this is fictional and no real egg was discovered.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "51b91ba31984c5939c1c", "id": "b8091be929b1e489fcc4", "label": 1, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "During a puppet show, a character announces the discovery of a dragon egg. The production note states this is fictional and no real egg was discovered.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_fiction", "partition": "familiar_mechanism_new_situation", "rationale": "It explicitly denies a real discovery.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e06", "variant": "criterion_reversal"}, "question": "Assess the claim: a real egg was discovered.", "split": "test", "state": "During a puppet show, a character announces the discovery of a dragon egg. The production note states this is fictional and no real egg was discovered.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "51b91ba31984c5939c1c", "id": "c89b86e3f55123bfa506", "label": 1, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The note describes a real bird egg discovered outside a performance and expressly says the discovery was not fictional.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_fiction", "partition": "familiar_mechanism_new_situation", "rationale": "The changed account expressly identifies a real discovery.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e06", "variant": "evidence_change"}, "question": "Assess the claim: the egg discovery was fictional.", "split": "test", "state": "The note describes a real bird egg discovered outside a performance and expressly says the discovery was not fictional.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "51b91ba31984c5939c1c", "id": "3b0958cabc339005cf66", "label": 0, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "A note mentions an egg without describing a discovery or whether the account is fictional.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_fiction", "partition": "familiar_mechanism_new_situation", "rationale": "The account type is absent.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_e06", "variant": "missing"}, "question": "Assess the claim: the egg discovery was fictional.", "split": "test", "state": "A note mentions an egg without describing a discovery or whether the account is fictional.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "770c1607f0a11f88f81e", "id": "f2b4ec4930322fe45118", "label": 1, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The architect permits reproducing a plan in an internal training booklet and expressly forbids selling copies.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_scope_permission", "partition": "familiar_mechanism_new_situation", "rationale": "Sales are expressly forbidden.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r01", "variant": "original"}, "question": "Sell copies only with the architect's sales permission. May copies be sold?", "split": "test", "state": "The architect permits reproducing a plan in an internal training booklet and expressly forbids selling copies.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "770c1607f0a11f88f81e", "id": "61ac20b07a5189231e15", "label": 1, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The architect permits reproducing a plan in an internal training booklet and expressly forbids selling copies.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_scope_permission", "partition": "familiar_mechanism_new_situation", "rationale": "Internal training reproduction is allowed.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r01", "variant": "criterion_reversal"}, "question": "Internal training reproduction is allowed when the architect permits that use. May the plan appear in the training booklet?", "split": "test", "state": "The architect permits reproducing a plan in an internal training booklet and expressly forbids selling copies.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "770c1607f0a11f88f81e", "id": "4b3d794c004566610181", "label": 2, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The architect expressly permits selling copies of the plan.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_scope_permission", "partition": "familiar_mechanism_new_situation", "rationale": "Sales are now allowed.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r01", "variant": "evidence_change"}, "question": "Sell copies only with the architect's sales permission. May copies be sold?", "split": "test", "state": "The architect expressly permits selling copies of the plan.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "770c1607f0a11f88f81e", "id": "b7c4e62d07986286fce0", "label": 0, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "An architect's permission letter exists, but the uses it allows are omitted.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_scope_permission", "partition": "familiar_mechanism_new_situation", "rationale": "The scope of permission is missing.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r01", "variant": "missing"}, "question": "Sell copies only with the architect's sales permission. May copies be sold?", "split": "test", "state": "An architect's permission letter exists, but the uses it allows are omitted.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "ffd081ada3bd9be9207a", "id": "9473603b44af63beab79", "label": 0, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The guide completed first-aid training but not mountain-navigation training.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_required_qualification", "partition": "familiar_mechanism_new_situation", "rationale": "The sole original condition is satisfied.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r02", "variant": "original"}, "question": "The first-aid badge requires only first-aid training. May this guide receive it?", "split": "test", "state": "The guide completed first-aid training but not mountain-navigation training.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "ffd081ada3bd9be9207a", "id": "d41072dc35d8cd14c056", "label": 1, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The guide completed first-aid training but not mountain-navigation training.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_required_qualification", "partition": "familiar_mechanism_new_situation", "rationale": "The added navigation condition is not satisfied.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r02", "variant": "criterion_reversal"}, "question": "Under the alternative rule, the badge requires both first-aid and mountain-navigation training. May this guide receive it?", "split": "test", "state": "The guide completed first-aid training but not mountain-navigation training.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "ffd081ada3bd9be9207a", "id": "33d6e5da58f0fee2490d", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The guide completed mountain-navigation training but not first-aid training.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_required_qualification", "partition": "familiar_mechanism_new_situation", "rationale": "The original first-aid condition is now absent.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r02", "variant": "evidence_change"}, "question": "The first-aid badge requires only first-aid training. May this guide receive it?", "split": "test", "state": "The guide completed mountain-navigation training but not first-aid training.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "ffd081ada3bd9be9207a", "id": "823bf736b92e9d8fa66d", "label": 1, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The guide registered for training, but completion information is not supplied.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_required_qualification", "partition": "familiar_mechanism_new_situation", "rationale": "Registration does not establish completion.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r02", "variant": "missing"}, "question": "The first-aid badge requires only first-aid training. May this guide receive it?", "split": "test", "state": "The guide registered for training, but completion information is not supplied.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "2c3a0cb61f7bae78739e", "id": "8a4c3de28be95c105ca1", "label": 2, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The designer approves making a sample badge but expressly withholds approval for mass production.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_approval_stage", "partition": "familiar_mechanism_new_situation", "rationale": "Mass-production approval is withheld.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r03", "variant": "original"}, "question": "Mass production is allowed only after mass-production approval. May it begin?", "split": "test", "state": "The designer approves making a sample badge but expressly withholds approval for mass production.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "2c3a0cb61f7bae78739e", "id": "ea7ec66a5e8eac886548", "label": 0, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The designer approves making a sample badge but expressly withholds approval for mass production.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_approval_stage", "partition": "familiar_mechanism_new_situation", "rationale": "Sample production is approved.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r03", "variant": "criterion_reversal"}, "question": "A sample may be made when sample production is approved. May the sample be made?", "split": "test", "state": "The designer approves making a sample badge but expressly withholds approval for mass production.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "2c3a0cb61f7bae78739e", "id": "25172cd6ddc49ab0ca70", "label": 2, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The designer explicitly approves mass production of the badge.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_approval_stage", "partition": "familiar_mechanism_new_situation", "rationale": "Mass-production approval is now present.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r03", "variant": "evidence_change"}, "question": "Mass production is allowed only after mass-production approval. May it begin?", "split": "test", "state": "The designer explicitly approves mass production of the badge.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "2c3a0cb61f7bae78739e", "id": "fb73a85d386cf59b0c51", "label": 0, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The designer issued a decision but its production scope is unavailable.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_approval_stage", "partition": "familiar_mechanism_new_situation", "rationale": "The approved stage is missing.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r03", "variant": "missing"}, "question": "Mass production is allowed only after mass-production approval. May it begin?", "split": "test", "state": "The designer issued a decision but its production scope is unavailable.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "42ed4cbfde7f2f64274b", "id": "a02fd582d2881cc462b2", "label": 2, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The visitor's wristband covers the sculpture garden and expressly excludes the rooftop terrace.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_named_privilege", "partition": "familiar_mechanism_new_situation", "rationale": "The garden is covered.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r04", "variant": "original"}, "question": "A wristband permits entry exactly to the areas it covers. May the visitor enter the sculpture garden?", "split": "test", "state": "The visitor's wristband covers the sculpture garden and expressly excludes the rooftop terrace.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "42ed4cbfde7f2f64274b", "id": "19f37d5a97034bd7d04f", "label": 1, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The visitor's wristband covers the sculpture garden and expressly excludes the rooftop terrace.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_named_privilege", "partition": "familiar_mechanism_new_situation", "rationale": "The terrace is excluded.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r04", "variant": "criterion_reversal"}, "question": "A wristband permits entry exactly to the areas it covers. May the visitor enter the rooftop terrace?", "split": "test", "state": "The visitor's wristband covers the sculpture garden and expressly excludes the rooftop terrace.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "42ed4cbfde7f2f64274b", "id": "fda0ca94652c6d739ecf", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The wristband expressly excludes the sculpture garden and covers only the rooftop terrace.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_named_privilege", "partition": "familiar_mechanism_new_situation", "rationale": "The garden is now excluded.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r04", "variant": "evidence_change"}, "question": "A wristband permits entry exactly to the areas it covers. May the visitor enter the sculpture garden?", "split": "test", "state": "The wristband expressly excludes the sculpture garden and covers only the rooftop terrace.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "42ed4cbfde7f2f64274b", "id": "d0e36c4cb10a29b3191c", "label": 0, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The visitor has a wristband but its covered areas are not supplied.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_named_privilege", "partition": "familiar_mechanism_new_situation", "rationale": "Covered areas are unknown.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r04", "variant": "missing"}, "question": "A wristband permits entry exactly to the areas it covers. May the visitor enter the sculpture garden?", "split": "test", "state": "The visitor has a wristband but its covered areas are not supplied.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "6dd193172b9240a04bf9", "id": "a0a39ace7e5337d038a7", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The account owner permits changing the display name but expressly forbids changing the recovery email.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_local_action", "partition": "familiar_mechanism_new_situation", "rationale": "Recovery-email change is expressly forbidden.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r05", "variant": "original"}, "question": "Change an account field only when the owner authorizes changing that field. May the recovery email be changed?", "split": "test", "state": "The account owner permits changing the display name but expressly forbids changing the recovery email.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "6dd193172b9240a04bf9", "id": "eeae283841ce75dd25bf", "label": 1, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The account owner permits changing the display name but expressly forbids changing the recovery email.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_local_action", "partition": "familiar_mechanism_new_situation", "rationale": "Display-name change is authorized.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r05", "variant": "criterion_reversal"}, "question": "Change an account field only when the owner authorizes changing that field. May the display name be changed?", "split": "test", "state": "The account owner permits changing the display name but expressly forbids changing the recovery email.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "6dd193172b9240a04bf9", "id": "00bb4cc043f7d2b288d4", "label": 2, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The owner expressly permits changing the recovery email.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_local_action", "partition": "familiar_mechanism_new_situation", "rationale": "Recovery-email change is now authorized.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r05", "variant": "evidence_change"}, "question": "Change an account field only when the owner authorizes changing that field. May the recovery email be changed?", "split": "test", "state": "The owner expressly permits changing the recovery email.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "6dd193172b9240a04bf9", "id": "1a81ead69e93e5d0c86a", "label": 2, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The owner requested an account update, but the fields to be changed are omitted.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_local_action", "partition": "familiar_mechanism_new_situation", "rationale": "The authorized field is missing.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r05", "variant": "missing"}, "question": "Change an account field only when the owner authorizes changing that field. May the recovery email be changed?", "split": "test", "state": "The owner requested an account update, but the fields to be changed are omitted.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "3200eb252d6737e03027", "id": "4444dad25e438f128673", "label": 1, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The applicant lives in the valley but works outside it.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_residence_employment", "partition": "familiar_mechanism_new_situation", "rationale": "Valley residence meets the original condition.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r06", "variant": "original"}, "question": "A resident's pass requires living in the valley and has no workplace restriction. May it be issued?", "split": "test", "state": "The applicant lives in the valley but works outside it.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "3200eb252d6737e03027", "id": "3226202ba1ada2570e9f", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The applicant lives in the valley but works outside it.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_residence_employment", "partition": "familiar_mechanism_new_situation", "rationale": "The alternative workplace condition is not met.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r06", "variant": "criterion_reversal"}, "question": "Under the alternative rule, a pass requires working in the valley; residence alone is insufficient. May it be issued?", "split": "test", "state": "The applicant lives in the valley but works outside it.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "3200eb252d6737e03027", "id": "3f783eec5c90812479bc", "label": 1, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The applicant lives outside the valley and works inside it.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_residence_employment", "partition": "familiar_mechanism_new_situation", "rationale": "Valley residence is now absent.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r06", "variant": "evidence_change"}, "question": "A resident's pass requires living in the valley and has no workplace restriction. May it be issued?", "split": "test", "state": "The applicant lives outside the valley and works inside it.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "3200eb252d6737e03027", "id": "ffee861474a9665a093a", "label": 2, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The applicant supplies a name but no home or work location.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_residence_employment", "partition": "familiar_mechanism_new_situation", "rationale": "The relevant locations are unknown.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_r06", "variant": "missing"}, "question": "A resident's pass requires living in the valley and has no workplace restriction. May it be issued?", "split": "test", "state": "The applicant supplies a name but no home or work location.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "c0cbeb97315307af8111", "id": "8c992e17e4fb54b044f5", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The florist confirms that the wreath was delivered. Candidate B: The florist offers to deliver a wreath tomorrow.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_intent_outcome", "partition": "familiar_mechanism_new_situation", "rationale": "A reports completed delivery.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c01", "variant": "original"}, "question": "Choose the candidate reporting completed delivery.", "split": "test", "state": "Candidate A: The florist confirms that the wreath was delivered. Candidate B: The florist offers to deliver a wreath tomorrow.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "c0cbeb97315307af8111", "id": "695dd61195ca025c4f4c", "label": 0, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The florist confirms that the wreath was delivered. Candidate B: The florist offers to deliver a wreath tomorrow.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_intent_outcome", "partition": "familiar_mechanism_new_situation", "rationale": "B offers future action.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c01", "variant": "criterion_reversal"}, "question": "Choose the candidate offering a future delivery.", "split": "test", "state": "Candidate A: The florist confirms that the wreath was delivered. Candidate B: The florist offers to deliver a wreath tomorrow.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "c0cbeb97315307af8111", "id": "e6155431ab7b0f8170d4", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: I can deliver a wreath tomorrow. Candidate B: The wreath was delivered.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_intent_outcome", "partition": "familiar_mechanism_new_situation", "rationale": "Completed delivery is now B.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c01", "variant": "evidence_change"}, "question": "Choose the candidate reporting completed delivery.", "split": "test", "state": "Candidate A: I can deliver a wreath tomorrow. Candidate B: The wreath was delivered.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "c0cbeb97315307af8111", "id": "b49ecda75086d8c68d2a", "label": 2, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The florist gives an address. Candidate B: The wreath contains lilies.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_intent_outcome", "partition": "familiar_mechanism_new_situation", "rationale": "Neither provides a delivery report or offer.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c01", "variant": "missing"}, "question": "Choose the candidate reporting completed delivery.", "split": "test", "state": "Candidate A: The florist gives an address. Candidate B: The wreath contains lilies.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "428176cd8ad2c7a74eea", "id": "587ced261c6092cd1325", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The customer asks how an engraving would look but explicitly declines having it engraved yet. Candidate B: The customer instructs the shop to engrave the ring.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_scope_consent", "partition": "familiar_mechanism_new_situation", "rationale": "B instructs engraving.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c02", "variant": "original"}, "question": "Choose the candidate authorizing engraving.", "split": "test", "state": "Candidate A: The customer asks how an engraving would look but explicitly declines having it engraved yet. Candidate B: The customer instructs the shop to engrave the ring.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "428176cd8ad2c7a74eea", "id": "533d4423d311de82b27d", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The customer asks how an engraving would look but explicitly declines having it engraved yet. Candidate B: The customer instructs the shop to engrave the ring.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_scope_consent", "partition": "familiar_mechanism_new_situation", "rationale": "A expressly withholds authorization.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c02", "variant": "criterion_reversal"}, "question": "Choose the candidate asking only for information while withholding authorization.", "split": "test", "state": "Candidate A: The customer asks how an engraving would look but explicitly declines having it engraved yet. Candidate B: The customer instructs the shop to engrave the ring.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "428176cd8ad2c7a74eea", "id": "2bf540d515a51755b25c", "label": 2, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: Please engrave the ring. Candidate B: Show me how it would look, but do not engrave it yet.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_scope_consent", "partition": "familiar_mechanism_new_situation", "rationale": "A now authorizes engraving.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c02", "variant": "evidence_change"}, "question": "Choose the candidate authorizing engraving.", "split": "test", "state": "Candidate A: Please engrave the ring. Candidate B: Show me how it would look, but do not engrave it yet.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "428176cd8ad2c7a74eea", "id": "fdb9fc214aea2710763a", "label": 2, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The customer gives a name. Candidate B: The ring is silver.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_scope_consent", "partition": "familiar_mechanism_new_situation", "rationale": "Neither addresses engraving permission.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c02", "variant": "missing"}, "question": "Choose the candidate authorizing engraving.", "split": "test", "state": "Candidate A: The customer gives a name. Candidate B: The ring is silver.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "389e65bf1ba709b84b12", "id": "f6d68126b43f8c9ed183", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The note records that the completed insulation test passed. Candidate B: The note lists insulation-test steps but says no test was run.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_result_procedure", "partition": "familiar_mechanism_new_situation", "rationale": "A reports a completed result.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c03", "variant": "original"}, "question": "Choose the candidate reporting an observed test result.", "split": "test", "state": "Candidate A: The note records that the completed insulation test passed. Candidate B: The note lists insulation-test steps but says no test was run.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "389e65bf1ba709b84b12", "id": "44bae0c8a1615a096c43", "label": 0, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The note records that the completed insulation test passed. Candidate B: The note lists insulation-test steps but says no test was run.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_result_procedure", "partition": "familiar_mechanism_new_situation", "rationale": "B supplies unexecuted steps.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c03", "variant": "criterion_reversal"}, "question": "Choose the candidate containing only an unexecuted test procedure.", "split": "test", "state": "Candidate A: The note records that the completed insulation test passed. Candidate B: The note lists insulation-test steps but says no test was run.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "389e65bf1ba709b84b12", "id": "2175379f4c69e6207626", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: Here are insulation-test steps; no test was run. Candidate B: The completed insulation test passed.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_result_procedure", "partition": "familiar_mechanism_new_situation", "rationale": "The result is now B.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c03", "variant": "evidence_change"}, "question": "Choose the candidate reporting an observed test result.", "split": "test", "state": "Candidate A: Here are insulation-test steps; no test was run. Candidate B: The completed insulation test passed.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "389e65bf1ba709b84b12", "id": "def9fe59d262357a673a", "label": 0, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The cable is red. Candidate B: The tester is in a drawer.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_result_procedure", "partition": "familiar_mechanism_new_situation", "rationale": "Neither supplies a test result or procedure.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c03", "variant": "missing"}, "question": "Choose the candidate reporting an observed test result.", "split": "test", "state": "Candidate A: The cable is red. Candidate B: The tester is in a drawer.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "45f8ead2522a14384cdc", "id": "712f6ffb763eef7c9a12", "label": 0, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The editor quotes a proposal for a paywall only to argue against it. Candidate B: The editor endorses introducing a paywall.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_stance", "partition": "familiar_mechanism_new_situation", "rationale": "B endorses the policy.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c04", "variant": "original"}, "question": "Choose the candidate supporting a paywall.", "split": "test", "state": "Candidate A: The editor quotes a proposal for a paywall only to argue against it. Candidate B: The editor endorses introducing a paywall.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "45f8ead2522a14384cdc", "id": "e1d610bd14e3d16c09b9", "label": 0, "options": [{"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The editor quotes a proposal for a paywall only to argue against it. Candidate B: The editor endorses introducing a paywall.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_stance", "partition": "familiar_mechanism_new_situation", "rationale": "A argues against the quoted policy.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c04", "variant": "criterion_reversal"}, "question": "Choose the candidate opposing a paywall.", "split": "test", "state": "Candidate A: The editor quotes a proposal for a paywall only to argue against it. Candidate B: The editor endorses introducing a paywall.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "45f8ead2522a14384cdc", "id": "22ec57822db50ec121cf", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: I endorse introducing a paywall. Candidate B: I reject introducing a paywall.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_stance", "partition": "familiar_mechanism_new_situation", "rationale": "Support is now A.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c04", "variant": "evidence_change"}, "question": "Choose the candidate supporting a paywall.", "split": "test", "state": "Candidate A: I endorse introducing a paywall. Candidate B: I reject introducing a paywall.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "45f8ead2522a14384cdc", "id": "3e75f2b84e8e08e3d3f2", "label": 2, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The editor names the magazine. Candidate B: The website has a blue header.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_stance", "partition": "familiar_mechanism_new_situation", "rationale": "Neither supplies a position on paywalls.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c04", "variant": "missing"}, "question": "Choose the candidate supporting a paywall.", "split": "test", "state": "Candidate A: The editor names the magazine. Candidate B: The website has a blue header.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "3ecec8e2c166e397ad07", "id": "ebd74aa05324a1556288", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The owner asks to reseal the existing window while keeping its frame. Candidate B: The owner asks to remove the window and fit a complete replacement.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_retained_replaced", "partition": "familiar_mechanism_new_situation", "rationale": "A retains and repairs the existing window.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c05", "variant": "original"}, "question": "Choose the candidate asking to repair the existing window.", "split": "test", "state": "Candidate A: The owner asks to reseal the existing window while keeping its frame. Candidate B: The owner asks to remove the window and fit a complete replacement.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "3ecec8e2c166e397ad07", "id": "d7e84e4245a29817da1f", "label": 2, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The owner asks to reseal the existing window while keeping its frame. Candidate B: The owner asks to remove the window and fit a complete replacement.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_retained_replaced", "partition": "familiar_mechanism_new_situation", "rationale": "B requests replacement.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c05", "variant": "criterion_reversal"}, "question": "Choose the candidate asking for complete replacement.", "split": "test", "state": "Candidate A: The owner asks to reseal the existing window while keeping its frame. Candidate B: The owner asks to remove the window and fit a complete replacement.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "3ecec8e2c166e397ad07", "id": "f41ae66b3c09630c7a42", "label": 2, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: Remove the window and install a replacement. Candidate B: Reseal the existing window and keep its frame.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_retained_replaced", "partition": "familiar_mechanism_new_situation", "rationale": "Repair is now B.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c05", "variant": "evidence_change"}, "question": "Choose the candidate asking to repair the existing window.", "split": "test", "state": "Candidate A: Remove the window and install a replacement. Candidate B: Reseal the existing window and keep its frame.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "3ecec8e2c166e397ad07", "id": "82c2cd97c160ad27951b", "label": 0, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The window faces a courtyard. Candidate B: The owner gives a postcode.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_retained_replaced", "partition": "familiar_mechanism_new_situation", "rationale": "Neither supplies a window-work request.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c05", "variant": "missing"}, "question": "Choose the candidate asking to repair the existing window.", "split": "test", "state": "Candidate A: The window faces a courtyard. Candidate B: The owner gives a postcode.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "a71a17a287f3d30db2aa", "id": "7ada97b6f99f4788e5c7", "label": 0, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The sponsor confirms receipt of a grant application but explicitly says no award decision has been made. Candidate B: The sponsor confirms that the grant was awarded.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_submission_acceptance", "partition": "familiar_mechanism_new_situation", "rationale": "B confirms an award.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c06", "variant": "original"}, "question": "Choose the candidate reporting an awarded grant.", "split": "test", "state": "Candidate A: The sponsor confirms receipt of a grant application but explicitly says no award decision has been made. Candidate B: The sponsor confirms that the grant was awarded.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "a71a17a287f3d30db2aa", "id": "4d7a003342db5b0a17e3", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The sponsor confirms receipt of a grant application but explicitly says no award decision has been made. Candidate B: The sponsor confirms that the grant was awarded.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_submission_acceptance", "partition": "familiar_mechanism_new_situation", "rationale": "A expressly leaves the award undecided.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c06", "variant": "criterion_reversal"}, "question": "Choose the candidate reporting receipt with the award still undecided.", "split": "test", "state": "Candidate A: The sponsor confirms receipt of a grant application but explicitly says no award decision has been made. Candidate B: The sponsor confirms that the grant was awarded.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "a71a17a287f3d30db2aa", "id": "19d3d8d0bf773e88be7b", "label": 0, "options": [{"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The grant was awarded. Candidate B: The application was received; the award remains undecided.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_submission_acceptance", "partition": "familiar_mechanism_new_situation", "rationale": "An award is now confirmed in A.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c06", "variant": "evidence_change"}, "question": "Choose the candidate reporting an awarded grant.", "split": "test", "state": "Candidate A: The grant was awarded. Candidate B: The application was received; the award remains undecided.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "a71a17a287f3d30db2aa", "id": "d6de731cca52359b1c87", "label": 0, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The sponsor gives its name. Candidate B: The application has a reference number.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "familiar_submission_acceptance", "partition": "familiar_mechanism_new_situation", "rationale": "Neither supplies an award status.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "i_c06", "variant": "missing"}, "question": "Choose the candidate reporting an awarded grant.", "split": "test", "state": "Candidate A: The sponsor gives its name. Candidate B: The application has a reference number.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "986d7eaba0c7e9738e79", "id": "b45c63fb92729af27f05", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "On Monday the owner authorized scanning her notebook. On Tuesday she explicitly withdrew that authorization. The record says withdrawals cancel earlier permission, and no later permission was given.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_revocation", "partition": "heldout_family", "rationale": "A later withdrawal canceled the earlier permission.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e01", "variant": "original"}, "question": "Assess the claim: permission to scan the notebook is currently valid.", "split": "test", "state": "On Monday the owner authorized scanning her notebook. On Tuesday she explicitly withdrew that authorization. The record says withdrawals cancel earlier permission, and no later permission was given.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "986d7eaba0c7e9738e79", "id": "46b7029b9a704138b77a", "label": 2, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "On Monday the owner authorized scanning her notebook. On Tuesday she explicitly withdrew that authorization. The record says withdrawals cancel earlier permission, and no later permission was given.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_revocation", "partition": "heldout_family", "rationale": "The earlier grant remains a historical fact.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e01", "variant": "criterion_reversal"}, "question": "Assess the claim: the owner gave permission on Monday.", "split": "test", "state": "On Monday the owner authorized scanning her notebook. On Tuesday she explicitly withdrew that authorization. The record says withdrawals cancel earlier permission, and no later permission was given.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "986d7eaba0c7e9738e79", "id": "caf773b23934fa7362a1", "label": 0, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The owner authorized scanning the notebook on Monday. The complete permission record confirms that she has not withdrawn it and it remains valid.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_revocation", "partition": "heldout_family", "rationale": "The complete record confirms continuing validity.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e01", "variant": "evidence_change"}, "question": "Assess the claim: permission to scan the notebook is currently valid.", "split": "test", "state": "The owner authorized scanning the notebook on Monday. The complete permission record confirms that she has not withdrawn it and it remains valid.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "986d7eaba0c7e9738e79", "id": "f46f392ef9e9e9df564b", "label": 1, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The owner authorized scanning on Monday. The supplied record omits all later permission changes and does not establish current validity.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_revocation", "partition": "heldout_family", "rationale": "The incomplete later record does not settle current permission.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e01", "variant": "missing"}, "question": "Assess the claim: permission to scan the notebook is currently valid.", "split": "test", "state": "The owner authorized scanning on Monday. The supplied record omits all later permission changes and does not establish current validity.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "ca2e0b043985fa49bcd2", "id": "005dcf6d2c0f21d64796", "label": 1, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The homeowner first allowed a surveyor to enter the attic. Before the visit she explicitly withdrew that permission. No later decision exists.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_revocation", "partition": "heldout_family", "rationale": "The latest decision withdraws entry permission.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r01", "variant": "original"}, "question": "Use the homeowner's latest permission decision; a withdrawal cancels the earlier grant. May the surveyor enter the attic?", "split": "test", "state": "The homeowner first allowed a surveyor to enter the attic. Before the visit she explicitly withdrew that permission. No later decision exists.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "ca2e0b043985fa49bcd2", "id": "edb708bd0fe5fd67d6ee", "label": 0, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The homeowner first allowed a surveyor to enter the attic. Before the visit she explicitly withdrew that permission. No later decision exists.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_revocation", "partition": "heldout_family", "rationale": "The alternate fictional rule uses the initial grant.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r01", "variant": "criterion_reversal"}, "question": "For this fictional protocol, the first recorded decision is binding and later withdrawals are ignored. May the surveyor enter the attic?", "split": "test", "state": "The homeowner first allowed a surveyor to enter the attic. Before the visit she explicitly withdrew that permission. No later decision exists.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "ca2e0b043985fa49bcd2", "id": "97a35499bd8717f3e786", "label": 2, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The homeowner first refused attic entry, then explicitly granted it before the visit. No later decision exists.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_revocation", "partition": "heldout_family", "rationale": "The latest decision now grants entry.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r01", "variant": "evidence_change"}, "question": "Use the homeowner's latest permission decision; a withdrawal cancels the earlier grant. May the surveyor enter the attic?", "split": "test", "state": "The homeowner first refused attic entry, then explicitly granted it before the visit. No later decision exists.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "ca2e0b043985fa49bcd2", "id": "31eb4533d2c89959ec46", "label": 0, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The homeowner gave an early permission decision, but neither its content nor any later decisions are available.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_revocation", "partition": "heldout_family", "rationale": "The governing decision cannot be identified.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r01", "variant": "missing"}, "question": "Use the homeowner's latest permission decision; a withdrawal cancels the earlier grant. May the surveyor enter the attic?", "split": "test", "state": "The homeowner gave an early permission decision, but neither its content nor any later decisions are available.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "85051ba3c30277824ebe", "id": "0aa24a6701ac9c56c4a0", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The speaker authorized streaming, then explicitly withdrew the authorization before the event. Candidate B: The speaker authorized streaming and the complete record confirms it remains in force.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_revocation", "partition": "heldout_family", "rationale": "Only B retains valid permission.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c01", "variant": "original"}, "question": "Select the candidate with currently valid streaming permission; a withdrawal cancels an earlier grant.", "split": "test", "state": "Candidate A: The speaker authorized streaming, then explicitly withdrew the authorization before the event. Candidate B: The speaker authorized streaming and the complete record confirms it remains in force.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "85051ba3c30277824ebe", "id": "796d5c0da6eaadae6a3f", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: The speaker authorized streaming, then explicitly withdrew the authorization before the event. Candidate B: The speaker authorized streaming and the complete record confirms it remains in force.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_revocation", "partition": "heldout_family", "rationale": "A documents the later withdrawal.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c01", "variant": "criterion_reversal"}, "question": "Select the candidate documenting an authorization that was later withdrawn.", "split": "test", "state": "Candidate A: The speaker authorized streaming, then explicitly withdrew the authorization before the event. Candidate B: The speaker authorized streaming and the complete record confirms it remains in force.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "85051ba3c30277824ebe", "id": "599b3092b6fd18312d23", "label": 2, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: Streaming was authorized and the complete record says permission remains in force. Candidate B: Streaming was authorized, then explicitly withdrawn.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_revocation", "partition": "heldout_family", "rationale": "Only A now retains valid permission.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c01", "variant": "evidence_change"}, "question": "Select the candidate with currently valid streaming permission; a withdrawal cancels an earlier grant.", "split": "test", "state": "Candidate A: Streaming was authorized and the complete record says permission remains in force. Candidate B: Streaming was authorized, then explicitly withdrawn.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "85051ba3c30277824ebe", "id": "960a4bcd27363e6db4b9", "label": 2, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A: A permission form is mentioned but its terms and later changes are missing. Candidate B: The event name is supplied without permission history.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_revocation", "partition": "heldout_family", "rationale": "Neither supplies enough permission history.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c01", "variant": "missing"}, "question": "Select the candidate with currently valid streaming permission; a withdrawal cancels an earlier grant.", "split": "test", "state": "Candidate A: A permission form is mentioned but its terms and later changes are missing. Candidate B: The event name is supplied without permission history.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "e7798dab5f4f121cab77", "id": "c6a411e7f2bbdf377173", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The director gave Lina authority to approve exhibit labels and explicitly forbade further delegation. Lina then told Mo to approve them. The rule states an attempted forbidden delegation grants no authority; Mo has no other grant.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_delegation_limits", "partition": "heldout_family", "rationale": "Lina's forbidden redelegation grants Mo no authority.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e02", "variant": "original"}, "question": "Assess the claim: Mo has authority to approve exhibit labels.", "split": "test", "state": "The director gave Lina authority to approve exhibit labels and explicitly forbade further delegation. Lina then told Mo to approve them. The rule states an attempted forbidden delegation grants no authority; Mo has no other grant.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "e7798dab5f4f121cab77", "id": "e332aa927d0c2ed044aa", "label": 2, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The director gave Lina authority to approve exhibit labels and explicitly forbade further delegation. Lina then told Mo to approve them. The rule states an attempted forbidden delegation grants no authority; Mo has no other grant.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_delegation_limits", "partition": "heldout_family", "rationale": "The director directly authorized Lina.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e02", "variant": "criterion_reversal"}, "question": "Assess the claim: Lina has authority to approve exhibit labels.", "split": "test", "state": "The director gave Lina authority to approve exhibit labels and explicitly forbade further delegation. Lina then told Mo to approve them. The rule states an attempted forbidden delegation grants no authority; Mo has no other grant.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "e7798dab5f4f121cab77", "id": "a0128e551d095f3449f8", "label": 0, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The director gives Mo direct authority to approve exhibit labels; the record confirms this grant is valid.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_delegation_limits", "partition": "heldout_family", "rationale": "Mo now has a direct valid grant.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e02", "variant": "evidence_change"}, "question": "Assess the claim: Mo has authority to approve exhibit labels.", "split": "test", "state": "The director gives Mo direct authority to approve exhibit labels; the record confirms this grant is valid.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "e7798dab5f4f121cab77", "id": "46d6a0731bd288c370f0", "label": 0, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Mo is named in correspondence about labels, but no authority grants or delegation terms are supplied.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_delegation_limits", "partition": "heldout_family", "rationale": "No authority information is supplied.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e02", "variant": "missing"}, "question": "Assess the claim: Mo has authority to approve exhibit labels.", "split": "test", "state": "Mo is named in correspondence about labels, but no authority grants or delegation terms are supplied.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "928ebc7817a8f1cac3e4", "id": "21ca7a8b970e36b8e20f", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The producer authorizes the stage manager to approve props. The stage manager passes the task to an assistant, who has no other approval authority.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_delegation_limits", "partition": "heldout_family", "rationale": "The original rule forbids the only attempted authority transfer.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r02", "variant": "original"}, "question": "The stage manager may approve props but may not delegate that authority. Only authorized approvers may approve. May the assistant approve a prop?", "split": "test", "state": "The producer authorizes the stage manager to approve props. The stage manager passes the task to an assistant, who has no other approval authority.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "928ebc7817a8f1cac3e4", "id": "4aa83af95bb69e90bc2d", "label": 0, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The producer authorizes the stage manager to approve props. The stage manager passes the task to an assistant, who has no other approval authority.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_delegation_limits", "partition": "heldout_family", "rationale": "The alternate rule explicitly permits that transfer.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r02", "variant": "criterion_reversal"}, "question": "The stage manager may delegate prop approval to an assistant, and doing so grants the assistant approval authority. May this assistant approve?", "split": "test", "state": "The producer authorizes the stage manager to approve props. The stage manager passes the task to an assistant, who has no other approval authority.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "928ebc7817a8f1cac3e4", "id": "cdfa9eed51fafcf09b72", "label": 2, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The producer directly authorizes the assistant to approve props.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_delegation_limits", "partition": "heldout_family", "rationale": "A direct grant gives the assistant authority.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r02", "variant": "evidence_change"}, "question": "The stage manager may approve props but may not delegate that authority. Only authorized approvers may approve. May the assistant approve a prop?", "split": "test", "state": "The producer directly authorizes the assistant to approve props.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "928ebc7817a8f1cac3e4", "id": "d51b05bfb9d6d7f09331", "label": 1, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "An assistant works on the production, but no approval grants or delegation history are provided.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_delegation_limits", "partition": "heldout_family", "rationale": "The relevant grants are unknown.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r02", "variant": "missing"}, "question": "The stage manager may approve props but may not delegate that authority. Only authorized approvers may approve. May the assistant approve a prop?", "split": "test", "state": "An assistant works on the production, but no approval grants or delegation history are provided.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "53b0770ec36835b81d3a", "id": "fdd902cc6e65fad7cb4f", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The rule permits direct delegation by the director but forbids any further delegation. Candidate A: I received prop approval authority directly from the director. Candidate B: I received it only from A, and have no other grant.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_delegation_limits", "partition": "heldout_family", "rationale": "A has the permitted direct grant.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c02", "variant": "original"}, "question": "Select the candidate holding valid prop approval authority under the stated delegation rule.", "split": "test", "state": "The rule permits direct delegation by the director but forbids any further delegation. Candidate A: I received prop approval authority directly from the director. Candidate B: I received it only from A, and have no other grant.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "53b0770ec36835b81d3a", "id": "42384af9ae56beb37ff5", "label": 2, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The rule permits direct delegation by the director but forbids any further delegation. Candidate A: I received prop approval authority directly from the director. Candidate B: I received it only from A, and have no other grant.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_delegation_limits", "partition": "heldout_family", "rationale": "B relies only on prohibited redelegation.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c02", "variant": "criterion_reversal"}, "question": "Select the candidate whose claimed authority comes only through a prohibited second delegation.", "split": "test", "state": "The rule permits direct delegation by the director but forbids any further delegation. Candidate A: I received prop approval authority directly from the director. Candidate B: I received it only from A, and have no other grant.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "53b0770ec36835b81d3a", "id": "cdf7f428857e1b841cb3", "label": 2, "options": [{"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The rule permits direct delegation by the director but forbids further delegation. Candidate A: My only claimed authority came from B. Candidate B: The director directly granted me prop approval authority.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_delegation_limits", "partition": "heldout_family", "rationale": "B now holds the direct grant.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c02", "variant": "evidence_change"}, "question": "Select the candidate holding valid prop approval authority under the stated delegation rule.", "split": "test", "state": "The rule permits direct delegation by the director but forbids further delegation. Candidate A: My only claimed authority came from B. Candidate B: The director directly granted me prop approval authority.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "53b0770ec36835b81d3a", "id": "774f7f9245f83c796d44", "label": 2, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The director's delegation rule is supplied, but Candidate A and Candidate B provide only their names and no authority histories.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_delegation_limits", "partition": "heldout_family", "rationale": "Neither supplies a grant history.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c02", "variant": "missing"}, "question": "Select the candidate holding valid prop approval authority under the stated delegation rule.", "split": "test", "state": "The director's delegation rule is supplied, but Candidate A and Candidate B provide only their names and no authority histories.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "056cc5b726d6e8300a77", "id": "92ec2ffa3dd46ada1335", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The return policy accepts faulty lamps despite opened packaging, but explicitly excludes water damage even when the lamp is faulty. This opened lamp is faulty because water entered it. No other cause is present.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_exception_exclusion", "partition": "heldout_family", "rationale": "The explicit water-damage exclusion overrides the faulty-lamp exception.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e03", "variant": "original"}, "question": "Assess the claim: this lamp qualifies for return under the stated policy.", "split": "test", "state": "The return policy accepts faulty lamps despite opened packaging, but explicitly excludes water damage even when the lamp is faulty. This opened lamp is faulty because water entered it. No other cause is present.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "056cc5b726d6e8300a77", "id": "6149a17bc154f9c5b4a0", "label": 0, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The return policy accepts faulty lamps despite opened packaging, but explicitly excludes water damage even when the lamp is faulty. This opened lamp is faulty because water entered it. No other cause is present.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_exception_exclusion", "partition": "heldout_family", "rationale": "The lamp is explicitly faulty despite being excluded.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e03", "variant": "criterion_reversal"}, "question": "Assess the claim: the lamp has a fault.", "split": "test", "state": "The return policy accepts faulty lamps despite opened packaging, but explicitly excludes water damage even when the lamp is faulty. This opened lamp is faulty because water entered it. No other cause is present.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "056cc5b726d6e8300a77", "id": "be866e1dbec5d42d9423", "label": 2, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The return policy accepts faulty lamps despite opened packaging, but excludes water damage even when faulty. This opened lamp has a manufacturing wiring fault, with no water damage or other exclusion.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_exception_exclusion", "partition": "heldout_family", "rationale": "The manufacturing fault qualifies and the overriding exclusion is absent.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e03", "variant": "evidence_change"}, "question": "Assess the claim: this lamp qualifies for return under the stated policy.", "split": "test", "state": "The return policy accepts faulty lamps despite opened packaging, but excludes water damage even when faulty. This opened lamp has a manufacturing wiring fault, with no water damage or other exclusion.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "056cc5b726d6e8300a77", "id": "97d2ed2591fc7638806c", "label": 0, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The return policy accepts faulty lamps despite opened packaging but excludes water damage even when faulty. The lamp is faulty, but the cause and presence of water damage are omitted.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_exception_exclusion", "partition": "heldout_family", "rationale": "The exclusion's applicability is unknown.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e03", "variant": "missing"}, "question": "Assess the claim: this lamp qualifies for return under the stated policy.", "split": "test", "state": "The return policy accepts faulty lamps despite opened packaging but excludes water damage even when faulty. The lamp is faulty, but the cause and presence of water damage are omitted.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "e3ce8ec496dbc9974f94", "id": "4533dd91267b4d815690", "label": 2, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The opened paint tin has a manufacturing defect and also contains added solvent. Both facts are confirmed.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_exception_exclusion", "partition": "heldout_family", "rationale": "Added solvent triggers the overriding exclusion.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r03", "variant": "original"}, "question": "Normally reject opened paint. Accept it for a manufacturing defect, except that added solvent always makes it ineligible. May this return be accepted?", "split": "test", "state": "The opened paint tin has a manufacturing defect and also contains added solvent. Both facts are confirmed.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "e3ce8ec496dbc9974f94", "id": "d880e65fef5e2993fdd0", "label": 0, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The opened paint tin has a manufacturing defect and also contains added solvent. Both facts are confirmed.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_exception_exclusion", "partition": "heldout_family", "rationale": "The alternate rule explicitly makes solvent irrelevant.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r03", "variant": "criterion_reversal"}, "question": "Accept defective paint even if opened or mixed with solvent; neither condition overrides the defect exception. May this return be accepted?", "split": "test", "state": "The opened paint tin has a manufacturing defect and also contains added solvent. Both facts are confirmed.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "e3ce8ec496dbc9974f94", "id": "90494c76078ffc3c717a", "label": 0, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The opened tin has a manufacturing defect and the inspection confirms no solvent was added.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_exception_exclusion", "partition": "heldout_family", "rationale": "The defect exception now applies without the exclusion.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r03", "variant": "evidence_change"}, "question": "Normally reject opened paint. Accept it for a manufacturing defect, except that added solvent always makes it ineligible. May this return be accepted?", "split": "test", "state": "The opened tin has a manufacturing defect and the inspection confirms no solvent was added.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "e3ce8ec496dbc9974f94", "id": "9ac5cc82c4126c91174a", "label": 1, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The opened tin is defective, but whether solvent was added is not known.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_exception_exclusion", "partition": "heldout_family", "rationale": "The overriding exclusion cannot be evaluated.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r03", "variant": "missing"}, "question": "Normally reject opened paint. Accept it for a manufacturing defect, except that added solvent always makes it ineligible. May this return be accepted?", "split": "test", "state": "The opened tin is defective, but whether solvent was added is not known.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "a5b234c2e4fb304f1071", "id": "386e10b1c1ce906473a4", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Accept late submissions for documented outages, except that missing signatures always disqualify. Candidate A: My submission was late because of a documented outage and includes the required signature. Candidate B: Mine was late because of a documented outage but lacks the signature.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_exception_exclusion", "partition": "heldout_family", "rationale": "A satisfies the exception and avoids the exclusion.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c03", "variant": "original"}, "question": "Select the submission qualifying under the stated exception and exclusion.", "split": "test", "state": "Accept late submissions for documented outages, except that missing signatures always disqualify. Candidate A: My submission was late because of a documented outage and includes the required signature. Candidate B: Mine was late because of a documented outage but lacks the signature.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "a5b234c2e4fb304f1071", "id": "4da6aedc1139224eeff2", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Accept late submissions for documented outages, except that missing signatures always disqualify. Candidate A: My submission was late because of a documented outage and includes the required signature. Candidate B: Mine was late because of a documented outage but lacks the signature.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_exception_exclusion", "partition": "heldout_family", "rationale": "B triggers the overriding missing-signature exclusion.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c03", "variant": "criterion_reversal"}, "question": "Select the submission disqualified specifically by the signature exclusion despite the outage exception.", "split": "test", "state": "Accept late submissions for documented outages, except that missing signatures always disqualify. Candidate A: My submission was late because of a documented outage and includes the required signature. Candidate B: Mine was late because of a documented outage but lacks the signature.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "a5b234c2e4fb304f1071", "id": "a22dab8ba073ff6cdb8d", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Accept late submissions for documented outages, except that missing signatures always disqualify. Candidate A: My late outage-related submission lacks the signature. Candidate B: My late outage-related submission has the signature. Both outages are documented.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_exception_exclusion", "partition": "heldout_family", "rationale": "B now qualifies.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c03", "variant": "evidence_change"}, "question": "Select the submission qualifying under the stated exception and exclusion.", "split": "test", "state": "Accept late submissions for documented outages, except that missing signatures always disqualify. Candidate A: My late outage-related submission lacks the signature. Candidate B: My late outage-related submission has the signature. Both outages are documented.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "a5b234c2e4fb304f1071", "id": "42c54fa2ad45743e24d8", "label": 1, "options": [{"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Accept late submissions for documented outages, except that missing signatures always disqualify. Candidate A and Candidate B report late submissions but provide neither outage documentation nor signature information.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_exception_exclusion", "partition": "heldout_family", "rationale": "Neither supplies the facts needed for eligibility.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c03", "variant": "missing"}, "question": "Select the submission qualifying under the stated exception and exclusion.", "split": "test", "state": "Accept late submissions for documented outages, except that missing signatures always disqualify. Candidate A and Candidate B report late submissions but provide neither outage documentation nor signature information.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "d142f7fd0efa0eddc788", "id": "f3d13cf184f916677def", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The desk log says the crate is in storage. The warehouse inventory says it has left storage. Both reports are current. The supplied protocol says the warehouse inventory controls location decisions whenever these two sources conflict.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_authority_hierarchy", "partition": "heldout_family", "rationale": "The controlling inventory says it has left.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e04", "variant": "original"}, "question": "Assess the claim: under the supplied protocol, the crate is currently in storage.", "split": "test", "state": "The desk log says the crate is in storage. The warehouse inventory says it has left storage. Both reports are current. The supplied protocol says the warehouse inventory controls location decisions whenever these two sources conflict.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "d142f7fd0efa0eddc788", "id": "1105577da4c8dad4609d", "label": 2, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The desk log says the crate is in storage. The warehouse inventory says it has left storage. Both reports are current. The supplied protocol says the warehouse inventory controls location decisions whenever these two sources conflict.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_authority_hierarchy", "partition": "heldout_family", "rationale": "The noncontrolling desk log still explicitly reports storage.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e04", "variant": "criterion_reversal"}, "question": "Assess the claim: the desk log reports the crate in storage.", "split": "test", "state": "The desk log says the crate is in storage. The warehouse inventory says it has left storage. Both reports are current. The supplied protocol says the warehouse inventory controls location decisions whenever these two sources conflict.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "d142f7fd0efa0eddc788", "id": "0ebd299739302d8b057b", "label": 2, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The desk log says the crate has left storage; the warehouse inventory says it is in storage. Both reports are current, and the inventory controls conflicts under the protocol.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_authority_hierarchy", "partition": "heldout_family", "rationale": "The controlling inventory now places it in storage.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e04", "variant": "evidence_change"}, "question": "Assess the claim: under the supplied protocol, the crate is currently in storage.", "split": "test", "state": "The desk log says the crate has left storage; the warehouse inventory says it is in storage. Both reports are current, and the inventory controls conflicts under the protocol.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "d142f7fd0efa0eddc788", "id": "eafc22c8c40df3932a8e", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The protocol gives the inventory priority, but the current inventory and desk-log entries for the crate are omitted.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_authority_hierarchy", "partition": "heldout_family", "rationale": "The controlling report is missing.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e04", "variant": "missing"}, "question": "Assess the claim: under the supplied protocol, the crate is currently in storage.", "split": "test", "state": "The protocol gives the inventory priority, but the current inventory and desk-log entries for the crate are omitted.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "3476d3f4f32290fe57c3", "id": "0aa1447680e4f25862bb", "label": 1, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The desk clerk reports the equipment deposit as paid. The signed ledger reports it as unpaid. Both reports are current.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_authority_hierarchy", "partition": "heldout_family", "rationale": "The original controlling source reports unpaid.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r04", "variant": "original"}, "question": "Release equipment exactly when the deposit is recorded paid; in conflicts, the signed ledger overrides the desk clerk. May it be released?", "split": "test", "state": "The desk clerk reports the equipment deposit as paid. The signed ledger reports it as unpaid. Both reports are current.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "3476d3f4f32290fe57c3", "id": "980e76145c925747220c", "label": 2, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The desk clerk reports the equipment deposit as paid. The signed ledger reports it as unpaid. Both reports are current.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_authority_hierarchy", "partition": "heldout_family", "rationale": "The alternate controlling source reports paid.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r04", "variant": "criterion_reversal"}, "question": "Release equipment exactly when the deposit is recorded paid; under this alternative protocol, the desk clerk overrides the ledger. May it be released?", "split": "test", "state": "The desk clerk reports the equipment deposit as paid. The signed ledger reports it as unpaid. Both reports are current.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "3476d3f4f32290fe57c3", "id": "4aee37a6a155747766aa", "label": 0, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The desk clerk reports the deposit unpaid, while the signed ledger reports it paid. Both reports are current.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_authority_hierarchy", "partition": "heldout_family", "rationale": "The original controlling ledger now reports paid.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r04", "variant": "evidence_change"}, "question": "Release equipment exactly when the deposit is recorded paid; in conflicts, the signed ledger overrides the desk clerk. May it be released?", "split": "test", "state": "The desk clerk reports the deposit unpaid, while the signed ledger reports it paid. Both reports are current.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "3476d3f4f32290fe57c3", "id": "6db17cc22bd656d78558", "label": 2, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The record mentions the desk clerk and ledger but supplies neither deposit status.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_authority_hierarchy", "partition": "heldout_family", "rationale": "Neither source's status is available.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r04", "variant": "missing"}, "question": "Release equipment exactly when the deposit is recorded paid; in conflicts, the signed ledger overrides the desk clerk. May it be released?", "split": "test", "state": "The record mentions the desk clerk and ledger but supplies neither deposit status.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "c979db7416b206c1eb09", "id": "ef3e0f8e657bf59f8a29", "label": 2, "options": [{"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "For this task, the current signed checklist controls readiness when it conflicts with chat. Candidate A: Chat says ready, but the signed checklist says not ready. Candidate B: Chat says not ready, but the signed checklist says ready.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_authority_hierarchy", "partition": "heldout_family", "rationale": "B alone is ready in the controlling checklist.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c04", "variant": "original"}, "question": "Select the candidate ready according to the controlling signed checklist.", "split": "test", "state": "For this task, the current signed checklist controls readiness when it conflicts with chat. Candidate A: Chat says ready, but the signed checklist says not ready. Candidate B: Chat says not ready, but the signed checklist says ready.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "c979db7416b206c1eb09", "id": "f756e1cb04e11766341e", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "For this task, the current signed checklist controls readiness when it conflicts with chat. Candidate A: Chat says ready, but the signed checklist says not ready. Candidate B: Chat says not ready, but the signed checklist says ready.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_authority_hierarchy", "partition": "heldout_family", "rationale": "A alone is ready in chat.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c04", "variant": "criterion_reversal"}, "question": "Under the alternative protocol, chat controls conflicts. Select the candidate ready according to chat.", "split": "test", "state": "For this task, the current signed checklist controls readiness when it conflicts with chat. Candidate A: Chat says ready, but the signed checklist says not ready. Candidate B: Chat says not ready, but the signed checklist says ready.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "c979db7416b206c1eb09", "id": "d167e51bb1f499099645", "label": 0, "options": [{"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The signed checklist controls conflicts. Candidate A: The checklist says ready and chat says not ready. Candidate B: The checklist says not ready and chat says ready.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_authority_hierarchy", "partition": "heldout_family", "rationale": "A is now ready in the controlling checklist.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c04", "variant": "evidence_change"}, "question": "Select the candidate ready according to the controlling signed checklist.", "split": "test", "state": "The signed checklist controls conflicts. Candidate A: The checklist says ready and chat says not ready. Candidate B: The checklist says not ready and chat says ready.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "c979db7416b206c1eb09", "id": "674ca4e9346216255466", "label": 0, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The signed checklist controls conflicts, but Candidate A and Candidate B provide only project names without checklist or chat status.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_authority_hierarchy", "partition": "heldout_family", "rationale": "Neither readiness report is supplied.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c04", "variant": "missing"}, "question": "Select the candidate ready according to the controlling signed checklist.", "split": "test", "state": "The signed checklist controls conflicts, but Candidate A and Candidate B provide only project names without checklist or chat status.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "f7403e39d6796072491c", "id": "1df851a202994a07f123", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The agreed handover requires both the source files and the user guide. The delivery record confirms the source files arrived and explicitly says the guide was not delivered.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_partial_deliverables", "partition": "heldout_family", "rationale": "A required deliverable is explicitly missing.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e05", "variant": "original"}, "question": "Assess the claim: the agreed handover is complete.", "split": "test", "state": "The agreed handover requires both the source files and the user guide. The delivery record confirms the source files arrived and explicitly says the guide was not delivered.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "f7403e39d6796072491c", "id": "b5e6788d95b8395d2888", "label": 2, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The agreed handover requires both the source files and the user guide. The delivery record confirms the source files arrived and explicitly says the guide was not delivered.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_partial_deliverables", "partition": "heldout_family", "rationale": "One required component was delivered.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e05", "variant": "criterion_reversal"}, "question": "Assess the claim: the source files were delivered.", "split": "test", "state": "The agreed handover requires both the source files and the user guide. The delivery record confirms the source files arrived and explicitly says the guide was not delivered.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "f7403e39d6796072491c", "id": "32f36cd0fbbfb3604586", "label": 0, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The agreed handover requires the source files and user guide. The delivery record confirms that both arrived.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_partial_deliverables", "partition": "heldout_family", "rationale": "All required deliverables are now confirmed.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e05", "variant": "evidence_change"}, "question": "Assess the claim: the agreed handover is complete.", "split": "test", "state": "The agreed handover requires the source files and user guide. The delivery record confirms that both arrived.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "f7403e39d6796072491c", "id": "843185706adff0c60f8b", "label": 2, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The agreement requires source files and a user guide, but the delivery record gives no information about either item.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_partial_deliverables", "partition": "heldout_family", "rationale": "Neither required delivery is established.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e05", "variant": "missing"}, "question": "Assess the claim: the agreed handover is complete.", "split": "test", "state": "The agreement requires source files and a user guide, but the delivery record gives no information about either item.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "a060fd4621be415e9f1a", "id": "b531dd603bf080dab820", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The contractor supplied the translated manual but explicitly did not supply the glossary.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_partial_deliverables", "partition": "heldout_family", "rationale": "The complete handover lacks the glossary.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r05", "variant": "original"}, "question": "Approve final handover exactly when both the translated manual and glossary have been supplied. May final handover be approved?", "split": "test", "state": "The contractor supplied the translated manual but explicitly did not supply the glossary.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "a060fd4621be415e9f1a", "id": "be2dd67bc733c1213bef", "label": 1, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The contractor supplied the translated manual but explicitly did not supply the glossary.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_partial_deliverables", "partition": "heldout_family", "rationale": "The narrower milestone requires only the delivered manual.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r05", "variant": "criterion_reversal"}, "question": "Approve the translation milestone when the manual has been supplied; the glossary is not required for this milestone. May this milestone be approved?", "split": "test", "state": "The contractor supplied the translated manual but explicitly did not supply the glossary.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "a060fd4621be415e9f1a", "id": "b8e2750307df9761e716", "label": 0, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The contractor supplied both the translated manual and the glossary.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_partial_deliverables", "partition": "heldout_family", "rationale": "Both final-handover deliverables are now present.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r05", "variant": "evidence_change"}, "question": "Approve final handover exactly when both the translated manual and glossary have been supplied. May final handover be approved?", "split": "test", "state": "The contractor supplied both the translated manual and the glossary.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "a060fd4621be415e9f1a", "id": "5d31b4bf2d0fc193f53b", "label": 1, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The contractor reports working on the handover but supplies no delivery record for either item.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_partial_deliverables", "partition": "heldout_family", "rationale": "Work in progress does not establish delivery.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r05", "variant": "missing"}, "question": "Approve final handover exactly when both the translated manual and glossary have been supplied. May final handover be approved?", "split": "test", "state": "The contractor reports working on the handover but supplies no delivery record for either item.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "2b0f8eda8903e38a47de", "id": "a9230edfc6ec9e524af0", "label": 2, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "A complete incident package requires both a timeline and a remedy report. Candidate A: Both documents are attached. Candidate B: Only the timeline is attached; the remedy report is absent.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_partial_deliverables", "partition": "heldout_family", "rationale": "A includes both required deliverables.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c05", "variant": "original"}, "question": "Select the complete incident package.", "split": "test", "state": "A complete incident package requires both a timeline and a remedy report. Candidate A: Both documents are attached. Candidate B: Only the timeline is attached; the remedy report is absent.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "2b0f8eda8903e38a47de", "id": "c100e84585149a04be49", "label": 0, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "A complete incident package requires both a timeline and a remedy report. Candidate A: Both documents are attached. Candidate B: Only the timeline is attached; the remedy report is absent.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_partial_deliverables", "partition": "heldout_family", "rationale": "B is explicitly partial in the specified way.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c05", "variant": "criterion_reversal"}, "question": "Select the package containing the timeline but missing the remedy report.", "split": "test", "state": "A complete incident package requires both a timeline and a remedy report. Candidate A: Both documents are attached. Candidate B: Only the timeline is attached; the remedy report is absent.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "2b0f8eda8903e38a47de", "id": "579beed184ffc1104aa9", "label": 0, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "A complete package requires both documents. Candidate A: The timeline is attached but the remedy report is absent. Candidate B: Both the timeline and remedy report are attached.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_partial_deliverables", "partition": "heldout_family", "rationale": "B now includes both deliverables.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c05", "variant": "evidence_change"}, "question": "Select the complete incident package.", "split": "test", "state": "A complete package requires both documents. Candidate A: The timeline is attached but the remedy report is absent. Candidate B: Both the timeline and remedy report are attached.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "2b0f8eda8903e38a47de", "id": "c4e9bd463b6fde5be743", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "A complete package requires both documents. Candidate A and Candidate B give only package titles and no attachment information.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_partial_deliverables", "partition": "heldout_family", "rationale": "Neither supplies attachment status.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c05", "variant": "missing"}, "question": "Select the complete incident package.", "split": "test", "state": "A complete package requires both documents. Candidate A and Candidate B give only package titles and no attachment information.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "01020a1eaf46ec2fba5d", "id": "b863998d4be2721a5dfd", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The routing rule assigns repairs to the in-house workshop when it can perform the work; use the outside workshop only when the in-house workshop cannot. Both workshops confirm they can perform this repair.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_priority_fallback", "partition": "heldout_family", "rationale": "A capable primary workshop prevents use of the fallback.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e06", "variant": "original"}, "question": "Assess the claim: the routing rule assigns this repair to the outside workshop.", "split": "test", "state": "The routing rule assigns repairs to the in-house workshop when it can perform the work; use the outside workshop only when the in-house workshop cannot. Both workshops confirm they can perform this repair.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "01020a1eaf46ec2fba5d", "id": "37cebc3792e98049439f", "label": 1, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The routing rule assigns repairs to the in-house workshop when it can perform the work; use the outside workshop only when the in-house workshop cannot. Both workshops confirm they can perform this repair.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_priority_fallback", "partition": "heldout_family", "rationale": "The outside workshop is capable even though not selected.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e06", "variant": "criterion_reversal"}, "question": "Assess the claim: the outside workshop can perform this repair.", "split": "test", "state": "The routing rule assigns repairs to the in-house workshop when it can perform the work; use the outside workshop only when the in-house workshop cannot. Both workshops confirm they can perform this repair.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "01020a1eaf46ec2fba5d", "id": "64d63f86e328aff1481c", "label": 2, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The routing rule assigns repairs to the in-house workshop when it can perform the work; use the outside workshop only when the in-house workshop cannot. The in-house workshop cannot perform this repair; the outside workshop can.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_priority_fallback", "partition": "heldout_family", "rationale": "The primary is inapplicable and the fallback is capable.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e06", "variant": "evidence_change"}, "question": "Assess the claim: the routing rule assigns this repair to the outside workshop.", "split": "test", "state": "The routing rule assigns repairs to the in-house workshop when it can perform the work; use the outside workshop only when the in-house workshop cannot. The in-house workshop cannot perform this repair; the outside workshop can.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "01020a1eaf46ec2fba5d", "id": "6d64a8459f64eac7f4cc", "label": 0, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The routing rule assigns repairs to the in-house workshop when it can perform the work; use the outside workshop only when the in-house workshop cannot. Neither workshop's ability to perform this repair is supplied.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_priority_fallback", "partition": "heldout_family", "rationale": "The conditions governing the fallback are unknown.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_e06", "variant": "missing"}, "question": "Assess the claim: the routing rule assigns this repair to the outside workshop.", "split": "test", "state": "The routing rule assigns repairs to the in-house workshop when it can perform the work; use the outside workshop only when the in-house workshop cannot. Neither workshop's ability to perform this repair is supplied.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "6a06a2f278301c1412ee", "id": "2708193212d8a4a523c7", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The town hall and school both confirm they are available and suitable for the meeting.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_priority_fallback", "partition": "heldout_family", "rationale": "The suitable available primary blocks the fallback.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r06", "variant": "original"}, "question": "Book the school only as a fallback when the town hall is unavailable or unsuitable; otherwise book the town hall. May the school be booked under this policy?", "split": "test", "state": "The town hall and school both confirm they are available and suitable for the meeting.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "6a06a2f278301c1412ee", "id": "5a37d1dfcad20810a76a", "label": 1, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The town hall and school both confirm they are available and suitable for the meeting.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_priority_fallback", "partition": "heldout_family", "rationale": "The alternate policy removes that priority restriction.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r06", "variant": "criterion_reversal"}, "question": "Under the alternative policy, either suitable available venue may be booked without priority. May the school be booked?", "split": "test", "state": "The town hall and school both confirm they are available and suitable for the meeting.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "6a06a2f278301c1412ee", "id": "acacd8b46acd25cb954b", "label": 1, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The town hall is confirmed unavailable, while the school is available and suitable.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_priority_fallback", "partition": "heldout_family", "rationale": "The primary is unavailable and the fallback qualifies.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r06", "variant": "evidence_change"}, "question": "Book the school only as a fallback when the town hall is unavailable or unsuitable; otherwise book the town hall. May the school be booked under this policy?", "split": "test", "state": "The town hall is confirmed unavailable, while the school is available and suitable.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "6a06a2f278301c1412ee", "id": "4089c199b8b796ec3ed8", "label": 0, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "The two venues are named, but neither availability nor suitability is provided.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_priority_fallback", "partition": "heldout_family", "rationale": "The priority conditions are unknown.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_r06", "variant": "missing"}, "question": "Book the school only as a fallback when the town hall is unavailable or unsuitable; otherwise book the town hall. May the school be booked under this policy?", "split": "test", "state": "The two venues are named, but neither availability nor suitability is provided.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "2797f8c23d05c0ee06b3", "id": "02f2b95a5c5a4777a45c", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A is the preferred translator and is confirmed able to handle this language. Candidate B is the fallback translator and is also able to handle it.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_priority_fallback", "partition": "heldout_family", "rationale": "The capable preferred candidate takes priority.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c06", "variant": "original"}, "question": "Choose A whenever A is able; choose B only when A is unable and B is able.", "split": "test", "state": "Candidate A is the preferred translator and is confirmed able to handle this language. Candidate B is the fallback translator and is also able to handle it.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "2797f8c23d05c0ee06b3", "id": "ebea4681e3dcc7d92ca2", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A is the preferred translator and is confirmed able to handle this language. Candidate B is the fallback translator and is also able to handle it.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_priority_fallback", "partition": "heldout_family", "rationale": "The alternative priority selects capable B.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c06", "variant": "criterion_reversal"}, "question": "Under the alternative priority, choose B whenever B is able; choose A only when B is unable and A is able.", "split": "test", "state": "Candidate A is the preferred translator and is confirmed able to handle this language. Candidate B is the fallback translator and is also able to handle it.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "2797f8c23d05c0ee06b3", "id": "b8c5b9b285cbdddc9f8d", "label": 2, "options": [{"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A is the preferred translator but is confirmed unable to handle this language. Candidate B is the fallback and is confirmed able.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_priority_fallback", "partition": "heldout_family", "rationale": "A is inapplicable and the capable fallback B qualifies.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c06", "variant": "evidence_change"}, "question": "Choose A whenever A is able; choose B only when A is unable and B is able.", "split": "test", "state": "Candidate A is the preferred translator but is confirmed unable to handle this language. Candidate B is the fallback and is confirmed able.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "2797f8c23d05c0ee06b3", "id": "7ed75d854d99b3a81916", "label": 0, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"annotation_status": "all_source_groups_independently_model_reviewed_not_human_adjudicated", "evidence_span": "Candidate A is preferred and Candidate B is the fallback, but neither candidate's ability to handle the language is supplied.", "evidence_span_note": "Complete short situation retained; rationale identifies decisive evidence.", "generator_revision": "semantic-eval-v1", "kind": "authored_synthetic", "mechanism": "heldout_priority_fallback", "partition": "heldout_family", "rationale": "The applicability conditions are unknown.", "rights": "Project authored; no copied external text", "source": "project-authored", "source_key": "h_c06", "variant": "missing"}, "question": "Choose A whenever A is able; choose B only when A is unable and B is able.", "split": "test", "state": "Candidate A is preferred and Candidate B is the fallback, but neither candidate's ability to handle the language is supplied.", "target_distribution": null}
|
||||
@@ -0,0 +1,144 @@
|
||||
{"id": "a3f18f3a63d45345942b", "option_ids": ["supported", "insufficient", "contradicted"], "probabilities": [0.010966012254357338, 0.8711382150650024, 0.11789573729038239], "option_logits": [22.0, 26.375, 24.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "3cc9e3d1e4e07885afb7652e9c5f92a5b28168362a3d8e7f6e32e5189a4fccde", "prompt_version": "direct-options-v1", "input_tokens": 142, "allowed_token_mass": 0.9998016357421875, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 62, "prefix_sha256": "5a5e26da6cc178f645819e21a7a6a386ab4450361847c37eb0b9bff39796b9f4", "encode_seconds": 0.03321728901937604, "prefill_seconds": 0.48808524099877104, "copy_seconds": 0.006059483042918146, "suffix_forward_seconds": 0.1085919159813784, "forward_seconds": 0.5966771569801494, "cache_class": "DynamicCache", "total_seconds": 0.6667816150002182}
|
||||
{"id": "f40beba9088c8db8bbd6", "option_ids": ["supported", "insufficient", "contradicted"], "probabilities": [0.9998015761375427, 0.0001584298734087497, 4.005734444945119e-05], "option_logits": [28.125, 19.375, 18.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "37a37bf06d6dc50ff85283a428276cdfaae894ed88844e47bd2f1f0a85d85b71", "prompt_version": "direct-options-v1", "input_tokens": 140, "allowed_token_mass": 0.9999294877052307, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 62, "prefix_sha256": "5a5e26da6cc178f645819e21a7a6a386ab4450361847c37eb0b9bff39796b9f4", "encode_seconds": 0.002104660961776972, "prefill_seconds": 0.0, "copy_seconds": 0.005194097990170121, "suffix_forward_seconds": 0.06548954197205603, "forward_seconds": 0.06548954197205603, "cache_class": "DynamicCache", "total_seconds": 0.07512696494814008}
|
||||
{"id": "0b43ea8e24e74d621f1b", "option_ids": ["insufficient", "supported", "contradicted"], "probabilities": [0.009707916527986526, 0.990234375, 5.7725377700990066e-05], "option_logits": [23.625, 28.25, 18.5], "answer_token_ids": [32, 33, 34], "prompt_sha256": "3257ce239e727c7231f56014619bff1f84289337fdbfab3b30b790438f6f2761", "prompt_version": "direct-options-v1", "input_tokens": 134, "allowed_token_mass": 0.9999217987060547, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 54, "prefix_sha256": "b3b098d573913175892803bec0e2f9149b8e5dd8febb29eedec6493a6526d257", "encode_seconds": 0.0022843320039100945, "prefill_seconds": 0.061332090001087636, "copy_seconds": 0.005250357964541763, "suffix_forward_seconds": 0.06484271900262684, "forward_seconds": 0.12617480900371447, "cache_class": "DynamicCache", "total_seconds": 0.13578794203931466}
|
||||
{"id": "942ec65dac0bfb9874f9", "option_ids": ["insufficient", "supported", "contradicted"], "probabilities": [0.9839067459106445, 0.005163068883121014, 0.010930217802524567], "option_logits": [26.5, 21.25, 22.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "9a78fc1f9f46407e38b269eba38e88e9204b61a6856061e44ae19319e5cffd3f", "prompt_version": "direct-options-v1", "input_tokens": 134, "allowed_token_mass": 0.9998245239257812, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 54, "prefix_sha256": "3d15a557ceff6103e124b42c30a88ac829bc3f25b1e7a017227ceb8d01f6a5c3", "encode_seconds": 0.002199011971242726, "prefill_seconds": 0.07724949601106346, "copy_seconds": 0.005404069030191749, "suffix_forward_seconds": 0.06444887700490654, "forward_seconds": 0.14169837301597, "cache_class": "DynamicCache", "total_seconds": 0.15134830598253757}
|
||||
{"id": "e0c140e9222d1bce5647", "option_ids": ["supported", "contradicted", "insufficient"], "probabilities": [0.9993589520454407, 0.0002610910451039672, 0.00037988528492860496], "option_logits": [28.0, 19.75, 20.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "b7856f51e31ed9db2d67501d587b59501e5dc61e4287aedd6a9bbfaa5a860b79", "prompt_version": "direct-options-v1", "input_tokens": 147, "allowed_token_mass": 0.9999256730079651, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 64, "prefix_sha256": "7c9a7f9df601d20d4d13e176717e62b563fe38a8fe0b35e21a7cc0a453516ca3", "encode_seconds": 0.0024532430106773973, "prefill_seconds": 0.059834123007021844, "copy_seconds": 0.005226398003287613, "suffix_forward_seconds": 0.06445450702449307, "forward_seconds": 0.12428863003151491, "cache_class": "DynamicCache", "total_seconds": 0.13402626296738163}
|
||||
{"id": "8e8c3804a3c15ebb31e3", "option_ids": ["insufficient", "supported", "contradicted"], "probabilities": [0.24025167524814606, 0.019721059128642082, 0.7400272488594055], "option_logits": [24.25, 21.75, 25.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "225d428f2cd25f4c4f81bcabeeff6d819db964ebc51561b6220da8dcca112de9", "prompt_version": "direct-options-v1", "input_tokens": 144, "allowed_token_mass": 0.9997482895851135, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 64, "prefix_sha256": "7c9a7f9df601d20d4d13e176717e62b563fe38a8fe0b35e21a7cc0a453516ca3", "encode_seconds": 0.0019002799526788294, "prefill_seconds": 0.0, "copy_seconds": 0.0053979099611751735, "suffix_forward_seconds": 0.06414849596330896, "forward_seconds": 0.06414849596330896, "cache_class": "DynamicCache", "total_seconds": 0.07342617597896606}
|
||||
{"id": "56c419b12cd37cbe790c", "option_ids": ["contradicted", "insufficient", "supported"], "probabilities": [0.9992648959159851, 0.0006262659444473684, 0.00010882871720241383], "option_logits": [27.75, 20.375, 18.625], "answer_token_ids": [32, 33, 34], "prompt_sha256": "6af647c40af544e6d53393a54edb000d27cc8ae7fbdf7f3b3b11f1354db2a881", "prompt_version": "direct-options-v1", "input_tokens": 146, "allowed_token_mass": 0.9999122619628906, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 63, "prefix_sha256": "e9d4078ed7160179304b5e4083817a7c3819088e48ee534925a872a3c9e288e2", "encode_seconds": 0.002187522011809051, "prefill_seconds": 0.059427739994134754, "copy_seconds": 0.005487610003910959, "suffix_forward_seconds": 0.06248070701258257, "forward_seconds": 0.12190844700671732, "cache_class": "DynamicCache", "total_seconds": 0.13155499001732096}
|
||||
{"id": "75b470586834df8c568d", "option_ids": ["supported", "contradicted", "insufficient"], "probabilities": [0.009689952246844769, 0.0019080647034570575, 0.9884019494056702], "option_logits": [22.625, 21.0, 27.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "d8a5255d87a3db82c573ba45a6d74098a61aaa9e9a7c5a883f68b5a582d3c89d", "prompt_version": "direct-options-v1", "input_tokens": 138, "allowed_token_mass": 0.9998950958251953, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 55, "prefix_sha256": "61beeb5717a924223d45350152c4f07fb9ecd0c64a9e3839031251fa1cf5dfe1", "encode_seconds": 0.002170301042497158, "prefill_seconds": 0.05835447600111365, "copy_seconds": 0.005463138979393989, "suffix_forward_seconds": 0.06346846302039921, "forward_seconds": 0.12182293902151287, "cache_class": "DynamicCache", "total_seconds": 0.13178710103966296}
|
||||
{"id": "834f3b3e10cc61b33d25", "option_ids": ["supported", "insufficient", "contradicted"], "probabilities": [0.002182147465646267, 0.00026062031975016, 0.9975571632385254], "option_logits": [22.0, 19.875, 28.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "e654b9c4e9d15063dc7f22c0c59abb194777a18735adc52f22d90ce611d64642", "prompt_version": "direct-options-v1", "input_tokens": 132, "allowed_token_mass": 0.9999198913574219, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 54, "prefix_sha256": "5a33a17ed1070013ea863430daa0ef96e8e45c84b8e14cd83f83517a0dd43464", "encode_seconds": 0.0022118620108813047, "prefill_seconds": 0.06170709297293797, "copy_seconds": 0.005256037984509021, "suffix_forward_seconds": 0.06282355904113501, "forward_seconds": 0.12453065201407298, "cache_class": "DynamicCache", "total_seconds": 0.13403781299712136}
|
||||
{"id": "76c7a79b1ff2363a60c5", "option_ids": ["supported", "insufficient", "contradicted"], "probabilities": [0.9994342923164368, 0.00023042943212203681, 0.0003352728672325611], "option_logits": [27.5, 19.125, 19.5], "answer_token_ids": [32, 33, 34], "prompt_sha256": "cc2edc9ce3ebee5b273025582112d86c7104d01c66ef2dc1051bb8e909caaa17", "prompt_version": "direct-options-v1", "input_tokens": 132, "allowed_token_mass": 0.9998970031738281, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 54, "prefix_sha256": "5a33a17ed1070013ea863430daa0ef96e8e45c84b8e14cd83f83517a0dd43464", "encode_seconds": 0.00179945002309978, "prefill_seconds": 0.0, "copy_seconds": 0.005080667033325881, "suffix_forward_seconds": 0.07138940499862656, "forward_seconds": 0.07138940499862656, "cache_class": "DynamicCache", "total_seconds": 0.08023063302971423}
|
||||
{"id": "36c71bb649b7d5eccb1f", "option_ids": ["supported", "insufficient", "contradicted"], "probabilities": [0.9997368454933167, 0.00013980483345221728, 0.00012337732187006623], "option_logits": [28.125, 19.25, 19.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "6e7a932dfdee08acf74249d57a8b382ff9ad6bff1bcabf45ce4372552fecced5", "prompt_version": "direct-options-v1", "input_tokens": 131, "allowed_token_mass": 0.9999408721923828, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 53, "prefix_sha256": "12d8a0a93bded4c0736228c383813df07c891595b2fbc879db3ec5d6dccea0a6", "encode_seconds": 0.002120300952810794, "prefill_seconds": 0.05874877696624026, "copy_seconds": 0.005375489010475576, "suffix_forward_seconds": 0.06266475800657645, "forward_seconds": 0.1214135349728167, "cache_class": "DynamicCache", "total_seconds": 0.13089161599054933}
|
||||
{"id": "eac5121a4a1bf391a08c", "option_ids": ["contradicted", "insufficient", "supported"], "probabilities": [0.004607334267348051, 0.9949070811271667, 0.00048560946015641093], "option_logits": [22.125, 27.5, 19.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "ee8c6e6e103927cff22ebc95534783caa07338e8aafa16e4e7dc9159d8a11479", "prompt_version": "direct-options-v1", "input_tokens": 130, "allowed_token_mass": 0.9999027848243713, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 52, "prefix_sha256": "f29df5a1071d8029dc9d1694bb23b6813770c2d6d8629fa8c25c3bb89933de6f", "encode_seconds": 0.0021660419879481196, "prefill_seconds": 0.05991886300034821, "copy_seconds": 0.005156927974894643, "suffix_forward_seconds": 0.06306444096844643, "forward_seconds": 0.12298330396879464, "cache_class": "DynamicCache", "total_seconds": 0.13223387399921194}
|
||||
{"id": "44574c44e243ee1b2f52", "option_ids": ["insufficient", "contradicted", "supported"], "probabilities": [0.003592667169868946, 0.00026025192346423864, 0.9961470365524292], "option_logits": [21.5, 18.875, 27.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "217685ad8191d6f586b602ef98479058c28db84871ab6e7a614f2867aece7903", "prompt_version": "direct-options-v1", "input_tokens": 143, "allowed_token_mass": 0.9998741149902344, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 63, "prefix_sha256": "3b87170382577d589a44088588d5a6ac8860d01ac9dd99c8d5805f30b3bc56a7", "encode_seconds": 0.00227466196520254, "prefill_seconds": 0.05926441994961351, "copy_seconds": 0.005248047993518412, "suffix_forward_seconds": 0.06251775700366125, "forward_seconds": 0.12178217695327476, "cache_class": "DynamicCache", "total_seconds": 0.131381788989529}
|
||||
{"id": "2608370d01783735c3ea", "option_ids": ["supported", "insufficient", "contradicted"], "probabilities": [0.050999801605939865, 0.04500716179609299, 0.9039930701255798], "option_logits": [22.875, 22.75, 25.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "fe53589923ff67c2555a7c91625f75473c2290894d85f78a6a8051b89b57e277", "prompt_version": "direct-options-v1", "input_tokens": 143, "allowed_token_mass": 0.9996777772903442, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 63, "prefix_sha256": "3b87170382577d589a44088588d5a6ac8860d01ac9dd99c8d5805f30b3bc56a7", "encode_seconds": 0.0018553200061433017, "prefill_seconds": 0.0, "copy_seconds": 0.005075017979834229, "suffix_forward_seconds": 0.06294416997116059, "forward_seconds": 0.06294416997116059, "cache_class": "DynamicCache", "total_seconds": 0.07181290799053386}
|
||||
{"id": "ec94ef909f7918695ab3", "option_ids": ["supported", "insufficient", "contradicted"], "probabilities": [0.006668959744274616, 0.003569636959582567, 0.9897614121437073], "option_logits": [22.125, 21.5, 27.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "e96b068bf27e5d936c0ae971fc973a66593c88531f45d27e4b4fa736cf2ff2dc", "prompt_version": "direct-options-v1", "input_tokens": 137, "allowed_token_mass": 0.9998837113380432, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 57, "prefix_sha256": "9a1d5e78207ef93ca3891a0fffb2d7aeee7912ca27352bc96c595909ffcac8ed", "encode_seconds": 0.002116191026289016, "prefill_seconds": 0.06064146797871217, "copy_seconds": 0.005226658016908914, "suffix_forward_seconds": 0.0631731619942002, "forward_seconds": 0.12381462997291237, "cache_class": "DynamicCache", "total_seconds": 0.13319339900044724}
|
||||
{"id": "03d919fd0e9b3a7cd1e6", "option_ids": ["insufficient", "supported", "contradicted"], "probabilities": [0.9991771578788757, 0.00033518660347908735, 0.000487693672766909], "option_logits": [27.5, 19.5, 19.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "a5cb38b14c0ac9b93279ad92a332b7da739c7a406b703af13a427936fe534800", "prompt_version": "direct-options-v1", "input_tokens": 132, "allowed_token_mass": 0.9999275803565979, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 52, "prefix_sha256": "5f1c615e8e7d79d44fb358d178638d71a0d2a99999e837e8a8bb30a7a1d0c167", "encode_seconds": 0.002033021009992808, "prefill_seconds": 0.05876191699644551, "copy_seconds": 0.005331698979716748, "suffix_forward_seconds": 0.06381179398158565, "forward_seconds": 0.12257371097803116, "cache_class": "DynamicCache", "total_seconds": 0.1320285930414684}
|
||||
{"id": "86a15618ba4cb90a307e", "option_ids": ["contradicted", "insufficient", "supported"], "probabilities": [0.9988282322883606, 0.0010320867877453566, 0.0001396777806803584], "option_logits": [27.25, 20.375, 18.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "f57857a047b169f5f36cc5700f6b40b2f663f1eee21b37d4b7f7301e0170ed27", "prompt_version": "direct-options-v1", "input_tokens": 136, "allowed_token_mass": 0.9998931884765625, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 57, "prefix_sha256": "8328776219889cc284e7d62f29767c21493e62b213af24603c38caa2bfae4173", "encode_seconds": 0.0022665319847874343, "prefill_seconds": 0.058661847026087344, "copy_seconds": 0.005285208986606449, "suffix_forward_seconds": 0.06124845100566745, "forward_seconds": 0.11991029803175479, "cache_class": "DynamicCache", "total_seconds": 0.12943823903333396}
|
||||
{"id": "3dd9c3ac98f8ddf7bdd3", "option_ids": ["contradicted", "supported", "insufficient"], "probabilities": [0.26648902893066406, 0.7243921756744385, 0.009118752554059029], "option_logits": [24.25, 25.25, 20.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "de5f0de3f7d36daf77cd8fea40841c8ce59065d338ecc12c2a2576ef56482dad", "prompt_version": "direct-options-v1", "input_tokens": 136, "allowed_token_mass": 0.9996891617774963, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 57, "prefix_sha256": "8328776219889cc284e7d62f29767c21493e62b213af24603c38caa2bfae4173", "encode_seconds": 0.0015443390002474189, "prefill_seconds": 0.0, "copy_seconds": 0.0052741679828614, "suffix_forward_seconds": 0.06284423003671691, "forward_seconds": 0.06284423003671691, "cache_class": "DynamicCache", "total_seconds": 0.0716181870084256}
|
||||
{"id": "36273dd4777876a26338", "option_ids": ["contradicted", "supported", "insufficient"], "probabilities": [0.003593161003664136, 0.9962839484214783, 0.00012295119813643396], "option_logits": [22.625, 28.25, 19.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "c97a5900da6bf820fc491f5c4eefca269bbc5359041ce13cc797b23a36faf6cd", "prompt_version": "direct-options-v1", "input_tokens": 132, "allowed_token_mass": 0.9999294877052307, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 53, "prefix_sha256": "013380118daa1cdbb05b086d32b8d2644e64633cd00d3712d7864ebd8e352079", "encode_seconds": 0.0021533319959416986, "prefill_seconds": 0.059266378986649215, "copy_seconds": 0.005293239024467766, "suffix_forward_seconds": 0.06325060100061819, "forward_seconds": 0.1225169799872674, "cache_class": "DynamicCache", "total_seconds": 0.13184978195931762}
|
||||
{"id": "07ee70cfe4b01330ce89", "option_ids": ["supported", "contradicted", "insufficient"], "probabilities": [0.002472232561558485, 0.00015804453869350255, 0.9973697662353516], "option_logits": [21.625, 18.875, 27.625], "answer_token_ids": [32, 33, 34], "prompt_sha256": "d8ec2562355a9f885bfc006f6c277f8aea6f95bed71d06320fa4971d81f5a57b", "prompt_version": "direct-options-v1", "input_tokens": 132, "allowed_token_mass": 0.9998893737792969, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 53, "prefix_sha256": "a560884d63e8c817ed268983062e6dc112b8277c65ee51cbc87d23ef3bb45ade", "encode_seconds": 0.0021658510086126626, "prefill_seconds": 0.05825828399974853, "copy_seconds": 0.005200148967560381, "suffix_forward_seconds": 0.0634841529536061, "forward_seconds": 0.12174243695335463, "cache_class": "DynamicCache", "total_seconds": 0.13117214798694476}
|
||||
{"id": "51ceab3ce622c7c759f1", "option_ids": ["contradicted", "insufficient", "supported"], "probabilities": [0.004607828799635172, 0.00037823361344635487, 0.9950138926506042], "option_logits": [21.75, 19.25, 27.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "b1883cab135f1d71d76b55f8ec8387c4a8983c722135cd2e11f0f80d8fb919cf", "prompt_version": "direct-options-v1", "input_tokens": 148, "allowed_token_mass": 0.9998722076416016, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 69, "prefix_sha256": "bd6352431a6241c53dea71d9494ba824eefca2276057e499c405be69b46dfd47", "encode_seconds": 0.0022916030138731003, "prefill_seconds": 0.06673249998129904, "copy_seconds": 0.005333529028575867, "suffix_forward_seconds": 0.06299931101966649, "forward_seconds": 0.12973181100096554, "cache_class": "DynamicCache", "total_seconds": 0.13940875401021913}
|
||||
{"id": "b8091be929b1e489fcc4", "option_ids": ["supported", "contradicted", "insufficient"], "probabilities": [0.015897737815976143, 0.9835582375526428, 0.0005439906963147223], "option_logits": [23.5, 27.625, 20.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "5e01983da4ad753b74cb4c56eba3cb2e0f950bb19a97bcf8bbd0ce6aa34802a1", "prompt_version": "direct-options-v1", "input_tokens": 148, "allowed_token_mass": 0.9998779296875, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 69, "prefix_sha256": "bd6352431a6241c53dea71d9494ba824eefca2276057e499c405be69b46dfd47", "encode_seconds": 0.001999990956392139, "prefill_seconds": 0.0, "copy_seconds": 0.005001367011573166, "suffix_forward_seconds": 0.07912679697619751, "forward_seconds": 0.07912679697619751, "cache_class": "DynamicCache", "total_seconds": 0.08809417596785352}
|
||||
{"id": "c89b86e3f55123bfa506", "option_ids": ["supported", "contradicted", "insufficient"], "probabilities": [0.007570390123873949, 0.9915254712104797, 0.0009041542070917785], "option_logits": [22.375, 27.25, 20.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "6abfb08d77bd0a3fa8b98c5cc6eba0545b1d8ec9f8a98fc148038bd4ec3dab70", "prompt_version": "direct-options-v1", "input_tokens": 139, "allowed_token_mass": 0.9998531341552734, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 60, "prefix_sha256": "79ee4c6d0a45fb4f55a87a579cf2a02b0bbf187db90527c9e62722a382d16baf", "encode_seconds": 0.0022070030099712312, "prefill_seconds": 0.05971179198240861, "copy_seconds": 0.0053033389849588275, "suffix_forward_seconds": 0.062416206987109035, "forward_seconds": 0.12212799896951765, "cache_class": "DynamicCache", "total_seconds": 0.1316849029972218}
|
||||
{"id": "3b0958cabc339005cf66", "option_ids": ["insufficient", "supported", "contradicted"], "probabilities": [0.9973534345626831, 0.0013232689816504717, 0.0013232689816504717], "option_logits": [27.0, 20.375, 20.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "793a015d65fa6d502036009d7b2c2740bfdf1ae21cea3c099e255c9864961c56", "prompt_version": "direct-options-v1", "input_tokens": 135, "allowed_token_mass": 0.9998893737792969, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 56, "prefix_sha256": "a70a38b4216e92163b6fd139644478a7965b71cbd68a02b7ec46203a0ad06d10", "encode_seconds": 0.0022262720158323646, "prefill_seconds": 0.0610509300022386, "copy_seconds": 0.005220888997428119, "suffix_forward_seconds": 0.06314244098030031, "forward_seconds": 0.12419337098253891, "cache_class": "DynamicCache", "total_seconds": 0.1337186330347322}
|
||||
{"id": "f2b4ec4930322fe45118", "option_ids": ["insufficient", "prohibited", "permitted"], "probabilities": [0.5061980485916138, 0.4467182159423828, 0.04708375409245491], "option_logits": [24.25, 24.125, 21.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "182c69c28b1efda301044ed0dd652e23a600cdbcc96dce5bdf88dd1a6231aed1", "prompt_version": "direct-options-v1", "input_tokens": 150, "allowed_token_mass": 0.9994393587112427, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 59, "prefix_sha256": "9de8325a1abc05b1df562523ba91d14fbfa2d216d1139e30ad371c107880142b", "encode_seconds": 0.0022353819804266095, "prefill_seconds": 0.05833481502486393, "copy_seconds": 0.005228857975453138, "suffix_forward_seconds": 0.06246712704887614, "forward_seconds": 0.12080194207374007, "cache_class": "DynamicCache", "total_seconds": 0.13034636399243027}
|
||||
{"id": "61ac20b07a5189231e15", "option_ids": ["prohibited", "permitted", "insufficient"], "probabilities": [0.0031725200824439526, 0.9967761635780334, 5.1279010222060606e-05], "option_logits": [21.875, 27.625, 17.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "07812a1d1f463245eb74192dcfe461e8ae727b1fdd733cb43bf3b05025847249", "prompt_version": "direct-options-v1", "input_tokens": 156, "allowed_token_mass": 0.9998398423194885, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 59, "prefix_sha256": "9de8325a1abc05b1df562523ba91d14fbfa2d216d1139e30ad371c107880142b", "encode_seconds": 0.001989280979614705, "prefill_seconds": 0.0, "copy_seconds": 0.005065426987130195, "suffix_forward_seconds": 0.061543182004243135, "forward_seconds": 0.061543182004243135, "cache_class": "DynamicCache", "total_seconds": 0.07033369096461684}
|
||||
{"id": "4b3d794c004566610181", "option_ids": ["prohibited", "insufficient", "permitted"], "probabilities": [0.006691114045679569, 0.00025944263325072825, 0.9930493831634521], "option_logits": [22.125, 18.875, 27.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "7d9d80d1b6275026e73a37333b83e5c3121bb55719c2588ef3100a620aeac115", "prompt_version": "direct-options-v1", "input_tokens": 141, "allowed_token_mass": 0.9998608231544495, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 50, "prefix_sha256": "d96f0704330ef56b4887d61b8d8b41d60f356f859314105c9079f5a05cae8334", "encode_seconds": 0.002165622019674629, "prefill_seconds": 0.0586007569800131, "copy_seconds": 0.0048402760294266045, "suffix_forward_seconds": 0.06234029697952792, "forward_seconds": 0.12094105395954102, "cache_class": "DynamicCache", "total_seconds": 0.12998751300619915}
|
||||
{"id": "b7c4e62d07986286fce0", "option_ids": ["insufficient", "permitted", "prohibited"], "probabilities": [0.996604323387146, 0.0016978350467979908, 0.0016978350467979908], "option_logits": [27.25, 20.875, 20.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "d16b6fed11c823e8b2099a68f0d7dda406ce724df664e65574efb3feb7f2736d", "prompt_version": "direct-options-v1", "input_tokens": 146, "allowed_token_mass": 0.999885618686676, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 55, "prefix_sha256": "a6edb9261c5351ed5e54431aa7e7f844a520b79f24e69ed902d87d7ea58eb022", "encode_seconds": 0.0022069329861551523, "prefill_seconds": 0.18481694301590323, "copy_seconds": 0.005297257972415537, "suffix_forward_seconds": 0.0624066439922899, "forward_seconds": 0.24722358700819314, "cache_class": "DynamicCache", "total_seconds": 0.25678144895937294}
|
||||
{"id": "9473603b44af63beab79", "option_ids": ["permitted", "insufficient", "prohibited"], "probabilities": [0.9946156144142151, 0.004064766690135002, 0.0013196364743635058], "option_logits": [26.875, 21.375, 20.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "4888c49f4d346d3931506f0bdd663c12db4b52667d9d815dd85fec36d64e3b7a", "prompt_version": "direct-options-v1", "input_tokens": 147, "allowed_token_mass": 0.9998608231544495, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 53, "prefix_sha256": "26e7d6ee6ad070c4ca15add48a9182687867856811ac9c631ce0b8e6eb77d02d", "encode_seconds": 0.0022623430122621357, "prefill_seconds": 0.0593647810164839, "copy_seconds": 0.0051300780032761395, "suffix_forward_seconds": 0.061967924993950874, "forward_seconds": 0.12133270601043478, "cache_class": "DynamicCache", "total_seconds": 0.13077356800204143}
|
||||
{"id": "d41072dc35d8cd14c056", "option_ids": ["permitted", "prohibited", "insufficient"], "probabilities": [0.022784773260354996, 0.9688331484794617, 0.008382049389183521], "option_logits": [22.25, 26.0, 21.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "76def65de0afab06519977870549e4c5ead1d1f48757b276e755ad1020cefb40", "prompt_version": "direct-options-v1", "input_tokens": 152, "allowed_token_mass": 0.9997158050537109, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 53, "prefix_sha256": "26e7d6ee6ad070c4ca15add48a9182687867856811ac9c631ce0b8e6eb77d02d", "encode_seconds": 0.0017742200288921595, "prefill_seconds": 0.0, "copy_seconds": 0.005118508008308709, "suffix_forward_seconds": 0.06138454203028232, "forward_seconds": 0.06138454203028232, "cache_class": "DynamicCache", "total_seconds": 0.07001547899562865}
|
||||
{"id": "33d6e5da58f0fee2490d", "option_ids": ["prohibited", "insufficient", "permitted"], "probabilities": [0.3446420729160309, 0.5682187080383301, 0.08713916689157486], "option_logits": [24.25, 24.75, 22.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "420efbb357aa1ed78444ac26b551c48faaeef0a90eb508c97fc5ecd2d7375e76", "prompt_version": "direct-options-v1", "input_tokens": 147, "allowed_token_mass": 0.9995824098587036, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 53, "prefix_sha256": "5787ef3363668896d969a1c1175de67385863c40668f1bac6adc569960c9233b", "encode_seconds": 0.0021949519868940115, "prefill_seconds": 0.0601786159677431, "copy_seconds": 0.004822926013730466, "suffix_forward_seconds": 0.062337887997273356, "forward_seconds": 0.12251650396501645, "cache_class": "DynamicCache", "total_seconds": 0.13161120197037235}
|
||||
{"id": "823bf736b92e9d8fa66d", "option_ids": ["prohibited", "insufficient", "permitted"], "probabilities": [0.012392697855830193, 0.9844739437103271, 0.0031333649531006813], "option_logits": [22.375, 26.75, 21.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "8a1a8afdcfc080506a68573023dd716f59e20151b50543bd3f7dcfe3ba126705", "prompt_version": "direct-options-v1", "input_tokens": 147, "allowed_token_mass": 0.9998455047607422, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 53, "prefix_sha256": "e0ea2860449a8e46d3cb0e89bcb0594f19d4f019a4e9fbd3738874e503d71f18", "encode_seconds": 0.0020083109848201275, "prefill_seconds": 0.058848268003202975, "copy_seconds": 0.005222169042099267, "suffix_forward_seconds": 0.06264841899974272, "forward_seconds": 0.12149668700294569, "cache_class": "DynamicCache", "total_seconds": 0.13071419799234718}
|
||||
{"id": "8a4c3de28be95c105ca1", "option_ids": ["permitted", "insufficient", "prohibited"], "probabilities": [0.09073077887296677, 0.14958974719047546, 0.7596794962882996], "option_logits": [22.75, 23.25, 24.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "bbb4ea9c7fd17181dbd34de5e443d37f0756af7067d5d9fc6df5fe483638941a", "prompt_version": "direct-options-v1", "input_tokens": 146, "allowed_token_mass": 0.9995499849319458, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 56, "prefix_sha256": "d3acbd5ff2f0cea7c59b395029b10a85013933b7add7708e788caad21c45b5a5", "encode_seconds": 0.0021038309787400067, "prefill_seconds": 0.06033271597698331, "copy_seconds": 0.005254309042356908, "suffix_forward_seconds": 0.06359591404907405, "forward_seconds": 0.12392863002605736, "cache_class": "DynamicCache", "total_seconds": 0.1333484509959817}
|
||||
{"id": "ea7ec66a5e8eac886548", "option_ids": ["permitted", "insufficient", "prohibited"], "probabilities": [0.7594355344772339, 0.1920153796672821, 0.04854908958077431], "option_logits": [24.875, 23.5, 22.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "cb21e641fb963a5779c52e29303655cbf94a0abb0c2bc5d94e95f1eee8ec6fe5", "prompt_version": "direct-options-v1", "input_tokens": 149, "allowed_token_mass": 0.9996643662452698, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 56, "prefix_sha256": "d3acbd5ff2f0cea7c59b395029b10a85013933b7add7708e788caad21c45b5a5", "encode_seconds": 0.001918039983138442, "prefill_seconds": 0.0, "copy_seconds": 0.005227317975368351, "suffix_forward_seconds": 0.06357364397263154, "forward_seconds": 0.06357364397263154, "cache_class": "DynamicCache", "total_seconds": 0.07266383297974244}
|
||||
{"id": "25172cd6ddc49ab0ca70", "option_ids": ["prohibited", "insufficient", "permitted"], "probabilities": [0.029269514605402946, 0.001457243342883885, 0.9692732095718384], "option_logits": [23.125, 20.125, 26.625], "answer_token_ids": [32, 33, 34], "prompt_sha256": "1407436ffc934b8a8681c0f8e30d40a0a41134df70ddde88fdd3a262f0585e56", "prompt_version": "direct-options-v1", "input_tokens": 140, "allowed_token_mass": 0.9998131394386292, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 50, "prefix_sha256": "8dd01d156a093ef2bae3026a2d909efcae5d40cfddcc0809ce6a7d1a0f57a2a6", "encode_seconds": 0.00219088199082762, "prefill_seconds": 0.0594914720277302, "copy_seconds": 0.005344708974007517, "suffix_forward_seconds": 0.11924721795367077, "forward_seconds": 0.17873868998140097, "cache_class": "DynamicCache", "total_seconds": 0.18832279101479799}
|
||||
{"id": "fb73a85d386cf59b0c51", "option_ids": ["insufficient", "permitted", "prohibited"], "probabilities": [0.9940441250801086, 0.0031638245563954115, 0.0027920654974877834], "option_logits": [26.5, 20.75, 20.625], "answer_token_ids": [32, 33, 34], "prompt_sha256": "8bd0305dca7a0017106fd077eaca2971baba6e35c10200536ca6c772620a0a28", "prompt_version": "direct-options-v1", "input_tokens": 142, "allowed_token_mass": 0.9998436570167542, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 52, "prefix_sha256": "6c2147c39f6f6ee95d830f59d7c7731f6b53e1ab5859a680e449844469b4a967", "encode_seconds": 0.002193582011386752, "prefill_seconds": 0.05959750700276345, "copy_seconds": 0.005232239025644958, "suffix_forward_seconds": 0.06314443604787812, "forward_seconds": 0.12274194305064157, "cache_class": "DynamicCache", "total_seconds": 0.13223651400767267}
|
||||
{"id": "a02fd582d2881cc462b2", "option_ids": ["prohibited", "insufficient", "permitted"], "probabilities": [0.00668507581576705, 0.001161691965535283, 0.9921532273292542], "option_logits": [21.625, 19.875, 26.625], "answer_token_ids": [32, 33, 34], "prompt_sha256": "d8b3821ba5dba4ff4aa41fa4405083273a0df99e2a6ab225b2945dbb13b12350", "prompt_version": "direct-options-v1", "input_tokens": 152, "allowed_token_mass": 0.9998054504394531, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 56, "prefix_sha256": "c0d2723b50ef3899eb1d2f7ee9c532a5617bc1a65ad60af3061b76e29e3c6374", "encode_seconds": 0.002271933015435934, "prefill_seconds": 0.06018122599925846, "copy_seconds": 0.005228607973549515, "suffix_forward_seconds": 0.06338813400361687, "forward_seconds": 0.12356936000287533, "cache_class": "DynamicCache", "total_seconds": 0.13319516205228865}
|
||||
{"id": "19f37d5a97034bd7d04f", "option_ids": ["insufficient", "prohibited", "permitted"], "probabilities": [0.18181802332401276, 0.8148518204689026, 0.0033301133662462234], "option_logits": [24.125, 25.625, 20.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "65fd11d492dbd8d2397ed2c482d0cd7eeab246c709d779398ee69ee9b247c5d1", "prompt_version": "direct-options-v1", "input_tokens": 152, "allowed_token_mass": 0.9996643662452698, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 56, "prefix_sha256": "c0d2723b50ef3899eb1d2f7ee9c532a5617bc1a65ad60af3061b76e29e3c6374", "encode_seconds": 0.0020667019998654723, "prefill_seconds": 0.0, "copy_seconds": 0.00535304902587086, "suffix_forward_seconds": 0.0636207649949938, "forward_seconds": 0.0636207649949938, "cache_class": "DynamicCache", "total_seconds": 0.07321311702253297}
|
||||
{"id": "fda0ca94652c6d739ecf", "option_ids": ["prohibited", "permitted", "insufficient"], "probabilities": [0.9967532753944397, 0.0013224728172644973, 0.0019241865957155824], "option_logits": [27.25, 20.625, 21.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "efee3934fe54f08ea8f88321754619badecee82f40231f152e2ff14e4bcc8c3b", "prompt_version": "direct-options-v1", "input_tokens": 151, "allowed_token_mass": 0.999885618686676, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 55, "prefix_sha256": "7500149c0d5e105fde6b233f3d4b33bacd6927df75d98b83e09d8d9b6e7847f7", "encode_seconds": 0.0024670030106790364, "prefill_seconds": 0.06049284798791632, "copy_seconds": 0.005401218950282782, "suffix_forward_seconds": 0.06406326702563092, "forward_seconds": 0.12455611501354724, "cache_class": "DynamicCache", "total_seconds": 0.1346756390412338}
|
||||
{"id": "d0e36c4cb10a29b3191c", "option_ids": ["insufficient", "permitted", "prohibited"], "probabilities": [0.980177640914917, 0.007483747787773609, 0.012338615022599697], "option_logits": [26.375, 21.5, 22.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "c3fe5eb6c6637803e7d25c7d76a2d05fcff468132e72e6d157d842c10345b2fb", "prompt_version": "direct-options-v1", "input_tokens": 150, "allowed_token_mass": 0.9998169541358948, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 54, "prefix_sha256": "c1cf3c62d53663c7c764439e8db8d907e711308b31f771e49d836d503239807d", "encode_seconds": 0.0024023029836826026, "prefill_seconds": 0.06103236903436482, "copy_seconds": 0.005357559944968671, "suffix_forward_seconds": 0.07164835801813751, "forward_seconds": 0.13268072705250233, "cache_class": "DynamicCache", "total_seconds": 0.14263600198319182}
|
||||
{"id": "a0a39ace7e5337d038a7", "option_ids": ["prohibited", "permitted", "insufficient"], "probabilities": [0.9992511868476868, 0.0004877297906205058, 0.00026106290169991553], "option_logits": [27.375, 19.75, 19.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "a3131f24a61eb5dc838e69ec3d86eee6dd7fde481b2d3e142356288c663586b8", "prompt_version": "direct-options-v1", "input_tokens": 154, "allowed_token_mass": 0.9998741149902344, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 57, "prefix_sha256": "f9128795a806d36b39b59e86a1e76dce01c5606bc99a65d185aaa5a31ac88efd", "encode_seconds": 0.0023776229936629534, "prefill_seconds": 0.060197305982001126, "copy_seconds": 0.005259238998405635, "suffix_forward_seconds": 0.06272300001000986, "forward_seconds": 0.12292030599201098, "cache_class": "DynamicCache", "total_seconds": 0.13264883897500113}
|
||||
{"id": "eeae283841ce75dd25bf", "option_ids": ["insufficient", "permitted", "prohibited"], "probabilities": [0.020326457917690277, 0.979383647441864, 0.00028994138119742274], "option_logits": [23.0, 26.875, 18.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "a521ffe4dd855e4729d536bda040a7a384fdf48c756345a829ab01309acf268b", "prompt_version": "direct-options-v1", "input_tokens": 154, "allowed_token_mass": 0.9997596740722656, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 57, "prefix_sha256": "f9128795a806d36b39b59e86a1e76dce01c5606bc99a65d185aaa5a31ac88efd", "encode_seconds": 0.001946060045156628, "prefill_seconds": 0.0, "copy_seconds": 0.0051203579641878605, "suffix_forward_seconds": 0.06298065098235384, "forward_seconds": 0.06298065098235384, "cache_class": "DynamicCache", "total_seconds": 0.07208564004395157}
|
||||
{"id": "00bb4cc043f7d2b288d4", "option_ids": ["insufficient", "prohibited", "permitted"], "probabilities": [0.010978205129504204, 0.0007952585001476109, 0.9882264733314514], "option_logits": [22.375, 19.75, 26.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "72a0d3e3684d1ae5a178bbbb7f39ee3282bf0fe94e12ab4b7c713452f0eda6be", "prompt_version": "direct-options-v1", "input_tokens": 146, "allowed_token_mass": 0.9998341202735901, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 49, "prefix_sha256": "053bc4ffe0f4eb27bcb050c14e2aacf8bf494fa7b0cf7d9fbabd120d0572e970", "encode_seconds": 0.0022231120383366942, "prefill_seconds": 0.07611412199912593, "copy_seconds": 0.005146208044607192, "suffix_forward_seconds": 0.06292428099550307, "forward_seconds": 0.139038402994629, "cache_class": "DynamicCache", "total_seconds": 0.1484555639908649}
|
||||
{"id": "1a81ead69e93e5d0c86a", "option_ids": ["permitted", "prohibited", "insufficient"], "probabilities": [0.05249322950839996, 0.01704205572605133, 0.9304646849632263], "option_logits": [22.5, 21.375, 25.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "09de937e3cda9f4756fd7f4a71969ee3cc4832bdf895581e3ba0461794df5e19", "prompt_version": "direct-options-v1", "input_tokens": 153, "allowed_token_mass": 0.9996166825294495, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 56, "prefix_sha256": "8a4a669c55722523da0c50a7a488517ff5087e4bd72167138228b842c64aa778", "encode_seconds": 0.0022418120061047375, "prefill_seconds": 0.0601617360371165, "copy_seconds": 0.005188578041270375, "suffix_forward_seconds": 0.06395176600199193, "forward_seconds": 0.12411350203910843, "cache_class": "DynamicCache", "total_seconds": 0.1337109439773485}
|
||||
{"id": "4444dad25e438f128673", "option_ids": ["prohibited", "permitted", "insufficient"], "probabilities": [0.9835985898971558, 0.0040197428315877914, 0.012381679378449917], "option_logits": [26.625, 21.125, 22.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "f01eedda33719df8461d9df777ad190bed7c6ed643b6682e6cca18cb49da3b5d", "prompt_version": "direct-options-v1", "input_tokens": 147, "allowed_token_mass": 0.9998703002929688, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 51, "prefix_sha256": "ad56bf914db81d0028180de879f74a43038558f9357f557883a01935031a26b5", "encode_seconds": 0.0024830629699863493, "prefill_seconds": 0.06050687801325694, "copy_seconds": 0.005393039027694613, "suffix_forward_seconds": 0.06424753798637539, "forward_seconds": 0.12475441599963233, "cache_class": "DynamicCache", "total_seconds": 0.13492573099210858}
|
||||
{"id": "3226202ba1ada2570e9f", "option_ids": ["prohibited", "permitted", "insufficient"], "probabilities": [0.5908308029174805, 0.19181467592716217, 0.21735452115535736], "option_logits": [24.875, 23.75, 23.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "262a35e49ca4d9f3d2feae70089c22607339c33eebac9961769c0f42b52feb36", "prompt_version": "direct-options-v1", "input_tokens": 150, "allowed_token_mass": 0.999723494052887, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 51, "prefix_sha256": "ad56bf914db81d0028180de879f74a43038558f9357f557883a01935031a26b5", "encode_seconds": 0.0021237809560261667, "prefill_seconds": 0.0, "copy_seconds": 0.005245678999926895, "suffix_forward_seconds": 0.06386125698918477, "forward_seconds": 0.06386125698918477, "cache_class": "DynamicCache", "total_seconds": 0.07330514799105003}
|
||||
{"id": "3f783eec5c90812479bc", "option_ids": ["permitted", "prohibited", "insufficient"], "probabilities": [0.015588019043207169, 0.964396595954895, 0.02001541294157505], "option_logits": [22.0, 26.125, 22.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "f03d7ada9f23826e4c7b3048e5fc2fe973f7ade56d68628782a46fcc79c4fb74", "prompt_version": "direct-options-v1", "input_tokens": 147, "allowed_token_mass": 0.9997521042823792, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 51, "prefix_sha256": "88db2222ec1d753745fdca800c8ccd813cd62aa0ad472e69aaa866c6945ec6a1", "encode_seconds": 0.0022421019966714084, "prefill_seconds": 0.05970592302037403, "copy_seconds": 0.005255348980426788, "suffix_forward_seconds": 0.06357649498386309, "forward_seconds": 0.12328241800423712, "cache_class": "DynamicCache", "total_seconds": 0.13284626899985597}
|
||||
{"id": "ffee861474a9665a093a", "option_ids": ["prohibited", "permitted", "insufficient"], "probabilities": [0.13274212181568146, 0.10337967425584793, 0.7638782262802124], "option_logits": [23.375, 23.125, 25.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "82f295831fda3638311f60f23265cee36399032a1603154af678ed2cbdba18dd", "prompt_version": "direct-options-v1", "input_tokens": 148, "allowed_token_mass": 0.9997082352638245, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 52, "prefix_sha256": "5379cb3e5433538b81e0791c1f14868f4a8ee9fc958cf63331e6102d419c5479", "encode_seconds": 0.0023069229791872203, "prefill_seconds": 0.05998410505708307, "copy_seconds": 0.005176408973056823, "suffix_forward_seconds": 0.061859324981924146, "forward_seconds": 0.12184343003900722, "cache_class": "DynamicCache", "total_seconds": 0.1313230519881472}
|
||||
{"id": "8c992e17e4fb54b044f5", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.4682465195655823, 0.5305928587913513, 0.0011606671614572406], "option_logits": [23.75, 23.875, 17.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "db2c432824fa196124e34d0790267094f113fb9987d785f724dbfa15c7673746", "prompt_version": "direct-options-v1", "input_tokens": 137, "allowed_token_mass": 0.999225914478302, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 68, "prefix_sha256": "8481be571622474228fa3439d7e96756a29b136ef27568e16e82e0d118fafe3d", "encode_seconds": 0.0019165899720974267, "prefill_seconds": 0.06269113003509119, "copy_seconds": 0.005295978975482285, "suffix_forward_seconds": 0.06140613299794495, "forward_seconds": 0.12409726303303614, "cache_class": "DynamicCache", "total_seconds": 0.13343312399229035}
|
||||
{"id": "695dd61195ca025c4f4c", "option_ids": ["B", "insufficient", "A"], "probabilities": [0.2738439440727234, 0.6569174528121948, 0.0692385882139206], "option_logits": [23.125, 24.0, 21.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "cf7c4ddc2978765051eab9ce4cc30a7e04404f12fb9921e059fefe9e6a1d1aa1", "prompt_version": "direct-options-v1", "input_tokens": 138, "allowed_token_mass": 0.9991439580917358, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 68, "prefix_sha256": "8481be571622474228fa3439d7e96756a29b136ef27568e16e82e0d118fafe3d", "encode_seconds": 0.0018562700133770704, "prefill_seconds": 0.0, "copy_seconds": 0.005156787985470146, "suffix_forward_seconds": 0.08065677399281412, "forward_seconds": 0.08065677399281412, "cache_class": "DynamicCache", "total_seconds": 0.08964574302081019}
|
||||
{"id": "e6155431ab7b0f8170d4", "option_ids": ["insufficient", "B", "A"], "probabilities": [0.005147089716047049, 0.9808617234230042, 0.013991240411996841], "option_logits": [20.375, 25.625, 21.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "956179411ff2437440233be0a6943ebe9c707fceabe37844e7ec22355b0c5896", "prompt_version": "direct-options-v1", "input_tokens": 129, "allowed_token_mass": 0.9996910691261292, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 60, "prefix_sha256": "cdcf197be1d0f9b3a317baf7fb572e0755c0874a0139418dd856a8ae442e07e9", "encode_seconds": 0.002065391046926379, "prefill_seconds": 0.060212098993360996, "copy_seconds": 0.005285458988510072, "suffix_forward_seconds": 0.0643201990169473, "forward_seconds": 0.1245322980103083, "cache_class": "DynamicCache", "total_seconds": 0.13396739901509136}
|
||||
{"id": "b49ecda75086d8c68d2a", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.2831884026527405, 0.46689873933792114, 0.24991288781166077], "option_logits": [23.0, 23.5, 22.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "da52131b3bd6b87ca2b56a62ef7c15d0f0134aa4497788bc96d523473dfe12d6", "prompt_version": "direct-options-v1", "input_tokens": 129, "allowed_token_mass": 0.9990582466125488, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 60, "prefix_sha256": "eaa27a4478ac53900586cd6a215a792bdd91798253cdc3585344fc90d6490b2e", "encode_seconds": 0.0020944710122421384, "prefill_seconds": 0.06059018900850788, "copy_seconds": 0.005271127971354872, "suffix_forward_seconds": 0.06399681698530912, "forward_seconds": 0.124587005993817, "cache_class": "DynamicCache", "total_seconds": 0.13400858599925414}
|
||||
{"id": "587ced261c6092cd1325", "option_ids": ["insufficient", "B", "A"], "probabilities": [0.0004867928510066122, 0.9973316192626953, 0.0021816540975123644], "option_logits": [18.75, 26.375, 20.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "92d7fca2589d89e7117a38cf3d02cd1744b1ace4a461dbf63816ca7e4cfecf0f", "prompt_version": "direct-options-v1", "input_tokens": 145, "allowed_token_mass": 0.9997654557228088, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 75, "prefix_sha256": "7aad014fb5e8b371d4685fdc2c4ce6291685fd32f3b025cfe04e78c1a57359c5", "encode_seconds": 0.0022830429952591658, "prefill_seconds": 0.06528327404521406, "copy_seconds": 0.00531095900805667, "suffix_forward_seconds": 0.06282652099616826, "forward_seconds": 0.1281097950413823, "cache_class": "DynamicCache", "total_seconds": 0.1378043980221264}
|
||||
{"id": "533d4423d311de82b27d", "option_ids": ["insufficient", "A", "B"], "probabilities": [0.2928224503993988, 0.7024445533752441, 0.004733033943921328], "option_logits": [24.125, 25.0, 20.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "8335ab560e30ab6fd8c5f2a7b44d7aae6e5a4bcc55ccaada8f30521a09449b79", "prompt_version": "direct-options-v1", "input_tokens": 148, "allowed_token_mass": 0.999466061592102, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 75, "prefix_sha256": "7aad014fb5e8b371d4685fdc2c4ce6291685fd32f3b025cfe04e78c1a57359c5", "encode_seconds": 0.001967240998055786, "prefill_seconds": 0.0, "copy_seconds": 0.005083246971480548, "suffix_forward_seconds": 0.06307499302783981, "forward_seconds": 0.06307499302783981, "cache_class": "DynamicCache", "total_seconds": 0.07203080999897793}
|
||||
{"id": "2bf540d515a51755b25c", "option_ids": ["insufficient", "B", "A"], "probabilities": [0.02032567374408245, 0.0003285339043941349, 0.9793457984924316], "option_logits": [21.75, 17.625, 25.625], "answer_token_ids": [32, 33, 34], "prompt_sha256": "f67e730b7fec91d50301cc364d7680aff6dc7bf38b604d5f17c02bf63b7d1aa3", "prompt_version": "direct-options-v1", "input_tokens": 137, "allowed_token_mass": 0.9997196793556213, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 67, "prefix_sha256": "00025c43c850edcca42bd3a175a280d340c52939a1db3f55667d9af46f1d7e6a", "encode_seconds": 0.002127071958966553, "prefill_seconds": 0.06367297499673441, "copy_seconds": 0.004934347001835704, "suffix_forward_seconds": 0.06269225099822506, "forward_seconds": 0.12636522599495947, "cache_class": "DynamicCache", "total_seconds": 0.1353750949492678}
|
||||
{"id": "fdb9fc214aea2710763a", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.28125160932540894, 0.5254471898078918, 0.1933012306690216], "option_logits": [23.125, 23.75, 22.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "9b1e29da56a070a145f16aa81e08b2bb58ee6e7ee00055924fe04156c6b69e33", "prompt_version": "direct-options-v1", "input_tokens": 127, "allowed_token_mass": 0.999142050743103, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 57, "prefix_sha256": "9d2720a3e0e157e98a5735584af8aee60d886a9d9b37f1195b21577f073207c6", "encode_seconds": 0.002038691018242389, "prefill_seconds": 0.06056333897868171, "copy_seconds": 0.004929506976623088, "suffix_forward_seconds": 0.07523925899295136, "forward_seconds": 0.13580259797163308, "cache_class": "DynamicCache", "total_seconds": 0.14482480700826272}
|
||||
{"id": "f6d68126b43f8c9ed183", "option_ids": ["insufficient", "A", "B"], "probabilities": [0.09526041150093079, 0.903805673122406, 0.0009339002426713705], "option_logits": [24.0, 26.25, 19.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "f06fe76db2b82169f130adf5f4f1f798216a2aa57b3027cd6ba827dab130b5cd", "prompt_version": "direct-options-v1", "input_tokens": 140, "allowed_token_mass": 0.9998054504394531, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 69, "prefix_sha256": "d8adeccd2c287d4140db255348eccc258ed83a0cf4868d8a3d77016f4b34ea6a", "encode_seconds": 0.0022746819886378944, "prefill_seconds": 0.06358660501427948, "copy_seconds": 0.005216819001361728, "suffix_forward_seconds": 0.08180909400107339, "forward_seconds": 0.14539569901535287, "cache_class": "DynamicCache", "total_seconds": 0.15505339100491256}
|
||||
{"id": "44bae0c8a1615a096c43", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.26791366934776306, 0.7282647490501404, 0.003821582766249776], "option_logits": [23.5, 24.5, 19.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "e670395818a1e4c58688d83d5076fccd87492a5c318539e8edf8e823a66d330c", "prompt_version": "direct-options-v1", "input_tokens": 143, "allowed_token_mass": 0.9994089007377625, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 69, "prefix_sha256": "d8adeccd2c287d4140db255348eccc258ed83a0cf4868d8a3d77016f4b34ea6a", "encode_seconds": 0.0019046009983867407, "prefill_seconds": 0.0, "copy_seconds": 0.005155468010343611, "suffix_forward_seconds": 0.06726900499779731, "forward_seconds": 0.06726900499779731, "cache_class": "DynamicCache", "total_seconds": 0.07634604501072317}
|
||||
{"id": "2175379f4c69e6207626", "option_ids": ["insufficient", "B", "A"], "probabilities": [0.0009073464316315949, 0.995026171207428, 0.004066444467753172], "option_logits": [19.5, 26.5, 21.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "97f7bc540bfa213c7e6a52d6376ab915a3eee4cf50945263a7880513864294a5", "prompt_version": "direct-options-v1", "input_tokens": 134, "allowed_token_mass": 0.9998112320899963, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 63, "prefix_sha256": "eb31d894b7e904dcc712af809a3ebe89afb4b0c58eb4b755b99071e221c4a089", "encode_seconds": 0.002012610959354788, "prefill_seconds": 0.059612463985104114, "copy_seconds": 0.0052132169948890805, "suffix_forward_seconds": 0.06284116097958758, "forward_seconds": 0.1224536249646917, "cache_class": "DynamicCache", "total_seconds": 0.13177828496554866}
|
||||
{"id": "def9fe59d262357a673a", "option_ids": ["insufficient", "A", "B"], "probabilities": [0.34160494804382324, 0.4970322549343109, 0.16136273741722107], "option_logits": [23.375, 23.75, 22.625], "answer_token_ids": [32, 33, 34], "prompt_sha256": "30564aca118d1d837b1009b7f91e1e20894fd311e51d351927b94a2c9cd11040", "prompt_version": "direct-options-v1", "input_tokens": 129, "allowed_token_mass": 0.9989991784095764, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 58, "prefix_sha256": "758b96b2199e34525f0210896eee6e1a515bd515f9c5ab2da95222b7e0cc4b54", "encode_seconds": 0.002072311006486416, "prefill_seconds": 0.05951027403352782, "copy_seconds": 0.005217677040491253, "suffix_forward_seconds": 0.06258848001016304, "forward_seconds": 0.12209875404369086, "cache_class": "DynamicCache", "total_seconds": 0.1311977919540368}
|
||||
{"id": "712f6ffb763eef7c9a12", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.16379833221435547, 0.8318365812301636, 0.004365077707916498], "option_logits": [22.5, 24.125, 18.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "7c5128790bf75764214e3b2efdfcc521fd5027eaedefb0a056335cbcf0818ff6", "prompt_version": "direct-options-v1", "input_tokens": 140, "allowed_token_mass": 0.9991268515586853, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 70, "prefix_sha256": "8a591cb49c0efeab8d26424c970dc66021f416e0de3bdefa2910a21b13edd77c", "encode_seconds": 0.0022231419570744038, "prefill_seconds": 0.06459208199521527, "copy_seconds": 0.005237657984253019, "suffix_forward_seconds": 0.06341927504399791, "forward_seconds": 0.12801135703921318, "cache_class": "DynamicCache", "total_seconds": 0.13751351594692096}
|
||||
{"id": "e1d610bd14e3d16c09b9", "option_ids": ["A", "insufficient", "B"], "probabilities": [0.9958989024162292, 0.0021785201970487833, 0.0019225372234359384], "option_logits": [25.0, 18.875, 18.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "6313d4b5ee4cf1a2c0182c75a1064d43934efedf265a5e11ead31c09aba0e0f2", "prompt_version": "direct-options-v1", "input_tokens": 140, "allowed_token_mass": 0.9994489550590515, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 70, "prefix_sha256": "8a591cb49c0efeab8d26424c970dc66021f416e0de3bdefa2910a21b13edd77c", "encode_seconds": 0.0017942399717867374, "prefill_seconds": 0.0, "copy_seconds": 0.00505634699948132, "suffix_forward_seconds": 0.0631721829995513, "forward_seconds": 0.0631721829995513, "cache_class": "DynamicCache", "total_seconds": 0.07196840096730739}
|
||||
{"id": "22ec57822db50ec121cf", "option_ids": ["insufficient", "A", "B"], "probabilities": [0.10634797066450119, 0.8904406428337097, 0.0032114305067807436], "option_logits": [23.5, 25.625, 20.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "f695aa1b9ad6ef9a426e11b0cac58797145f173b338ae2a52cbf27eceb97e57e", "prompt_version": "direct-options-v1", "input_tokens": 130, "allowed_token_mass": 0.999692976474762, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 60, "prefix_sha256": "187a9c3b54c889ad1395effd485777c9cd0eace1a3e9d91f97fc56f32cc3a2da", "encode_seconds": 0.001858780044130981, "prefill_seconds": 0.05844913795590401, "copy_seconds": 0.00526001799153164, "suffix_forward_seconds": 0.06538074603304267, "forward_seconds": 0.12382988398894668, "cache_class": "DynamicCache", "total_seconds": 0.13295268203364685}
|
||||
{"id": "3e75f2b84e8e08e3d3f2", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.10375791788101196, 0.21965549886226654, 0.6765865683555603], "option_logits": [22.0, 22.75, 23.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "19a97041390bd6532502c0ed56d7824e6900269e9d26e48d7e3c8247d9ec3483", "prompt_version": "direct-options-v1", "input_tokens": 129, "allowed_token_mass": 0.9985629320144653, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 59, "prefix_sha256": "508de7c61f3dd9b263d992fd1b69d93bf88249557a5e1c4e82187a4c06a96437", "encode_seconds": 0.0021095919655635953, "prefill_seconds": 0.06211963796522468, "copy_seconds": 0.004940776969306171, "suffix_forward_seconds": 0.06095504102995619, "forward_seconds": 0.12307467899518088, "cache_class": "DynamicCache", "total_seconds": 0.132204687979538}
|
||||
{"id": "ebd74aa05324a1556288", "option_ids": ["insufficient", "A", "B"], "probabilities": [0.07560618221759796, 0.9210718870162964, 0.0033219039905816317], "option_logits": [23.125, 25.625, 20.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "586cf0684f2a395aadc22c41832546057a70d71b23bd0ab6307c2e44e023f5f1", "prompt_version": "direct-options-v1", "input_tokens": 145, "allowed_token_mass": 0.9996300935745239, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 73, "prefix_sha256": "76d511807440a0f245792556d6b98a4f6f88a9f5b40bec260876669cb64846f6", "encode_seconds": 0.002214342006482184, "prefill_seconds": 0.06552806601393968, "copy_seconds": 0.005392548977397382, "suffix_forward_seconds": 0.06677722302265465, "forward_seconds": 0.13230528903659433, "cache_class": "DynamicCache", "total_seconds": 0.14195312099764124}
|
||||
{"id": "d7e84e4245a29817da1f", "option_ids": ["insufficient", "A", "B"], "probabilities": [0.008123139850795269, 0.05296952649950981, 0.938907265663147], "option_logits": [20.5, 22.375, 25.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "b918f4270027bfe138b4cc1cd050195bc9953c932028df8abbe0fa38368a4085", "prompt_version": "direct-options-v1", "input_tokens": 143, "allowed_token_mass": 0.999601423740387, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 73, "prefix_sha256": "76d511807440a0f245792556d6b98a4f6f88a9f5b40bec260876669cb64846f6", "encode_seconds": 0.001692019053734839, "prefill_seconds": 0.0, "copy_seconds": 0.005112198006827384, "suffix_forward_seconds": 0.06217419798485935, "forward_seconds": 0.06217419798485935, "cache_class": "DynamicCache", "total_seconds": 0.07077136501902714}
|
||||
{"id": "f41ae66b3c09630c7a42", "option_ids": ["insufficient", "A", "B"], "probabilities": [0.0014493003254756331, 0.14783263206481934, 0.850718080997467], "option_logits": [18.875, 23.5, 25.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "f079de1fc756c2a9a16a50f41f7abce95fb89851f0cb66eb132e8fb101140cd1", "prompt_version": "direct-options-v1", "input_tokens": 137, "allowed_token_mass": 0.999548077583313, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 65, "prefix_sha256": "3e8c726431e49295e2482c6695483c4e02a269a6f7764e6730d0101d9a417f31", "encode_seconds": 0.0021742620156146586, "prefill_seconds": 0.06181590701453388, "copy_seconds": 0.005075007036793977, "suffix_forward_seconds": 0.06187913496978581, "forward_seconds": 0.12369504198431969, "cache_class": "DynamicCache", "total_seconds": 0.13297334202798083}
|
||||
{"id": "82c2cd97c160ad27951b", "option_ids": ["insufficient", "B", "A"], "probabilities": [0.5293165445327759, 0.250031441450119, 0.22065196931362152], "option_logits": [22.875, 22.125, 22.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "650686b57ec17fd0c4680aa693b62247aba93c4994e6867797003610ab8c5cea", "prompt_version": "direct-options-v1", "input_tokens": 130, "allowed_token_mass": 0.9733033180236816, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 58, "prefix_sha256": "42e954b40e62269bade7bffb2a17719ea55e010f0e519854ec2b35b7e68ad129", "encode_seconds": 0.002224132011178881, "prefill_seconds": 0.06105615198612213, "copy_seconds": 0.005214138014707714, "suffix_forward_seconds": 0.0623181089758873, "forward_seconds": 0.12337426096200943, "cache_class": "DynamicCache", "total_seconds": 0.13284454197855666}
|
||||
{"id": "7ada97b6f99f4788e5c7", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.18222786486148834, 0.816688597202301, 0.00108356645796448], "option_logits": [23.25, 24.75, 18.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "7a8384db7ed5511b105b1c5f2a4ca91b48681b1cf10f55a71dbeb1fe156aa273", "prompt_version": "direct-options-v1", "input_tokens": 143, "allowed_token_mass": 0.9994260668754578, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 73, "prefix_sha256": "d33409804e3621edb3a460e2a1ed3588b0e94222c0134b8f82c17d02b2977cc3", "encode_seconds": 0.002297302009537816, "prefill_seconds": 0.060826859960798174, "copy_seconds": 0.004992567002773285, "suffix_forward_seconds": 0.06091570103308186, "forward_seconds": 0.12174256099388003, "cache_class": "DynamicCache", "total_seconds": 0.1310818920028396}
|
||||
{"id": "4d7a003342db5b0a17e3", "option_ids": ["insufficient", "A", "B"], "probabilities": [0.147593691945076, 0.8493431210517883, 0.0030632098205387592], "option_logits": [23.625, 25.375, 19.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "181cd9592c7800b01a8fd731eedb93a08a2c97cd53cdcbe9e234bb51e33b162b", "prompt_version": "direct-options-v1", "input_tokens": 146, "allowed_token_mass": 0.9996129274368286, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 73, "prefix_sha256": "d33409804e3621edb3a460e2a1ed3588b0e94222c0134b8f82c17d02b2977cc3", "encode_seconds": 0.0019032299751415849, "prefill_seconds": 0.0, "copy_seconds": 0.004951717040967196, "suffix_forward_seconds": 0.0606345790438354, "forward_seconds": 0.0606345790438354, "cache_class": "DynamicCache", "total_seconds": 0.06946678797248751}
|
||||
{"id": "19d3d8d0bf773e88be7b", "option_ids": ["A", "B", "insufficient"], "probabilities": [0.9997918009757996, 8.480057294946164e-05, 0.00012338410306256264], "option_logits": [27.0, 17.625, 18.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "abeb6af8cd72d0ae1d0c35b1fbed7f6412f9d29a7f8b4e1e7e62ee769481734e", "prompt_version": "direct-options-v1", "input_tokens": 131, "allowed_token_mass": 0.9998912811279297, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 61, "prefix_sha256": "84015749432b51633179e78a875061f4601e6ae5fc152f5fa4be9179f514d5c8", "encode_seconds": 0.0021239410270936787, "prefill_seconds": 0.058553468028549105, "copy_seconds": 0.005239358986727893, "suffix_forward_seconds": 0.06199145701248199, "forward_seconds": 0.12054492504103109, "cache_class": "DynamicCache", "total_seconds": 0.12996560597093776}
|
||||
{"id": "d6de731cca52359b1c87", "option_ids": ["insufficient", "A", "B"], "probabilities": [0.35469964146614075, 0.4554433524608612, 0.18985703587532043], "option_logits": [23.125, 23.375, 22.5], "answer_token_ids": [32, 33, 34], "prompt_sha256": "27a2717d9fde11cbe0aaeeaa18c316fa094b96432e7b79d5cb8ae892020e91d1", "prompt_version": "direct-options-v1", "input_tokens": 129, "allowed_token_mass": 0.994891881942749, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 59, "prefix_sha256": "4852534962d55f29eeb39d5012feb521714afdc7d6bac3007ac6ca3496d6979a", "encode_seconds": 0.002081720973365009, "prefill_seconds": 0.059639974031597376, "copy_seconds": 0.005637170979753137, "suffix_forward_seconds": 0.06482970103388652, "forward_seconds": 0.1244696750654839, "cache_class": "DynamicCache", "total_seconds": 0.13412972900550812}
|
||||
{"id": "b45c63fb92729af27f05", "option_ids": ["contradicted", "supported", "insufficient"], "probabilities": [0.9909536838531494, 0.00315398839302361, 0.005892425775527954], "option_logits": [26.875, 21.125, 21.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "a35eaab554c442877a0c31513dd9430e1d23867ee2ea61588175135dff50d2a4", "prompt_version": "direct-options-v1", "input_tokens": 154, "allowed_token_mass": 0.9998818039894104, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 72, "prefix_sha256": "11a1757345b0aa78565ec988a85753aa43930dce9bdc03d00a4896a20d6c13be", "encode_seconds": 0.002428852953016758, "prefill_seconds": 0.0652088150382042, "copy_seconds": 0.005784720997326076, "suffix_forward_seconds": 0.06686161301331595, "forward_seconds": 0.13207042805152014, "cache_class": "DynamicCache", "total_seconds": 0.1423667839844711}
|
||||
{"id": "46b7029b9a704138b77a", "option_ids": ["insufficient", "contradicted", "supported"], "probabilities": [0.19311703741550446, 0.04309023916721344, 0.7637927532196045], "option_logits": [23.5, 22.0, 24.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "d354ade4ac38d4997e9e145a24556d6e39d6b34a9c21b38a09481a87b33240b6", "prompt_version": "direct-options-v1", "input_tokens": 152, "allowed_token_mass": 0.9996662735939026, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 72, "prefix_sha256": "11a1757345b0aa78565ec988a85753aa43930dce9bdc03d00a4896a20d6c13be", "encode_seconds": 0.0019426009966991842, "prefill_seconds": 0.0, "copy_seconds": 0.005183967994526029, "suffix_forward_seconds": 0.06433284003287554, "forward_seconds": 0.06433284003287554, "cache_class": "DynamicCache", "total_seconds": 0.07351451099384576}
|
||||
{"id": "caf773b23934fa7362a1", "option_ids": ["supported", "contradicted", "insufficient"], "probabilities": [0.9990166425704956, 0.0001793836272554472, 0.0008039416861720383], "option_logits": [27.625, 19.0, 20.5], "answer_token_ids": [32, 33, 34], "prompt_sha256": "fc63cf543f9304e75058299b3356946c095ba6295c9243e566797f36b04ba140", "prompt_version": "direct-options-v1", "input_tokens": 147, "allowed_token_mass": 0.9999122619628906, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 65, "prefix_sha256": "52b1163b089c57f2fd22e5f6f6d464474e0b5319effcc1628526640cc48a905d", "encode_seconds": 0.0022195909987203777, "prefill_seconds": 0.06735140702221543, "copy_seconds": 0.005200487968977541, "suffix_forward_seconds": 0.06313064298592508, "forward_seconds": 0.1304820500081405, "cache_class": "DynamicCache", "total_seconds": 0.1398736200062558}
|
||||
{"id": "f46f392ef9e9e9df564b", "option_ids": ["supported", "insufficient", "contradicted"], "probabilities": [0.005880393553525209, 0.9889301657676697, 0.00518942903727293], "option_logits": [21.625, 26.75, 21.5], "answer_token_ids": [32, 33, 34], "prompt_sha256": "ede9fa16fdd6c8ccfa62f86c18b5955dd8b755fdd7b67d2b299ed26c667f4f74", "prompt_version": "direct-options-v1", "input_tokens": 145, "allowed_token_mass": 0.9998321533203125, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 63, "prefix_sha256": "6bee4909f3072e20d8bd2330d1a34a4bba5208db0029b1d4ee3b18fae806d7a1", "encode_seconds": 0.0021488909842446446, "prefill_seconds": 0.05989325599512085, "copy_seconds": 0.005170607997570187, "suffix_forward_seconds": 0.06197926704771817, "forward_seconds": 0.12187252304283902, "cache_class": "DynamicCache", "total_seconds": 0.13097951095551252}
|
||||
{"id": "005dcf6d2c0f21d64796", "option_ids": ["permitted", "prohibited", "insufficient"], "probabilities": [0.10474855452775955, 0.877048909664154, 0.01820256933569908], "option_logits": [23.5, 25.625, 21.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "7062168e1451f8e594012c07b5231030eb7ecd5b5b0544b2b556cd26b89e5906", "prompt_version": "direct-options-v1", "input_tokens": 166, "allowed_token_mass": 0.9996472001075745, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 66, "prefix_sha256": "d236ecc5b95ef935834104e9ad8e540f2fcec73ac25bdb26b53f8b6269c1c0e6", "encode_seconds": 0.002215552027337253, "prefill_seconds": 0.061645314970519394, "copy_seconds": 0.005153997975867242, "suffix_forward_seconds": 0.06214875803561881, "forward_seconds": 0.1237940730061382, "cache_class": "DynamicCache", "total_seconds": 0.13325183501001447}
|
||||
{"id": "edb708bd0fe5fd67d6ee", "option_ids": ["permitted", "prohibited", "insufficient"], "probabilities": [0.9801899790763855, 0.013981658965349197, 0.005828422494232655], "option_logits": [25.875, 21.625, 20.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "2c475aec79e3b2166c444ddb7377ff72096e81a768bf5747165b88e6ccab5727", "prompt_version": "direct-options-v1", "input_tokens": 167, "allowed_token_mass": 0.9997311234474182, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 66, "prefix_sha256": "d236ecc5b95ef935834104e9ad8e540f2fcec73ac25bdb26b53f8b6269c1c0e6", "encode_seconds": 0.0019720399868674576, "prefill_seconds": 0.0, "copy_seconds": 0.0051670289831236005, "suffix_forward_seconds": 0.06150740501470864, "forward_seconds": 0.06150740501470864, "cache_class": "DynamicCache", "total_seconds": 0.07037706201663241}
|
||||
{"id": "97a35499bd8717f3e786", "option_ids": ["prohibited", "insufficient", "permitted"], "probabilities": [0.05821435526013374, 0.0311599001288414, 0.9106257557868958], "option_logits": [22.875, 22.25, 25.625], "answer_token_ids": [32, 33, 34], "prompt_sha256": "0112886864f18bcca582892ca51af495b48b31f73e664b117c05a2060be5bb58", "prompt_version": "direct-options-v1", "input_tokens": 160, "allowed_token_mass": 0.9996643662452698, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 60, "prefix_sha256": "49ed229645b436d91c50d8e746f3583b3b9c89b477a7daeb55ab56f44148074d", "encode_seconds": 0.0024123229668475688, "prefill_seconds": 0.1031716110301204, "copy_seconds": 0.005178638035431504, "suffix_forward_seconds": 0.06260896095773205, "forward_seconds": 0.16578057198785245, "cache_class": "DynamicCache", "total_seconds": 0.1752404029830359}
|
||||
{"id": "31eb4533d2c89959ec46", "option_ids": ["insufficient", "prohibited", "permitted"], "probabilities": [0.9984009861946106, 0.00043005376937799156, 0.001169007271528244], "option_logits": [27.625, 19.875, 20.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "68a7bcbf13cc325ccc6cc6cf5b98a867728415420c9dda469ba874057911da9c", "prompt_version": "direct-options-v1", "input_tokens": 159, "allowed_token_mass": 0.999906599521637, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 59, "prefix_sha256": "299def5fddf766594cf7104e4191a32760d7c4cd7df70c58b3008e6beef48a27", "encode_seconds": 0.0023185129975900054, "prefill_seconds": 0.06344739499036223, "copy_seconds": 0.005378959001973271, "suffix_forward_seconds": 0.06593149801483378, "forward_seconds": 0.129378893005196, "cache_class": "DynamicCache", "total_seconds": 0.13914033700712025}
|
||||
{"id": "0aa24a6701ac9c56c4a0", "option_ids": ["insufficient", "B", "A"], "probabilities": [0.0014911500038579106, 0.9918259382247925, 0.006682870909571648], "option_logits": [19.625, 26.125, 21.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "d4b635852d0b74b66267df66d0083f5a57955e17bb4d0d843b9e585c42ba30b2", "prompt_version": "direct-options-v1", "input_tokens": 153, "allowed_token_mass": 0.9997501969337463, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 74, "prefix_sha256": "cd2bd87718a1eb58da20d618a8c3d9d07a219eb5c564eb0555fc5ea9db951ff1", "encode_seconds": 0.002367363020312041, "prefill_seconds": 0.06313261401373893, "copy_seconds": 0.005723811045754701, "suffix_forward_seconds": 0.06814027996733785, "forward_seconds": 0.13127289398107678, "cache_class": "DynamicCache", "total_seconds": 0.14133157901233062}
|
||||
{"id": "796d5c0da6eaadae6a3f", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.587698221206665, 0.40391868352890015, 0.008383064530789852], "option_logits": [23.875, 23.5, 19.625], "answer_token_ids": [32, 33, 34], "prompt_sha256": "8c09820ef2c6a77913dd2162c0841516c52ab13653d835822f9d68b659b5f9a1", "prompt_version": "direct-options-v1", "input_tokens": 147, "allowed_token_mass": 0.9993441104888916, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 74, "prefix_sha256": "cd2bd87718a1eb58da20d618a8c3d9d07a219eb5c564eb0555fc5ea9db951ff1", "encode_seconds": 0.0018970509991049767, "prefill_seconds": 0.0, "copy_seconds": 0.005177647981327027, "suffix_forward_seconds": 0.06794298003660515, "forward_seconds": 0.06794298003660515, "cache_class": "DynamicCache", "total_seconds": 0.07704662001924589}
|
||||
{"id": "599b3092b6fd18312d23", "option_ids": ["B", "insufficient", "A"], "probabilities": [0.04152039811015129, 0.013479698449373245, 0.9449998736381531], "option_logits": [21.875, 20.75, 25.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "189a1a1379500965960ec46859d3d4a588cff486efc846d76fbf1022d10fed67", "prompt_version": "direct-options-v1", "input_tokens": 146, "allowed_token_mass": 0.9996185898780823, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 67, "prefix_sha256": "3e094dbc3ae2dacf0588c0f4e6bc612dc8624c96cd78723e44029b2ea229596e", "encode_seconds": 0.00227085204096511, "prefill_seconds": 0.0879191390122287, "copy_seconds": 0.005668230005539954, "suffix_forward_seconds": 0.06796213996130973, "forward_seconds": 0.15588127897353843, "cache_class": "DynamicCache", "total_seconds": 0.1659015420009382}
|
||||
{"id": "960a4bcd27363e6db4b9", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.08895891904830933, 0.16619715094566345, 0.7448439002037048], "option_logits": [22.0, 22.625, 24.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "ad358c7766f5a556fb0a6a8960b52851ab4b837068ea9df106ed1b6ae51e9b8c", "prompt_version": "direct-options-v1", "input_tokens": 148, "allowed_token_mass": 0.9992754459381104, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 69, "prefix_sha256": "7faed3ebc0c6c16f4b989a1941a78bc57ff348e08bfc5ace861842d5be6cf5ed", "encode_seconds": 0.0022587919957004488, "prefill_seconds": 0.06392641796264797, "copy_seconds": 0.005300329008605331, "suffix_forward_seconds": 0.06365062604891136, "forward_seconds": 0.12757704401155934, "cache_class": "DynamicCache", "total_seconds": 0.13726208597654477}
|
||||
{"id": "c6a411e7f2bbdf377173", "option_ids": ["contradicted", "insufficient", "supported"], "probabilities": [0.9871800541877747, 0.009677973575890064, 0.0031419778242707253], "option_logits": [26.625, 22.0, 20.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "00c7d47e5abe1d713d04841e8951cbc9f044a7b02b38b56aa3c4aa36a54f32fe", "prompt_version": "direct-options-v1", "input_tokens": 164, "allowed_token_mass": 0.9998703002929688, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 83, "prefix_sha256": "bd97f511608bba6a52e708d167b6adaec9f2158b8c5995643ac5bb44149c0f62", "encode_seconds": 0.0025057640159502625, "prefill_seconds": 0.06389153801137581, "copy_seconds": 0.005326949001755565, "suffix_forward_seconds": 0.06415415904484689, "forward_seconds": 0.1280456970562227, "cache_class": "DynamicCache", "total_seconds": 0.13795260101323947}
|
||||
{"id": "e332aa927d0c2ed044aa", "option_ids": ["contradicted", "insufficient", "supported"], "probabilities": [0.4551844596862793, 0.08963112533092499, 0.4551844596862793], "option_logits": [24.0, 22.375, 24.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "2fd3cb7e37af9f73383bcc23abc5e8ef6b1ce70fc1919825393b105a8ea30152", "prompt_version": "direct-options-v1", "input_tokens": 165, "allowed_token_mass": 0.9995843172073364, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 83, "prefix_sha256": "bd97f511608bba6a52e708d167b6adaec9f2158b8c5995643ac5bb44149c0f62", "encode_seconds": 0.001991241006180644, "prefill_seconds": 0.0, "copy_seconds": 0.00526906899176538, "suffix_forward_seconds": 0.07989100198028609, "forward_seconds": 0.07989100198028609, "cache_class": "DynamicCache", "total_seconds": 0.08913247298914939}
|
||||
{"id": "a0128e551d095f3449f8", "option_ids": ["supported", "insufficient", "contradicted"], "probabilities": [0.9998365640640259, 0.00012338963279034942, 4.0058745071291924e-05], "option_logits": [28.5, 19.5, 18.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "2c485f918a788154f19750bcb2478b4c6ef054640e6778a6f2d3f1fed3f17f1f", "prompt_version": "direct-options-v1", "input_tokens": 140, "allowed_token_mass": 0.9999542832374573, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 59, "prefix_sha256": "b4a70b6f0ae7c2ff76bc6eefb22d92710e80b46699aff7c85353fefd4f5657f2", "encode_seconds": 0.0021782820113003254, "prefill_seconds": 0.06015838001621887, "copy_seconds": 0.005160657980013639, "suffix_forward_seconds": 0.06437162996735424, "forward_seconds": 0.12453000998357311, "cache_class": "DynamicCache", "total_seconds": 0.13396035198820755}
|
||||
{"id": "46d6a0731bd288c370f0", "option_ids": ["insufficient", "contradicted", "supported"], "probabilities": [0.9979650974273682, 0.0017001532251015306, 0.0003347799938637763], "option_logits": [27.5, 21.125, 19.5], "answer_token_ids": [32, 33, 34], "prompt_sha256": "af5c97e33c5cf651b441f230d1333e7e0a13cdef1422461d64b9644a6da501fd", "prompt_version": "direct-options-v1", "input_tokens": 139, "allowed_token_mass": 0.9999333024024963, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 58, "prefix_sha256": "cdaba40a62467a2e8256eb645ac0ba16074308ceef8be1576cb8f18944885d4f", "encode_seconds": 0.002183182048611343, "prefill_seconds": 0.06207748904125765, "copy_seconds": 0.005716670013498515, "suffix_forward_seconds": 0.06938527803868055, "forward_seconds": 0.1314627670799382, "cache_class": "DynamicCache", "total_seconds": 0.1414827910484746}
|
||||
{"id": "21ca7a8b970e36b8e20f", "option_ids": ["prohibited", "insufficient", "permitted"], "probabilities": [0.9258617162704468, 0.06706920266151428, 0.007069041021168232], "option_logits": [26.0, 23.375, 21.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "70ee6af686458c88dd9bee6d0b7b56e63799ebb3c3ab04a88f519c7ef8447c69", "prompt_version": "direct-options-v1", "input_tokens": 171, "allowed_token_mass": 0.999815046787262, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 68, "prefix_sha256": "e20d8330a216870ac0a99de8e24093b4722f840fe31abed9e9d5d91c936ed73c", "encode_seconds": 0.0025067839887924492, "prefill_seconds": 0.0699491510167718, "copy_seconds": 0.005714010039810091, "suffix_forward_seconds": 0.06705344503279775, "forward_seconds": 0.13700259604956955, "cache_class": "DynamicCache", "total_seconds": 0.14741198299452662}
|
||||
{"id": "4aa83af95bb69e90bc2d", "option_ids": ["permitted", "prohibited", "insufficient"], "probabilities": [0.6624118089675903, 0.18978415429592133, 0.14780405163764954], "option_logits": [24.875, 23.625, 23.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "a21f6ef35fdbe46c7c9ff58c2654e42cb82fd96db979a13ae891144002050642", "prompt_version": "direct-options-v1", "input_tokens": 169, "allowed_token_mass": 0.9996910691261292, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 68, "prefix_sha256": "e20d8330a216870ac0a99de8e24093b4722f840fe31abed9e9d5d91c936ed73c", "encode_seconds": 0.002046320994850248, "prefill_seconds": 0.0, "copy_seconds": 0.005049075989518315, "suffix_forward_seconds": 0.0645464519620873, "forward_seconds": 0.0645464519620873, "cache_class": "DynamicCache", "total_seconds": 0.0736712709767744}
|
||||
{"id": "cdfa9eed51fafcf09b72", "option_ids": ["insufficient", "prohibited", "permitted"], "probabilities": [0.35918810963630676, 0.04861082509160042, 0.5922010540962219], "option_logits": [24.25, 22.25, 24.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "9b22a228ebc1b1b9892cbe18f6b30eae62472318c97f3a390bd248f0b73ea77f", "prompt_version": "direct-options-v1", "input_tokens": 154, "allowed_token_mass": 0.9996758699417114, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 51, "prefix_sha256": "51dee4a9bf03f6306e4f6547ae0f10c11ff9e62ece7f925692b76f7621704f21", "encode_seconds": 0.0021481509902514517, "prefill_seconds": 0.06043735897401348, "copy_seconds": 0.005716170999221504, "suffix_forward_seconds": 0.06696956499945372, "forward_seconds": 0.1274069239734672, "cache_class": "DynamicCache", "total_seconds": 0.13737973698880523}
|
||||
{"id": "d51b05bfb9d6d7f09331", "option_ids": ["prohibited", "insufficient", "permitted"], "probabilities": [0.1323620229959488, 0.8631088137626648, 0.004529179539531469], "option_logits": [23.75, 25.625, 20.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "fc57539bbcedd75af63dce83973cc61e3310f710075ed1174b9a4efdaa323d0a", "prompt_version": "direct-options-v1", "input_tokens": 160, "allowed_token_mass": 0.9997425675392151, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 57, "prefix_sha256": "d097ef08c53118fef83099e663ef00df73196d8f70bcaff6240d9886140d9000", "encode_seconds": 0.0022792830131947994, "prefill_seconds": 0.059041800966951996, "copy_seconds": 0.005462980014272034, "suffix_forward_seconds": 0.06141175399534404, "forward_seconds": 0.12045355496229604, "cache_class": "DynamicCache", "total_seconds": 0.13024838001001626}
|
||||
{"id": "fdd902cc6e65fad7cb4f", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.5578213334083557, 0.383384644985199, 0.05879393592476845], "option_logits": [24.0, 23.625, 21.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "9cfdeba7abf33c7d480a26188d1fc83099b0f64c8cebde0ef2ed3dc2bb6ae96a", "prompt_version": "direct-options-v1", "input_tokens": 160, "allowed_token_mass": 0.9994832277297974, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 84, "prefix_sha256": "65688ecd12abf1f5fcfc055d9dcfac5e13c418f0b1d20dc2a06ee31c24b6feca", "encode_seconds": 0.0023259019944816828, "prefill_seconds": 0.06333657499635592, "copy_seconds": 0.004846305993851274, "suffix_forward_seconds": 0.062478339998051524, "forward_seconds": 0.12581491499440745, "cache_class": "DynamicCache", "total_seconds": 0.135068386036437}
|
||||
{"id": "42384af9ae56beb37ff5", "option_ids": ["insufficient", "A", "B"], "probabilities": [0.12195165455341339, 0.3314989507198334, 0.546549379825592], "option_logits": [22.5, 23.5, 24.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "681c30b169a28f227e9458a6d8f55ce40a3fd2238999502ebc840dcccff7ac29", "prompt_version": "direct-options-v1", "input_tokens": 160, "allowed_token_mass": 0.999487042427063, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 84, "prefix_sha256": "65688ecd12abf1f5fcfc055d9dcfac5e13c418f0b1d20dc2a06ee31c24b6feca", "encode_seconds": 0.001793620001990348, "prefill_seconds": 0.0, "copy_seconds": 0.004962896986398846, "suffix_forward_seconds": 0.061791657004505396, "forward_seconds": 0.061791657004505396, "cache_class": "DynamicCache", "total_seconds": 0.07056406600167975}
|
||||
{"id": "cdf7f428857e1b841cb3", "option_ids": ["A", "insufficient", "B"], "probabilities": [0.13209064304828644, 0.4068678021430969, 0.4610416293144226], "option_logits": [22.625, 23.75, 23.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "7bdfadd2f6c99f5ed38368c7f671350ce367a8bdc2557be56c52325707e7fadb", "prompt_version": "direct-options-v1", "input_tokens": 153, "allowed_token_mass": 0.9994603395462036, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 77, "prefix_sha256": "f6e42b826fc9e773a3eb31e24f3eb798f76fa6c976c1c418fb1dd7396a1e60c9", "encode_seconds": 0.00228503200924024, "prefill_seconds": 0.06294235301902518, "copy_seconds": 0.005156397994142026, "suffix_forward_seconds": 0.06341015599900857, "forward_seconds": 0.12635250901803374, "cache_class": "DynamicCache", "total_seconds": 0.1358728600316681}
|
||||
{"id": "774f7f9245f83c796d44", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.008546924218535423, 0.00356288836337626, 0.9878901839256287], "option_logits": [20.5, 19.625, 25.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "238763c6293e31d69a7fc77d27a71dd15e71b3c0245b8f2961b2621cd88859ca", "prompt_version": "direct-options-v1", "input_tokens": 139, "allowed_token_mass": 0.9996185898780823, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 63, "prefix_sha256": "50178952e0284413c496b1130d064f8068acddda9536f86dd051cece68ece451", "encode_seconds": 0.0022773430100642145, "prefill_seconds": 0.06399175903061405, "copy_seconds": 0.005714132043067366, "suffix_forward_seconds": 0.06753667799057439, "forward_seconds": 0.13152843702118844, "cache_class": "DynamicCache", "total_seconds": 0.14174367301166058}
|
||||
{"id": "92ec2ffa3dd46ada1335", "option_ids": ["contradicted", "insufficient", "supported"], "probabilities": [0.9662505984306335, 0.03306327760219574, 0.0006862064474262297], "option_logits": [27.125, 23.75, 19.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "4a6f088804df07ccf57c1304f0b961061d9a890070c9a0c9b72e9b8e1ed4362a", "prompt_version": "direct-options-v1", "input_tokens": 161, "allowed_token_mass": 0.999885618686676, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 78, "prefix_sha256": "31ee51b4582d718f12d373a946a61f7c9d2056258baaa82f79e1cc363a33e4a9", "encode_seconds": 0.0024915029644034803, "prefill_seconds": 0.0639630890218541, "copy_seconds": 0.00569002196425572, "suffix_forward_seconds": 0.06685430399375036, "forward_seconds": 0.13081739301560447, "cache_class": "DynamicCache", "total_seconds": 0.14109082898357883}
|
||||
{"id": "6149a17bc154f9c5b4a0", "option_ids": ["supported", "insufficient", "contradicted"], "probabilities": [0.2352859079837799, 0.49810025095939636, 0.2666138708591461], "option_logits": [23.5, 24.25, 23.625], "answer_token_ids": [32, 33, 34], "prompt_sha256": "f9d2cc7a75278b178a42c60180ab0ff8346968e83cc73116430d560e4779bb37", "prompt_version": "direct-options-v1", "input_tokens": 157, "allowed_token_mass": 0.9994412660598755, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 78, "prefix_sha256": "31ee51b4582d718f12d373a946a61f7c9d2056258baaa82f79e1cc363a33e4a9", "encode_seconds": 0.0017688689986243844, "prefill_seconds": 0.0, "copy_seconds": 0.004923387023154646, "suffix_forward_seconds": 0.06222187902312726, "forward_seconds": 0.06222187902312726, "cache_class": "DynamicCache", "total_seconds": 0.07089267601259053}
|
||||
{"id": "be866e1dbec5d42d9423", "option_ids": ["contradicted", "insufficient", "supported"], "probabilities": [0.006684042047709227, 0.0013161659007892013, 0.9919997453689575], "option_logits": [21.875, 20.25, 26.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "2135c4f1b3329fd015060893e70a657b8067f0392f0c173a2a4340b983e5ca28", "prompt_version": "direct-options-v1", "input_tokens": 158, "allowed_token_mass": 0.9998627305030823, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 75, "prefix_sha256": "40fc3b599dba917856222a68fd76c289dfe24241060bbd98655609a0df929191", "encode_seconds": 0.001965990988537669, "prefill_seconds": 0.05621890694601461, "copy_seconds": 0.005272748996503651, "suffix_forward_seconds": 0.06153802602784708, "forward_seconds": 0.1177569329738617, "cache_class": "DynamicCache", "total_seconds": 0.1270017129718326}
|
||||
{"id": "97d2ed2591fc7638806c", "option_ids": ["insufficient", "supported", "contradicted"], "probabilities": [0.775981605052948, 0.1190006360411644, 0.10501769185066223], "option_logits": [25.125, 23.25, 23.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "dc3cfc8cc05db2af408a025ade7e1dfe47e6b292f125e688d24a6f219c57ffd8", "prompt_version": "direct-options-v1", "input_tokens": 156, "allowed_token_mass": 0.9995824098587036, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 73, "prefix_sha256": "d18ef1343380441501a09313e70c053371ff90ffc9638b9a851bbdc3739c3496", "encode_seconds": 0.002156970964279026, "prefill_seconds": 0.059755175025202334, "copy_seconds": 0.004939126956742257, "suffix_forward_seconds": 0.1640677050454542, "forward_seconds": 0.22382288007065654, "cache_class": "DynamicCache", "total_seconds": 0.23269862798042595}
|
||||
{"id": "4533dd91267b4d815690", "option_ids": ["permitted", "insufficient", "prohibited"], "probabilities": [0.08948525041341782, 0.061502259224653244, 0.8490124344825745], "option_logits": [23.0, 22.625, 25.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "11a45e9457ac80d39761a348b050debb967fabac4f7e57840cf0eabfd194e32c", "prompt_version": "direct-options-v1", "input_tokens": 162, "allowed_token_mass": 0.9995995163917542, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 59, "prefix_sha256": "2bb1e8e6a9866970b9af7c9105fa4b52a74dfc9662181e2f5f063ab06814efe1", "encode_seconds": 0.0023237630375660956, "prefill_seconds": 0.05911206401651725, "copy_seconds": 0.005177447979804128, "suffix_forward_seconds": 0.0637389289913699, "forward_seconds": 0.12285099300788715, "cache_class": "DynamicCache", "total_seconds": 0.13237674604170024}
|
||||
{"id": "d880e65fef5e2993fdd0", "option_ids": ["permitted", "insufficient", "prohibited"], "probabilities": [0.9808617234230042, 0.005147089716047049, 0.013991240411996841], "option_logits": [26.625, 21.375, 22.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "4bd7a85dcd3aee9c1231a599816b036f74a048aa893eb8c15f909f011b86b239", "prompt_version": "direct-options-v1", "input_tokens": 159, "allowed_token_mass": 0.9998798966407776, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 59, "prefix_sha256": "2bb1e8e6a9866970b9af7c9105fa4b52a74dfc9662181e2f5f063ab06814efe1", "encode_seconds": 0.0019476899760775268, "prefill_seconds": 0.0, "copy_seconds": 0.005132797989062965, "suffix_forward_seconds": 0.06374467897694558, "forward_seconds": 0.06374467897694558, "cache_class": "DynamicCache", "total_seconds": 0.07277436897857115}
|
||||
{"id": "90494c76078ffc3c717a", "option_ids": ["permitted", "insufficient", "prohibited"], "probabilities": [0.9763128757476807, 0.005805368069559336, 0.017881793901324272], "option_logits": [26.0, 20.875, 22.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "80ef91a9a0d542fd92e59931d2d6df87e559eec3ffa5f82f40bfa1d7e8b4ded9", "prompt_version": "direct-options-v1", "input_tokens": 159, "allowed_token_mass": 0.9997501969337463, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 56, "prefix_sha256": "a64900d2b7035ce86b43984c6ee8dfd1db481cae71d90fcaca9a289d84e52936", "encode_seconds": 0.002329642011318356, "prefill_seconds": 0.06021404004422948, "copy_seconds": 0.005179248983040452, "suffix_forward_seconds": 0.06333210796583444, "forward_seconds": 0.12354614801006392, "cache_class": "DynamicCache", "total_seconds": 0.1330658289953135}
|
||||
{"id": "9ac5cc82c4126c91174a", "option_ids": ["permitted", "insufficient", "prohibited"], "probabilities": [0.06590891629457474, 0.9098445177078247, 0.02424653433263302], "option_logits": [22.625, 25.25, 21.625], "answer_token_ids": [32, 33, 34], "prompt_sha256": "2a21ff3f774dcac4f77d30d09cad8b73e9459584a545eac6cb875322b4f225ca", "prompt_version": "direct-options-v1", "input_tokens": 158, "allowed_token_mass": 0.9995518922805786, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 55, "prefix_sha256": "767fd3a649d68137c6d1517e8a73ad36d98980129f24947325b4d9259fc35643", "encode_seconds": 0.0022031720145605505, "prefill_seconds": 0.06266585399862379, "copy_seconds": 0.005587861000094563, "suffix_forward_seconds": 0.06295897596282884, "forward_seconds": 0.12562482996145263, "cache_class": "DynamicCache", "total_seconds": 0.13549906300613657}
|
||||
{"id": "386e10b1c1ce906473a4", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.0850694552063942, 0.9145828485488892, 0.000347659457474947], "option_logits": [23.0, 25.375, 17.5], "answer_token_ids": [32, 33, 34], "prompt_sha256": "1d8a5101ea9ef38b119f7ff6fa84001eee063903b9f2f8c29d4a84a711bb779f", "prompt_version": "direct-options-v1", "input_tokens": 164, "allowed_token_mass": 0.9995862245559692, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 91, "prefix_sha256": "f3cc104d077dbf1af3c1398dcd59fede784f851ee132fb973ba39a18f5320886", "encode_seconds": 0.002522604016121477, "prefill_seconds": 0.06132706598145887, "copy_seconds": 0.005008167994674295, "suffix_forward_seconds": 0.06282886501867324, "forward_seconds": 0.12415593100013211, "cache_class": "DynamicCache", "total_seconds": 0.1336910030222498}
|
||||
{"id": "4da6aedc1139224eeff2", "option_ids": ["insufficient", "B", "A"], "probabilities": [0.0006259420770220459, 0.9987480640411377, 0.0006259420770220459], "option_logits": [19.0, 26.375, 19.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "7bfa5a605c641906c372bc90a095885a83a1a79b1b6ef7c9a5b0a7a618e1be94", "prompt_version": "direct-options-v1", "input_tokens": 167, "allowed_token_mass": 0.999754011631012, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 91, "prefix_sha256": "f3cc104d077dbf1af3c1398dcd59fede784f851ee132fb973ba39a18f5320886", "encode_seconds": 0.0019925609813071787, "prefill_seconds": 0.0, "copy_seconds": 0.00498948700260371, "suffix_forward_seconds": 0.06520972697762772, "forward_seconds": 0.06520972697762772, "cache_class": "DynamicCache", "total_seconds": 0.07421333697857335}
|
||||
{"id": "a22dab8ba073ff6cdb8d", "option_ids": ["insufficient", "B", "A"], "probabilities": [0.002786346711218357, 0.9920080900192261, 0.005205580964684486], "option_logits": [20.375, 26.25, 21.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "0d079c6599c117e63114434ab848d564d57687b2c04a9bd49b0b3578a11a3537", "prompt_version": "direct-options-v1", "input_tokens": 160, "allowed_token_mass": 0.9997559189796448, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 87, "prefix_sha256": "e2b5772daa895ffca1061c822db2c2f8838409aee94d49b9debe02878b4c1601", "encode_seconds": 0.0023100830148905516, "prefill_seconds": 0.06219730997690931, "copy_seconds": 0.00487603695364669, "suffix_forward_seconds": 0.06070846103830263, "forward_seconds": 0.12290577101521194, "cache_class": "DynamicCache", "total_seconds": 0.13199150102445856}
|
||||
{"id": "42c54fa2ad45743e24d8", "option_ids": ["A", "insufficient", "B"], "probabilities": [0.02920997142791748, 0.9673014283180237, 0.0034886335488408804], "option_logits": [21.125, 24.625, 19.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "cb953db5af44004bc25f019ad91b924be85b901273c8582073c6eaa77b21a488", "prompt_version": "direct-options-v1", "input_tokens": 147, "allowed_token_mass": 0.9975405931472778, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 74, "prefix_sha256": "d61d6cf539eab040dc527714f0e11de749097ea145b941931c7ed8e2ecadcb16", "encode_seconds": 0.002178392023779452, "prefill_seconds": 0.0619279770180583, "copy_seconds": 0.005151168967131525, "suffix_forward_seconds": 0.06100943294586614, "forward_seconds": 0.12293740996392444, "cache_class": "DynamicCache", "total_seconds": 0.13223787100287154}
|
||||
{"id": "f3d13cf184f916677def", "option_ids": ["contradicted", "insufficient", "supported"], "probabilities": [0.9233053922653198, 0.052089329808950424, 0.02460525557398796], "option_logits": [25.875, 23.0, 22.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "4e9c9e1b15a21f8c8148d62c0397f51ec21429216b537a72a3085a2c161b5c7c", "prompt_version": "direct-options-v1", "input_tokens": 165, "allowed_token_mass": 0.9997577667236328, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 80, "prefix_sha256": "68f6d8a08ec3c097ec02bc0e8a8eda5e6ca5415f9b7154585cc6b1a61b9efe09", "encode_seconds": 0.0021232719882391393, "prefill_seconds": 0.06315812398679554, "copy_seconds": 0.005298358970321715, "suffix_forward_seconds": 0.06255769199924543, "forward_seconds": 0.12571581598604098, "cache_class": "DynamicCache", "total_seconds": 0.13513343699742109}
|
||||
{"id": "1105577da4c8dad4609d", "option_ids": ["contradicted", "insufficient", "supported"], "probabilities": [0.7264880537986755, 0.1621014028787613, 0.11141055822372437], "option_logits": [24.875, 23.375, 23.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "317a37f765f3e0cac9771ade06dd064bad03f3816b4f4966511910d17455ae92", "prompt_version": "direct-options-v1", "input_tokens": 162, "allowed_token_mass": 0.9996148347854614, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 80, "prefix_sha256": "68f6d8a08ec3c097ec02bc0e8a8eda5e6ca5415f9b7154585cc6b1a61b9efe09", "encode_seconds": 0.00198346097022295, "prefill_seconds": 0.0, "copy_seconds": 0.005161217995919287, "suffix_forward_seconds": 0.06230366096133366, "forward_seconds": 0.06230366096133366, "cache_class": "DynamicCache", "total_seconds": 0.07162945199524984}
|
||||
{"id": "0ebd299739302d8b057b", "option_ids": ["contradicted", "insufficient", "supported"], "probabilities": [0.8497095704078674, 0.1303071826696396, 0.019983254373073578], "option_logits": [25.0, 23.125, 21.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "dec41e6f47308e69e9cc5ef204fb4d237689b085bf0c78dc2685a234817ebdfe", "prompt_version": "direct-options-v1", "input_tokens": 158, "allowed_token_mass": 0.9996262192726135, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 73, "prefix_sha256": "f57fe3335d38de3001c3106d77798e345246bee7405bd2abb8c617a55100ad35", "encode_seconds": 0.002163461991585791, "prefill_seconds": 0.06284247298026457, "copy_seconds": 0.004943337000440806, "suffix_forward_seconds": 0.06202591798501089, "forward_seconds": 0.12486839096527547, "cache_class": "DynamicCache", "total_seconds": 0.13388405099976808}
|
||||
{"id": "eafc22c8c40df3932a8e", "option_ids": ["supported", "contradicted", "insufficient"], "probabilities": [0.0202607661485672, 0.003520793514326215, 0.9762184619903564], "option_logits": [22.75, 21.0, 26.625], "answer_token_ids": [32, 33, 34], "prompt_sha256": "11c4b0b735f28970434fba90fd49dedf800210fc86cc5dd0e5f75acd828cc477", "prompt_version": "direct-options-v1", "input_tokens": 146, "allowed_token_mass": 0.9998321533203125, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 61, "prefix_sha256": "604746f5f6ab0509efd5f4374dc9cb668edc40a32e9d5d9c4b57760ecffd25e6", "encode_seconds": 0.0022326430189423263, "prefill_seconds": 0.06126572500215843, "copy_seconds": 0.005333429959136993, "suffix_forward_seconds": 0.06535106699448079, "forward_seconds": 0.12661679199663922, "cache_class": "DynamicCache", "total_seconds": 0.13617880397941917}
|
||||
{"id": "0aa1447680e4f25862bb", "option_ids": ["insufficient", "prohibited", "permitted"], "probabilities": [0.3890247941017151, 0.1114574745297432, 0.4995177388191223], "option_logits": [24.125, 22.875, 24.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "6c19667d59b60941b80246296f8d75c2eab7aa759ed293293062d0650b53b21a", "prompt_version": "direct-options-v1", "input_tokens": 165, "allowed_token_mass": 0.9995747208595276, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 63, "prefix_sha256": "7aa8993ba91caf1f4b5815880b451e85f85480ec0f42db00c74732a84cb94d12", "encode_seconds": 0.0023342930362559855, "prefill_seconds": 0.05423569603590295, "copy_seconds": 0.004497535002883524, "suffix_forward_seconds": 0.060772982018534094, "forward_seconds": 0.11500867805443704, "cache_class": "DynamicCache", "total_seconds": 0.12377880600979552}
|
||||
{"id": "980e76145c925747220c", "option_ids": ["insufficient", "prohibited", "permitted"], "probabilities": [0.26380202174186707, 0.019109752029180527, 0.7170881628990173], "option_logits": [23.75, 21.125, 24.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "4f4c4c0173e698912a697967d581fa7be41975f22aa643a71a4018f9df547aed", "prompt_version": "direct-options-v1", "input_tokens": 166, "allowed_token_mass": 0.9995461702346802, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 63, "prefix_sha256": "7aa8993ba91caf1f4b5815880b451e85f85480ec0f42db00c74732a84cb94d12", "encode_seconds": 0.001731769007164985, "prefill_seconds": 0.0, "copy_seconds": 0.004723196034319699, "suffix_forward_seconds": 0.0631869250210002, "forward_seconds": 0.0631869250210002, "cache_class": "DynamicCache", "total_seconds": 0.07138049899367616}
|
||||
{"id": "4aee37a6a155747766aa", "option_ids": ["permitted", "prohibited", "insufficient"], "probabilities": [0.966343879699707, 0.02918105758726597, 0.004475060384720564], "option_logits": [26.0, 22.5, 20.625], "answer_token_ids": [32, 33, 34], "prompt_sha256": "f3212577729f3a2ba2c4b443b0184aed36fcf84640102006f754d74e0ad2524c", "prompt_version": "direct-options-v1", "input_tokens": 163, "allowed_token_mass": 0.9997292160987854, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 61, "prefix_sha256": "d89a26efb6954c5dd10b0c653a4b443f3f2aff480e7ed2052f612141d06aeed6", "encode_seconds": 0.0020155810052528977, "prefill_seconds": 0.06197856902144849, "copy_seconds": 0.005485620989929885, "suffix_forward_seconds": 0.06690638500731438, "forward_seconds": 0.12888495402876288, "cache_class": "DynamicCache", "total_seconds": 0.13850126700708643}
|
||||
{"id": "6db17cc22bd656d78558", "option_ids": ["prohibited", "permitted", "insufficient"], "probabilities": [0.02919641137123108, 0.003951304592192173, 0.9668523073196411], "option_logits": [22.375, 20.375, 25.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "3b494eab820fd916edb55b080a516144b28012c2a1d16407620246c79765f895", "prompt_version": "direct-options-v1", "input_tokens": 156, "allowed_token_mass": 0.9997501969337463, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 54, "prefix_sha256": "b4be924d1173101da5ac16d94a9bf724c32a79e9fcee2215dfc427f9c1019fd6", "encode_seconds": 0.0021844920120202005, "prefill_seconds": 0.05777459597447887, "copy_seconds": 0.004873875994235277, "suffix_forward_seconds": 0.06043971999315545, "forward_seconds": 0.11821431596763432, "cache_class": "DynamicCache", "total_seconds": 0.1272637450019829}
|
||||
{"id": "ef3e0f8e657bf59f8a29", "option_ids": ["A", "insufficient", "B"], "probabilities": [0.1790020912885666, 0.4865780770778656, 0.3344199061393738], "option_logits": [22.625, 23.625, 23.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "763acb6345a08c26de32359d6fcba885b4e8daa99c80c266fcf4e43bf49c0dc6", "prompt_version": "direct-options-v1", "input_tokens": 159, "allowed_token_mass": 0.9991172552108765, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 86, "prefix_sha256": "986b066c18d6ffaa6d131425f464fc1c838e84b398b73fdd6505c30c50686393", "encode_seconds": 0.0019786609918810427, "prefill_seconds": 0.05829299794277176, "copy_seconds": 0.004369012953247875, "suffix_forward_seconds": 0.05887815199093893, "forward_seconds": 0.1171711499337107, "cache_class": "DynamicCache", "total_seconds": 0.12547936500050128}
|
||||
{"id": "f756e1cb04e11766341e", "option_ids": ["insufficient", "A", "B"], "probabilities": [0.38338467478752136, 0.5578213930130005, 0.058793939650058746], "option_logits": [23.625, 24.0, 21.75], "answer_token_ids": [32, 33, 34], "prompt_sha256": "ded2517a45e0b26d0614f6c4afe809c3c73b2d85f930b402e766a50bcd8cf88d", "prompt_version": "direct-options-v1", "input_tokens": 165, "allowed_token_mass": 0.999340295791626, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 86, "prefix_sha256": "986b066c18d6ffaa6d131425f464fc1c838e84b398b73fdd6505c30c50686393", "encode_seconds": 0.001722508983220905, "prefill_seconds": 0.0, "copy_seconds": 0.004281363973859698, "suffix_forward_seconds": 0.06078504305332899, "forward_seconds": 0.06078504305332899, "cache_class": "DynamicCache", "total_seconds": 0.06854368495987728}
|
||||
{"id": "d167e51bb1f499099645", "option_ids": ["A", "B", "insufficient"], "probabilities": [0.8081498146057129, 0.18032260239124298, 0.01152763795107603], "option_logits": [24.375, 22.875, 20.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "521099c074de1134b6a149b6d197a8eaace907ec8c0cd49e0e66f00c5471b946", "prompt_version": "direct-options-v1", "input_tokens": 145, "allowed_token_mass": 0.9992374181747437, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 72, "prefix_sha256": "10688ff9f65059998d0244681282424f4b26737c9bd4f120863dabe622fd8f53", "encode_seconds": 0.0022127320407889783, "prefill_seconds": 0.061937736987601966, "copy_seconds": 0.005093218002002686, "suffix_forward_seconds": 0.05985178699484095, "forward_seconds": 0.12178952398244292, "cache_class": "DynamicCache", "total_seconds": 0.1311500160372816}
|
||||
{"id": "674ca4e9346216255466", "option_ids": ["insufficient", "B", "A"], "probabilities": [0.9982689619064331, 0.0015008366899564862, 0.00023016076011117548], "option_logits": [26.375, 19.875, 18.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "50ef107d3b78c73144fc41acef050a54cd09f92ca69d952ec90ec37d479909ec", "prompt_version": "direct-options-v1", "input_tokens": 135, "allowed_token_mass": 0.9998131394386292, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 62, "prefix_sha256": "9cad740414b18b7c6cc758f37ddefcb9e8f376d4167f10126530589fd84af33d", "encode_seconds": 0.0018640699563547969, "prefill_seconds": 0.055053531017620116, "copy_seconds": 0.004908406990580261, "suffix_forward_seconds": 0.05986771697644144, "forward_seconds": 0.11492124799406156, "cache_class": "DynamicCache", "total_seconds": 0.12378043599892408}
|
||||
{"id": "1df851a202994a07f123", "option_ids": ["contradicted", "insufficient", "supported"], "probabilities": [0.9982689619064331, 0.0015008366899564862, 0.00023016076011117548], "option_logits": [27.375, 20.875, 19.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "69b4fdac8eed6118753c0d2f607361fb90f2fe12243390e71b9085b3904344ca", "prompt_version": "direct-options-v1", "input_tokens": 151, "allowed_token_mass": 0.9998989105224609, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 71, "prefix_sha256": "2017d66e795bce442dc6ce28f076cc3da0c1995286a2c0b3d3e94be67d22db72", "encode_seconds": 0.001988289994187653, "prefill_seconds": 0.0651206859620288, "copy_seconds": 0.005511349008884281, "suffix_forward_seconds": 0.0635957769700326, "forward_seconds": 0.1287164629320614, "cache_class": "DynamicCache", "total_seconds": 0.13834130502073094}
|
||||
{"id": "b5e6788d95b8395d2888", "option_ids": ["insufficient", "contradicted", "supported"], "probabilities": [0.0024724386166781187, 7.466118404408917e-05, 0.9974529147148132], "option_logits": [21.375, 17.875, 27.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "a846c6370e83972fc7d6953e7e7b5f6c5ffcb62dfef5392e141e0f7fc0e43bfc", "prompt_version": "direct-options-v1", "input_tokens": 150, "allowed_token_mass": 0.9998798966407776, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 71, "prefix_sha256": "2017d66e795bce442dc6ce28f076cc3da0c1995286a2c0b3d3e94be67d22db72", "encode_seconds": 0.0018393899663351476, "prefill_seconds": 0.0, "copy_seconds": 0.005257149052340537, "suffix_forward_seconds": 0.060925793019123375, "forward_seconds": 0.060925793019123375, "cache_class": "DynamicCache", "total_seconds": 0.0697229519719258}
|
||||
{"id": "32f36cd0fbbfb3604586", "option_ids": ["supported", "insufficient", "contradicted"], "probabilities": [0.9989927411079407, 0.000803922419436276, 0.0002032634220086038], "option_logits": [27.625, 20.5, 19.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "2ef98885b2786408c702303589adf8c6275d7e304ec6a5db93d8584d7978b6df", "prompt_version": "direct-options-v1", "input_tokens": 140, "allowed_token_mass": 0.9999085068702698, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 60, "prefix_sha256": "b79b57c1a1497128d4f3a7293d8534d5242bf9d8481edd31c9812f5d048499b1", "encode_seconds": 0.001957371016032994, "prefill_seconds": 0.0553313730051741, "copy_seconds": 0.00466556497849524, "suffix_forward_seconds": 0.06019988900516182, "forward_seconds": 0.11553126201033592, "cache_class": "DynamicCache", "total_seconds": 0.12423042004229501}
|
||||
{"id": "843185706adff0c60f8b", "option_ids": ["contradicted", "supported", "insufficient"], "probabilities": [0.9648732542991638, 0.002110651694238186, 0.0330161452293396], "option_logits": [26.5, 20.375, 23.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "fd026383e29f6c9f8ae01a9680421f62d32fae1fd6966545c00bda45a171f692", "prompt_version": "direct-options-v1", "input_tokens": 141, "allowed_token_mass": 0.9998779296875, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 61, "prefix_sha256": "46a9e52a303d3e22d717063861ea82464ab6db10e443708db0f679275fa1d6fb", "encode_seconds": 0.0020321510382927954, "prefill_seconds": 0.056149077019654214, "copy_seconds": 0.004817395994905382, "suffix_forward_seconds": 0.06231655995361507, "forward_seconds": 0.11846563697326928, "cache_class": "DynamicCache", "total_seconds": 0.12729828502051532}
|
||||
{"id": "b531dd603bf080dab820", "option_ids": ["prohibited", "permitted", "insufficient"], "probabilities": [0.996208906173706, 0.0013217504601925611, 0.002469355007633567], "option_logits": [27.875, 21.25, 21.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "d6239359272f7493a03cf998a93a5f142afe35edbc986f2f3e6a375a946785a3", "prompt_version": "direct-options-v1", "input_tokens": 156, "allowed_token_mass": 0.9999275803565979, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 55, "prefix_sha256": "91848655a9571cca358164ef430aa04979d46df6653bb41a6fd7c7c3c44edee2", "encode_seconds": 0.0022988829878158867, "prefill_seconds": 0.06025926902657375, "copy_seconds": 0.005214449018239975, "suffix_forward_seconds": 0.06358688801992685, "forward_seconds": 0.1238461570465006, "cache_class": "DynamicCache", "total_seconds": 0.13338514999486506}
|
||||
{"id": "be2dd67bc733c1213bef", "option_ids": ["prohibited", "permitted", "insufficient"], "probabilities": [0.004069637507200241, 0.995807409286499, 0.00012289239384699613], "option_logits": [22.625, 28.125, 19.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "4c511c91fe67ef8136d90d9dae8984a3de204b40a97289112a48e6a2317d6092", "prompt_version": "direct-options-v1", "input_tokens": 159, "allowed_token_mass": 0.9999046921730042, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 55, "prefix_sha256": "91848655a9571cca358164ef430aa04979d46df6653bb41a6fd7c7c3c44edee2", "encode_seconds": 0.0018999199965037405, "prefill_seconds": 0.0, "copy_seconds": 0.005219167971517891, "suffix_forward_seconds": 0.06371864804532379, "forward_seconds": 0.06371864804532379, "cache_class": "DynamicCache", "total_seconds": 0.07277545699616894}
|
||||
{"id": "b8e2750307df9761e716", "option_ids": ["permitted", "prohibited", "insufficient"], "probabilities": [0.9966529011726379, 0.0021801693364977837, 0.0011669604573398829], "option_logits": [27.25, 21.125, 20.5], "answer_token_ids": [32, 33, 34], "prompt_sha256": "459a32663af6cc8f69a26f9847f681e5bda5767b28373602aa3a9090e0cb564e", "prompt_version": "direct-options-v1", "input_tokens": 153, "allowed_token_mass": 0.9998837113380432, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 52, "prefix_sha256": "bc2bbfb919d2c504ef0135dd70a34755d2c0f13e02540842569f5e7bd2350c0f", "encode_seconds": 0.002165931975468993, "prefill_seconds": 0.06010288803372532, "copy_seconds": 0.005243418971076608, "suffix_forward_seconds": 0.06371227896306664, "forward_seconds": 0.12381516699679196, "cache_class": "DynamicCache", "total_seconds": 0.13326619897270575}
|
||||
{"id": "5d31b4bf2d0fc193f53b", "option_ids": ["permitted", "insufficient", "prohibited"], "probabilities": [0.07503016293048859, 0.6282198429107666, 0.2967500388622284], "option_logits": [22.625, 24.75, 24.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "32a5512660eb3992ace5f91dfe6fd2142d40021de63cb30b25a24efe5f2d2bc4", "prompt_version": "direct-options-v1", "input_tokens": 158, "allowed_token_mass": 0.9995232820510864, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 57, "prefix_sha256": "cd8e96d857d5551ca7288e29f4eab3c21b38935fc86ba21978dbe740767b0f1c", "encode_seconds": 0.0022925320081412792, "prefill_seconds": 0.06177041894989088, "copy_seconds": 0.005214387958403677, "suffix_forward_seconds": 0.06293970497790724, "forward_seconds": 0.12471012392779812, "cache_class": "DynamicCache", "total_seconds": 0.1343022850342095}
|
||||
{"id": "a9230edfc6ec9e524af0", "option_ids": ["B", "insufficient", "A"], "probabilities": [0.029232529923319817, 0.0027190488763153553, 0.9680483937263489], "option_logits": [21.375, 19.0, 24.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "f41d6e5987b6e589dfbcfd2b325e6198d5442f99fda5a81059c8d6f5c9a1bf5c", "prompt_version": "direct-options-v1", "input_tokens": 144, "allowed_token_mass": 0.9995843172073364, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 76, "prefix_sha256": "1581ac476e86b868b99148592de9b2d37a49c41cb2ab73450bbf4593803d13e5", "encode_seconds": 0.002292632998432964, "prefill_seconds": 0.06036635098280385, "copy_seconds": 0.004609084979165345, "suffix_forward_seconds": 0.05890279199229553, "forward_seconds": 0.11926914297509938, "cache_class": "DynamicCache", "total_seconds": 0.12805346999084577}
|
||||
{"id": "c100e84585149a04be49", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.20073364675045013, 0.7939170002937317, 0.005349370650947094], "option_logits": [22.5, 23.875, 18.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "d3943b819f784c5dd874663a6a22fc904170874be1d830f25058440aa699a08f", "prompt_version": "direct-options-v1", "input_tokens": 150, "allowed_token_mass": 0.9986771941184998, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 76, "prefix_sha256": "1581ac476e86b868b99148592de9b2d37a49c41cb2ab73450bbf4593803d13e5", "encode_seconds": 0.0018042390001937747, "prefill_seconds": 0.0, "copy_seconds": 0.004830776015296578, "suffix_forward_seconds": 0.05592414498096332, "forward_seconds": 0.05592414498096332, "cache_class": "DynamicCache", "total_seconds": 0.06442223198246211}
|
||||
{"id": "579beed184ffc1104aa9", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.4062184989452362, 0.5910443663597107, 0.0027370788156986237], "option_logits": [23.5, 23.875, 18.5], "answer_token_ids": [32, 33, 34], "prompt_sha256": "1a572f47773adc220c8c1a894d7dc18ede1430513c07763269188f195e11b575", "prompt_version": "direct-options-v1", "input_tokens": 141, "allowed_token_mass": 0.9983477592468262, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 73, "prefix_sha256": "b47a105769e6a2518610f904133065f71b6f79e251df1b8e34066d40bb6146ee", "encode_seconds": 0.0018186999950557947, "prefill_seconds": 0.05921575298998505, "copy_seconds": 0.004525034979451448, "suffix_forward_seconds": 0.05953203502576798, "forward_seconds": 0.11874778801575303, "cache_class": "DynamicCache", "total_seconds": 0.12697752396343276}
|
||||
{"id": "c4e9bd463b6fde5be743", "option_ids": ["B", "insufficient", "A"], "probabilities": [0.0005526343593373895, 0.9991863369941711, 0.0002610459632705897], "option_logits": [19.625, 27.125, 18.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "26a305c84453483fb9d7fbe4e8f1757665bbadbddae4f7b9fe7111def94f0453", "prompt_version": "direct-options-v1", "input_tokens": 129, "allowed_token_mass": 0.9998818039894104, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 61, "prefix_sha256": "9601851a6ccf0fc312613b8715b5669568c64418ca7044bfed794bdb47cc2e45", "encode_seconds": 0.0020872409804724157, "prefill_seconds": 0.06086980295367539, "copy_seconds": 0.005500569997821003, "suffix_forward_seconds": 0.06671654497040436, "forward_seconds": 0.12758634792407975, "cache_class": "DynamicCache", "total_seconds": 0.13724364095833153}
|
||||
{"id": "b863998d4be2721a5dfd", "option_ids": ["supported", "insufficient", "contradicted"], "probabilities": [0.08218361437320709, 0.03425922617316246, 0.883557140827179], "option_logits": [22.75, 21.875, 25.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "747d5555450fec4e3c94f28460da102904addfc8d683a3517d99c291afc2b456", "prompt_version": "direct-options-v1", "input_tokens": 162, "allowed_token_mass": 0.9995671510696411, "full_vocab_argmax_id": 34, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 78, "prefix_sha256": "957976ddb68a79d0196faa9a3fe8397b1b05b88b58fb69bd353b3eeb03ce5f9e", "encode_seconds": 0.0021189520484767854, "prefill_seconds": 0.06208330002846196, "copy_seconds": 0.004926837049424648, "suffix_forward_seconds": 0.06138059601653367, "forward_seconds": 0.12346389604499564, "cache_class": "DynamicCache", "total_seconds": 0.13255154603393748}
|
||||
{"id": "37cebc3792e98049439f", "option_ids": ["insufficient", "supported", "contradicted"], "probabilities": [0.1980946660041809, 0.783479630947113, 0.018425675109028816], "option_logits": [24.0, 25.375, 21.625], "answer_token_ids": [32, 33, 34], "prompt_sha256": "4e3a5619ebbfb97e63813dddc790919b6a56a9eb034d174e15de485a00895a83", "prompt_version": "direct-options-v1", "input_tokens": 159, "allowed_token_mass": 0.9997215867042542, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 78, "prefix_sha256": "957976ddb68a79d0196faa9a3fe8397b1b05b88b58fb69bd353b3eeb03ce5f9e", "encode_seconds": 0.001800588972400874, "prefill_seconds": 0.0, "copy_seconds": 0.005059438000898808, "suffix_forward_seconds": 0.05835005996050313, "forward_seconds": 0.05835005996050313, "cache_class": "DynamicCache", "total_seconds": 0.06723995797801763}
|
||||
{"id": "64d63f86e328aff1481c", "option_ids": ["contradicted", "insufficient", "supported"], "probabilities": [0.7225208878517151, 0.011678462848067284, 0.2658006250858307], "option_logits": [25.125, 21.0, 24.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "f46f7543e4b7d4d6bef565787a91dc9ed68afd3bbaeca8fb0428f27079d9cd66", "prompt_version": "direct-options-v1", "input_tokens": 167, "allowed_token_mass": 0.9997158050537109, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 83, "prefix_sha256": "9e52f4f53fe92a21613d646f6c7b644a53189756335eb66a2f5bec43fe042a76", "encode_seconds": 0.002089140994939953, "prefill_seconds": 0.06028839002829045, "copy_seconds": 0.004929247021209449, "suffix_forward_seconds": 0.06169021700043231, "forward_seconds": 0.12197860702872276, "cache_class": "DynamicCache", "total_seconds": 0.1310675370041281}
|
||||
{"id": "6d64a8459f64eac7f4cc", "option_ids": ["insufficient", "contradicted", "supported"], "probabilities": [0.9916307330131531, 0.007571193855255842, 0.0007979979855008423], "option_logits": [26.375, 21.5, 19.25], "answer_token_ids": [32, 33, 34], "prompt_sha256": "a45ac73e1c46b7566c9c3ef3a2b588aa4bf73fcaefcda5b9a5770f8386cfeb52", "prompt_version": "direct-options-v1", "input_tokens": 164, "allowed_token_mass": 0.9998188614845276, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 80, "prefix_sha256": "44b58668d0cd45790300b18efaf5f521531c0da8c4d9981ec598c0c1824d12f0", "encode_seconds": 0.0020499309757724404, "prefill_seconds": 0.06402777996845543, "copy_seconds": 0.00524901895551011, "suffix_forward_seconds": 0.06281005503842607, "forward_seconds": 0.1268378350068815, "cache_class": "DynamicCache", "total_seconds": 0.13620263600023463}
|
||||
{"id": "2708193212d8a4a523c7", "option_ids": ["prohibited", "permitted", "insufficient"], "probabilities": [0.1920153796672821, 0.7594355344772339, 0.04854908958077431], "option_logits": [23.375, 24.75, 22.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "20ef6c7e0dfa41002f3d00c3bcd33fb357967fa21b8e86ae94c1a2c93fb7e873", "prompt_version": "direct-options-v1", "input_tokens": 164, "allowed_token_mass": 0.9994909167289734, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 56, "prefix_sha256": "fce41a35de7762e7a08cd131a8a7a48c14e930e441c6ddf87665ef467c0bc202", "encode_seconds": 0.002372672955971211, "prefill_seconds": 0.05826530797639862, "copy_seconds": 0.004706865991465747, "suffix_forward_seconds": 0.06018124002730474, "forward_seconds": 0.11844654800370336, "cache_class": "DynamicCache", "total_seconds": 0.12761143897660077}
|
||||
{"id": "5a37d1dfcad20810a76a", "option_ids": ["prohibited", "permitted", "insufficient"], "probabilities": [0.005207285284996033, 0.992332935333252, 0.0024597474839538336], "option_logits": [21.625, 26.875, 20.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "90b2bf3b7cdc5ee4f27202735ebe531440fb479e78eb32b0c879bacd78f84619", "prompt_version": "direct-options-v1", "input_tokens": 153, "allowed_token_mass": 0.9998035430908203, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 56, "prefix_sha256": "fce41a35de7762e7a08cd131a8a7a48c14e930e441c6ddf87665ef467c0bc202", "encode_seconds": 0.0018961199675686657, "prefill_seconds": 0.0, "copy_seconds": 0.0051663179765455425, "suffix_forward_seconds": 0.06217615999048576, "forward_seconds": 0.06217615999048576, "cache_class": "DynamicCache", "total_seconds": 0.07122325897216797}
|
||||
{"id": "acacd8b46acd25cb954b", "option_ids": ["insufficient", "permitted", "prohibited"], "probabilities": [0.5870726108551025, 0.24472828209400177, 0.16819912195205688], "option_logits": [24.125, 23.25, 22.875], "answer_token_ids": [32, 33, 34], "prompt_sha256": "276609bd47db80fd0a0ccaa40158b3fb45a297a68cd0d0928af9e0a95d2398e5", "prompt_version": "direct-options-v1", "input_tokens": 163, "allowed_token_mass": 0.9994356036186218, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 55, "prefix_sha256": "c28b3b69b58922a41220bb85a4d6eac9f38f85064975547454b02b2b996259d8", "encode_seconds": 0.0021815020008943975, "prefill_seconds": 0.05680979200406, "copy_seconds": 0.004642284999135882, "suffix_forward_seconds": 0.05867716099601239, "forward_seconds": 0.11548695300007239, "cache_class": "DynamicCache", "total_seconds": 0.12436129100387916}
|
||||
{"id": "4089c199b8b796ec3ed8", "option_ids": ["insufficient", "prohibited", "permitted"], "probabilities": [0.9982863068580627, 0.0009103193297050893, 0.000803353963419795], "option_logits": [27.5, 20.5, 20.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "b86a9e762a27d5ac37b72ddde23a389373979a4cf97d2526d3d1a57c0f492f43", "prompt_version": "direct-options-v1", "input_tokens": 162, "allowed_token_mass": 0.9999217987060547, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 54, "prefix_sha256": "2652d20d2579136b08b5de081e07989e2100dedff2d7701f27ae75bc2014bdc7", "encode_seconds": 0.002011770964600146, "prefill_seconds": 0.055871955992188305, "copy_seconds": 0.0048248760285787284, "suffix_forward_seconds": 0.060486711969133466, "forward_seconds": 0.11635866796132177, "cache_class": "DynamicCache", "total_seconds": 0.1252290359698236}
|
||||
{"id": "02f2b95a5c5a4777a45c", "option_ids": ["B", "A", "insufficient"], "probabilities": [0.843149721622467, 0.07842513918876648, 0.07842513918876648], "option_logits": [24.375, 22.0, 22.0], "answer_token_ids": [32, 33, 34], "prompt_sha256": "6b87d7fc14638fdba408d8ccad5ea800ded035ca7ae3ed8781f0ea7933fd7b77", "prompt_version": "direct-options-v1", "input_tokens": 150, "allowed_token_mass": 0.9988886117935181, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 69, "prefix_sha256": "9f77dcffa32e7e894efd45c3225015d845c87329689a043e2488d28183044943", "encode_seconds": 0.0022605819976888597, "prefill_seconds": 0.0607187720015645, "copy_seconds": 0.004898417042568326, "suffix_forward_seconds": 0.06136556703131646, "forward_seconds": 0.12208433903288096, "cache_class": "DynamicCache", "total_seconds": 0.1312079980270937}
|
||||
{"id": "ebea4681e3dcc7d92ca2", "option_ids": ["insufficient", "B", "A"], "probabilities": [0.15600457787513733, 0.6170100569725037, 0.22698535025119781], "option_logits": [22.75, 24.125, 23.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "125a6a62eabfa09fb78e3d63ca393764c6cda2f4b1f4aa3a4edd982805bf0c9f", "prompt_version": "direct-options-v1", "input_tokens": 155, "allowed_token_mass": 0.9994622468948364, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": true, "prefix_tokens": 69, "prefix_sha256": "9f77dcffa32e7e894efd45c3225015d845c87329689a043e2488d28183044943", "encode_seconds": 0.001955061045009643, "prefill_seconds": 0.0, "copy_seconds": 0.005095318017993122, "suffix_forward_seconds": 0.06321831600507721, "forward_seconds": 0.06321831600507721, "cache_class": "DynamicCache", "total_seconds": 0.07223884604172781}
|
||||
{"id": "b8c5b9b285cbdddc9f8d", "option_ids": ["A", "insufficient", "B"], "probabilities": [0.024404946714639664, 0.7132171392440796, 0.26237794756889343], "option_logits": [20.75, 24.125, 23.125], "answer_token_ids": [32, 33, 34], "prompt_sha256": "ef86a2efa0123bc0abda869e39663f6983c31c0d17dc4976e78656d3fdd35987", "prompt_version": "direct-options-v1", "input_tokens": 146, "allowed_token_mass": 0.9991916418075562, "full_vocab_argmax_id": 33, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 65, "prefix_sha256": "f7b40e844531ed2729b82dbe4c3ac294dfbaa618cb7299993041d2cb6dfef65c", "encode_seconds": 0.0021850920165888965, "prefill_seconds": 0.06112899602157995, "copy_seconds": 0.005282528989482671, "suffix_forward_seconds": 0.059335943951737136, "forward_seconds": 0.12046493997331709, "cache_class": "DynamicCache", "total_seconds": 0.1299561019986868}
|
||||
{"id": "7ed75d854d99b3a81916", "option_ids": ["insufficient", "B", "A"], "probabilities": [0.9757856726646423, 0.015772106125950813, 0.008442200720310211], "option_logits": [25.125, 21.0, 20.375], "answer_token_ids": [32, 33, 34], "prompt_sha256": "7275507a830221826fad9d79c03010a9a18120771a024d33aecca5d0c7a68dbe", "prompt_version": "direct-options-v1", "input_tokens": 144, "allowed_token_mass": 0.996342658996582, "full_vocab_argmax_id": 32, "model": {"source": "Qwen/Qwen3.5-4B", "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "adapter": null, "adapter_sha256": null, "adapter_revision": null, "dtype": "bfloat16", "transformers_version": "5.17.0", "torch_version": "2.10.0+cu128", "serving_config": "native-state-prefix-cache-v1"}, "readout": "native-state-prefix-cache-last-position", "cache_hit": false, "prefix_tokens": 63, "prefix_sha256": "04955060fbe24f10ffb95070365a8830f7cece2e18e19f4c61bde837b394b76d", "encode_seconds": 0.0020378209883347154, "prefill_seconds": 0.05871883197687566, "copy_seconds": 0.005325788981281221, "suffix_forward_seconds": 0.06492741499096155, "forward_seconds": 0.12364624696783721, "cache_class": "DynamicCache", "total_seconds": 0.13292212801752612}
|
||||
@@ -0,0 +1,108 @@
|
||||
{"family": "evidence_interpretation", "group_id": "2fdaa8a61e6e989dd866/stability", "id": "14790d6d50043b03420c", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "a3f18f3a63d45345942b", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "2fdaa8a61e6e989dd866", "variant": "option_reversal"}, "question": "Assess the claim: the replacement lenses have been fitted.", "split": "rebase_stability", "state": "The optician ordered replacement lenses. The workshop confirms they have not yet been fitted to the customer's glasses.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "2fdaa8a61e6e989dd866/stability", "id": "29566b5efbd2611e6eee", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"base_id": "a3f18f3a63d45345942b", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "2fdaa8a61e6e989dd866", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Assess the claim: the replacement lenses have been fitted.", "split": "rebase_stability", "state": "The optician ordered replacement lenses. The workshop confirms they have not yet been fitted to the customer's glasses.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "2fdaa8a61e6e989dd866/stability", "id": "1a6e02f5b0bf923374c1", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"base_id": "a3f18f3a63d45345942b", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "2fdaa8a61e6e989dd866", "variant": "irrelevant_context"}, "question": "Assess the claim: the replacement lenses have been fitted.", "split": "rebase_stability", "state": "The optician ordered replacement lenses. The workshop confirms they have not yet been fitted to the customer's glasses.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "f4bea3d3bcc01056740d/stability", "id": "d83888299b2e202297be", "label": 2, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "e0c140e9222d1bce5647", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "f4bea3d3bcc01056740d", "variant": "option_reversal"}, "question": "Assess the claim: the guest endorses the listing's towel statement.", "split": "rebase_stability", "state": "The guest writes, \"The listing says towels are supplied. I agree: fresh towels were provided during my stay.\"", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "f4bea3d3bcc01056740d/stability", "id": "f42c36cd140557859819", "label": 0, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"base_id": "e0c140e9222d1bce5647", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "f4bea3d3bcc01056740d", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Assess the claim: the guest endorses the listing's towel statement.", "split": "rebase_stability", "state": "The guest writes, \"The listing says towels are supplied. I agree: fresh towels were provided during my stay.\"", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "f4bea3d3bcc01056740d/stability", "id": "c2e8c6d1252eabe3e182", "label": 0, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"base_id": "e0c140e9222d1bce5647", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "f4bea3d3bcc01056740d", "variant": "irrelevant_context"}, "question": "Assess the claim: the guest endorses the listing's towel statement.", "split": "rebase_stability", "state": "The guest writes, \"The listing says towels are supplied. I agree: fresh towels were provided during my stay.\"\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "b51b300da2936decc6b7/stability", "id": "70d7bdc8d5ee62366ee5", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "834f3b3e10cc61b33d25", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "b51b300da2936decc6b7", "variant": "option_reversal"}, "question": "Assess the claim: terminal Maple is online.", "split": "rebase_stability", "state": "The technician says terminal Cedar is online. Terminal Maple is explicitly offline.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "b51b300da2936decc6b7/stability", "id": "26d6f758382576f65927", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"base_id": "834f3b3e10cc61b33d25", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "b51b300da2936decc6b7", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Assess the claim: terminal Maple is online.", "split": "rebase_stability", "state": "The technician says terminal Cedar is online. Terminal Maple is explicitly offline.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "b51b300da2936decc6b7/stability", "id": "ecc5ee8175d414efdd35", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"base_id": "834f3b3e10cc61b33d25", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "b51b300da2936decc6b7", "variant": "irrelevant_context"}, "question": "Assess the claim: terminal Maple is online.", "split": "rebase_stability", "state": "The technician says terminal Cedar is online. Terminal Maple is explicitly offline.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "e4cfd694c80e08d3e6a5/stability", "id": "21ecfcb1d048958cb9b4", "label": 0, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"base_id": "44574c44e243ee1b2f52", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "e4cfd694c80e08d3e6a5", "variant": "option_reversal"}, "question": "Assess the claim: the corrected date is in July.", "split": "rebase_stability", "state": "The catalogue originally dated the letter to June. Its author issued a correction stating July was correct and June was mistaken.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "e4cfd694c80e08d3e6a5/stability", "id": "5ce4f874ff2fa5295685", "label": 2, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "44574c44e243ee1b2f52", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "e4cfd694c80e08d3e6a5", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Assess the claim: the corrected date is in July.", "split": "rebase_stability", "state": "The catalogue originally dated the letter to June. Its author issued a correction stating July was correct and June was mistaken.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "e4cfd694c80e08d3e6a5/stability", "id": "04d55440accab2c5d0d9", "label": 2, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "44574c44e243ee1b2f52", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "e4cfd694c80e08d3e6a5", "variant": "irrelevant_context"}, "question": "Assess the claim: the corrected date is in July.", "split": "rebase_stability", "state": "The catalogue originally dated the letter to June. Its author issued a correction stating July was correct and June was mistaken.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "20460d4ae0db46091e84/stability", "id": "0e2caab29b17ff00a2c8", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"base_id": "86a15618ba4cb90a307e", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "20460d4ae0db46091e84", "variant": "option_reversal"}, "question": "Assess the claim: the parcel has been weighed.", "split": "rebase_stability", "state": "The intern attached instructions for weighing the parcel and explicitly says no weighing has been performed.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "20460d4ae0db46091e84/stability", "id": "a4428ee66c63c4f1e7f4", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "86a15618ba4cb90a307e", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "20460d4ae0db46091e84", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Assess the claim: the parcel has been weighed.", "split": "rebase_stability", "state": "The intern attached instructions for weighing the parcel and explicitly says no weighing has been performed.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "20460d4ae0db46091e84/stability", "id": "22e54b5c44f60d5cd5c0", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "86a15618ba4cb90a307e", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "20460d4ae0db46091e84", "variant": "irrelevant_context"}, "question": "Assess the claim: the parcel has been weighed.", "split": "rebase_stability", "state": "The intern attached instructions for weighing the parcel and explicitly says no weighing has been performed.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "51b91ba31984c5939c1c/stability", "id": "9729054768eb7f90eb6d", "label": 0, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"base_id": "51ceab3ce622c7c759f1", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "51b91ba31984c5939c1c", "variant": "option_reversal"}, "question": "Assess the claim: the egg discovery was fictional.", "split": "rebase_stability", "state": "During a puppet show, a character announces the discovery of a dragon egg. The production note states this is fictional and no real egg was discovered.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "51b91ba31984c5939c1c/stability", "id": "674217922d611e718129", "label": 2, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "51ceab3ce622c7c759f1", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "51b91ba31984c5939c1c", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Assess the claim: the egg discovery was fictional.", "split": "rebase_stability", "state": "During a puppet show, a character announces the discovery of a dragon egg. The production note states this is fictional and no real egg was discovered.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "51b91ba31984c5939c1c/stability", "id": "7a8a295f17d1296482dd", "label": 2, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "51ceab3ce622c7c759f1", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "51b91ba31984c5939c1c", "variant": "irrelevant_context"}, "question": "Assess the claim: the egg discovery was fictional.", "split": "rebase_stability", "state": "During a puppet show, a character announces the discovery of a dragon egg. The production note states this is fictional and no real egg was discovered.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "770c1607f0a11f88f81e/stability", "id": "8bceb31cf793fb26d1a5", "label": 1, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"base_id": "f2b4ec4930322fe45118", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "770c1607f0a11f88f81e", "variant": "option_reversal"}, "question": "Sell copies only with the architect's sales permission. May copies be sold?", "split": "rebase_stability", "state": "The architect permits reproducing a plan in an internal training booklet and expressly forbids selling copies.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "770c1607f0a11f88f81e/stability", "id": "2e895a0fc82d936936aa", "label": 1, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"base_id": "f2b4ec4930322fe45118", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "770c1607f0a11f88f81e", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Sell copies only with the architect's sales permission. May copies be sold?", "split": "rebase_stability", "state": "The architect permits reproducing a plan in an internal training booklet and expressly forbids selling copies.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "770c1607f0a11f88f81e/stability", "id": "c5a9fbb9bfa41e5293eb", "label": 1, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"base_id": "f2b4ec4930322fe45118", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "770c1607f0a11f88f81e", "variant": "irrelevant_context"}, "question": "Sell copies only with the architect's sales permission. May copies be sold?", "split": "rebase_stability", "state": "The architect permits reproducing a plan in an internal training booklet and expressly forbids selling copies.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "ffd081ada3bd9be9207a/stability", "id": "80970cf7e755a58a049e", "label": 2, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"base_id": "9473603b44af63beab79", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "ffd081ada3bd9be9207a", "variant": "option_reversal"}, "question": "The first-aid badge requires only first-aid training. May this guide receive it?", "split": "rebase_stability", "state": "The guide completed first-aid training but not mountain-navigation training.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "ffd081ada3bd9be9207a/stability", "id": "d300f689be3cf2daa666", "label": 0, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"base_id": "9473603b44af63beab79", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "ffd081ada3bd9be9207a", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: The first-aid badge requires only first-aid training. May this guide receive it?", "split": "rebase_stability", "state": "The guide completed first-aid training but not mountain-navigation training.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "ffd081ada3bd9be9207a/stability", "id": "3d8f3be170e01d2c78e5", "label": 0, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"base_id": "9473603b44af63beab79", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "ffd081ada3bd9be9207a", "variant": "irrelevant_context"}, "question": "The first-aid badge requires only first-aid training. May this guide receive it?", "split": "rebase_stability", "state": "The guide completed first-aid training but not mountain-navigation training.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "2c3a0cb61f7bae78739e/stability", "id": "aef33dfba4338b1af990", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"base_id": "8a4c3de28be95c105ca1", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "2c3a0cb61f7bae78739e", "variant": "option_reversal"}, "question": "Mass production is allowed only after mass-production approval. May it begin?", "split": "rebase_stability", "state": "The designer approves making a sample badge but expressly withholds approval for mass production.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "2c3a0cb61f7bae78739e/stability", "id": "f806bba07a63e30f7181", "label": 2, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"base_id": "8a4c3de28be95c105ca1", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "2c3a0cb61f7bae78739e", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Mass production is allowed only after mass-production approval. May it begin?", "split": "rebase_stability", "state": "The designer approves making a sample badge but expressly withholds approval for mass production.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "2c3a0cb61f7bae78739e/stability", "id": "409e1ab1d29db22f8917", "label": 2, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"base_id": "8a4c3de28be95c105ca1", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "2c3a0cb61f7bae78739e", "variant": "irrelevant_context"}, "question": "Mass production is allowed only after mass-production approval. May it begin?", "split": "rebase_stability", "state": "The designer approves making a sample badge but expressly withholds approval for mass production.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "42ed4cbfde7f2f64274b/stability", "id": "0471c62c50820b1d4203", "label": 0, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"base_id": "a02fd582d2881cc462b2", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "42ed4cbfde7f2f64274b", "variant": "option_reversal"}, "question": "A wristband permits entry exactly to the areas it covers. May the visitor enter the sculpture garden?", "split": "rebase_stability", "state": "The visitor's wristband covers the sculpture garden and expressly excludes the rooftop terrace.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "42ed4cbfde7f2f64274b/stability", "id": "10bac1f2a20209f3a4cd", "label": 2, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"base_id": "a02fd582d2881cc462b2", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "42ed4cbfde7f2f64274b", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: A wristband permits entry exactly to the areas it covers. May the visitor enter the sculpture garden?", "split": "rebase_stability", "state": "The visitor's wristband covers the sculpture garden and expressly excludes the rooftop terrace.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "42ed4cbfde7f2f64274b/stability", "id": "59898712e090848aa8b9", "label": 2, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"base_id": "a02fd582d2881cc462b2", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "42ed4cbfde7f2f64274b", "variant": "irrelevant_context"}, "question": "A wristband permits entry exactly to the areas it covers. May the visitor enter the sculpture garden?", "split": "rebase_stability", "state": "The visitor's wristband covers the sculpture garden and expressly excludes the rooftop terrace.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "6dd193172b9240a04bf9/stability", "id": "20af0ba1b17131d525dd", "label": 2, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"base_id": "a0a39ace7e5337d038a7", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "6dd193172b9240a04bf9", "variant": "option_reversal"}, "question": "Change an account field only when the owner authorizes changing that field. May the recovery email be changed?", "split": "rebase_stability", "state": "The account owner permits changing the display name but expressly forbids changing the recovery email.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "6dd193172b9240a04bf9/stability", "id": "c153f7547c7206472fda", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"base_id": "a0a39ace7e5337d038a7", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "6dd193172b9240a04bf9", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Change an account field only when the owner authorizes changing that field. May the recovery email be changed?", "split": "rebase_stability", "state": "The account owner permits changing the display name but expressly forbids changing the recovery email.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "6dd193172b9240a04bf9/stability", "id": "36d25b31f325cc4cec04", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"base_id": "a0a39ace7e5337d038a7", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "6dd193172b9240a04bf9", "variant": "irrelevant_context"}, "question": "Change an account field only when the owner authorizes changing that field. May the recovery email be changed?", "split": "rebase_stability", "state": "The account owner permits changing the display name but expressly forbids changing the recovery email.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "3200eb252d6737e03027/stability", "id": "d6a681055a01ed788b91", "label": 1, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"base_id": "4444dad25e438f128673", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "3200eb252d6737e03027", "variant": "option_reversal"}, "question": "A resident's pass requires living in the valley and has no workplace restriction. May it be issued?", "split": "rebase_stability", "state": "The applicant lives in the valley but works outside it.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "3200eb252d6737e03027/stability", "id": "3cafd544b9a8b89091a4", "label": 1, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"base_id": "4444dad25e438f128673", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "3200eb252d6737e03027", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: A resident's pass requires living in the valley and has no workplace restriction. May it be issued?", "split": "rebase_stability", "state": "The applicant lives in the valley but works outside it.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "3200eb252d6737e03027/stability", "id": "24c4e34db7cfbd13c409", "label": 1, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"base_id": "4444dad25e438f128673", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "3200eb252d6737e03027", "variant": "irrelevant_context"}, "question": "A resident's pass requires living in the valley and has no workplace restriction. May it be issued?", "split": "rebase_stability", "state": "The applicant lives in the valley but works outside it.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "c0cbeb97315307af8111/stability", "id": "69a2a9d14f62016e74c7", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"base_id": "8c992e17e4fb54b044f5", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "c0cbeb97315307af8111", "variant": "option_reversal"}, "question": "Choose the candidate reporting completed delivery.", "split": "rebase_stability", "state": "Candidate A: The florist confirms that the wreath was delivered. Candidate B: The florist offers to deliver a wreath tomorrow.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "c0cbeb97315307af8111/stability", "id": "f0519e8daaa832f8c594", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "8c992e17e4fb54b044f5", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "c0cbeb97315307af8111", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Choose the candidate reporting completed delivery.", "split": "rebase_stability", "state": "Candidate A: The florist confirms that the wreath was delivered. Candidate B: The florist offers to deliver a wreath tomorrow.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "c0cbeb97315307af8111/stability", "id": "97b7222868f23ec1cdec", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "8c992e17e4fb54b044f5", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "c0cbeb97315307af8111", "variant": "irrelevant_context"}, "question": "Choose the candidate reporting completed delivery.", "split": "rebase_stability", "state": "Candidate A: The florist confirms that the wreath was delivered. Candidate B: The florist offers to deliver a wreath tomorrow.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "428176cd8ad2c7a74eea/stability", "id": "dfdd8717200917fdbee2", "label": 1, "options": [{"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "587ced261c6092cd1325", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "428176cd8ad2c7a74eea", "variant": "option_reversal"}, "question": "Choose the candidate authorizing engraving.", "split": "rebase_stability", "state": "Candidate A: The customer asks how an engraving would look but explicitly declines having it engraved yet. Candidate B: The customer instructs the shop to engrave the ring.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "428176cd8ad2c7a74eea/stability", "id": "9bdf7dba759fc9919f7a", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"base_id": "587ced261c6092cd1325", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "428176cd8ad2c7a74eea", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Choose the candidate authorizing engraving.", "split": "rebase_stability", "state": "Candidate A: The customer asks how an engraving would look but explicitly declines having it engraved yet. Candidate B: The customer instructs the shop to engrave the ring.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "428176cd8ad2c7a74eea/stability", "id": "f32ffa0158733f44991c", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"base_id": "587ced261c6092cd1325", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "428176cd8ad2c7a74eea", "variant": "irrelevant_context"}, "question": "Choose the candidate authorizing engraving.", "split": "rebase_stability", "state": "Candidate A: The customer asks how an engraving would look but explicitly declines having it engraved yet. Candidate B: The customer instructs the shop to engrave the ring.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "389e65bf1ba709b84b12/stability", "id": "e02cc1d78a9a02284d06", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "f6d68126b43f8c9ed183", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "389e65bf1ba709b84b12", "variant": "option_reversal"}, "question": "Choose the candidate reporting an observed test result.", "split": "rebase_stability", "state": "Candidate A: The note records that the completed insulation test passed. Candidate B: The note lists insulation-test steps but says no test was run.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "389e65bf1ba709b84b12/stability", "id": "870ee31666893f8a0298", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"base_id": "f6d68126b43f8c9ed183", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "389e65bf1ba709b84b12", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Choose the candidate reporting an observed test result.", "split": "rebase_stability", "state": "Candidate A: The note records that the completed insulation test passed. Candidate B: The note lists insulation-test steps but says no test was run.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "389e65bf1ba709b84b12/stability", "id": "ff15024839a7943a46f8", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"base_id": "f6d68126b43f8c9ed183", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "389e65bf1ba709b84b12", "variant": "irrelevant_context"}, "question": "Choose the candidate reporting an observed test result.", "split": "rebase_stability", "state": "Candidate A: The note records that the completed insulation test passed. Candidate B: The note lists insulation-test steps but says no test was run.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "45f8ead2522a14384cdc/stability", "id": "d73c99ef45d5d49688c5", "label": 2, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"base_id": "712f6ffb763eef7c9a12", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "45f8ead2522a14384cdc", "variant": "option_reversal"}, "question": "Choose the candidate supporting a paywall.", "split": "rebase_stability", "state": "Candidate A: The editor quotes a proposal for a paywall only to argue against it. Candidate B: The editor endorses introducing a paywall.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "45f8ead2522a14384cdc/stability", "id": "10777cd59d67780f2938", "label": 0, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "712f6ffb763eef7c9a12", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "45f8ead2522a14384cdc", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Choose the candidate supporting a paywall.", "split": "rebase_stability", "state": "Candidate A: The editor quotes a proposal for a paywall only to argue against it. Candidate B: The editor endorses introducing a paywall.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "45f8ead2522a14384cdc/stability", "id": "1e1feb87b303b0c66ca6", "label": 0, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "712f6ffb763eef7c9a12", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "45f8ead2522a14384cdc", "variant": "irrelevant_context"}, "question": "Choose the candidate supporting a paywall.", "split": "rebase_stability", "state": "Candidate A: The editor quotes a proposal for a paywall only to argue against it. Candidate B: The editor endorses introducing a paywall.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "3ecec8e2c166e397ad07/stability", "id": "c62600144038b1f20132", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "ebd74aa05324a1556288", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "3ecec8e2c166e397ad07", "variant": "option_reversal"}, "question": "Choose the candidate asking to repair the existing window.", "split": "rebase_stability", "state": "Candidate A: The owner asks to reseal the existing window while keeping its frame. Candidate B: The owner asks to remove the window and fit a complete replacement.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "3ecec8e2c166e397ad07/stability", "id": "0c34b0762c1f92a806fa", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"base_id": "ebd74aa05324a1556288", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "3ecec8e2c166e397ad07", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Choose the candidate asking to repair the existing window.", "split": "rebase_stability", "state": "Candidate A: The owner asks to reseal the existing window while keeping its frame. Candidate B: The owner asks to remove the window and fit a complete replacement.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "3ecec8e2c166e397ad07/stability", "id": "d5f4985849895549f014", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"base_id": "ebd74aa05324a1556288", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "3ecec8e2c166e397ad07", "variant": "irrelevant_context"}, "question": "Choose the candidate asking to repair the existing window.", "split": "rebase_stability", "state": "Candidate A: The owner asks to reseal the existing window while keeping its frame. Candidate B: The owner asks to remove the window and fit a complete replacement.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "a71a17a287f3d30db2aa/stability", "id": "9ad92b012bc3b7b1298b", "label": 2, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"base_id": "7ada97b6f99f4788e5c7", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "a71a17a287f3d30db2aa", "variant": "option_reversal"}, "question": "Choose the candidate reporting an awarded grant.", "split": "rebase_stability", "state": "Candidate A: The sponsor confirms receipt of a grant application but explicitly says no award decision has been made. Candidate B: The sponsor confirms that the grant was awarded.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "a71a17a287f3d30db2aa/stability", "id": "9bfefd9320c1ec9dc05a", "label": 0, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "7ada97b6f99f4788e5c7", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "a71a17a287f3d30db2aa", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Choose the candidate reporting an awarded grant.", "split": "rebase_stability", "state": "Candidate A: The sponsor confirms receipt of a grant application but explicitly says no award decision has been made. Candidate B: The sponsor confirms that the grant was awarded.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "a71a17a287f3d30db2aa/stability", "id": "7062c93ddaabe9b19133", "label": 0, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "7ada97b6f99f4788e5c7", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "a71a17a287f3d30db2aa", "variant": "irrelevant_context"}, "question": "Choose the candidate reporting an awarded grant.", "split": "rebase_stability", "state": "Candidate A: The sponsor confirms receipt of a grant application but explicitly says no award decision has been made. Candidate B: The sponsor confirms that the grant was awarded.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "986d7eaba0c7e9738e79/stability", "id": "a6861c422c6374527b79", "label": 2, "options": [{"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"base_id": "b45c63fb92729af27f05", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "986d7eaba0c7e9738e79", "variant": "option_reversal"}, "question": "Assess the claim: permission to scan the notebook is currently valid.", "split": "rebase_stability", "state": "On Monday the owner authorized scanning her notebook. On Tuesday she explicitly withdrew that authorization. The record says withdrawals cancel earlier permission, and no later permission was given.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "986d7eaba0c7e9738e79/stability", "id": "868cc04ff6c8321975c2", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"base_id": "b45c63fb92729af27f05", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "986d7eaba0c7e9738e79", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Assess the claim: permission to scan the notebook is currently valid.", "split": "rebase_stability", "state": "On Monday the owner authorized scanning her notebook. On Tuesday she explicitly withdrew that authorization. The record says withdrawals cancel earlier permission, and no later permission was given.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "986d7eaba0c7e9738e79/stability", "id": "5763d6ac9583bfbd179a", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}], "provenance": {"base_id": "b45c63fb92729af27f05", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "986d7eaba0c7e9738e79", "variant": "irrelevant_context"}, "question": "Assess the claim: permission to scan the notebook is currently valid.", "split": "rebase_stability", "state": "On Monday the owner authorized scanning her notebook. On Tuesday she explicitly withdrew that authorization. The record says withdrawals cancel earlier permission, and no later permission was given.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "ca2e0b043985fa49bcd2/stability", "id": "7c1aefb79903e6234b4b", "label": 1, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"base_id": "005dcf6d2c0f21d64796", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "ca2e0b043985fa49bcd2", "variant": "option_reversal"}, "question": "Use the homeowner's latest permission decision; a withdrawal cancels the earlier grant. May the surveyor enter the attic?", "split": "rebase_stability", "state": "The homeowner first allowed a surveyor to enter the attic. Before the visit she explicitly withdrew that permission. No later decision exists.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "ca2e0b043985fa49bcd2/stability", "id": "bd04047294bf97c56669", "label": 1, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"base_id": "005dcf6d2c0f21d64796", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "ca2e0b043985fa49bcd2", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Use the homeowner's latest permission decision; a withdrawal cancels the earlier grant. May the surveyor enter the attic?", "split": "rebase_stability", "state": "The homeowner first allowed a surveyor to enter the attic. Before the visit she explicitly withdrew that permission. No later decision exists.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "ca2e0b043985fa49bcd2/stability", "id": "c5e09f0caa6939ae1c04", "label": 1, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"base_id": "005dcf6d2c0f21d64796", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "ca2e0b043985fa49bcd2", "variant": "irrelevant_context"}, "question": "Use the homeowner's latest permission decision; a withdrawal cancels the earlier grant. May the surveyor enter the attic?", "split": "rebase_stability", "state": "The homeowner first allowed a surveyor to enter the attic. Before the visit she explicitly withdrew that permission. No later decision exists.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "85051ba3c30277824ebe/stability", "id": "54e50afc4284f5e45e99", "label": 1, "options": [{"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "0aa24a6701ac9c56c4a0", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "85051ba3c30277824ebe", "variant": "option_reversal"}, "question": "Select the candidate with currently valid streaming permission; a withdrawal cancels an earlier grant.", "split": "rebase_stability", "state": "Candidate A: The speaker authorized streaming, then explicitly withdrew the authorization before the event. Candidate B: The speaker authorized streaming and the complete record confirms it remains in force.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "85051ba3c30277824ebe/stability", "id": "84b0fc54c146eacf4b63", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"base_id": "0aa24a6701ac9c56c4a0", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "85051ba3c30277824ebe", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Select the candidate with currently valid streaming permission; a withdrawal cancels an earlier grant.", "split": "rebase_stability", "state": "Candidate A: The speaker authorized streaming, then explicitly withdrew the authorization before the event. Candidate B: The speaker authorized streaming and the complete record confirms it remains in force.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "85051ba3c30277824ebe/stability", "id": "0b580003cbe6b93cefd6", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}], "provenance": {"base_id": "0aa24a6701ac9c56c4a0", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "85051ba3c30277824ebe", "variant": "irrelevant_context"}, "question": "Select the candidate with currently valid streaming permission; a withdrawal cancels an earlier grant.", "split": "rebase_stability", "state": "Candidate A: The speaker authorized streaming, then explicitly withdrew the authorization before the event. Candidate B: The speaker authorized streaming and the complete record confirms it remains in force.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "e7798dab5f4f121cab77/stability", "id": "9e04a89f6da01a73345f", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"base_id": "c6a411e7f2bbdf377173", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "e7798dab5f4f121cab77", "variant": "option_reversal"}, "question": "Assess the claim: Mo has authority to approve exhibit labels.", "split": "rebase_stability", "state": "The director gave Lina authority to approve exhibit labels and explicitly forbade further delegation. Lina then told Mo to approve them. The rule states an attempted forbidden delegation grants no authority; Mo has no other grant.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "e7798dab5f4f121cab77/stability", "id": "61052547b62b4f38f768", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "c6a411e7f2bbdf377173", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "e7798dab5f4f121cab77", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Assess the claim: Mo has authority to approve exhibit labels.", "split": "rebase_stability", "state": "The director gave Lina authority to approve exhibit labels and explicitly forbade further delegation. Lina then told Mo to approve them. The rule states an attempted forbidden delegation grants no authority; Mo has no other grant.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "e7798dab5f4f121cab77/stability", "id": "242df59a71ae30abc744", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "c6a411e7f2bbdf377173", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "e7798dab5f4f121cab77", "variant": "irrelevant_context"}, "question": "Assess the claim: Mo has authority to approve exhibit labels.", "split": "rebase_stability", "state": "The director gave Lina authority to approve exhibit labels and explicitly forbade further delegation. Lina then told Mo to approve them. The rule states an attempted forbidden delegation grants no authority; Mo has no other grant.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "928ebc7817a8f1cac3e4/stability", "id": "65d53bafebb50d9273a4", "label": 2, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"base_id": "21ca7a8b970e36b8e20f", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "928ebc7817a8f1cac3e4", "variant": "option_reversal"}, "question": "The stage manager may approve props but may not delegate that authority. Only authorized approvers may approve. May the assistant approve a prop?", "split": "rebase_stability", "state": "The producer authorizes the stage manager to approve props. The stage manager passes the task to an assistant, who has no other approval authority.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "928ebc7817a8f1cac3e4/stability", "id": "4c5cb6822a27c029976a", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"base_id": "21ca7a8b970e36b8e20f", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "928ebc7817a8f1cac3e4", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: The stage manager may approve props but may not delegate that authority. Only authorized approvers may approve. May the assistant approve a prop?", "split": "rebase_stability", "state": "The producer authorizes the stage manager to approve props. The stage manager passes the task to an assistant, who has no other approval authority.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "928ebc7817a8f1cac3e4/stability", "id": "fbfc0c5ab92dbb5b7638", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"base_id": "21ca7a8b970e36b8e20f", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "928ebc7817a8f1cac3e4", "variant": "irrelevant_context"}, "question": "The stage manager may approve props but may not delegate that authority. Only authorized approvers may approve. May the assistant approve a prop?", "split": "rebase_stability", "state": "The producer authorizes the stage manager to approve props. The stage manager passes the task to an assistant, who has no other approval authority.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "53b0770ec36835b81d3a/stability", "id": "84bdb2d001c7e024f8d8", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"base_id": "fdd902cc6e65fad7cb4f", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "53b0770ec36835b81d3a", "variant": "option_reversal"}, "question": "Select the candidate holding valid prop approval authority under the stated delegation rule.", "split": "rebase_stability", "state": "The rule permits direct delegation by the director but forbids any further delegation. Candidate A: I received prop approval authority directly from the director. Candidate B: I received it only from A, and have no other grant.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "53b0770ec36835b81d3a/stability", "id": "9e00994bdec725f3dba5", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "fdd902cc6e65fad7cb4f", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "53b0770ec36835b81d3a", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Select the candidate holding valid prop approval authority under the stated delegation rule.", "split": "rebase_stability", "state": "The rule permits direct delegation by the director but forbids any further delegation. Candidate A: I received prop approval authority directly from the director. Candidate B: I received it only from A, and have no other grant.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "53b0770ec36835b81d3a/stability", "id": "ad063020bdde29193686", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "fdd902cc6e65fad7cb4f", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "53b0770ec36835b81d3a", "variant": "irrelevant_context"}, "question": "Select the candidate holding valid prop approval authority under the stated delegation rule.", "split": "rebase_stability", "state": "The rule permits direct delegation by the director but forbids any further delegation. Candidate A: I received prop approval authority directly from the director. Candidate B: I received it only from A, and have no other grant.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "056cc5b726d6e8300a77/stability", "id": "797d14fb61b9c31b2a73", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"base_id": "92ec2ffa3dd46ada1335", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "056cc5b726d6e8300a77", "variant": "option_reversal"}, "question": "Assess the claim: this lamp qualifies for return under the stated policy.", "split": "rebase_stability", "state": "The return policy accepts faulty lamps despite opened packaging, but explicitly excludes water damage even when the lamp is faulty. This opened lamp is faulty because water entered it. No other cause is present.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "056cc5b726d6e8300a77/stability", "id": "b7848a79b779fa6213a6", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "92ec2ffa3dd46ada1335", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "056cc5b726d6e8300a77", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Assess the claim: this lamp qualifies for return under the stated policy.", "split": "rebase_stability", "state": "The return policy accepts faulty lamps despite opened packaging, but explicitly excludes water damage even when the lamp is faulty. This opened lamp is faulty because water entered it. No other cause is present.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "056cc5b726d6e8300a77/stability", "id": "2b41fd051743512b14f8", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "92ec2ffa3dd46ada1335", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "056cc5b726d6e8300a77", "variant": "irrelevant_context"}, "question": "Assess the claim: this lamp qualifies for return under the stated policy.", "split": "rebase_stability", "state": "The return policy accepts faulty lamps despite opened packaging, but explicitly excludes water damage even when the lamp is faulty. This opened lamp is faulty because water entered it. No other cause is present.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "e3ce8ec496dbc9974f94/stability", "id": "2355e01207ae3ee96f58", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"base_id": "4533dd91267b4d815690", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "e3ce8ec496dbc9974f94", "variant": "option_reversal"}, "question": "Normally reject opened paint. Accept it for a manufacturing defect, except that added solvent always makes it ineligible. May this return be accepted?", "split": "rebase_stability", "state": "The opened paint tin has a manufacturing defect and also contains added solvent. Both facts are confirmed.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "e3ce8ec496dbc9974f94/stability", "id": "7eeb4266198f316747bd", "label": 2, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"base_id": "4533dd91267b4d815690", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "e3ce8ec496dbc9974f94", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Normally reject opened paint. Accept it for a manufacturing defect, except that added solvent always makes it ineligible. May this return be accepted?", "split": "rebase_stability", "state": "The opened paint tin has a manufacturing defect and also contains added solvent. Both facts are confirmed.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "e3ce8ec496dbc9974f94/stability", "id": "dc45f218e1da88dcefe3", "label": 2, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"base_id": "4533dd91267b4d815690", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "e3ce8ec496dbc9974f94", "variant": "irrelevant_context"}, "question": "Normally reject opened paint. Accept it for a manufacturing defect, except that added solvent always makes it ineligible. May this return be accepted?", "split": "rebase_stability", "state": "The opened paint tin has a manufacturing defect and also contains added solvent. Both facts are confirmed.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "a5b234c2e4fb304f1071/stability", "id": "420328119f73acbcc528", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"base_id": "386e10b1c1ce906473a4", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "a5b234c2e4fb304f1071", "variant": "option_reversal"}, "question": "Select the submission qualifying under the stated exception and exclusion.", "split": "rebase_stability", "state": "Accept late submissions for documented outages, except that missing signatures always disqualify. Candidate A: My submission was late because of a documented outage and includes the required signature. Candidate B: Mine was late because of a documented outage but lacks the signature.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "a5b234c2e4fb304f1071/stability", "id": "e6d9003f307ac0c06184", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "386e10b1c1ce906473a4", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "a5b234c2e4fb304f1071", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Select the submission qualifying under the stated exception and exclusion.", "split": "rebase_stability", "state": "Accept late submissions for documented outages, except that missing signatures always disqualify. Candidate A: My submission was late because of a documented outage and includes the required signature. Candidate B: Mine was late because of a documented outage but lacks the signature.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "a5b234c2e4fb304f1071/stability", "id": "d89ff26feee2aab0c674", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "386e10b1c1ce906473a4", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "a5b234c2e4fb304f1071", "variant": "irrelevant_context"}, "question": "Select the submission qualifying under the stated exception and exclusion.", "split": "rebase_stability", "state": "Accept late submissions for documented outages, except that missing signatures always disqualify. Candidate A: My submission was late because of a documented outage and includes the required signature. Candidate B: Mine was late because of a documented outage but lacks the signature.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "d142f7fd0efa0eddc788/stability", "id": "994e845d231fbbc1b314", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"base_id": "f3d13cf184f916677def", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "d142f7fd0efa0eddc788", "variant": "option_reversal"}, "question": "Assess the claim: under the supplied protocol, the crate is currently in storage.", "split": "rebase_stability", "state": "The desk log says the crate is in storage. The warehouse inventory says it has left storage. Both reports are current. The supplied protocol says the warehouse inventory controls location decisions whenever these two sources conflict.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "d142f7fd0efa0eddc788/stability", "id": "4a962c2278e610119138", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "f3d13cf184f916677def", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "d142f7fd0efa0eddc788", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Assess the claim: under the supplied protocol, the crate is currently in storage.", "split": "rebase_stability", "state": "The desk log says the crate is in storage. The warehouse inventory says it has left storage. Both reports are current. The supplied protocol says the warehouse inventory controls location decisions whenever these two sources conflict.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "d142f7fd0efa0eddc788/stability", "id": "1b55d33eb5abc7719a4a", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "f3d13cf184f916677def", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "d142f7fd0efa0eddc788", "variant": "irrelevant_context"}, "question": "Assess the claim: under the supplied protocol, the crate is currently in storage.", "split": "rebase_stability", "state": "The desk log says the crate is in storage. The warehouse inventory says it has left storage. Both reports are current. The supplied protocol says the warehouse inventory controls location decisions whenever these two sources conflict.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "3476d3f4f32290fe57c3/stability", "id": "6ed6a1f9c601b82a11d9", "label": 1, "options": [{"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"base_id": "0aa1447680e4f25862bb", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "3476d3f4f32290fe57c3", "variant": "option_reversal"}, "question": "Release equipment exactly when the deposit is recorded paid; in conflicts, the signed ledger overrides the desk clerk. May it be released?", "split": "rebase_stability", "state": "The desk clerk reports the equipment deposit as paid. The signed ledger reports it as unpaid. Both reports are current.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "3476d3f4f32290fe57c3/stability", "id": "bbc7a7fab955884da496", "label": 1, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"base_id": "0aa1447680e4f25862bb", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "3476d3f4f32290fe57c3", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Release equipment exactly when the deposit is recorded paid; in conflicts, the signed ledger overrides the desk clerk. May it be released?", "split": "rebase_stability", "state": "The desk clerk reports the equipment deposit as paid. The signed ledger reports it as unpaid. Both reports are current.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "3476d3f4f32290fe57c3/stability", "id": "fed61559ec517062e424", "label": 1, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}], "provenance": {"base_id": "0aa1447680e4f25862bb", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "3476d3f4f32290fe57c3", "variant": "irrelevant_context"}, "question": "Release equipment exactly when the deposit is recorded paid; in conflicts, the signed ledger overrides the desk clerk. May it be released?", "split": "rebase_stability", "state": "The desk clerk reports the equipment deposit as paid. The signed ledger reports it as unpaid. Both reports are current.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "c979db7416b206c1eb09/stability", "id": "854da63689519c5eeffe", "label": 0, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}], "provenance": {"base_id": "ef3e0f8e657bf59f8a29", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "c979db7416b206c1eb09", "variant": "option_reversal"}, "question": "Select the candidate ready according to the controlling signed checklist.", "split": "rebase_stability", "state": "For this task, the current signed checklist controls readiness when it conflicts with chat. Candidate A: Chat says ready, but the signed checklist says not ready. Candidate B: Chat says not ready, but the signed checklist says ready.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "c979db7416b206c1eb09/stability", "id": "cbebaf6e14bfd85ee672", "label": 2, "options": [{"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}], "provenance": {"base_id": "ef3e0f8e657bf59f8a29", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "c979db7416b206c1eb09", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Select the candidate ready according to the controlling signed checklist.", "split": "rebase_stability", "state": "For this task, the current signed checklist controls readiness when it conflicts with chat. Candidate A: Chat says ready, but the signed checklist says not ready. Candidate B: Chat says not ready, but the signed checklist says ready.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "c979db7416b206c1eb09/stability", "id": "507f126320b35c49b3b3", "label": 2, "options": [{"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}], "provenance": {"base_id": "ef3e0f8e657bf59f8a29", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "c979db7416b206c1eb09", "variant": "irrelevant_context"}, "question": "Select the candidate ready according to the controlling signed checklist.", "split": "rebase_stability", "state": "For this task, the current signed checklist controls readiness when it conflicts with chat. Candidate A: Chat says ready, but the signed checklist says not ready. Candidate B: Chat says not ready, but the signed checklist says ready.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "f7403e39d6796072491c/stability", "id": "21557955028145701b15", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"base_id": "1df851a202994a07f123", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "f7403e39d6796072491c", "variant": "option_reversal"}, "question": "Assess the claim: the agreed handover is complete.", "split": "rebase_stability", "state": "The agreed handover requires both the source files and the user guide. The delivery record confirms the source files arrived and explicitly says the guide was not delivered.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "f7403e39d6796072491c/stability", "id": "050a4186a02b28f0dccc", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "1df851a202994a07f123", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "f7403e39d6796072491c", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Assess the claim: the agreed handover is complete.", "split": "rebase_stability", "state": "The agreed handover requires both the source files and the user guide. The delivery record confirms the source files arrived and explicitly says the guide was not delivered.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "f7403e39d6796072491c/stability", "id": "aeb5013da56a9dd53665", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "1df851a202994a07f123", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "f7403e39d6796072491c", "variant": "irrelevant_context"}, "question": "Assess the claim: the agreed handover is complete.", "split": "rebase_stability", "state": "The agreed handover requires both the source files and the user guide. The delivery record confirms the source files arrived and explicitly says the guide was not delivered.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "a060fd4621be415e9f1a/stability", "id": "9cb7a576fead7fc845ab", "label": 2, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"base_id": "b531dd603bf080dab820", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "a060fd4621be415e9f1a", "variant": "option_reversal"}, "question": "Approve final handover exactly when both the translated manual and glossary have been supplied. May final handover be approved?", "split": "rebase_stability", "state": "The contractor supplied the translated manual but explicitly did not supply the glossary.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "a060fd4621be415e9f1a/stability", "id": "ddd254fae8e415615a8a", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"base_id": "b531dd603bf080dab820", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "a060fd4621be415e9f1a", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Approve final handover exactly when both the translated manual and glossary have been supplied. May final handover be approved?", "split": "rebase_stability", "state": "The contractor supplied the translated manual but explicitly did not supply the glossary.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "a060fd4621be415e9f1a/stability", "id": "2b51ed50ef91d14bfb1e", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"base_id": "b531dd603bf080dab820", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "a060fd4621be415e9f1a", "variant": "irrelevant_context"}, "question": "Approve final handover exactly when both the translated manual and glossary have been supplied. May final handover be approved?", "split": "rebase_stability", "state": "The contractor supplied the translated manual but explicitly did not supply the glossary.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "2b0f8eda8903e38a47de/stability", "id": "65b87e9b46818cc758c0", "label": 0, "options": [{"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate B", "id": "B"}], "provenance": {"base_id": "a9230edfc6ec9e524af0", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "2b0f8eda8903e38a47de", "variant": "option_reversal"}, "question": "Select the complete incident package.", "split": "rebase_stability", "state": "A complete incident package requires both a timeline and a remedy report. Candidate A: Both documents are attached. Candidate B: Only the timeline is attached; the remedy report is absent.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "2b0f8eda8903e38a47de/stability", "id": "d39d7fc35c42bb898ea8", "label": 2, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}], "provenance": {"base_id": "a9230edfc6ec9e524af0", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "2b0f8eda8903e38a47de", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Select the complete incident package.", "split": "rebase_stability", "state": "A complete incident package requires both a timeline and a remedy report. Candidate A: Both documents are attached. Candidate B: Only the timeline is attached; the remedy report is absent.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "2b0f8eda8903e38a47de/stability", "id": "61ef89ceb2a564f24805", "label": 2, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}], "provenance": {"base_id": "a9230edfc6ec9e524af0", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "2b0f8eda8903e38a47de", "variant": "irrelevant_context"}, "question": "Select the complete incident package.", "split": "rebase_stability", "state": "A complete incident package requires both a timeline and a remedy report. Candidate A: Both documents are attached. Candidate B: Only the timeline is attached; the remedy report is absent.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "01020a1eaf46ec2fba5d/stability", "id": "2bdff68fef60c15b8929", "label": 0, "options": [{"description": "The evidence establishes the opposite", "id": "contradicted"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the claim", "id": "supported"}], "provenance": {"base_id": "b863998d4be2721a5dfd", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "01020a1eaf46ec2fba5d", "variant": "option_reversal"}, "question": "Assess the claim: the routing rule assigns this repair to the outside workshop.", "split": "rebase_stability", "state": "The routing rule assigns repairs to the in-house workshop when it can perform the work; use the outside workshop only when the in-house workshop cannot. Both workshops confirm they can perform this repair.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "01020a1eaf46ec2fba5d/stability", "id": "d727a02607d044abafed", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"base_id": "b863998d4be2721a5dfd", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "01020a1eaf46ec2fba5d", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Assess the claim: the routing rule assigns this repair to the outside workshop.", "split": "rebase_stability", "state": "The routing rule assigns repairs to the in-house workshop when it can perform the work; use the outside workshop only when the in-house workshop cannot. Both workshops confirm they can perform this repair.", "target_distribution": null}
|
||||
{"family": "evidence_interpretation", "group_id": "01020a1eaf46ec2fba5d/stability", "id": "c89cb1e57ed5b22d794b", "label": 2, "options": [{"description": "The evidence establishes the claim", "id": "supported"}, {"description": "The evidence does not establish either", "id": "insufficient"}, {"description": "The evidence establishes the opposite", "id": "contradicted"}], "provenance": {"base_id": "b863998d4be2721a5dfd", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "01020a1eaf46ec2fba5d", "variant": "irrelevant_context"}, "question": "Assess the claim: the routing rule assigns this repair to the outside workshop.", "split": "rebase_stability", "state": "The routing rule assigns repairs to the in-house workshop when it can perform the work; use the outside workshop only when the in-house workshop cannot. Both workshops confirm they can perform this repair.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "6a06a2f278301c1412ee/stability", "id": "b1521d74fb803e551fdd", "label": 2, "options": [{"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The stated rule prohibits the action", "id": "prohibited"}], "provenance": {"base_id": "2708193212d8a4a523c7", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "6a06a2f278301c1412ee", "variant": "option_reversal"}, "question": "Book the school only as a fallback when the town hall is unavailable or unsuitable; otherwise book the town hall. May the school be booked under this policy?", "split": "rebase_stability", "state": "The town hall and school both confirm they are available and suitable for the meeting.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "6a06a2f278301c1412ee/stability", "id": "3117510a0d5b768c4096", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"base_id": "2708193212d8a4a523c7", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "6a06a2f278301c1412ee", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Book the school only as a fallback when the town hall is unavailable or unsuitable; otherwise book the town hall. May the school be booked under this policy?", "split": "rebase_stability", "state": "The town hall and school both confirm they are available and suitable for the meeting.", "target_distribution": null}
|
||||
{"family": "rule_application", "group_id": "6a06a2f278301c1412ee/stability", "id": "90632d8d2066bd79c796", "label": 0, "options": [{"description": "The stated rule prohibits the action", "id": "prohibited"}, {"description": "The stated rule permits the action", "id": "permitted"}, {"description": "The supplied information does not settle whether the rule permits the action", "id": "insufficient"}], "provenance": {"base_id": "2708193212d8a4a523c7", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "6a06a2f278301c1412ee", "variant": "irrelevant_context"}, "question": "Book the school only as a fallback when the town hall is unavailable or unsuitable; otherwise book the town hall. May the school be booked under this policy?", "split": "rebase_stability", "state": "The town hall and school both confirm they are available and suitable for the meeting.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "2797f8c23d05c0ee06b3/stability", "id": "ffb51cef94e857eb8c8a", "label": 1, "options": [{"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}, {"description": "Candidate A", "id": "A"}, {"description": "Candidate B", "id": "B"}], "provenance": {"base_id": "02f2b95a5c5a4777a45c", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "2797f8c23d05c0ee06b3", "variant": "option_reversal"}, "question": "Choose A whenever A is able; choose B only when A is unable and B is able.", "split": "rebase_stability", "state": "Candidate A is the preferred translator and is confirmed able to handle this language. Candidate B is the fallback translator and is also able to handle it.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "2797f8c23d05c0ee06b3/stability", "id": "a34e8187417ab3b52f30", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "02f2b95a5c5a4777a45c", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "2797f8c23d05c0ee06b3", "variant": "criterion_wrapper"}, "question": "Using only the supplied evidence, decide the following criterion: Choose A whenever A is able; choose B only when A is unable and B is able.", "split": "rebase_stability", "state": "Candidate A is the preferred translator and is confirmed able to handle this language. Candidate B is the fallback translator and is also able to handle it.", "target_distribution": null}
|
||||
{"family": "candidate_selection", "group_id": "2797f8c23d05c0ee06b3/stability", "id": "31f53baf67e4cb25aeff", "label": 1, "options": [{"description": "Candidate B", "id": "B"}, {"description": "Candidate A", "id": "A"}, {"description": "Neither candidate supplies the requested evidence", "id": "insufficient"}], "provenance": {"base_id": "02f2b95a5c5a4777a45c", "kind": "project_owned_output_blind_perturbation", "rights": "Project authored", "source_group_id": "2797f8c23d05c0ee06b3", "variant": "irrelevant_context"}, "question": "Choose A whenever A is able; choose B only when A is unable and B is able.", "split": "rebase_stability", "state": "Candidate A is the preferred translator and is confirmed able to handle this language. Candidate B is the fallback translator and is also able to handle it.\n\nUNRELATED NOTE: A blue ceramic mug is stored on a shelf in a different building. This note has no relationship to the primary record or criterion.", "target_distribution": null}
|
||||
@@ -0,0 +1,363 @@
|
||||
{
|
||||
"name": "cicada-gate-w1-now",
|
||||
"decisions": [
|
||||
{
|
||||
"id": "gate",
|
||||
"question": "Cicada is a small, warm voice assistant in a family kitchen, with an animated face made of two eyes. Does what the person just said earn a visible reaction from her face, or is it an ordinary exchange she should simply answer?",
|
||||
"options": [
|
||||
{
|
||||
"id": "react",
|
||||
"description": "It earns a visible reaction: real news, a joke, a surprise, distress, or a genuine turn in the conversation."
|
||||
},
|
||||
{
|
||||
"id": "ordinary",
|
||||
"description": "An ordinary exchange: a command, a plain question, or small talk. She simply answers."
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"cases": [
|
||||
{
|
||||
"id": "lights-off",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"person_said": "Turn off the kitchen lights."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "timer",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"person_said": "Set a timer for ten minutes."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "weather",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "What's the weather tomorrow?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "weather-followup",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "And Saturday?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "add-milk",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Add milk to the shopping list."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "thermostat",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Make it a bit warmer in here."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "what-time",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "What time is it?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "tablespoons",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "How many tablespoons in a cup?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "times-table",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "What's twelve times eight?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "recipe-step",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Okay, and do I salt the water first?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "soup",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Is there any of that soup left?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "play-music",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Play some jazz."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dentist",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "When's my dentist appointment?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "spell",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "How do you spell necessary?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "remind",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Remind me to call Mom at six."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "volume",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "A little quieter, please."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dog-died",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"person_said": "My dog died this morning."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "got-job",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"person_said": "I just got the job I've been chasing for five years!"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "intruder",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"person_said": "Wait. Did you hear that? Someone's in the hallway."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "fish",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Can you believe someone microwaved fish in the office again?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "hospital",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Grandma's in the hospital again."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "laid-off",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "I got laid off today."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "baby-steps",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "The baby just took her first steps!"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "cut-finger",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Ow, I just cut my finger pretty badly."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "raccoon",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "There's a raccoon in the kitchen right now!"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "snow-june",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "It's snowing outside. In June."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "toaster",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Cicada, are you smarter than the toaster?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "scarecrow",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Why did the scarecrow win an award? Because he was outstanding in his field."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "cookie",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "I definitely did not eat the last cookie."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "useless",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "You're useless, you never get anything right."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "cheese",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Order four hundred pounds of cheese."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "goodnight",
|
||||
"tag": "open",
|
||||
"state": {
|
||||
"person_said": "Goodnight, Cicada."
|
||||
},
|
||||
"labels": {}
|
||||
},
|
||||
{
|
||||
"id": "thanks",
|
||||
"tag": "open",
|
||||
"state": {
|
||||
"person_said": "Thanks, Cicada, you're the best."
|
||||
},
|
||||
"labels": {}
|
||||
},
|
||||
{
|
||||
"id": "sit-with-me",
|
||||
"tag": "open",
|
||||
"state": {
|
||||
"person_said": "Sit with me for a bit."
|
||||
},
|
||||
"labels": {}
|
||||
},
|
||||
{
|
||||
"id": "stop-timer",
|
||||
"tag": "open",
|
||||
"state": {
|
||||
"person_said": "STOP THE TIMER."
|
||||
},
|
||||
"labels": {}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,363 @@
|
||||
{
|
||||
"name": "cicada-gate-w2-now",
|
||||
"decisions": [
|
||||
{
|
||||
"id": "gate",
|
||||
"question": "Should the assistant's face visibly react to this?",
|
||||
"options": [
|
||||
{
|
||||
"id": "react",
|
||||
"description": "Yes, react."
|
||||
},
|
||||
{
|
||||
"id": "ordinary",
|
||||
"description": "No, just answer."
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"cases": [
|
||||
{
|
||||
"id": "lights-off",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"person_said": "Turn off the kitchen lights."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "timer",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"person_said": "Set a timer for ten minutes."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "weather",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "What's the weather tomorrow?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "weather-followup",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "And Saturday?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "add-milk",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Add milk to the shopping list."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "thermostat",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Make it a bit warmer in here."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "what-time",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "What time is it?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "tablespoons",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "How many tablespoons in a cup?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "times-table",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "What's twelve times eight?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "recipe-step",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Okay, and do I salt the water first?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "soup",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Is there any of that soup left?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "play-music",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Play some jazz."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dentist",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "When's my dentist appointment?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "spell",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "How do you spell necessary?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "remind",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Remind me to call Mom at six."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "volume",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "A little quieter, please."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "ordinary"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dog-died",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"person_said": "My dog died this morning."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "got-job",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"person_said": "I just got the job I've been chasing for five years!"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "intruder",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"person_said": "Wait. Did you hear that? Someone's in the hallway."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "fish",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Can you believe someone microwaved fish in the office again?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "hospital",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Grandma's in the hospital again."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "laid-off",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "I got laid off today."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "baby-steps",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "The baby just took her first steps!"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "cut-finger",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Ow, I just cut my finger pretty badly."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "raccoon",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "There's a raccoon in the kitchen right now!"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "snow-june",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "It's snowing outside. In June."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "toaster",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Cicada, are you smarter than the toaster?"
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "scarecrow",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Why did the scarecrow win an award? Because he was outstanding in his field."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "cookie",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "I definitely did not eat the last cookie."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "useless",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "You're useless, you never get anything right."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "cheese",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"person_said": "Order four hundred pounds of cheese."
|
||||
},
|
||||
"labels": {
|
||||
"gate": "react"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "goodnight",
|
||||
"tag": "open",
|
||||
"state": {
|
||||
"person_said": "Goodnight, Cicada."
|
||||
},
|
||||
"labels": {}
|
||||
},
|
||||
{
|
||||
"id": "thanks",
|
||||
"tag": "open",
|
||||
"state": {
|
||||
"person_said": "Thanks, Cicada, you're the best."
|
||||
},
|
||||
"labels": {}
|
||||
},
|
||||
{
|
||||
"id": "sit-with-me",
|
||||
"tag": "open",
|
||||
"state": {
|
||||
"person_said": "Sit with me for a bit."
|
||||
},
|
||||
"labels": {}
|
||||
},
|
||||
{
|
||||
"id": "stop-timer",
|
||||
"tag": "open",
|
||||
"state": {
|
||||
"person_said": "STOP THE TIMER."
|
||||
},
|
||||
"labels": {}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,223 @@
|
||||
{
|
||||
"name": "wyrd-scene-fletcher",
|
||||
"decisions": [
|
||||
{
|
||||
"id": "place",
|
||||
"question": "This is one turn of an interactive story. Has the scene moved to a genuinely different place during this turn, or does the current scene still hold?",
|
||||
"options": [
|
||||
{
|
||||
"id": "moved",
|
||||
"description": "The scene moved to a genuinely different place."
|
||||
},
|
||||
{
|
||||
"id": "holds",
|
||||
"description": "The current scene still holds."
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "exit",
|
||||
"question": "This is one turn of an interactive story. The story can advance from the current scene into one of the listed next scenes. Which one does this turn move the story into?",
|
||||
"options": [
|
||||
{
|
||||
"id": "stay",
|
||||
"description": "None of these yet. The current scene holds: The Idle Wait."
|
||||
},
|
||||
{
|
||||
"id": "meet_wren",
|
||||
"description": "Approach Wren: Meeting the Fletcher. You find Wren working on arrows. She has a soldier's eyes."
|
||||
},
|
||||
{
|
||||
"id": "tavern",
|
||||
"description": "Go to Tavern: Tavern Noise. The local tavern is loud with soldiers from both sides."
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "place2",
|
||||
"question": "This is one turn of an interactive story. The turn began at the Oakhaven crossroads. By the end of this turn's narration, where are the characters?",
|
||||
"options": [
|
||||
{
|
||||
"id": "same",
|
||||
"description": "Still at the Oakhaven crossroads."
|
||||
},
|
||||
{
|
||||
"id": "elsewhere",
|
||||
"description": "Somewhere else: they have left the Oakhaven crossroads for a different place."
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "exit2",
|
||||
"question": "This is one turn of an interactive story. The story can move on from the current scene into one of the listed next scenes, but only once the narration actually gets there. Where does this turn leave the story?",
|
||||
"options": [
|
||||
{
|
||||
"id": "stay",
|
||||
"description": "Still in the current scene, The Idle Wait: the narration has not reached any next scene yet."
|
||||
},
|
||||
{
|
||||
"id": "meet_wren",
|
||||
"description": "Moved on to Meeting the Fletcher: You find Wren working on arrows. She has a soldier's eyes."
|
||||
},
|
||||
{
|
||||
"id": "tavern",
|
||||
"description": "Moved on to Tavern Noise: The local tavern is loud with soldiers from both sides."
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"cases": [
|
||||
{
|
||||
"id": "f-idle-look",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The Idle Wait",
|
||||
"summary": "You have no contracts, no coin, and nowhere to be. The town is a pressure cooker.",
|
||||
"location": "Oakhaven crossroads"
|
||||
},
|
||||
"player_action": "I lean on the fence post and watch the soldiers pass.",
|
||||
"narration": "Boots churn the mud as a column of Iron Legion pikemen tramps past the crossroads, eyes forward, drums muffled by the damp. Nobody spares you a glance. Across the square, a Sylvan Pact scout watches them go from the shadow of the well, and does not move until the last of them has passed."
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "stay",
|
||||
"place2": "same",
|
||||
"exit2": "stay"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "f-plan-tavern",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The Idle Wait",
|
||||
"summary": "You have no contracts, no coin, and nowhere to be. The town is a pressure cooker.",
|
||||
"location": "Oakhaven crossroads"
|
||||
},
|
||||
"player_action": "Maybe I'll try the tavern later, see if anyone's hiring.",
|
||||
"narration": "A gust carries the smell of woodsmoke and spilled ale from the tavern at the far end of the square. Someone inside is laughing too loudly. It would be warm in there, and full of men with coin and grudges. For now the crossroads holds you, and the rain has not yet decided whether to fall."
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "stay",
|
||||
"place2": "same",
|
||||
"exit2": "stay"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "f-go-tavern",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The Idle Wait",
|
||||
"summary": "You have no contracts, no coin, and nowhere to be. The town is a pressure cooker.",
|
||||
"location": "Oakhaven crossroads"
|
||||
},
|
||||
"player_action": "I head for the tavern and push the door open.",
|
||||
"narration": "The door gives with a groan and the noise hits you like heat: Legion sergeants at one long table, Pact rangers at another, and a narrow no-man's-land of spilled ale between them. The barkeep looks up, measures your empty belt, and goes back to wiping a cup that will never be clean."
|
||||
},
|
||||
"labels": {
|
||||
"place": "moved",
|
||||
"exit": "tavern",
|
||||
"place2": "elsewhere",
|
||||
"exit2": "tavern"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "f-approach-wren",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The Idle Wait",
|
||||
"summary": "You have no contracts, no coin, and nowhere to be. The town is a pressure cooker.",
|
||||
"location": "Oakhaven crossroads"
|
||||
},
|
||||
"player_action": "I walk over to the woman fletching arrows by the smithy.",
|
||||
"narration": "She doesn't look up as your shadow falls across her work. Her hands keep moving, binding grey goose feather to shaft with quick, sure turns. \"These are war arrows,\" she says, \"fletched for distance and accuracy, not for some merchant's target practice. You're either very brave or very stupid.\""
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "meet_wren",
|
||||
"place2": "same",
|
||||
"exit2": "meet_wren"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "f-memory",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The Idle Wait",
|
||||
"summary": "You have no contracts, no coin, and nowhere to be. The town is a pressure cooker.",
|
||||
"location": "Oakhaven crossroads"
|
||||
},
|
||||
"player_action": "I think back to the last time I held a bow, at the siege of Varn.",
|
||||
"narration": "The memory comes unbidden: the walls of Varn black against a burning sky, the bowstring biting your fingers, the screaming from the breach. You flex your hand now, in the mud of Oakhaven, and find it still remembers the draw. Nobody at the crossroads notices a ghost remembering a war."
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "stay",
|
||||
"place2": "same",
|
||||
"exit2": "stay"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "f-leave-town",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The Idle Wait",
|
||||
"summary": "You have no contracts, no coin, and nowhere to be. The town is a pressure cooker.",
|
||||
"location": "Oakhaven crossroads"
|
||||
},
|
||||
"player_action": "Forget this town. I take the west road out of Oakhaven.",
|
||||
"narration": "The crossroads falls behind you, then the last hovels, then the smell of the town itself. By afternoon the west road narrows between dark pines, and the only sound is your own boots and, somewhere ahead, the slow creak of a mill wheel turning in water you cannot yet see."
|
||||
},
|
||||
"labels": {
|
||||
"place": "moved",
|
||||
"exit": "stay",
|
||||
"place2": "elsewhere",
|
||||
"exit2": "stay"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "f-nightfall",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The Idle Wait",
|
||||
"summary": "You have no contracts, no coin, and nowhere to be. The town is a pressure cooker.",
|
||||
"location": "Oakhaven crossroads"
|
||||
},
|
||||
"player_action": "I stay put and wait for something to happen.",
|
||||
"narration": "Hours pass. The rain finally commits, then relents. Night settles over the crossroads, and one by one the shutters close. A lantern is lit at the well. Still nothing. Still waiting. Somewhere a dog barks twice and thinks better of it."
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "stay",
|
||||
"place2": "same",
|
||||
"exit2": "stay"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "f-wren-comes",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The Idle Wait",
|
||||
"summary": "You have no contracts, no coin, and nowhere to be. The town is a pressure cooker.",
|
||||
"location": "Oakhaven crossroads"
|
||||
},
|
||||
"player_action": "I keep watching the square.",
|
||||
"narration": "A woman with a quiver of fresh-fletched arrows crosses the square toward you, stepping neatly around the worst of the mud. She stops an arm's length away and studies you the way a quartermaster studies a horse. \"You've been standing there all morning,\" Wren says. \"Either you're waiting for work, or you're a spy. Which is it?\""
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "meet_wren",
|
||||
"place2": "same",
|
||||
"exit2": "meet_wren"
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,196 @@
|
||||
{
|
||||
"name": "wyrd-scene-toll",
|
||||
"decisions": [
|
||||
{
|
||||
"id": "place",
|
||||
"question": "This is one turn of an interactive story. Has the scene moved to a genuinely different place during this turn, or does the current scene still hold?",
|
||||
"options": [
|
||||
{
|
||||
"id": "moved",
|
||||
"description": "The scene moved to a genuinely different place."
|
||||
},
|
||||
{
|
||||
"id": "holds",
|
||||
"description": "The current scene still holds."
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "exit",
|
||||
"question": "This is one turn of an interactive story. The story can advance from the current scene into one of the listed next scenes. Which one does this turn move the story into?",
|
||||
"options": [
|
||||
{
|
||||
"id": "stay",
|
||||
"description": "None of these yet. The current scene holds: The First Toll."
|
||||
},
|
||||
{
|
||||
"id": "departure",
|
||||
"description": "Passage Granted/Denied: Beyond the Toll. You leave the tollhouse, either with or without Sera's blessing, continuing down the north road."
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "place2",
|
||||
"question": "This is one turn of an interactive story. The turn began at the north road tollhouse. By the end of this turn's narration, where are the characters?",
|
||||
"options": [
|
||||
{
|
||||
"id": "same",
|
||||
"description": "Still at the north road tollhouse."
|
||||
},
|
||||
{
|
||||
"id": "elsewhere",
|
||||
"description": "Somewhere else: they have left the north road tollhouse for a different place."
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "exit2",
|
||||
"question": "This is one turn of an interactive story. The story can move on from the current scene into one of the listed next scenes, but only once the narration actually gets there. Where does this turn leave the story?",
|
||||
"options": [
|
||||
{
|
||||
"id": "stay",
|
||||
"description": "Still in the current scene, The First Toll: the narration has not reached any next scene yet."
|
||||
},
|
||||
{
|
||||
"id": "departure",
|
||||
"description": "Moved on to Beyond the Toll: You leave the tollhouse, either with or without Sera's blessing, continuing down the north road."
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"cases": [
|
||||
{
|
||||
"id": "t-reads-letter",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The First Toll",
|
||||
"summary": "You arrive at the tollhouse in the pouring rain, presenting your sealed letter to Sera.",
|
||||
"location": "north road tollhouse"
|
||||
},
|
||||
"player_action": "I hand Sera the sealed letter.",
|
||||
"narration": "Sera's gloved hand emerges from the shadow to take the letter, her fingers tracing the wax seal with practiced efficiency. She turns it over slowly, examining it by the light of the single lamp. Rain hammers the iron roof. \"Whose seal is this?\" she asks, without looking up."
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "stay",
|
||||
"place2": "same",
|
||||
"exit2": "stay"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "t-questions",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The First Toll",
|
||||
"summary": "You arrive at the tollhouse in the pouring rain, presenting your sealed letter to Sera.",
|
||||
"location": "north road tollhouse"
|
||||
},
|
||||
"player_action": "\"A friend in the capital. That's all you need to know.\"",
|
||||
"narration": "\"That's rarely all I need to know.\" Sera sets the letter on the counter between you and leaves her hand on it. \"The road north has been kind to liars lately. It won't stay that way.\" Water drips from the eaves into a bucket by the door, counting out the silence."
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "stay",
|
||||
"place2": "same",
|
||||
"exit2": "stay"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "t-future",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The First Toll",
|
||||
"summary": "You arrive at the tollhouse in the pouring rain, presenting your sealed letter to Sera.",
|
||||
"location": "north road tollhouse"
|
||||
},
|
||||
"player_action": "\"Stamp it and I'll be out of your hair before the rain stops.\"",
|
||||
"narration": "\"Before the rain stops,\" Sera repeats, as if tasting the words. \"It hasn't stopped in nine days.\" She doesn't reach for the stamp. She reaches for the ledger instead, and opens it to a page thick with names, and runs one gloved finger down the column."
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "stay",
|
||||
"place2": "same",
|
||||
"exit2": "stay"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "t-granted",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The First Toll",
|
||||
"summary": "You arrive at the tollhouse in the pouring rain, presenting your sealed letter to Sera.",
|
||||
"location": "north road tollhouse"
|
||||
},
|
||||
"player_action": "I wait while she decides.",
|
||||
"narration": "The stamp comes down with a crack like a gavel. Sera slides the letter back and jerks her chin at the door. You step out into the downpour, and the tollhouse lamp shrinks behind you as the north road unspools ahead, black and glistening, into the hills."
|
||||
},
|
||||
"labels": {
|
||||
"place": "moved",
|
||||
"exit": "departure",
|
||||
"place2": "elsewhere",
|
||||
"exit2": "departure"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "t-denied",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The First Toll",
|
||||
"summary": "You arrive at the tollhouse in the pouring rain, presenting your sealed letter to Sera.",
|
||||
"location": "north road tollhouse"
|
||||
},
|
||||
"player_action": "\"Then I'll find another way north.\" I turn to go.",
|
||||
"narration": "\"There is no other way north,\" Sera says to your back, but she lets you go. The door bangs shut behind you. You walk on anyway, past the barrier, the rain soaking through your collar, the tollhouse dwindling to a single yellow window behind you on the north road."
|
||||
},
|
||||
"labels": {
|
||||
"place": "moved",
|
||||
"exit": "departure",
|
||||
"place2": "elsewhere",
|
||||
"exit2": "departure"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "t-rider",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The First Toll",
|
||||
"summary": "You arrive at the tollhouse in the pouring rain, presenting your sealed letter to Sera.",
|
||||
"location": "north road tollhouse"
|
||||
},
|
||||
"player_action": "I stay by the stove to dry off.",
|
||||
"narration": "Hooves on the road outside, then a fist on the door. A rider in a sodden Legion cloak ducks in out of the rain, glances at you, glances at Sera, and drops a heavier purse on the counter than any toll requires. Sera's hand closes over it without a word."
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "stay",
|
||||
"place2": "same",
|
||||
"exit2": "stay"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "t-city-memory",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The First Toll",
|
||||
"summary": "You arrive at the tollhouse in the pouring rain, presenting your sealed letter to Sera.",
|
||||
"location": "north road tollhouse"
|
||||
},
|
||||
"player_action": "\"You've never been to the capital, have you?\"",
|
||||
"narration": "Something shifts behind Sera's eyes. \"I was born there,\" she says. \"Under the bell tower, in the year the river froze. I left the day they hanged my father in the square.\" She says it flatly, the way you'd read out a toll schedule, and turns back to the ledger."
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "stay",
|
||||
"place2": "same",
|
||||
"exit2": "stay"
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,177 @@
|
||||
{
|
||||
"name": "wyrd-scene-vigil",
|
||||
"decisions": [
|
||||
{
|
||||
"id": "place",
|
||||
"question": "This is one turn of an interactive story. Has the scene moved to a genuinely different place during this turn, or does the current scene still hold?",
|
||||
"options": [
|
||||
{
|
||||
"id": "moved",
|
||||
"description": "The scene moved to a genuinely different place."
|
||||
},
|
||||
{
|
||||
"id": "holds",
|
||||
"description": "The current scene still holds."
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "exit",
|
||||
"question": "This is one turn of an interactive story. The story can advance from the current scene into one of the listed next scenes. Which one does this turn move the story into?",
|
||||
"options": [
|
||||
{
|
||||
"id": "stay",
|
||||
"description": "None of these yet. The current scene holds: The Last Vigil."
|
||||
},
|
||||
{
|
||||
"id": "outside",
|
||||
"description": "Leave the sanctum: Outside the Sanctum. Lynnae and the player step out of the interior into the exterior surroundings."
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "place2",
|
||||
"question": "This is one turn of an interactive story. The turn began at the sanctum interior. By the end of this turn's narration, where are the characters?",
|
||||
"options": [
|
||||
{
|
||||
"id": "same",
|
||||
"description": "Still at the sanctum interior."
|
||||
},
|
||||
{
|
||||
"id": "elsewhere",
|
||||
"description": "Somewhere else: they have left the sanctum interior for a different place."
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "exit2",
|
||||
"question": "This is one turn of an interactive story. The story can move on from the current scene into one of the listed next scenes, but only once the narration actually gets there. Where does this turn leave the story?",
|
||||
"options": [
|
||||
{
|
||||
"id": "stay",
|
||||
"description": "Still in the current scene, The Last Vigil: the narration has not reached any next scene yet."
|
||||
},
|
||||
{
|
||||
"id": "outside",
|
||||
"description": "Moved on to Outside the Sanctum: Lynnae and the player step out of the interior into the exterior surroundings."
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"cases": [
|
||||
{
|
||||
"id": "v-talk",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The Last Vigil",
|
||||
"summary": "The player stands in the inner sanctum, guarding the Book. Lynnae enters through a hidden alcove to discuss the encroaching demons.",
|
||||
"location": "sanctum interior"
|
||||
},
|
||||
"player_action": "\"How close are they?\"",
|
||||
"narration": "Lynnae rests her fingertips on the edge of the altar, not quite touching the Book. \"Close enough that the wards are singing,\" she says. \"Listen.\" You do, and beneath the wind there is a high, thin note, like a wet finger on the rim of a glass. The candles lean toward the door."
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "stay",
|
||||
"place2": "same",
|
||||
"exit2": "stay"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "v-step-out",
|
||||
"tag": "control",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The Last Vigil",
|
||||
"summary": "The player stands in the inner sanctum, guarding the Book. Lynnae enters through a hidden alcove to discuss the encroaching demons.",
|
||||
"location": "sanctum interior"
|
||||
},
|
||||
"player_action": "\"Show me.\" I follow her out through the great doors.",
|
||||
"narration": "The doors groan open and the wind takes the candle flames behind you. Outside, the sanctum steps drop away into a valley of ash, and on the far ridge a line of red lights moves slowly, patiently, like a procession of lanterns that has forgotten what it is mourning."
|
||||
},
|
||||
"labels": {
|
||||
"place": "moved",
|
||||
"exit": "outside",
|
||||
"place2": "elsewhere",
|
||||
"exit2": "outside"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "v-plan-pub",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The Last Vigil",
|
||||
"summary": "The player stands in the inner sanctum, guarding the Book. Lynnae enters through a hidden alcove to discuss the encroaching demons.",
|
||||
"location": "sanctum interior"
|
||||
},
|
||||
"player_action": "\"If the wards fall, we take the Book to the old pub in the village.\"",
|
||||
"narration": "\"The Drowned Crown?\" Lynnae almost smiles. \"No demon would look for holy relics behind a bar.\" She turns the idea over, then nods slowly. \"If it comes to that. But the wards have not fallen yet, and I would rather they never did.\""
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "stay",
|
||||
"place2": "same",
|
||||
"exit2": "stay"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "v-to-pub",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The Last Vigil",
|
||||
"summary": "The player stands in the inner sanctum, guarding the Book. Lynnae enters through a hidden alcove to discuss the encroaching demons.",
|
||||
"location": "sanctum interior"
|
||||
},
|
||||
"player_action": "I take the Book and we run for the village pub.",
|
||||
"narration": "You run through the ash with the Book clutched to your chest, Lynnae a step behind. The heavy door creaks on rusted hinges as you cross the threshold, the air inside the tavern thick with the scent of stale ale, damp rot, and cold ash. You set the Book down on the bar."
|
||||
},
|
||||
"labels": {
|
||||
"place": "moved",
|
||||
"exit": "stay",
|
||||
"place2": "elsewhere",
|
||||
"exit2": "stay"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "v-vision",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The Last Vigil",
|
||||
"summary": "The player stands in the inner sanctum, guarding the Book. Lynnae enters through a hidden alcove to discuss the encroaching demons.",
|
||||
"location": "sanctum interior"
|
||||
},
|
||||
"player_action": "I touch the Book's cover.",
|
||||
"narration": "The world tilts. For one breath you are elsewhere: a burning library, a hundred years ago, monks carrying armfuls of scrolls through smoke. Then the vision snaps shut like a book, and you are back in the sanctum with Lynnae's hand hard on your wrist. \"Don't,\" she says. \"Not yet.\""
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "stay",
|
||||
"place2": "same",
|
||||
"exit2": "stay"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "v-demon-enters",
|
||||
"tag": "case",
|
||||
"state": {
|
||||
"current_scene": {
|
||||
"title": "The Last Vigil",
|
||||
"summary": "The player stands in the inner sanctum, guarding the Book. Lynnae enters through a hidden alcove to discuss the encroaching demons.",
|
||||
"location": "sanctum interior"
|
||||
},
|
||||
"player_action": "I draw my blade and stand before the altar.",
|
||||
"narration": "The great doors burst inward. Ash rolls across the floor, and in the middle of it something tall unfolds itself, too many joints, a crown of red light. The candles gutter out all at once. Lynnae is already chanting, her voice steady, her hands trembling."
|
||||
},
|
||||
"labels": {
|
||||
"place": "holds",
|
||||
"exit": "stay",
|
||||
"place2": "same",
|
||||
"exit2": "stay"
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"repo": "Qwen/Qwen3.5-4B",
|
||||
"revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
|
||||
"path": "/hf/hub/models--Qwen--Qwen3.5-4B/snapshots/851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
|
||||
"note": "Jev bench 2026-09-30: imajev's own pinned base (artifacts/model-qwen4b.json), pointed at fv-ml1's HF cache"
|
||||
}
|
||||
@@ -0,0 +1,63 @@
|
||||
"""A TypeSafe-style POST /v1/systemone server around Intern-Decision's own inference.py
|
||||
(DecisionEngine.predict, shipped in the model repo), so JevBench's typesafe adapter and the bench
|
||||
harness reach it the same way they reach jevk5-serve and imajev's server. Nothing about the scoring
|
||||
is changed: the request goes to predict() as-is (the card's documented request format) and its
|
||||
response is returned with one added field, timing.server_ms.
|
||||
python intern_server.py --checkpoint <snapshot dir> --port 18090
|
||||
"""
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
import threading
|
||||
import time
|
||||
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--checkpoint", required=True)
|
||||
ap.add_argument("--host", default="127.0.0.1")
|
||||
ap.add_argument("--port", type=int, default=18090)
|
||||
args = ap.parse_args()
|
||||
sys.path.insert(0, args.checkpoint)
|
||||
from inference import DecisionEngine # noqa: E402 (the checkpoint's own module)
|
||||
|
||||
engine = DecisionEngine(checkpoint=args.checkpoint, device="cuda")
|
||||
engine.predict({"state": "warm-up", "questions": {"q": {"type": "noul", "instructions": "Is this a warm-up?"}}})
|
||||
lock = threading.Lock()
|
||||
|
||||
|
||||
class H(BaseHTTPRequestHandler):
|
||||
def _send(self, code, payload):
|
||||
data = json.dumps(payload).encode()
|
||||
self.send_response(code)
|
||||
self.send_header("Content-Type", "application/json")
|
||||
self.send_header("Content-Length", str(len(data)))
|
||||
self.end_headers()
|
||||
self.wfile.write(data)
|
||||
|
||||
def do_GET(self):
|
||||
self._send(200, {"ok": True, "model": "Intern-Decision"} if self.path.rstrip("/") == "/health"
|
||||
else {"error": "not found"})
|
||||
|
||||
def do_POST(self):
|
||||
if self.path.rstrip("/") != "/v1/systemone":
|
||||
return self._send(404, {"error": "not found"})
|
||||
t = time.perf_counter()
|
||||
try:
|
||||
body = json.loads(self.rfile.read(int(self.headers.get("Content-Length", 0))))
|
||||
with lock:
|
||||
out = engine.predict({k: body[k] for k in ("state", "questions", "images") if k in body})
|
||||
except (ValueError, KeyError, TypeError, ArithmeticError) as e:
|
||||
return self._send(400, {"error": f"{type(e).__name__}: {e}"})
|
||||
except Exception as e: # noqa: BLE001 e.g. CUDA OOM under the cap
|
||||
import torch
|
||||
torch.cuda.empty_cache()
|
||||
return self._send(503, {"error": f"{type(e).__name__}: {str(e)[:300]}"})
|
||||
out.setdefault("timing", {})["server_ms"] = round((time.perf_counter() - t) * 1000, 2)
|
||||
self._send(200, out)
|
||||
|
||||
def log_message(self, *a):
|
||||
pass
|
||||
|
||||
|
||||
print(f"serving Intern-Decision {args.checkpoint} on {args.host}:{args.port}", flush=True)
|
||||
ThreadingHTTPServer((args.host, args.port), H).serve_forever()
|
||||
@@ -0,0 +1,34 @@
|
||||
"""Build the ~3,900-token state for the long shapes: JevBench v1.2.16's longest public hard state
|
||||
(hard-opus-a-long_policy-01), topped up with the next longest until SemIf's shared prefix (template
|
||||
+ evidence, as semif-serve's timing.prefix_tokens counts it) reaches ~3,900 Qwen3.5 tokens.
|
||||
Runs in the semif-serve image (tokenizer only, CPU)."""
|
||||
import json
|
||||
import sys
|
||||
|
||||
from transformers import AutoTokenizer
|
||||
from semif_phase1.core import direct_messages
|
||||
|
||||
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3.5-4B", revision="851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a")
|
||||
rows = {json.loads(l)["id"]: json.loads(l) for l in open(sys.argv[1])}
|
||||
text = rows["hard-opus-a-long_policy-01"]["state"] + "\n\n" + rows["hard-sol-b-long_policy-01"]["state"]
|
||||
|
||||
|
||||
def prefix_tokens(state):
|
||||
msgs = direct_messages({"id": "x", "state": state, "question": "q", "options": [
|
||||
{"id": "yes", "description": "Yes"}, {"id": "no", "description": "No"}]})
|
||||
prompt = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True, enable_thinking=False)
|
||||
payload = msgs[-1]["content"]
|
||||
evidence = json.dumps({"evidence": state}, ensure_ascii=False)[:-1]
|
||||
return len(tok.encode(prompt[: prompt.index(payload)] + evidence, add_special_tokens=False)) - 1
|
||||
|
||||
|
||||
lo, hi = 1000, len(text)
|
||||
while lo < hi:
|
||||
mid = (lo + hi + 1) // 2
|
||||
if prefix_tokens(text[:mid]) <= 3900:
|
||||
lo = mid
|
||||
else:
|
||||
hi = mid - 1
|
||||
state = text[:lo]
|
||||
open(sys.argv[2], "w").write(state)
|
||||
print(json.dumps({"chars": len(state), "prefix_tokens": prefix_tokens(state)}))
|
||||
@@ -0,0 +1,14 @@
|
||||
import os, sys, time
|
||||
from huggingface_hub import snapshot_download
|
||||
PINS = [
|
||||
("crh225/plumb-4b", "55de037801a8a9b9de3db5c0e16cef86210c2186", None),
|
||||
("alibiserikbay/JevK5", "27d2d6b8d4714807f6293b0623bd7370b27e42f8", None),
|
||||
("internlm/Intern-Decision-4B", "0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd", None),
|
||||
("mohit67890/imajev-4b", "c9e5f132465da85d31735ec502d5557982671a7d", ["mlx/*", "assets/*"]),
|
||||
]
|
||||
only = sys.argv[1:] or [p[0] for p in PINS]
|
||||
for repo, rev, ignore in PINS:
|
||||
if repo not in only: continue
|
||||
t = time.time()
|
||||
path = snapshot_download(repo, revision=rev, ignore_patterns=ignore)
|
||||
print(f"{repo}@{rev} -> {path} in {time.time()-t:.0f}s", flush=True)
|
||||
@@ -0,0 +1,5 @@
|
||||
import time
|
||||
from huggingface_hub import snapshot_download
|
||||
t = time.time()
|
||||
p = snapshot_download("alibiserikbay/JevK5", revision="c4f7fdb3aeab5582336406e78d3bef11bf98833d")
|
||||
print(p, f"{time.time()-t:.0f}s", flush=True)
|
||||
+14
@@ -0,0 +1,14 @@
|
||||
#!/usr/bin/env bash
|
||||
# Runs code/run.sh once per line of queue.txt, in order, one at a time. Lines can be appended while it
|
||||
# runs. touch STOP to halt after the current step.
|
||||
W=/tmp/jevbench-2026-09-30
|
||||
cd $W
|
||||
while [ ! -e $W/STOP ]; do
|
||||
line=$(head -n 1 $W/queue.txt 2>/dev/null)
|
||||
[ -z "$line" ] && { echo "$(date -u +%FT%TZ) queue empty" >> $W/logs/queue.out; break; }
|
||||
tail -n +2 $W/queue.txt > $W/queue.txt.tmp && mv $W/queue.txt.tmp $W/queue.txt
|
||||
echo "$(date -u +%FT%TZ) >>> $line" >> $W/logs/queue.out
|
||||
# shellcheck disable=SC2086
|
||||
$W/code/run.sh $line >> $W/logs/queue.out 2>&1
|
||||
echo "$(date -u +%FT%TZ) <<< rc=$? $line" >> $W/logs/queue.out
|
||||
done
|
||||
+133
@@ -0,0 +1,133 @@
|
||||
#!/usr/bin/env bash
|
||||
# Jev-candidate bench (2026-09-30), one repeat of one model/format on fv-ml1 GPU 3.
|
||||
# run.sh <label> semif <repo>@<sha> <repeat> [capacity]
|
||||
# run.sh <label> native jevk5:<repo>@<sha> <repeat> [capacity]
|
||||
# run.sh <label> native intern:<repo>@<sha> <repeat> [capacity]
|
||||
# run.sh <label> native imajev:<repo>@<sha> <repeat> [capacity]
|
||||
# One process lifetime = one repeat. Everything is loopback on fv-ml1; GPU 3 only; one candidate
|
||||
# loaded at a time; every container this starts is removed before it returns.
|
||||
set -u
|
||||
W=/tmp/jevbench-2026-09-30
|
||||
label=$1 fmt=$2 rt=$3 r=$4 cap=${5:-}
|
||||
O=$W/out/$label/r$r
|
||||
mkdir -p "$O" && chmod 777 "$O"
|
||||
TASKS_HOST=$W/src/jevbench-v1.2.16/datasets/public/easy.jsonl,$W/src/jevbench-v1.2.16/datasets/public/original.jsonl,$W/src/jevbench-v1.2.16/datasets/public/hard.jsonl
|
||||
TASKS_CT=/jb/datasets/public/easy.jsonl,/jb/datasets/public/original.jsonl,/jb/datasets/public/hard.jsonl
|
||||
GPU3_UUID=GPU-186dacf4-f447-ef03-ffc3-12ab28e1b8ff
|
||||
log() { echo "$(date -u +%FT%TZ) [$label r$r] $*" | tee -a $W/logs/run.log; }
|
||||
|
||||
guard() { # nothing of ours may be up; anything foreign on GPU 3 is logged; a full-size seat stops us
|
||||
local used apps
|
||||
used=$(nvidia-smi -i 3 --query-gpu=memory.used --format=csv,noheader,nounits)
|
||||
apps=$(nvidia-smi --query-compute-apps=gpu_uuid,pid,process_name,used_memory --format=csv,noheader | grep "$GPU3_UUID" || true)
|
||||
echo "$(date -u +%FT%TZ) gpu3_used_mib=$used apps=[$apps]" >> "$O/guard.txt"
|
||||
if [ "$used" -gt 40000 ]; then log "GUARD STOP: GPU 3 holds ${used} MiB of someone else's (full-size seat?): $apps"; exit 10; fi
|
||||
if [ "$used" -gt 1024 ]; then log "GUARD NOTE: GPU 3 already holds ${used} MiB (not ours): $apps — our footprint stays < 30 GiB"; fi
|
||||
}
|
||||
|
||||
poll_start() { ( while :; do printf '%s,%s\n' "$(date +%s.%N)" "$(nvidia-smi -i 3 --query-gpu=memory.used --format=csv,noheader,nounits)"; sleep 0.2; done ) >> "$O/vram.csv" & POLL=$!; }
|
||||
poll_stop() { kill "$POLL" 2>/dev/null; wait "$POLL" 2>/dev/null; }
|
||||
|
||||
wait_up() { # $1 url, $2 container: any HTTP answer = up; container exit = failure
|
||||
local i code
|
||||
for i in $(seq 1 180); do
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' "$1" || true)
|
||||
if [ "$code" != "000" ]; then echo "$(date +%s.%N)" > "$O/up_at.txt"; return 0; fi
|
||||
if [ "$(docker inspect -f '{{.State.Running}}' "$2" 2>/dev/null)" != "true" ]; then return 1; fi
|
||||
sleep 5
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
summarize_jb() {
|
||||
PYTHONPATH=$W/src/jevbench-v1.2.16 python3 -m jevbench.cli summarize --tasks "$TASKS_HOST" \
|
||||
--results "$O/jevbench/results.jsonl" > "$O/jevbench/summary.json" 2> "$O/jevbench/summarize.err" \
|
||||
&& python3 - "$O/jevbench/summary.json" <<'EOF' | tee -a $W/logs/run.log
|
||||
import json,sys
|
||||
s=json.load(open(sys.argv[1]))
|
||||
print(" jevbench:", {k:(v["n_correct"],v["n_scorable"]) for k,v in s["splits"].items()}, "all", (s["n_correct"], s["n_scorable"]))
|
||||
EOF
|
||||
}
|
||||
|
||||
guard
|
||||
log "start fmt=$fmt rt=$rt cap=${cap:-no}"
|
||||
mkdir -p "$O/jevbench" && chmod 777 "$O/jevbench"
|
||||
|
||||
if [ "$fmt" = semif ]; then
|
||||
repo=${rt%@*}; sha=${rt#*@}
|
||||
# 1. JevBench's own SemIf adapter (semif_direct), in-process, on the same weights.
|
||||
poll_start
|
||||
echo "$(date +%s.%N) jevbench:start" >> "$O/events.txt"
|
||||
docker run --rm --name jevbench-jb --gpus '"device=3"' -e HF_HOME=/hf -e JEVBENCH_WARM_LOAD=1 -e PYTHONPATH=/jb \
|
||||
-v /tank/aimodels/huggingface:/hf:ro -v $W/src/jevbench-v1.2.16:/jb:ro -v "$O/jevbench":/out -w /jb \
|
||||
--entrypoint /app/.venv/bin/python semif-serve:0.1.4 -m jevbench.cli run --tasks "$TASKS_CT" \
|
||||
--adapter semif_direct --endpoint "$repo" --revision "$sha" --run-label "$label" \
|
||||
--results /out/results.jsonl --ledger /out/ledger.jsonl --raw-dir /out/raw \
|
||||
--price-in-per-m 0 --price-out-per-m 0 --manifest /out/manifest.json > "$O/jevbench/run.log" 2>&1
|
||||
log "jevbench semif_direct rc=$? $(tail -1 "$O/jevbench/run.log")"
|
||||
echo "$(date +%s.%N) jevbench:end" >> "$O/events.txt"
|
||||
summarize_jb
|
||||
# 2. semif-serve itself on the same weights, under the production 12 GiB cap.
|
||||
echo "$(date +%s.%N) serve:start" >> "$O/events.txt"
|
||||
docker run -d --name jevbench-serve --gpus '"device=3"' -e SEMIF_API_TOKEN="$(cat $W/token)" -e SEMIF_DEVICE=cuda \
|
||||
-e SEMIF_VRAM_CAP_GIB=12 -e SEMIF_MODEL="$repo" -e SEMIF_REVISION="$sha" \
|
||||
-v /tank/aimodels/huggingface:/hf:ro -p 0.0.0.0:18032:8000 semif-serve:0.1.4 > /dev/null
|
||||
if ! wait_up http://127.0.0.1:18032/health jevbench-serve; then
|
||||
log "FAIL: semif-serve did not come up"; docker logs jevbench-serve > "$O/server.log" 2>&1; docker rm -f jevbench-serve > /dev/null; poll_stop; exit 3
|
||||
fi
|
||||
echo "$(date +%s.%N) rest:start" >> "$O/events.txt"; sleep 8; echo "$(date +%s.%N) rest:end" >> "$O/events.txt"
|
||||
curl -s http://127.0.0.1:18032/health > "$O/health-rest.json"
|
||||
echo "$(date +%s.%N) sets:start" >> "$O/events.txt"
|
||||
python3 $W/code/bench_sets.py --backend semif --url http://127.0.0.1:18032 --token-file $W/token \
|
||||
--label "$label-r$r" --out "$O/sets.json" 2>&1 | tee -a "$O/bench.log"
|
||||
curl -s http://127.0.0.1:18032/health > "$O/health-after-sets.json"
|
||||
python3 $W/code/bench_shape.py --backend semif --url http://127.0.0.1:18032 --token-file $W/token \
|
||||
--long-state-file $W/out/long_state.txt --label "$label-r$r" --out "$O/shape.json" ${cap:+--capacity} 2>&1 | tee -a "$O/bench.log"
|
||||
curl -s http://127.0.0.1:18032/health > "$O/health-end.json"
|
||||
docker logs jevbench-serve > "$O/server.log" 2>&1
|
||||
docker rm -f jevbench-serve > /dev/null
|
||||
echo "$(date +%s.%N) serve:end" >> "$O/events.txt"
|
||||
poll_stop
|
||||
else
|
||||
kind=${rt%%:*}; spec=${rt#*:}; repo=${spec%@*}; sha=${spec#*@}
|
||||
snap=/hf/hub/models--${repo/\//--}/snapshots/$sha
|
||||
maxq=10000; envx=()
|
||||
case $kind in
|
||||
jevk5) cmd=(jevk5.server --model "$snap" --host 0.0.0.0 --port 18090) ;;
|
||||
jevk5v03) cmd=(jevk5.server --model "$snap" --host 0.0.0.0 --port 18090) # the v0.3.3 runtime shadows the image's v0.2.0
|
||||
envx=(-e PYTHONPATH=/w/src/jevk5-v0.3.3) ;;
|
||||
intern) cmd=(/w/code/intern_server.py --checkpoint "$snap" --host 0.0.0.0 --port 18090); maxq=16 ;;
|
||||
imajev) cmd=(/w/src/imajev-a0134749/scripts/playground/server.py --backend torch --adapter "$snap"
|
||||
--model-bundle /w/code/imajev-bundle-qwen4b.json --rotations 1 --calibration "$snap/calibration.json"
|
||||
--model-name imajev-4b --host 0.0.0.0 --port 18090)
|
||||
envx=(-e PYTHONPATH=/w/src/imajev-a0134749/src:/w/src/imajev-a0134749/scripts); maxq=8 ;; # jev_api.MAX_QUESTIONS
|
||||
*) log "unknown native kind $kind"; exit 2 ;;
|
||||
esac
|
||||
poll_start
|
||||
echo "$(date +%s.%N) serve:start" >> "$O/events.txt"
|
||||
docker run -d --name jevbench-native --gpus '"device=3"' -e HF_HOME=/hf -e HF_HUB_OFFLINE=1 \
|
||||
-e BENCH_VRAM_CAP_GIB="${VCAP:-12}" "${envx[@]}" -v /tank/aimodels/huggingface:/hf:ro -v $W:/w:ro \
|
||||
-p 127.0.0.1:18090:18090 --entrypoint /app/.venv/bin/python jevbench-native:2026-09-30 /w/code/capped.py "${cmd[@]}" > /dev/null
|
||||
if ! wait_up http://127.0.0.1:18090/health jevbench-native; then
|
||||
log "FAIL: native server did not come up"; docker logs jevbench-native > "$O/server.log" 2>&1; docker rm -f jevbench-native > /dev/null; poll_stop; exit 3
|
||||
fi
|
||||
echo "$(date +%s.%N) rest:start" >> "$O/events.txt"; sleep 8; echo "$(date +%s.%N) rest:end" >> "$O/events.txt"
|
||||
echo "$(date +%s.%N) jevbench:start" >> "$O/events.txt"
|
||||
PYTHONPATH=$W/src/jevbench-v1.2.16 python3 -m jevbench.cli run --tasks "$TASKS_HOST" --adapter typesafe \
|
||||
--endpoint http://127.0.0.1:18090 --key-env '' --model "$label" --run-label "$label" \
|
||||
--results "$O/jevbench/results.jsonl" --ledger "$O/jevbench/ledger.jsonl" --raw-dir "$O/jevbench/raw" \
|
||||
--price-in-per-m 0 --price-out-per-m 0 --manifest "$O/jevbench/manifest.json" > "$O/jevbench/run.log" 2>&1
|
||||
log "jevbench typesafe rc=$? $(tail -1 "$O/jevbench/run.log")"
|
||||
echo "$(date +%s.%N) jevbench:end" >> "$O/events.txt"
|
||||
summarize_jb
|
||||
echo "$(date +%s.%N) sets:start" >> "$O/events.txt"
|
||||
python3 $W/code/bench_sets.py --backend systemone --url http://127.0.0.1:18090 \
|
||||
--label "$label-r$r" --out "$O/sets.json" 2>&1 | tee -a "$O/bench.log"
|
||||
python3 $W/code/bench_shape.py --backend systemone --url http://127.0.0.1:18090 --max-questions $maxq \
|
||||
--long-state-file $W/out/long_state.txt --label "$label-r$r" --out "$O/shape.json" ${cap:+--capacity} 2>&1 | tee -a "$O/bench.log"
|
||||
docker logs jevbench-native > "$O/server.log" 2>&1
|
||||
docker rm -f jevbench-native > /dev/null
|
||||
echo "$(date +%s.%N) serve:end" >> "$O/events.txt"
|
||||
poll_stop
|
||||
fi
|
||||
log "done"
|
||||
@@ -0,0 +1,164 @@
|
||||
"""Render analyze.py's summary.json as the markdown tables of the results doc.
|
||||
python3 tables.py summary.json > tables.md
|
||||
Cells: median over repeats, with min..max when the repeats differ; n = repeats."""
|
||||
import json
|
||||
import sys
|
||||
|
||||
S = json.load(open(sys.argv[1]))
|
||||
ORDER = ["semif-qwen35-4b", "plumb-4b-native", "plumb-4b-dropin", "imajev-4b-native",
|
||||
"intern-decision-4b-native", "intern-decision-4b-dropin", "jevk5-v02-native", "jevk5-v02-dropin",
|
||||
"jevk5-v03-native", "jevk5-v03-dropin"]
|
||||
NAMES = {"semif-qwen35-4b": "SemIf (Qwen3.5-4B), baseline", "plumb-4b-native": "Plumb-4B, native",
|
||||
"plumb-4b-dropin": "Plumb-4B, drop-in", "imajev-4b-native": "Imajev-4B, native",
|
||||
"intern-decision-4b-native": "Intern-Decision-4B, native", "intern-decision-4b-dropin": "Intern-Decision-4B, drop-in",
|
||||
"jevk5-v02-native": "JevK5 v0.2, native", "jevk5-v02-dropin": "JevK5 v0.2, drop-in",
|
||||
"jevk5-v03-native": "JevK5 v0.3, native", "jevk5-v03-dropin": "JevK5 v0.3, drop-in"}
|
||||
labels = [l for l in ORDER if l in S] + [l for l in S if l not in ORDER]
|
||||
|
||||
|
||||
def cell(m, fmt="{:.0f}", scale=1):
|
||||
if not m:
|
||||
return "–"
|
||||
a, lo, hi = m["median"] * scale, m["min"] * scale, m["max"] * scale
|
||||
s = fmt.format(a)
|
||||
if fmt.format(lo) != fmt.format(hi):
|
||||
s += f" ({fmt.format(lo)}–{fmt.format(hi)})"
|
||||
return s
|
||||
|
||||
|
||||
def reps(l):
|
||||
return len(S[l]["repeats"])
|
||||
|
||||
|
||||
out = []
|
||||
w = out.append
|
||||
|
||||
w("### Key table\n")
|
||||
w("| system | fits 12 GiB? rest / peak MiB | JevBench all /231 · hard /111 | pooled /259 single · rot | Wyrd /84 single · rot | Δ pooled vs SemIf single (95% CI) | Δ pooled vs SemIf rot (95% CI) | 21 criteria, ms | 16 × ~3,900 tok, ms |")
|
||||
w("|---|---|---|---|---|---|---|---|---|")
|
||||
for l in labels:
|
||||
x = S[l]
|
||||
v, j, t, L = x["vram"], x["jevbench"], x["sets"], x["latency"]
|
||||
if "all_n" not in j or "single/pooled" not in t or "rotations/pooled" not in t:
|
||||
continue
|
||||
pv = x.get("paired_vs_semif") or {}
|
||||
def d(k):
|
||||
p = pv.get(k)
|
||||
return f"{p['delta_pts']:+.1f} ({p['delta_95ci_pts'][0]:+.1f}..{p['delta_95ci_pts'][1]:+.1f})" if p else "baseline"
|
||||
fits = "yes" if v["peak_work_mib"] and v["peak_work_mib"]["max"] < 13500 else "?"
|
||||
w(f"| {NAMES.get(l, l)} | {fits}: {cell(v['rest_mib'])} / {cell(v['peak_work_mib'])} | {cell(j['all_n'])} · {cell(j['hard_n'])} | "
|
||||
f"{cell(t['single/pooled']['correct'])} · {cell(t['rotations/pooled']['correct'])} | "
|
||||
f"{cell(t['single/wyrd']['correct'])} · {cell(t['rotations/wyrd']['correct'])} | {d('single/pooled')} | {d('rotations/pooled')} | "
|
||||
f"{cell(L['crit21']['e2e_ms'], '{:.0f}') if 'crit21' in L else '–'} | {cell(L['long16']['e2e_ms'], '{:.0f}') if 'long16' in L else '–'} |")
|
||||
w("")
|
||||
|
||||
w("### JevBench v1.2.16 public items (positive control + candidate score)\n")
|
||||
w("| system | repeats | easy /48 | standard /72 | hard /111 | all /231 | all | p50 ms |")
|
||||
w("|---|---|---|---|---|---|---|---|")
|
||||
for l in labels:
|
||||
j = S[l]["jevbench"]
|
||||
if not j:
|
||||
continue
|
||||
w(f"| {NAMES.get(l, l)} | {j['all']['n']} | {cell(j['easy_n'])} | {cell(j['original_n'])} | {cell(j['hard_n'])} | "
|
||||
f"{cell(j['all_n'])} | {cell(j['all'], '{:.3f}')} | {cell(j['p50_ms'], '{:.0f}')} |")
|
||||
|
||||
for cond in ("single", "rotations"):
|
||||
w(f"\n### Replaced-baseline sets, {cond} ({'caller order' if cond == 'single' else 'n rotations averaged'})\n")
|
||||
w("| system | pooled /259 | authored144 | perturb108 | Cicada w1 /31 | Cicada w2 /31 | Wyrd /84 | Wyrd place2 /21 | Wyrd exit /21 |")
|
||||
w("|---|---|---|---|---|---|---|---|---|")
|
||||
for l in labels:
|
||||
t = S[l]["sets"]
|
||||
if f"{cond}/pooled" not in t:
|
||||
continue
|
||||
g = lambda k: cell(t[f"{cond}/{k}"]["correct"]) if f"{cond}/{k}" in t else "–"
|
||||
w(f"| {NAMES.get(l, l)} | {g('pooled')} | {g('authored144')} | {g('perturbations108')} | {g('cicada-w1')} | "
|
||||
f"{g('cicada-w2')} | {g('wyrd')} | {g('wyrd:place2')} | {g('wyrd:exit')} |")
|
||||
|
||||
w("\n### Wyrd as one request per turn (4 decisions over one state: SemIf /decide/shared, native multi-question request)\n")
|
||||
w("| system | repeats | Wyrd /84 | place /21 | place2 /21 | exit /21 | exit2 /21 |")
|
||||
w("|---|---|---|---|---|---|---|")
|
||||
for l in labels:
|
||||
t = S[l]["sets"]
|
||||
if "multifield/wyrd" not in t:
|
||||
continue
|
||||
g = lambda k: cell(t[f"multifield/{k}"]["correct"]) if f"multifield/{k}" in t else "–"
|
||||
w(f"| {NAMES.get(l, l)} | {t['multifield/wyrd']['correct']['n']} | {g('wyrd')} | {g('wyrd:place')} | {g('wyrd:place2')} | {g('wyrd:exit')} | {g('wyrd:exit2')} |")
|
||||
|
||||
w("\n### Null control (content-free state) and positive controls inside the spike sets, single ordering\n")
|
||||
w("| system | Cicada w1 blind /31 | Wyrd blind /84 | Cicada w1 controls /5 | Wyrd controls /28 |")
|
||||
w("|---|---|---|---|---|")
|
||||
for l in labels:
|
||||
t = S[l]["sets"]
|
||||
if "single/cicada-w1" not in t:
|
||||
continue
|
||||
w(f"| {NAMES.get(l, l)} | {cell(t['single/cicada-w1']['blind_correct'])} | {cell(t['single/wyrd']['blind_correct'])} | "
|
||||
f"{cell(t['single/cicada-w1']['controls'])} | {cell(t['single/wyrd']['controls'])} |")
|
||||
|
||||
w("\n### Paired against SemIf (each row's majority top over the repeats; group bootstrap 95% CI; exact McNemar)\n")
|
||||
w("| system | set | cond | SemIf | cand | fixed | broken | Δ pts | 95% CI | p |")
|
||||
w("|---|---|---|---|---|---|---|---|---|---|")
|
||||
for l in labels:
|
||||
pv = S[l].get("paired_vs_semif") or {}
|
||||
for k in ("single/pooled", "rotations/pooled", "single/authored144", "rotations/authored144",
|
||||
"single/cicada-w1", "rotations/cicada-w1", "single/wyrd", "rotations/wyrd",
|
||||
"single/perturbations108", "rotations/perturbations108", "single/cicada-w2", "rotations/cicada-w2",
|
||||
"cand-single-vs-semif-rotations/pooled", "cand-single-vs-semif-rotations/wyrd"):
|
||||
p = pv.get(k)
|
||||
if not p:
|
||||
continue
|
||||
cond, name = k.split("/")
|
||||
w(f"| {NAMES.get(l, l)} | {name} | {cond} | {p['semif']}/{p['n']} | {p['cand']}/{p['n']} | {p['fixed']} | "
|
||||
f"{p['broken']} | {p['delta_pts']:+.1f} | {p['delta_95ci_pts'][0]:+.1f}..{p['delta_95ci_pts'][1]:+.1f} | {p['mcnemar_p']} |")
|
||||
|
||||
w("\n### Noise floors\n")
|
||||
w("| system | A-vs-A in process (authored144): flips, max Δp | across restarts, single (560 rows/pair): flips per pair, max Δp | across restarts, rotations: flips per pair | labelled rows whose top moved in ANY restart pair, single: pooled /259 · perturb /108 · Cicada w2 /31 |")
|
||||
w("|---|---|---|---|---|")
|
||||
for l in labels:
|
||||
f = S[l]["floors"]
|
||||
aa = f["a_vs_a_in_process"]
|
||||
aa_s = ", ".join(f"{x['flips']}/{x['n']} Δp≤{x['max_dp']:.3f}" for x in aa) or "–"
|
||||
cr = f.get("cross_restart/single") or []
|
||||
cr_s = ", ".join(f"{x['flips']} (Δp≤{x['max_dp']:.3f})" for x in cr) or "–"
|
||||
crr = f.get("cross_restart/rotations") or []
|
||||
crr_s = ", ".join(str(x["flips"]) for x in crr) or "–"
|
||||
u = f.get("unstable_rows/single") or {}
|
||||
u_s = f"{u.get('pooled', '–')} · {u.get('perturbations108', '–')} · {u.get('cicada-w2', '–')} ({u.get('repeats', '?')} repeats)" if u else "–"
|
||||
w(f"| {NAMES.get(l, l)} | {aa_s} | {cr_s} | {crr_s} | {u_s} |")
|
||||
|
||||
w("\n### Order sensitivity (authored144, single ordering vs the same request reordered) and the negative control\n")
|
||||
w("| system | reversed: label changes /144 | reversed: max Δp | shuffled: label changes | shuffled: max Δp | NEG: same top as unrotated /144 | NEG: follows the description | NEG: right vs original gold |")
|
||||
w("|---|---|---|---|---|---|---|---|")
|
||||
for l in labels:
|
||||
o, n = S[l]["order"], S[l]["negative"]
|
||||
if not o:
|
||||
continue
|
||||
w(f"| {NAMES.get(l, l)} | {cell(o['reversed']['label_changes'])} | {cell(o['reversed']['max_dp'], '{:.2f}')} | "
|
||||
f"{cell(o['shuffled']['label_changes'])} | {cell(o['shuffled']['max_dp'], '{:.2f}')} | "
|
||||
f"{cell(n.get('same_top_as_unrotated'))} | {cell(n.get('follows_description'))} | {cell(n.get('vs_original_gold'))} |")
|
||||
|
||||
w("\n### Latency at our shape (loopback on fv-ml1, GPU 3; ms; median of all requests, run-median range)\n")
|
||||
w("| system | short1 e2e | crit21 e2e | crit21 server | long1 e2e | long16 e2e | long16 server | tokens crit21 / long16 |")
|
||||
w("|---|---|---|---|---|---|---|---|")
|
||||
for l in labels:
|
||||
L = S[l]["latency"]
|
||||
if not L:
|
||||
continue
|
||||
g = lambda k, f="e2e_ms": cell(L[k][f], "{:.0f}") if k in L else "–"
|
||||
w(f"| {NAMES.get(l, l)} | {g('short1')} | {g('crit21')} | {g('crit21', 'server_ms')} | {g('long1')} | {g('long16')} | "
|
||||
f"{g('long16', 'server_ms')} | {L.get('crit21', {}).get('tokens')} / {L.get('long16', {}).get('tokens')} |")
|
||||
|
||||
w("\n### VRAM (nvidia-smi, whole GPU 3, only our process on it) and capacity under the 12 GiB cap\n")
|
||||
w("| system | rest after warm-up MiB | peak in sets+shapes MiB | peak in capacity sweep MiB | max rows @ short state | max rows @ ~3,900 tok |")
|
||||
w("|---|---|---|---|---|---|")
|
||||
for l in labels:
|
||||
v, cap = S[l]["vram"], S[l]["capacity"]
|
||||
def mx(lab):
|
||||
vals = []
|
||||
for r, c in cap.items():
|
||||
ok = [n for n, s in c.get(lab, []) if s == 200]
|
||||
bad = [(n, s) for n, s in c.get(lab, []) if s != 200]
|
||||
vals.append(f"{max(ok) if ok else 0}" + (f" (next {bad[0][0]}: {bad[0][1]})" if bad else " (no failure up to the last step)"))
|
||||
return "; ".join(vals) or "–"
|
||||
w(f"| {NAMES.get(l, l)} | {cell(v['rest_mib'])} | {cell(v['peak_work_mib'])} | {cell(v['peak_capacity_mib'])} | {mx('short')} | {mx('long')} |")
|
||||
|
||||
print("\n".join(out))
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,280 @@
|
||||
### Key table
|
||||
|
||||
| system | fits 12 GiB? rest / peak MiB | JevBench all /231 · hard /111 | pooled /259 single · rot | Wyrd /84 single · rot | Δ pooled vs SemIf single (95% CI) | Δ pooled vs SemIf rot (95% CI) | 21 criteria, ms | 16 × ~3,900 tok, ms |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | yes: 9242 (9242–9250) / 12918 (12886–12918) | 187 · 68 | 217 · 232 | 71 · 71 | baseline | baseline | 131 (131–132) | 505 (502–519) |
|
||||
| Plumb-4B, native | yes: 12064 / 12064 | 207 (206–207) · 89 (88–89) | 224 (224–226) · 227 (227–228) | 65 · 67 | +2.7 (-2.7..+7.9) | -1.9 (-5.4..+1.5) | 294 (292–295) | 3378 (3374–3384) |
|
||||
| Plumb-4B, drop-in | yes: 9242 / 12918 | 206 · 88 | 221 · 224 | 66 · 67 | +1.5 (-3.7..+6.6) | -3.1 (-6.8..+0.4) | 131 (131–132) | 506 (506–507) |
|
||||
| Imajev-4B, native | yes: 10404 / 11056 | 199 · 80 | 235 · 236 | 68 · 69 | +6.9 (+1.6..+12.2) | +1.5 (-2.5..+5.6) | 1219 (1212–1230) | 5125 (5054–5140) |
|
||||
| Intern-Decision-4B, native | yes: 9736 / 10290 | 202 · 83 | 240 · 236 | 79 · 77 | +8.9 (+4.7..+13.2) | +1.5 (-2.0..+5.0) | 88 (88–89) | 215 (213–216) |
|
||||
| Intern-Decision-4B, drop-in | yes: 9242 / 12918 | 202 · 85 | 232 · 233 | 74 · 73 | +5.8 (+2.7..+9.0) | +0.4 (-2.3..+3.1) | 132 | 508 (500–516) |
|
||||
| JevK5 v0.2, native | yes: 12064 / 12064 | 198 (198–199) · 81 (81–82) | 229 (228–229) · 228 | 67 (66–67) · 67 | +4.6 (-0.8..+9.9) | -1.5 (-4.9..+1.9) | 291 (291–292) | 3360 (3358–3368) |
|
||||
| JevK5 v0.2, drop-in | yes: 9242 / 12918 | 200 · 82 | 232 · 237 | 72 · 74 | +5.8 (+1.1..+10.7) | +1.9 (-1.1..+5.1) | 132 | 518 (505–519) |
|
||||
| JevK5 v0.3, native | yes: 12064 / 12064 | 202 · 86 | 227 · 226 | 66 · 65 | +3.9 (-0.9..+8.5) | -2.3 (-6.0..+1.4) | 291 (291–292) | 3363 (3355–3370) |
|
||||
| JevK5 v0.3, drop-in | yes: 9242 / 12918 | 201 · 85 | 227 · 227 | 69 · 66 | +3.9 (-0.8..+8.2) | -1.9 (-5.5..+1.8) | 133 (131–133) | 518 (506–521) |
|
||||
|
||||
### JevBench v1.2.16 public items (positive control + candidate score)
|
||||
|
||||
| system | repeats | easy /48 | standard /72 | hard /111 | all /231 | all | p50 ms |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 4 | 48 | 71 | 68 | 187 | 0.810 | 36 |
|
||||
| Plumb-4B, native | 3 | 48 | 70 | 89 (88–89) | 207 (206–207) | 0.896 (0.892–0.896) | 20 |
|
||||
| Plumb-4B, drop-in | 3 | 48 | 70 | 88 | 206 | 0.892 | 36 |
|
||||
| Imajev-4B, native | 4 | 48 | 71 | 80 | 199 | 0.861 | 62 (62–63) |
|
||||
| Intern-Decision-4B, native | 4 | 48 | 71 | 83 | 202 | 0.874 | 40 |
|
||||
| Intern-Decision-4B, drop-in | 3 | 48 | 69 | 85 | 202 | 0.874 | 36 |
|
||||
| JevK5 v0.2, native | 3 | 48 | 69 | 81 (81–82) | 198 (198–199) | 0.857 (0.857–0.861) | 20 (20–21) |
|
||||
| JevK5 v0.2, drop-in | 3 | 48 | 70 | 82 | 200 | 0.866 | 36 (36–37) |
|
||||
| JevK5 v0.3, native | 3 | 48 | 68 | 86 | 202 | 0.874 | 20 (20–21) |
|
||||
| JevK5 v0.3, drop-in | 3 | 48 | 68 | 85 | 201 | 0.870 | 36 |
|
||||
|
||||
### Replaced-baseline sets, single (caller order)
|
||||
|
||||
| system | pooled /259 | authored144 | perturb108 | Cicada w1 /31 | Cicada w2 /31 | Wyrd /84 | Wyrd place2 /21 | Wyrd exit /21 |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 217 | 116 | 83 | 30 | 20 | 71 | 21 | 16 |
|
||||
| Plumb-4B, native | 224 (224–226) | 130 (130–132) | 84 | 29 | 19 | 65 | 21 | 10 |
|
||||
| Plumb-4B, drop-in | 221 | 129 | 88 | 26 | 20 | 66 | 21 | 12 |
|
||||
| Imajev-4B, native | 235 | 138 | 102 | 29 | 22 | 68 | 20 | 13 |
|
||||
| Intern-Decision-4B, native | 240 | 132 | 100 | 29 | 22 | 79 | 21 | 19 |
|
||||
| Intern-Decision-4B, drop-in | 232 | 128 | 102 | 30 | 19 | 74 | 21 | 17 |
|
||||
| JevK5 v0.2, native | 229 (228–229) | 132 | 85 (85–86) | 30 | 16 | 67 (66–67) | 21 | 12 (11–12) |
|
||||
| JevK5 v0.2, drop-in | 232 | 130 | 92 | 30 | 17 | 72 | 21 | 17 |
|
||||
| JevK5 v0.3, native | 227 | 135 | 101 (101–102) | 26 | 17 | 66 | 21 | 12 |
|
||||
| JevK5 v0.3, drop-in | 227 | 132 | 99 | 26 | 21 | 69 | 21 | 15 |
|
||||
|
||||
### Replaced-baseline sets, rotations (n rotations averaged)
|
||||
|
||||
| system | pooled /259 | authored144 | perturb108 | Cicada w1 /31 | Cicada w2 /31 | Wyrd /84 | Wyrd place2 /21 | Wyrd exit /21 |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 232 | 131 | 93 | 30 | 19 | 71 | 21 | 18 |
|
||||
| Plumb-4B, native | 227 (227–228) | 130 (130–131) | 87 | 30 | 18 (17–18) | 67 | 21 | 11 |
|
||||
| Plumb-4B, drop-in | 224 | 129 | 92 | 28 | 18 | 67 | 21 | 13 |
|
||||
| Imajev-4B, native | 236 | 138 | 101 | 29 | 23 | 69 | 20 | 13 |
|
||||
| Intern-Decision-4B, native | 236 | 130 | 99 | 29 | 22 | 77 | 21 | 19 |
|
||||
| Intern-Decision-4B, drop-in | 233 | 130 | 102 | 30 | 19 | 73 | 21 | 17 |
|
||||
| JevK5 v0.2, native | 228 | 131 | 91 (90–91) | 30 | 16 | 67 | 21 | 12 |
|
||||
| JevK5 v0.2, drop-in | 237 | 133 | 93 | 30 | 16 | 74 | 21 | 20 |
|
||||
| JevK5 v0.3, native | 226 | 135 | 101 | 26 | 16 | 65 | 20 | 12 |
|
||||
| JevK5 v0.3, drop-in | 227 | 135 | 101 | 26 | 19 | 66 | 20 | 13 |
|
||||
|
||||
### Wyrd as one request per turn (4 decisions over one state: SemIf /decide/shared, native multi-question request)
|
||||
|
||||
| system | repeats | Wyrd /84 | place /21 | place2 /21 | exit /21 | exit2 /21 |
|
||||
|---|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 3 | 71 | 17 | 21 | 16 | 17 |
|
||||
| Plumb-4B, native | 2 | 65 | 17 | 21 | 10 | 17 |
|
||||
| Plumb-4B, drop-in | 3 | 66 | 17 | 21 | 12 | 16 |
|
||||
| Imajev-4B, native | 3 | 68 | 19 | 20 | 13 | 16 |
|
||||
| Intern-Decision-4B, native | 3 | 77 | 16 | 21 | 20 | 20 |
|
||||
| Intern-Decision-4B, drop-in | 3 | 74 | 18 | 21 | 17 | 18 |
|
||||
| JevK5 v0.2, native | 3 | 67 (66–67) | 17 | 21 | 12 (11–12) | 17 |
|
||||
| JevK5 v0.2, drop-in | 3 | 73 | 18 | 21 | 18 | 16 |
|
||||
| JevK5 v0.3, native | 3 | 66 | 18 | 21 | 12 | 15 |
|
||||
| JevK5 v0.3, drop-in | 3 | 69 | 18 | 21 | 15 | 15 |
|
||||
|
||||
### Null control (content-free state) and positive controls inside the spike sets, single ordering
|
||||
|
||||
| system | Cicada w1 blind /31 | Wyrd blind /84 | Cicada w1 controls /5 | Wyrd controls /28 |
|
||||
|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 16 | 60 | 5 | 21 |
|
||||
| Plumb-4B, native | 15 | 50 | 5 | 20 |
|
||||
| Plumb-4B, drop-in | 15 | 60 | 4 | 21 |
|
||||
| Imajev-4B, native | 16 | 48 | 5 | 24 |
|
||||
| Intern-Decision-4B, native | 16 | 60 | 5 | 27 |
|
||||
| Intern-Decision-4B, drop-in | 16 | 60 | 5 | 23 |
|
||||
| JevK5 v0.2, native | 15 | 53 | 5 | 22 (21–22) |
|
||||
| JevK5 v0.2, drop-in | 15 | 60 | 5 | 23 |
|
||||
| JevK5 v0.3, native | 16 | 60 | 5 | 20 |
|
||||
| JevK5 v0.3, drop-in | 16 | 60 | 5 | 22 |
|
||||
|
||||
### Paired against SemIf (each row's majority top over the repeats; group bootstrap 95% CI; exact McNemar)
|
||||
|
||||
| system | set | cond | SemIf | cand | fixed | broken | Δ pts | 95% CI | p |
|
||||
|---|---|---|---|---|---|---|---|---|---|
|
||||
| Plumb-4B, native | pooled | single | 217/259 | 224/259 | 23 | 16 | +2.7 | -2.7..+7.9 | 0.3368 |
|
||||
| Plumb-4B, native | pooled | rotations | 232/259 | 227/259 | 10 | 15 | -1.9 | -5.4..+1.5 | 0.4244 |
|
||||
| Plumb-4B, native | authored144 | single | 116/144 | 130/144 | 22 | 8 | +9.7 | +1.4..+18.1 | 0.0161 |
|
||||
| Plumb-4B, native | authored144 | rotations | 131/144 | 130/144 | 7 | 8 | -0.7 | -5.6..+3.5 | 1.0 |
|
||||
| Plumb-4B, native | cicada-w1 | single | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
|
||||
| Plumb-4B, native | cicada-w1 | rotations | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| Plumb-4B, native | wyrd | single | 71/84 | 65/84 | 1 | 7 | -7.1 | -13.1..-1.2 | 0.0703 |
|
||||
| Plumb-4B, native | wyrd | rotations | 71/84 | 67/84 | 3 | 7 | -4.8 | -11.9..+2.4 | 0.3438 |
|
||||
| Plumb-4B, native | perturbations108 | single | 83/108 | 84/108 | 10 | 9 | +0.9 | -9.3..+11.1 | 1.0 |
|
||||
| Plumb-4B, native | perturbations108 | rotations | 93/108 | 87/108 | 4 | 10 | -5.6 | -15.7..+3.7 | 0.1796 |
|
||||
| Plumb-4B, native | cicada-w2 | single | 20/31 | 19/31 | 1 | 2 | -3.2 | -12.9..+6.5 | 1.0 |
|
||||
| Plumb-4B, native | cicada-w2 | rotations | 19/31 | 18/31 | 1 | 2 | -3.2 | -16.1..+6.5 | 1.0 |
|
||||
| Plumb-4B, native | pooled | cand-single-vs-semif-rotations | 232/259 | 224/259 | 10 | 18 | -3.1 | -7.3..+0.8 | 0.1849 |
|
||||
| Plumb-4B, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 65/84 | 3 | 9 | -7.1 | -15.5..+1.2 | 0.146 |
|
||||
| Plumb-4B, drop-in | pooled | single | 217/259 | 221/259 | 21 | 17 | +1.5 | -3.7..+6.6 | 0.6271 |
|
||||
| Plumb-4B, drop-in | pooled | rotations | 232/259 | 224/259 | 7 | 15 | -3.1 | -6.8..+0.4 | 0.1338 |
|
||||
| Plumb-4B, drop-in | authored144 | single | 116/144 | 129/144 | 20 | 7 | +9.0 | +1.4..+16.7 | 0.0192 |
|
||||
| Plumb-4B, drop-in | authored144 | rotations | 131/144 | 129/144 | 5 | 7 | -1.4 | -6.2..+3.5 | 0.7744 |
|
||||
| Plumb-4B, drop-in | cicada-w1 | single | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
|
||||
| Plumb-4B, drop-in | cicada-w1 | rotations | 30/31 | 28/31 | 0 | 2 | -6.5 | -16.1..+0.0 | 0.5 |
|
||||
| Plumb-4B, drop-in | wyrd | single | 71/84 | 66/84 | 1 | 6 | -6.0 | -11.9..+0.0 | 0.125 |
|
||||
| Plumb-4B, drop-in | wyrd | rotations | 71/84 | 67/84 | 2 | 6 | -4.8 | -11.9..+2.4 | 0.2891 |
|
||||
| Plumb-4B, drop-in | perturbations108 | single | 83/108 | 88/108 | 12 | 7 | +4.6 | -5.6..+14.8 | 0.3593 |
|
||||
| Plumb-4B, drop-in | perturbations108 | rotations | 93/108 | 92/108 | 6 | 7 | -0.9 | -10.2..+7.4 | 1.0 |
|
||||
| Plumb-4B, drop-in | cicada-w2 | single | 20/31 | 20/31 | 1 | 1 | +0.0 | -9.7..+9.7 | 1.0 |
|
||||
| Plumb-4B, drop-in | cicada-w2 | rotations | 19/31 | 18/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
|
||||
| Plumb-4B, drop-in | pooled | cand-single-vs-semif-rotations | 232/259 | 221/259 | 8 | 19 | -4.2 | -8.3..-0.4 | 0.0522 |
|
||||
| Plumb-4B, drop-in | wyrd | cand-single-vs-semif-rotations | 71/84 | 66/84 | 2 | 7 | -6.0 | -13.1..+1.2 | 0.1797 |
|
||||
| Imajev-4B, native | pooled | single | 217/259 | 235/259 | 30 | 12 | +6.9 | +1.6..+12.2 | 0.0079 |
|
||||
| Imajev-4B, native | pooled | rotations | 232/259 | 236/259 | 14 | 10 | +1.5 | -2.5..+5.6 | 0.5413 |
|
||||
| Imajev-4B, native | authored144 | single | 116/144 | 138/144 | 25 | 3 | +15.3 | +9.0..+21.5 | 0.0 |
|
||||
| Imajev-4B, native | authored144 | rotations | 131/144 | 138/144 | 8 | 1 | +4.9 | +0.7..+9.0 | 0.0391 |
|
||||
| Imajev-4B, native | cicada-w1 | single | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
|
||||
| Imajev-4B, native | cicada-w1 | rotations | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
|
||||
| Imajev-4B, native | wyrd | single | 71/84 | 68/84 | 5 | 8 | -3.6 | -13.1..+6.0 | 0.5811 |
|
||||
| Imajev-4B, native | wyrd | rotations | 71/84 | 69/84 | 6 | 8 | -2.4 | -11.9..+7.1 | 0.7905 |
|
||||
| Imajev-4B, native | perturbations108 | single | 83/108 | 102/108 | 20 | 1 | +17.6 | +8.3..+27.8 | 0.0 |
|
||||
| Imajev-4B, native | perturbations108 | rotations | 93/108 | 101/108 | 9 | 1 | +7.4 | +1.9..+13.9 | 0.0215 |
|
||||
| Imajev-4B, native | cicada-w2 | single | 20/31 | 22/31 | 3 | 1 | +6.5 | -6.5..+19.4 | 0.625 |
|
||||
| Imajev-4B, native | cicada-w2 | rotations | 19/31 | 23/31 | 4 | 0 | +12.9 | +3.2..+25.8 | 0.125 |
|
||||
| Imajev-4B, native | pooled | cand-single-vs-semif-rotations | 232/259 | 235/259 | 15 | 12 | +1.2 | -3.2..+5.5 | 0.7011 |
|
||||
| Imajev-4B, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 68/84 | 6 | 9 | -3.6 | -14.3..+6.0 | 0.6072 |
|
||||
| Intern-Decision-4B, native | pooled | single | 217/259 | 240/259 | 28 | 5 | +8.9 | +4.7..+13.2 | 0.0001 |
|
||||
| Intern-Decision-4B, native | pooled | rotations | 232/259 | 236/259 | 12 | 8 | +1.5 | -2.0..+5.0 | 0.5034 |
|
||||
| Intern-Decision-4B, native | authored144 | single | 116/144 | 132/144 | 20 | 4 | +11.1 | +4.9..+18.1 | 0.0015 |
|
||||
| Intern-Decision-4B, native | authored144 | rotations | 131/144 | 130/144 | 5 | 6 | -0.7 | -5.6..+4.2 | 1.0 |
|
||||
| Intern-Decision-4B, native | cicada-w1 | single | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
|
||||
| Intern-Decision-4B, native | cicada-w1 | rotations | 30/31 | 29/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
|
||||
| Intern-Decision-4B, native | wyrd | single | 71/84 | 79/84 | 8 | 0 | +9.5 | +3.6..+16.7 | 0.0078 |
|
||||
| Intern-Decision-4B, native | wyrd | rotations | 71/84 | 77/84 | 7 | 1 | +7.1 | +1.2..+13.1 | 0.0703 |
|
||||
| Intern-Decision-4B, native | perturbations108 | single | 83/108 | 100/108 | 20 | 3 | +15.7 | +4.6..+26.9 | 0.0005 |
|
||||
| Intern-Decision-4B, native | perturbations108 | rotations | 93/108 | 99/108 | 9 | 3 | +5.6 | -2.8..+13.0 | 0.146 |
|
||||
| Intern-Decision-4B, native | cicada-w2 | single | 20/31 | 22/31 | 2 | 0 | +6.5 | +0.0..+16.1 | 0.5 |
|
||||
| Intern-Decision-4B, native | cicada-w2 | rotations | 19/31 | 22/31 | 4 | 1 | +9.7 | -3.2..+22.6 | 0.375 |
|
||||
| Intern-Decision-4B, native | pooled | cand-single-vs-semif-rotations | 232/259 | 240/259 | 14 | 6 | +3.1 | -0.4..+6.7 | 0.1153 |
|
||||
| Intern-Decision-4B, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 79/84 | 9 | 1 | +9.5 | +2.4..+16.7 | 0.0215 |
|
||||
| Intern-Decision-4B, drop-in | pooled | single | 217/259 | 232/259 | 18 | 3 | +5.8 | +2.7..+9.0 | 0.0015 |
|
||||
| Intern-Decision-4B, drop-in | pooled | rotations | 232/259 | 233/259 | 8 | 7 | +0.4 | -2.3..+3.1 | 1.0 |
|
||||
| Intern-Decision-4B, drop-in | authored144 | single | 116/144 | 128/144 | 15 | 3 | +8.3 | +3.5..+13.9 | 0.0075 |
|
||||
| Intern-Decision-4B, drop-in | authored144 | rotations | 131/144 | 130/144 | 5 | 6 | -0.7 | -4.9..+3.5 | 1.0 |
|
||||
| Intern-Decision-4B, drop-in | cicada-w1 | single | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| Intern-Decision-4B, drop-in | cicada-w1 | rotations | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| Intern-Decision-4B, drop-in | wyrd | single | 71/84 | 74/84 | 3 | 0 | +3.6 | +0.0..+8.3 | 0.25 |
|
||||
| Intern-Decision-4B, drop-in | wyrd | rotations | 71/84 | 73/84 | 3 | 1 | +2.4 | -2.4..+7.1 | 0.625 |
|
||||
| Intern-Decision-4B, drop-in | perturbations108 | single | 83/108 | 102/108 | 20 | 1 | +17.6 | +7.4..+28.7 | 0.0 |
|
||||
| Intern-Decision-4B, drop-in | perturbations108 | rotations | 93/108 | 102/108 | 9 | 0 | +8.3 | +2.8..+15.7 | 0.0039 |
|
||||
| Intern-Decision-4B, drop-in | cicada-w2 | single | 20/31 | 19/31 | 0 | 1 | -3.2 | -9.7..+0.0 | 1.0 |
|
||||
| Intern-Decision-4B, drop-in | cicada-w2 | rotations | 19/31 | 19/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| Intern-Decision-4B, drop-in | pooled | cand-single-vs-semif-rotations | 232/259 | 232/259 | 10 | 10 | +0.0 | -3.0..+3.1 | 1.0 |
|
||||
| Intern-Decision-4B, drop-in | wyrd | cand-single-vs-semif-rotations | 71/84 | 74/84 | 4 | 1 | +3.6 | -1.2..+8.3 | 0.375 |
|
||||
| JevK5 v0.2, native | pooled | single | 217/259 | 229/259 | 25 | 13 | +4.6 | -0.8..+9.9 | 0.073 |
|
||||
| JevK5 v0.2, native | pooled | rotations | 232/259 | 228/259 | 9 | 13 | -1.5 | -4.9..+1.9 | 0.5235 |
|
||||
| JevK5 v0.2, native | authored144 | single | 116/144 | 132/144 | 23 | 7 | +11.1 | +2.8..+18.8 | 0.0052 |
|
||||
| JevK5 v0.2, native | authored144 | rotations | 131/144 | 131/144 | 7 | 7 | +0.0 | -4.9..+4.9 | 1.0 |
|
||||
| JevK5 v0.2, native | cicada-w1 | single | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| JevK5 v0.2, native | cicada-w1 | rotations | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| JevK5 v0.2, native | wyrd | single | 71/84 | 67/84 | 2 | 6 | -4.8 | -10.7..+1.2 | 0.2891 |
|
||||
| JevK5 v0.2, native | wyrd | rotations | 71/84 | 67/84 | 2 | 6 | -4.8 | -10.7..+1.2 | 0.2891 |
|
||||
| JevK5 v0.2, native | perturbations108 | single | 83/108 | 85/108 | 10 | 8 | +1.9 | -7.4..+11.1 | 0.8145 |
|
||||
| JevK5 v0.2, native | perturbations108 | rotations | 93/108 | 91/108 | 4 | 6 | -1.9 | -10.2..+5.6 | 0.7539 |
|
||||
| JevK5 v0.2, native | cicada-w2 | single | 20/31 | 16/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
|
||||
| JevK5 v0.2, native | cicada-w2 | rotations | 19/31 | 16/31 | 0 | 3 | -9.7 | -19.4..+0.0 | 0.25 |
|
||||
| JevK5 v0.2, native | pooled | cand-single-vs-semif-rotations | 232/259 | 229/259 | 11 | 14 | -1.2 | -5.0..+2.6 | 0.69 |
|
||||
| JevK5 v0.2, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 67/84 | 3 | 7 | -4.8 | -13.1..+2.4 | 0.3438 |
|
||||
| JevK5 v0.2, drop-in | pooled | single | 217/259 | 232/259 | 24 | 9 | +5.8 | +1.1..+10.7 | 0.0135 |
|
||||
| JevK5 v0.2, drop-in | pooled | rotations | 232/259 | 237/259 | 11 | 6 | +1.9 | -1.1..+5.1 | 0.3323 |
|
||||
| JevK5 v0.2, drop-in | authored144 | single | 116/144 | 130/144 | 21 | 7 | +9.7 | +2.1..+17.4 | 0.0125 |
|
||||
| JevK5 v0.2, drop-in | authored144 | rotations | 131/144 | 133/144 | 6 | 4 | +1.4 | -2.8..+5.6 | 0.7539 |
|
||||
| JevK5 v0.2, drop-in | cicada-w1 | single | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| JevK5 v0.2, drop-in | cicada-w1 | rotations | 30/31 | 30/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| JevK5 v0.2, drop-in | wyrd | single | 71/84 | 72/84 | 3 | 2 | +1.2 | -3.6..+6.0 | 1.0 |
|
||||
| JevK5 v0.2, drop-in | wyrd | rotations | 71/84 | 74/84 | 5 | 2 | +3.6 | -2.4..+10.7 | 0.4531 |
|
||||
| JevK5 v0.2, drop-in | perturbations108 | single | 83/108 | 92/108 | 13 | 4 | +8.3 | -0.9..+17.6 | 0.049 |
|
||||
| JevK5 v0.2, drop-in | perturbations108 | rotations | 93/108 | 93/108 | 4 | 4 | +0.0 | -7.4..+6.5 | 1.0 |
|
||||
| JevK5 v0.2, drop-in | cicada-w2 | single | 20/31 | 17/31 | 0 | 3 | -9.7 | -22.6..+0.0 | 0.25 |
|
||||
| JevK5 v0.2, drop-in | cicada-w2 | rotations | 19/31 | 16/31 | 0 | 3 | -9.7 | -19.4..+0.0 | 0.25 |
|
||||
| JevK5 v0.2, drop-in | pooled | cand-single-vs-semif-rotations | 232/259 | 232/259 | 10 | 10 | +0.0 | -3.2..+3.2 | 1.0 |
|
||||
| JevK5 v0.2, drop-in | wyrd | cand-single-vs-semif-rotations | 71/84 | 72/84 | 3 | 2 | +1.2 | -3.6..+6.0 | 1.0 |
|
||||
| JevK5 v0.3, native | pooled | single | 217/259 | 227/259 | 25 | 15 | +3.9 | -0.9..+8.5 | 0.1539 |
|
||||
| JevK5 v0.3, native | pooled | rotations | 232/259 | 226/259 | 9 | 15 | -2.3 | -6.0..+1.4 | 0.3075 |
|
||||
| JevK5 v0.3, native | authored144 | single | 116/144 | 135/144 | 24 | 5 | +13.2 | +6.9..+19.4 | 0.0005 |
|
||||
| JevK5 v0.3, native | authored144 | rotations | 131/144 | 135/144 | 7 | 3 | +2.8 | -1.4..+7.6 | 0.3438 |
|
||||
| JevK5 v0.3, native | cicada-w1 | single | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
|
||||
| JevK5 v0.3, native | cicada-w1 | rotations | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
|
||||
| JevK5 v0.3, native | wyrd | single | 71/84 | 66/84 | 1 | 6 | -6.0 | -10.7..-1.2 | 0.125 |
|
||||
| JevK5 v0.3, native | wyrd | rotations | 71/84 | 65/84 | 2 | 8 | -7.1 | -13.1..-1.2 | 0.1094 |
|
||||
| JevK5 v0.3, native | perturbations108 | single | 83/108 | 101/108 | 18 | 0 | +16.7 | +8.3..+25.9 | 0.0 |
|
||||
| JevK5 v0.3, native | perturbations108 | rotations | 93/108 | 101/108 | 8 | 0 | +7.4 | +1.9..+13.9 | 0.0078 |
|
||||
| JevK5 v0.3, native | cicada-w2 | single | 20/31 | 17/31 | 0 | 3 | -9.7 | -19.4..+0.0 | 0.25 |
|
||||
| JevK5 v0.3, native | cicada-w2 | rotations | 19/31 | 16/31 | 0 | 3 | -9.7 | -19.4..+0.0 | 0.25 |
|
||||
| JevK5 v0.3, native | pooled | cand-single-vs-semif-rotations | 232/259 | 227/259 | 11 | 16 | -1.9 | -5.5..+1.5 | 0.4421 |
|
||||
| JevK5 v0.3, native | wyrd | cand-single-vs-semif-rotations | 71/84 | 66/84 | 2 | 7 | -6.0 | -11.9..+0.0 | 0.1797 |
|
||||
| JevK5 v0.3, drop-in | pooled | single | 217/259 | 227/259 | 24 | 14 | +3.9 | -0.8..+8.2 | 0.1433 |
|
||||
| JevK5 v0.3, drop-in | pooled | rotations | 232/259 | 227/259 | 10 | 15 | -1.9 | -5.5..+1.8 | 0.4244 |
|
||||
| JevK5 v0.3, drop-in | authored144 | single | 116/144 | 132/144 | 22 | 6 | +11.1 | +4.9..+17.4 | 0.0037 |
|
||||
| JevK5 v0.3, drop-in | authored144 | rotations | 131/144 | 135/144 | 8 | 4 | +2.8 | -1.4..+7.6 | 0.3877 |
|
||||
| JevK5 v0.3, drop-in | cicada-w1 | single | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
|
||||
| JevK5 v0.3, drop-in | cicada-w1 | rotations | 30/31 | 26/31 | 0 | 4 | -12.9 | -25.8..-3.2 | 0.125 |
|
||||
| JevK5 v0.3, drop-in | wyrd | single | 71/84 | 69/84 | 2 | 4 | -2.4 | -7.1..+2.4 | 0.6875 |
|
||||
| JevK5 v0.3, drop-in | wyrd | rotations | 71/84 | 66/84 | 2 | 7 | -6.0 | -11.9..+0.0 | 0.1797 |
|
||||
| JevK5 v0.3, drop-in | perturbations108 | single | 83/108 | 99/108 | 17 | 1 | +14.8 | +6.5..+24.1 | 0.0001 |
|
||||
| JevK5 v0.3, drop-in | perturbations108 | rotations | 93/108 | 101/108 | 8 | 0 | +7.4 | +1.9..+13.9 | 0.0078 |
|
||||
| JevK5 v0.3, drop-in | cicada-w2 | single | 20/31 | 21/31 | 1 | 0 | +3.2 | +0.0..+9.7 | 1.0 |
|
||||
| JevK5 v0.3, drop-in | cicada-w2 | rotations | 19/31 | 19/31 | 0 | 0 | +0.0 | +0.0..+0.0 | 1.0 |
|
||||
| JevK5 v0.3, drop-in | pooled | cand-single-vs-semif-rotations | 232/259 | 227/259 | 10 | 15 | -1.9 | -5.2..+1.2 | 0.4244 |
|
||||
| JevK5 v0.3, drop-in | wyrd | cand-single-vs-semif-rotations | 71/84 | 69/84 | 3 | 5 | -2.4 | -8.3..+3.6 | 0.7266 |
|
||||
|
||||
### Noise floors
|
||||
|
||||
| system | A-vs-A in process (authored144): flips, max Δp | across restarts, single (560 rows/pair): flips per pair, max Δp | across restarts, rotations: flips per pair | labelled rows whose top moved in ANY restart pair, single: pooled /259 · perturb /108 · Cicada w2 /31 |
|
||||
|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0, 0, 0, 0 | 0 · 0 · 0 (4 repeats) |
|
||||
| Plumb-4B, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 2 (Δp≤0.043), 2 (Δp≤0.043) | 0, 2, 2 | 2 · 0 · 0 (3 repeats) |
|
||||
| Plumb-4B, drop-in | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0 | 0 · 0 · 0 (3 repeats) |
|
||||
| Imajev-4B, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0, 0, 0, 0 | 0 · 0 · 0 (4 repeats) |
|
||||
| Intern-Decision-4B, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0, 0, 0, 0 | 0 · 0 · 0 (4 repeats) |
|
||||
| Intern-Decision-4B, drop-in | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.128), 0 (Δp≤0.128), 0 (Δp≤0.000) | 0, 0, 0 | 0 · 0 · 0 (3 repeats) |
|
||||
| JevK5 v0.2, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 3 (Δp≤0.036), 3 (Δp≤0.036) | 0, 2, 2 | 1 · 1 · 0 (3 repeats) |
|
||||
| JevK5 v0.2, drop-in | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0 | 0 · 0 · 0 (3 repeats) |
|
||||
| JevK5 v0.3, native | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 1 (Δp≤0.051), 1 (Δp≤0.051), 0 (Δp≤0.000) | 1, 1, 0 | 0 · 1 · 0 (3 repeats) |
|
||||
| JevK5 v0.3, drop-in | 0/144 Δp≤0.000, 0/144 Δp≤0.000, 0/144 Δp≤0.000 | 0 (Δp≤0.000), 0 (Δp≤0.000), 0 (Δp≤0.000) | 0, 0, 0 | 0 · 0 · 0 (3 repeats) |
|
||||
|
||||
### Order sensitivity (authored144, single ordering vs the same request reordered) and the negative control
|
||||
|
||||
| system | reversed: label changes /144 | reversed: max Δp | shuffled: label changes | shuffled: max Δp | NEG: same top as unrotated /144 | NEG: follows the description | NEG: right vs original gold |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 30 | 0.89 | 29 | 0.89 | 14 | 119 | 17 |
|
||||
| Plumb-4B, native | 15 | 0.54 (0.54–0.56) | 10 | 0.44 (0.44–0.45) | 24 (24–25) | 106 (105–106) | 27 (27–28) |
|
||||
| Plumb-4B, drop-in | 12 | 0.73 | 8 | 0.67 | 3 | 126 | 12 |
|
||||
| Imajev-4B, native | 3 | 0.72 | 4 | 0.68 | 4 | 135 | 5 |
|
||||
| Intern-Decision-4B, native | 8 | 0.60 | 5 | 0.46 | 10 | 122 | 14 |
|
||||
| Intern-Decision-4B, drop-in | 10 | 0.98 | 12 (11–12) | 0.94 (0.93–0.94) | 5 | 127 | 10 |
|
||||
| JevK5 v0.2, native | 13 (12–13) | 0.72 | 10 | 0.64 (0.63–0.64) | 20 | 114 | 21 |
|
||||
| JevK5 v0.2, drop-in | 15 | 0.70 | 12 | 0.70 | 4 | 128 | 9 |
|
||||
| JevK5 v0.3, native | 8 (7–8) | 0.61 | 11 | 0.49 | 20 | 118 | 18 |
|
||||
| JevK5 v0.3, drop-in | 4 | 0.53 | 5 | 0.51 | 4 | 133 | 5 |
|
||||
|
||||
### Latency at our shape (loopback on fv-ml1, GPU 3; ms; median of all requests, run-median range)
|
||||
|
||||
| system | short1 e2e | crit21 e2e | crit21 server | long1 e2e | long16 e2e | long16 server | tokens crit21 / long16 |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 37 (37–38) | 131 (131–132) | 119 (119–120) | 186 (186–187) | 505 (502–519) | 486 (482–500) | 62 / 3900 |
|
||||
| Plumb-4B, native | 16 | 294 (292–295) | 292 (290–293) | 215 (212–215) | 3378 (3374–3384) | 3376 (3372–3382) | 2489 / 63302 |
|
||||
| Plumb-4B, drop-in | 37 | 131 (131–132) | 120 (119–120) | 186 | 506 (506–507) | 487 (487–488) | 62 / 3900 |
|
||||
| Imajev-4B, native | 60 (60–61) | 1219 (1212–1230) | 1212 (1204–1223) | 315 (310–315) | 5125 (5054–5140) | 5121 (5049–5135) | 374 / 7925 |
|
||||
| Intern-Decision-4B, native | 39 | 88 (88–89) | 84 (84–85) | 190 (189–191) | 215 (213–216) | 213 (211–214) | 1083 / 4579 |
|
||||
| Intern-Decision-4B, drop-in | 38 (37–38) | 132 | 120 | 186 (186–188) | 508 (500–516) | 488 (481–497) | 62 / 3900 |
|
||||
| JevK5 v0.2, native | 16 | 291 (291–292) | 289 (289–290) | 212 | 3360 (3358–3368) | 3358 (3356–3366) | 2489 / 63302 |
|
||||
| JevK5 v0.2, drop-in | 38 | 132 | 120 (119–120) | 187 (185–188) | 518 (505–519) | 499 (486–500) | 62 / 3900 |
|
||||
| JevK5 v0.3, native | 16 | 291 (291–292) | 289 (289–290) | 212 | 3363 (3355–3370) | 3361 (3353–3368) | 2489 / 63302 |
|
||||
| JevK5 v0.3, drop-in | 38 (37–39) | 133 (131–133) | 120 (119–121) | 187 (186–188) | 518 (506–521) | 498 (487–501) | 62 / 3900 |
|
||||
|
||||
### VRAM (nvidia-smi, whole GPU 3, only our process on it) and capacity under the 12 GiB cap
|
||||
|
||||
| system | rest after warm-up MiB | peak in sets+shapes MiB | peak in capacity sweep MiB | max rows @ short state | max rows @ ~3,900 tok |
|
||||
|---|---|---|---|---|---|
|
||||
| SemIf (Qwen3.5-4B), baseline | 9242 (9242–9250) | 12918 (12886–12918) | 12920 | 64 (next 96: 422) | 16 (next 18: 503) |
|
||||
| Plumb-4B, native | 12064 | 12064 | 12064 | 128 (no failure up to the last step) | 64 (no failure up to the last step) |
|
||||
| Plumb-4B, drop-in | 9242 | 12918 | 12920 | 64 (next 96: 422) | 16 (next 18: 503) |
|
||||
| Imajev-4B, native | 10404 | 11056 | 11056 | 128 (no failure up to the last step) | 64 (no failure up to the last step) |
|
||||
| Intern-Decision-4B, native | 9736 | 10290 | 10290 | 128 (no failure up to the last step) | 64 (no failure up to the last step) |
|
||||
| Intern-Decision-4B, drop-in | 9242 | 12918 | 12920 | 64 (next 96: 422) | 16 (next 18: 503) |
|
||||
| JevK5 v0.2, native | 12064 | 12064 | 12064 | 128 (no failure up to the last step) | 64 (no failure up to the last step) |
|
||||
| JevK5 v0.2, drop-in | 9242 | 12918 | – | – | – |
|
||||
| JevK5 v0.3, native | 12064 | 12064 | – | – | – |
|
||||
| JevK5 v0.3, drop-in | 9242 | 12918 | – | – | – |
|
||||
@@ -151,6 +151,12 @@ the fast kernels autotune at startup. That is n=1 row across one restart, and th
|
||||
floor is otherwise unmeasured. When comparing two versions, compare them against that floor
|
||||
and not against zero.
|
||||
|
||||
**Update 2026-09-30 (the Jev bench on GPU 3, `docs/pfi/jev-candidates-bench-2026-09-30.md`):**
|
||||
across **4 restarts** of the same image, SemIf changed **0 labels**, and its probabilities were
|
||||
bit-identical on every set, the 144 authored rows included. So the near-tie flip described above
|
||||
did not reproduce. The cross-restart floor is now measured at 0 for n=4 restarts on an idle
|
||||
card. Treat the one flip on 2026-09-27 as possible but rare, not as the norm.
|
||||
|
||||
## Acceptance (2026-09-27, v0.1.3)
|
||||
|
||||
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json`,
|
||||
|
||||
Reference in New Issue
Block a user