fix(semif): 0.1.4 — object states ending in ) ; } no longer 422 (INV-7)
SemIf's shared scorer trims one token at the state boundary. When an object state's last value ends in ')', ';' or '}', the JSON that follows re-merges two tokens back, so score_shared refused the request with 422. The engine now wraps semif_phase1.shared._state_prefix to keep only the tokens the full prompts share. Each row scores the same token sequence; only the prefill/suffix split moves. Startup proves the fix is in effect, not just installed (heid bug hunt SKAL, folded). It checks that the hook is callable and is what score_shared resolves, that an ordinary state keeps upstream's whole prefix, and that a merge-prone state scores through the shared path. Real tokenizer: 154 states, 23 refused before and 0 after, with no ordinary or authored144 prefix changed. Acceptance: 144/144 parity. Shared vs direct 71/72; the miss is a bf16 tie that flipped across a plain restart (see README).
This commit is contained in:
@@ -104,6 +104,26 @@ Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from
|
||||
- **INV-5 no network at runtime.** Weights come from the mounted HF cache at the
|
||||
pinned revision. The entry point sets `HF_HUB_OFFLINE=1` itself before torch or
|
||||
transformers load, so this holds outside the image too (S10).
|
||||
- **INV-7 boundary-safe shared prefix (0.1.4).** At load, the engine wraps SemIf's
|
||||
module-global `semif_phase1.shared._state_prefix`, which `score_shared` calls. The wrapper
|
||||
keeps only the leading tokens that upstream's prefix shares with a real full prompt for
|
||||
the same state (a probe row: evidence, then a placeholder criterion).
|
||||
- **Why:** upstream drops just one token at the state boundary. An object state whose
|
||||
last value ends in `)`, `;` or `}` re-tokenises two tokens back once `, "criterion"`
|
||||
follows, so it was refused with 422 "The fixed state prefix does not match every full
|
||||
prompt" (found 2026-09-27).
|
||||
- **Effect:** each row still scores the same token sequence; only the prefill/suffix split
|
||||
moves, costing a few tokens of sharing. An ordinary state keeps upstream's prefix
|
||||
unchanged. `score_shared` still checks every real row and fails closed.
|
||||
- **Startup proves the fix is in effect, not just installed** (heid bug hunt SKAL,
|
||||
2026-09-27). Before the weights load, it refuses to start if `_state_prefix` is
|
||||
missing or not callable, or if `score_shared` does not resolve it from
|
||||
`semif_phase1.shared`'s globals (a SemIf bump that moves or re-exports it). After the
|
||||
warm-up, it checks two things. The wrapper must return upstream's exact prefix for
|
||||
the warm-up state, since a wrapper rendering the wrong prompt would silently drop all
|
||||
sharing. And a state upstream alone refuses, `{"person_said": "ok :)"}`, must score
|
||||
through `score_shared` itself; a hook bound before the patch would fail here. A
|
||||
reload wraps the original again rather than stacking wrappers.
|
||||
- **INV-6 constant-time auth.** Token comparison uses `hmac.compare_digest`. The
|
||||
token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything
|
||||
else, because a CR, LF or NUL in the token can never arrive in a header (S2).
|
||||
@@ -169,7 +189,13 @@ cancels a position bias exactly; `all` sends n! rows and is 422 above 4 options;
|
||||
agreement and spread are computed from the orderings; a mixed shared request
|
||||
(averaged + plain) is one engine call, with results in request order and plain
|
||||
results unchanged; expanded rows count toward the cap; `workload` + `orderings`
|
||||
→ 422.
|
||||
→ 422. **Prefix (INV-7):** against a tokenizer whose merge reaches two tokens back, the
|
||||
wrapped prefix is a token-prefix of every row's full prompt and gives up exactly the
|
||||
mismatched token; an ordinary state keeps the whole upstream prefix; load() installs
|
||||
one wrapper however many times it runs, and refuses to start when `_state_prefix` is
|
||||
gone or not callable, when `score_shared` binds it early or resolves its globals
|
||||
elsewhere, or when the wrapper renders the wrong prompt. The fake `score_shared` is
|
||||
compiled into the fake module, so that it resolves its globals as the real one does.
|
||||
|
||||
## Acceptance (on fv-ml1, real model; not unit tests)
|
||||
|
||||
@@ -180,3 +206,6 @@ results unchanged; expanded rows count toward the cap; `workload` + `orderings`
|
||||
4. **Shared vs direct:** the same rows agree within the A-vs-A floor.
|
||||
5. **Speed:** 21 binary criteria over one state, N ≥ 3, p50 + spread.
|
||||
6. **VRAM:** the peak at a 4096-token input sets `SEMIF_VRAM_CAP_GIB`.
|
||||
7. **Boundary (0.1.4):** states whose last value ends in `)`, `;` and `}`, as objects,
|
||||
are all answered by `/decide/shared`, with the same top choice as a string state
|
||||
holding the same text; parity (1) still holds.
|
||||
|
||||
Reference in New Issue
Block a user