fix(semif): 0.1.4 — object states ending in ) ; } no longer 422 (INV-7)

SemIf's shared scorer trims one token at the state boundary. When an object
state's last value ends in ')', ';' or '}', the JSON that follows re-merges two
tokens back, so score_shared refused the request with 422. The engine now wraps
semif_phase1.shared._state_prefix to keep only the tokens the full prompts
share. Each row scores the same token sequence; only the prefill/suffix split
moves.

Startup proves the fix is in effect, not just installed (heid bug hunt SKAL,
folded). It checks that the hook is callable and is what score_shared resolves,
that an ordinary state keeps upstream's whole prefix, and that a merge-prone
state scores through the shared path.

Real tokenizer: 154 states, 23 refused before and 0 after, with no ordinary or
authored144 prefix changed. Acceptance: 144/144 parity. Shared vs direct 71/72;
the miss is a bf16 tie that flipped across a plain restart (see README).
This commit is contained in:
vh
2026-09-27 10:23:14 -07:00
parent 7e11cf247b
commit 47cad33dd1
11 changed files with 384 additions and 20 deletions
@@ -0,0 +1,63 @@
{
"url": "http://10.251.50.54:8032",
"health": {
"status": "ok",
"semif_commit": "23cf1f39fc9534fe81437200959b6dfc7106e45a",
"model": {
"source": "Qwen/Qwen3.5-4B",
"revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
"dtype": "bfloat16",
"device": "cuda:0",
"torch_version": "2.10.0+cu128",
"transformers_version": "5.17.0",
"device_name": "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition",
"allocated_gib": 7.84,
"reserved_gib": 8.12
},
"vram_cap_gib": 12.0,
"max_tokens": 4096,
"max_decisions": 64,
"workloads": []
},
"1_parity_vs_upstream": {
"rows": 144,
"top_choice_agree": 144,
"max_abs_prob_gap": 0.0595381559570336
},
"1_prompt_sha256_equal": 144,
"2_noise_floor_a_vs_b": {
"rows": 144,
"top_choice_agree": 144,
"max_abs_prob_gap": 0.0
},
"3_negative_rotated_options": {
"rows": 144,
"top_choice_agree": 14,
"max_abs_prob_gap": 0.9987391090270772
},
"4_shared_vs_direct": {
"groups": 36,
"rows": 72,
"top_choice_agree": 71,
"max_abs_prob_gap": 0.05927145076874801
},
"5_speed_21_binary": {
"prefix_tokens": 62,
"shared_s": {
"runs": [
0.13865972400526516,
0.1382300010009203,
0.13749008599552326
],
"median": 0.1382300010009203
},
"sequential_decide_s": {
"runs": [
0.9489945390087087,
0.9463060130074155,
0.9458249929884914
],
"median": 0.9463060130074155
}
}
}
+1 -1
View File
@@ -1,6 +1,6 @@
[project]
name = "semif-serve"
version = "0.1.3"
version = "0.1.4"
description = "HTTP wrapper around SemIf's direct and shared option-logit scorers"
requires-python = ">=3.12"
dependencies = [
+30 -1
View File
@@ -104,6 +104,26 @@ Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from
- **INV-5 no network at runtime.** Weights come from the mounted HF cache at the
pinned revision. The entry point sets `HF_HUB_OFFLINE=1` itself before torch or
transformers load, so this holds outside the image too (S10).
- **INV-7 boundary-safe shared prefix (0.1.4).** At load, the engine wraps SemIf's
module-global `semif_phase1.shared._state_prefix`, which `score_shared` calls. The wrapper
keeps only the leading tokens that upstream's prefix shares with a real full prompt for
the same state (a probe row: evidence, then a placeholder criterion).
- **Why:** upstream drops just one token at the state boundary. An object state whose
last value ends in `)`, `;` or `}` re-tokenises two tokens back once `, "criterion"`
follows, so it was refused with 422 "The fixed state prefix does not match every full
prompt" (found 2026-09-27).
- **Effect:** each row still scores the same token sequence; only the prefill/suffix split
moves, costing a few tokens of sharing. An ordinary state keeps upstream's prefix
unchanged. `score_shared` still checks every real row and fails closed.
- **Startup proves the fix is in effect, not just installed** (heid bug hunt SKAL,
2026-09-27). Before the weights load, it refuses to start if `_state_prefix` is
missing or not callable, or if `score_shared` does not resolve it from
`semif_phase1.shared`'s globals (a SemIf bump that moves or re-exports it). After the
warm-up, it checks two things. The wrapper must return upstream's exact prefix for
the warm-up state, since a wrapper rendering the wrong prompt would silently drop all
sharing. And a state upstream alone refuses, `{"person_said": "ok :)"}`, must score
through `score_shared` itself; a hook bound before the patch would fail here. A
reload wraps the original again rather than stacking wrappers.
- **INV-6 constant-time auth.** Token comparison uses `hmac.compare_digest`. The
token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything
else, because a CR, LF or NUL in the token can never arrive in a header (S2).
@@ -169,7 +189,13 @@ cancels a position bias exactly; `all` sends n! rows and is 422 above 4 options;
agreement and spread are computed from the orderings; a mixed shared request
(averaged + plain) is one engine call, with results in request order and plain
results unchanged; expanded rows count toward the cap; `workload` + `orderings`
→ 422.
→ 422. **Prefix (INV-7):** against a tokenizer whose merge reaches two tokens back, the
wrapped prefix is a token-prefix of every row's full prompt and gives up exactly the
mismatched token; an ordinary state keeps the whole upstream prefix; load() installs
one wrapper however many times it runs, and refuses to start when `_state_prefix` is
gone or not callable, when `score_shared` binds it early or resolves its globals
elsewhere, or when the wrapper renders the wrong prompt. The fake `score_shared` is
compiled into the fake module, so that it resolves its globals as the real one does.
## Acceptance (on fv-ml1, real model; not unit tests)
@@ -180,3 +206,6 @@ results unchanged; expanded rows count toward the cap; `workload` + `orderings`
4. **Shared vs direct:** the same rows agree within the A-vs-A floor.
5. **Speed:** 21 binary criteria over one state, N ≥ 3, p50 + spread.
6. **VRAM:** the peak at a 4096-token input sets `SEMIF_VRAM_CAP_GIB`.
7. **Boundary (0.1.4):** states whose last value ends in `)`, `;` and `}`, as objects,
are all answered by `/decide/shared`, with the same top choice as a string state
holding the same text; parity (1) still holds.
+66 -1
View File
@@ -28,6 +28,44 @@ WARMUP_ROW = {
}
# INV-7 startup proof: a state SemIf's own prefix refuses (measured with the real tokenizer,
# 2026-09-27), sent through score_shared itself, so the fix is shown to be IN EFFECT, not just
# installed (bug hunt SKAL, R2/H1/H2).
BOUNDARY_ROW = {**WARMUP_ROW, "id": "semif-serve-boundary-check", "state": {"person_said": "ok :)"}}
PREFIX_PROBE_ROW = {
"id": "semif-serve-prefix-probe",
"question": "prefix boundary placeholder",
"options": [{"id": "yes", "description": "Yes"}, {"id": "no", "description": "No"}],
}
def boundary_safe_prefix(state_prefix: Callable, messages: Callable) -> Callable:
"""INV-7: wrap SemIf's `shared._state_prefix` so the prefill never covers a token that the full
prompts do not share.
Upstream encodes the prompt up to the end of the state and drops ONE token, because the JSON
punctuation that follows the state can merge with it. One is not always enough: an object
state whose last value ends in ")", ";" or "}" re-tokenises TWO tokens back once `, "criterion"`
follows, and score_shared then refused the request with 422 (found 2026-09-27). This keeps only
the leading tokens that upstream's prefix shares with a real full prompt for the same state.
The suffix starts that much earlier and scores the same token sequence; the cost is a few
tokens of lost sharing. Only the punctuation run at the boundary can merge, and it is the same
in every row, so a probe row stands for all of them. score_shared still checks every real row
and fails closed if that ever stops holding."""
def prefix(tokenizer, state):
ids = state_prefix(tokenizer, state)
full = tokenizer.encode(tokenizer.apply_chat_template(
messages({**PREFIX_PROBE_ROW, "state": state}), tokenize=False, add_generation_prompt=True,
enable_thinking=False), add_special_tokens=False)
shared = 0
while shared < min(len(ids), len(full)) and ids[shared] == full[shared]:
shared += 1
return ids[:shared]
prefix.semif_serve_wraps = state_prefix
return prefix
def _first_line(exc: BaseException) -> str:
lines = str(exc).splitlines()
return lines[0] if lines else ""
@@ -44,10 +82,24 @@ class TorchEngine:
@classmethod
def load(cls, settings: Settings) -> "TorchEngine":
import torch
from semif_phase1.core import load_causal_model
import semif_phase1.shared as upstream_shared
from semif_phase1.core import direct_messages, load_causal_model
from semif_phase1.direct import score
from semif_phase1.shared import score_shared
# INV-7: score_shared looks `_state_prefix` up as a module global, so the wrapper goes there.
# A SemIf bump that renames it, or moves score_shared so it resolves its globals elsewhere,
# must stop startup, not silently bring the 422 back.
original = getattr(upstream_shared, "_state_prefix", None)
if not callable(original):
raise RuntimeError("semif_phase1.shared._state_prefix is gone or not callable at this SemIf "
"commit: re-check the boundary-safe prefix (INV-7) before serving")
original = getattr(original, "semif_serve_wraps", original) # one wrapper, however many loads
upstream_shared._state_prefix = boundary_safe_prefix(original, direct_messages)
if getattr(score_shared, "__globals__", {}).get("_state_prefix") is not upstream_shared._state_prefix:
raise RuntimeError("INV-7: score_shared does not resolve semif_phase1.shared._state_prefix, "
"so the boundary-safe prefix would be inert")
if settings.device == "cuda":
if not torch.cuda.is_available():
raise RuntimeError("SEMIF_DEVICE=cuda but torch sees no CUDA device")
@@ -70,10 +122,23 @@ class TorchEngine:
raise RuntimeError(f"model landed on {placed}, expected {settings.device}")
engine = cls(torch, model, tokenizer, metadata, settings, direct_fn=score, shared_fn=score_shared)
engine.direct(WARMUP_ROW) # INV-3: one decision must score
engine._prove_prefix_hook(upstream_shared._state_prefix, original)
if settings.device == "cuda": # INV-4: the resting footprint
engine._release_above = torch.cuda.memory_reserved(0) + RELEASE_SLACK_BYTES
return engine
def _prove_prefix_hook(self, hook: Callable, original: Callable) -> None:
"""INV-7, at startup: the wrapper keeps upstream's whole prefix on an ordinary state (a wrapper
rendering the wrong prompt would silently give up all sharing), and a state upstream alone
refuses scores through score_shared itself (a hook score_shared never calls would not)."""
state = WARMUP_ROW["state"]
if hook(self._tokenizer, state) != original(self._tokenizer, state):
raise RuntimeError("INV-7: the boundary-safe prefix does not keep upstream's prefix on an ordinary state")
try:
self.shared([BOUNDARY_ROW])
except ValueError as exc:
raise RuntimeError(f"INV-7: shared scoring still refuses a merge-prone state: {exc}") from None
def health(self) -> dict:
info = dict(self._metadata)
if self._settings.device == "cuda":
@@ -0,0 +1,37 @@
"""Test doubles for SemIf's tokenizer seam (INV-7), shared by test_prefix and test_engine_load.
MergeTokenizer reproduces the real failure's shape without the Qwen vocabulary; the real
tokenizer is checked on the card, in acceptance."""
import json
class MergeTokenizer:
"""Char-level with three merges, in the shape of the real failure. ')"' and '"}' are single
tokens, but ')"},' splits as [')', '"},']. So a state tail ')"}' encodes as [')"', '}'] when
nothing follows it (the prefix), and the full prompt's ')"},' encodes as [')', '"},'].
Dropping one token from the prefix leaves ')"', which the full prompt does not contain. A
tail like '."}' stays ['.', '"}'] either way, which is the ordinary case."""
def apply_chat_template(self, messages, tokenize, add_generation_prompt, enable_thinking):
assert tokenize is False and add_generation_prompt is True and enable_thinking is False
return "<s>" + "|".join(m["content"] for m in messages) + "<a>"
def encode(self, text, add_special_tokens):
assert add_special_tokens is False
out, i = [], 0
while i < len(text):
if text.startswith(')"},', i):
out += [')', '"},']; i += 4
elif text.startswith(')"', i) or text.startswith('"}', i):
out.append(text[i:i + 2]); i += 2
else:
out.append(text[i]); i += 1
return out
def messages(row): # the shape of semif_phase1.core.direct_messages: evidence first, then the criterion
return [{"role": "user", "content": json.dumps(
{"evidence": row["state"], "criterion": row["question"], "options": row["options"]}, ensure_ascii=False)}]
def upstream_prefix(tokenizer, state): # what SemIf's _state_prefix does: through the state, minus one token
return tokenizer.encode("<s>" + json.dumps({"evidence": state}, ensure_ascii=False)[:-1], add_special_tokens=False)[:-1]
+100 -3
View File
@@ -6,6 +6,8 @@ import types
import pytest
from fake_tokenizer import MergeTokenizer, messages, upstream_prefix
from semif_serve.config import Settings
TOKEN = "t" * 40
@@ -24,6 +26,25 @@ class Model:
yield Param(self._device)
# The shape of SemIf's score_shared, compiled INTO the fake module so that, like the real one, it
# resolves _state_prefix and direct_messages as module globals at call time and refuses (ValueError)
# a prefix that is not a token-prefix of every row.
SHARED_SRC = """
def score_shared(model, tokenizer, rows, metadata, max_tokens=4096):
prefix = _state_prefix(tokenizer, rows[0]["state"])
for row in rows:
ids = tokenizer.encode(tokenizer.apply_chat_template(direct_messages(row), tokenize=False,
add_generation_prompt=True, enable_thinking=False), add_special_tokens=False)
if not prefix or ids[:len(prefix)] != prefix:
raise ValueError("The fixed state prefix does not match every full prompt")
CALLS.append(("shared", rows[0]["state"]))
return [{"id": row["id"]} for row in rows], {}
"""
# A SemIf revision that binds the helper when score_shared is defined: the module global can be
# replaced all day and scoring never sees it.
SHARED_SRC_EARLY_BOUND = SHARED_SRC.replace("max_tokens=4096):", "max_tokens=4096, _state_prefix=_state_prefix):")
@pytest.fixture
def fakes(monkeypatch):
calls = []
@@ -42,7 +63,7 @@ def fakes(monkeypatch):
def load_causal_model(model, revision, device, dtype):
calls.append(("load", model, revision, device, dtype))
return Model(state["device"]), object(), {"source": model}
return Model(state["device"]), MergeTokenizer(), {"source": model}
def score(model, tok, row, meta, max_tokens):
calls.append(("score", row["id"]))
@@ -52,10 +73,12 @@ def fakes(monkeypatch):
core = types.ModuleType("semif_phase1.core")
core.load_causal_model = load_causal_model
core.direct_messages = messages
direct = types.ModuleType("semif_phase1.direct")
direct.score = score
shared = types.ModuleType("semif_phase1.shared")
shared.score_shared = lambda *a: ([], {})
shared._state_prefix, shared.direct_messages, shared.CALLS = upstream_prefix, messages, calls
exec(SHARED_SRC, shared.__dict__)
pkg = types.ModuleType("semif_phase1")
for name, mod in {"torch": torch, "semif_phase1": pkg, "semif_phase1.core": core,
"semif_phase1.direct": direct, "semif_phase1.shared": shared}.items():
@@ -67,7 +90,7 @@ def test_load_caps_before_the_weights_land_then_warms_up(fakes):
from semif_serve.engine import TorchEngine
_torch, calls, _ = fakes
TorchEngine.load(Settings(api_token=TOKEN, vram_cap_gib=12.0))
assert [c[0] for c in calls] == ["cap", "load", "score"]
assert [c[0] for c in calls] == ["cap", "load", "score", "shared"]
assert calls[0] == ("cap", round(12 / 96, 4))
assert calls[2] == ("score", "semif-serve-warmup")
@@ -108,3 +131,77 @@ def test_the_entry_point_forces_offline_mode_before_the_engine_loads(fakes, monk
classmethod(lambda cls, s: seen.update(offline=os.environ.get("HF_HUB_OFFLINE")) or object()))
main.app_from_env()
assert seen["offline"] == "1"
def test_load_wraps_the_upstream_state_prefix_once_even_across_reloads(fakes):
"""INV-7: score_shared looks _state_prefix up as a module global, so the wrapper must replace
it there. A second load must wrap the ORIGINAL again, not stack wrapper on wrapper."""
import semif_phase1.shared as upstream
from semif_serve.engine import TorchEngine
for _ in range(2):
TorchEngine.load(Settings(api_token=TOKEN))
assert upstream._state_prefix is not upstream_prefix
assert upstream._state_prefix.semif_serve_wraps is upstream_prefix
def test_load_fails_closed_when_the_pinned_semif_no_longer_has_the_prefix_hook(fakes, monkeypatch):
import semif_phase1.shared as upstream
from semif_serve.engine import TorchEngine
_torch, calls, _ = fakes
monkeypatch.delattr(upstream, "_state_prefix")
with pytest.raises(RuntimeError, match="_state_prefix"):
TorchEngine.load(Settings(api_token=TOKEN))
assert not any(c[0] == "load" for c in calls)
def test_load_proves_the_hook_on_a_merge_prone_state_through_the_real_shared_path(fakes):
"""Bug hunt SKAL (R2/H1, H2): existence is not effect. Startup scores one state that upstream
alone refuses, through score_shared itself."""
import semif_phase1.shared as upstream
from semif_serve.engine import BOUNDARY_ROW, TorchEngine
_torch, calls, _ = fakes
with pytest.raises(ValueError): # the probe really is merge-prone here
upstream.score_shared(None, MergeTokenizer(), [BOUNDARY_ROW], {})
TorchEngine.load(Settings(api_token=TOKEN))
assert ("shared", BOUNDARY_ROW["state"]) in calls
def test_load_fails_closed_when_score_shared_binds_the_prefix_before_the_patch(fakes):
import semif_phase1.shared as upstream
from semif_serve.engine import TorchEngine
exec(SHARED_SRC_EARLY_BOUND, upstream.__dict__)
with pytest.raises(RuntimeError, match="INV-7"):
TorchEngine.load(Settings(api_token=TOKEN))
def test_load_fails_closed_when_score_shared_resolves_its_globals_elsewhere(fakes):
import semif_phase1.shared as upstream
from semif_serve.engine import TorchEngine
_torch, calls, _ = fakes
elsewhere = {"_state_prefix": upstream_prefix, "direct_messages": messages, "CALLS": calls}
exec(SHARED_SRC, elsewhere)
upstream.score_shared = elsewhere["score_shared"] # re-exported from another module
with pytest.raises(RuntimeError, match="INV-7"):
TorchEngine.load(Settings(api_token=TOKEN))
assert not any(c[0] == "load" for c in calls)
def test_load_fails_closed_when_the_hook_is_not_callable(fakes, monkeypatch):
import semif_phase1.shared as upstream
from semif_serve.engine import TorchEngine
_torch, calls, _ = fakes
monkeypatch.setattr(upstream, "_state_prefix", "not a function")
with pytest.raises(RuntimeError, match="_state_prefix"):
TorchEngine.load(Settings(api_token=TOKEN))
assert not any(c[0] == "load" for c in calls)
def test_load_fails_closed_when_the_wrapper_renders_the_wrong_prompt(fakes, monkeypatch):
"""The SURVIVED row of bug hunt SKAL: a wrapper that renders nothing still yields a valid
(degenerate) prefix, so scoring stays correct and silently loses all sharing. Startup
requires the wrapper to keep upstream's whole prefix on an ordinary state."""
import semif_phase1.core as core
from semif_serve.engine import TorchEngine
monkeypatch.setattr(core, "direct_messages", lambda row: [])
with pytest.raises(RuntimeError, match="INV-7"):
TorchEngine.load(Settings(api_token=TOKEN))
+39
View File
@@ -0,0 +1,39 @@
"""The shared prefix must never cover a token that the full prompts do not share (0.1.4, INV-7).
Found 2026-09-27 by the Cicada mood spike. An object state whose last value ends in ")", ";" or
"}" was refused with 422 "The fixed state prefix does not match every full prompt". SemIf's
_state_prefix encodes the prompt up to the end of the state and drops ONE token, because the
JSON punctuation that follows can merge with it. Here the merge reaches two tokens back.
MergeTokenizer reproduces that shape without the real Qwen vocabulary. The real tokenizer is
checked on the card, in acceptance.
"""
from fake_tokenizer import MergeTokenizer, messages, upstream_prefix
from semif_serve.engine import boundary_safe_prefix
ROW = {"question": "Which pose?", "options": [{"id": "a", "description": "A"}, {"id": "b", "description": "B"}]}
def full(tok, state, row=ROW):
return tok.encode(tok.apply_chat_template(messages({**row, "state": state}), tokenize=False,
add_generation_prompt=True, enable_thinking=False), add_special_tokens=False)
def test_the_merge_that_broke_upstream_is_reproduced():
tok, state = MergeTokenizer(), {"person_said": "ok :)"}
ids = upstream_prefix(tok, state)
assert full(tok, state)[:len(ids)] != ids # the 422: upstream's prefix is not a prefix
def test_a_state_ending_in_a_merging_character_gets_a_prefix_every_row_shares():
tok, state = MergeTokenizer(), {"person_said": "ok :)"}
ids = boundary_safe_prefix(upstream_prefix, messages)(tok, state)
for row in (ROW, {"question": "A different criterion entirely?", "options": ROW["options"][::-1]}):
assert full(tok, state, row)[:len(ids)] == ids
assert ids == upstream_prefix(tok, state)[:-1] # it gave up exactly the one mismatched token
def test_an_ordinary_state_keeps_the_whole_upstream_prefix():
tok = MergeTokenizer()
for state in ({"person_said": "Turn off the lights."}, "ok :)", "Nothing was said."):
assert boundary_safe_prefix(upstream_prefix, messages)(tok, state) == upstream_prefix(tok, state)
+1 -1
View File
@@ -991,7 +991,7 @@ dependencies = [
[[package]]
name = "semif-serve"
version = "0.1.3"
version = "0.1.4"
source = { editable = "." }
dependencies = [
{ name = "fastapi" },