fix(semif): 0.1.4 — object states ending in ) ; } no longer 422 (INV-7)
SemIf's shared scorer trims one token at the state boundary. When an object state's last value ends in ')', ';' or '}', the JSON that follows re-merges two tokens back, so score_shared refused the request with 422. The engine now wraps semif_phase1.shared._state_prefix to keep only the tokens the full prompts share. Each row scores the same token sequence; only the prefill/suffix split moves. Startup proves the fix is in effect, not just installed (heid bug hunt SKAL, folded). It checks that the hook is callable and is what score_shared resolves, that an ordinary state keeps upstream's whole prefix, and that a merge-prone state scores through the shared path. Real tokenizer: 154 states, 23 refused before and 0 after, with no ordinary or authored144 prefix changed. Acceptance: 144/144 parity. Shared vs direct 71/72; the miss is a bf16 tie that flipped across a plain restart (see README).
This commit is contained in:
@@ -0,0 +1,39 @@
|
||||
"""The shared prefix must never cover a token that the full prompts do not share (0.1.4, INV-7).
|
||||
|
||||
Found 2026-09-27 by the Cicada mood spike. An object state whose last value ends in ")", ";" or
|
||||
"}" was refused with 422 "The fixed state prefix does not match every full prompt". SemIf's
|
||||
_state_prefix encodes the prompt up to the end of the state and drops ONE token, because the
|
||||
JSON punctuation that follows can merge with it. Here the merge reaches two tokens back.
|
||||
MergeTokenizer reproduces that shape without the real Qwen vocabulary. The real tokenizer is
|
||||
checked on the card, in acceptance.
|
||||
"""
|
||||
from fake_tokenizer import MergeTokenizer, messages, upstream_prefix
|
||||
|
||||
from semif_serve.engine import boundary_safe_prefix
|
||||
|
||||
ROW = {"question": "Which pose?", "options": [{"id": "a", "description": "A"}, {"id": "b", "description": "B"}]}
|
||||
|
||||
|
||||
def full(tok, state, row=ROW):
|
||||
return tok.encode(tok.apply_chat_template(messages({**row, "state": state}), tokenize=False,
|
||||
add_generation_prompt=True, enable_thinking=False), add_special_tokens=False)
|
||||
|
||||
|
||||
def test_the_merge_that_broke_upstream_is_reproduced():
|
||||
tok, state = MergeTokenizer(), {"person_said": "ok :)"}
|
||||
ids = upstream_prefix(tok, state)
|
||||
assert full(tok, state)[:len(ids)] != ids # the 422: upstream's prefix is not a prefix
|
||||
|
||||
|
||||
def test_a_state_ending_in_a_merging_character_gets_a_prefix_every_row_shares():
|
||||
tok, state = MergeTokenizer(), {"person_said": "ok :)"}
|
||||
ids = boundary_safe_prefix(upstream_prefix, messages)(tok, state)
|
||||
for row in (ROW, {"question": "A different criterion entirely?", "options": ROW["options"][::-1]}):
|
||||
assert full(tok, state, row)[:len(ids)] == ids
|
||||
assert ids == upstream_prefix(tok, state)[:-1] # it gave up exactly the one mismatched token
|
||||
|
||||
|
||||
def test_an_ordinary_state_keeps_the_whole_upstream_prefix():
|
||||
tok = MergeTokenizer()
|
||||
for state in ({"person_said": "Turn off the lights."}, "ok :)", "Nothing was said."):
|
||||
assert boundary_safe_prefix(upstream_prefix, messages)(tok, state) == upstream_prefix(tok, state)
|
||||
Reference in New Issue
Block a user