fix(semif): 0.1.4 — object states ending in ) ; } no longer 422 (INV-7)
SemIf's shared scorer trims one token at the state boundary. When an object state's last value ends in ')', ';' or '}', the JSON that follows re-merges two tokens back, so score_shared refused the request with 422. The engine now wraps semif_phase1.shared._state_prefix to keep only the tokens the full prompts share. Each row scores the same token sequence; only the prefill/suffix split moves. Startup proves the fix is in effect, not just installed (heid bug hunt SKAL, folded). It checks that the hook is callable and is what score_shared resolves, that an ordinary state keeps upstream's whole prefix, and that a merge-prone state scores through the shared path. Real tokenizer: 154 states, 23 refused before and 0 after, with no ordinary or authored144 prefix changed. Acceptance: 144/144 parity. Shared vs direct 71/72; the miss is a bf16 tie that flipped across a plain restart (see README).
This commit is contained in:
@@ -0,0 +1,37 @@
|
||||
"""Test doubles for SemIf's tokenizer seam (INV-7), shared by test_prefix and test_engine_load.
|
||||
MergeTokenizer reproduces the real failure's shape without the Qwen vocabulary; the real
|
||||
tokenizer is checked on the card, in acceptance."""
|
||||
import json
|
||||
|
||||
|
||||
class MergeTokenizer:
|
||||
"""Char-level with three merges, in the shape of the real failure. ')"' and '"}' are single
|
||||
tokens, but ')"},' splits as [')', '"},']. So a state tail ')"}' encodes as [')"', '}'] when
|
||||
nothing follows it (the prefix), and the full prompt's ')"},' encodes as [')', '"},'].
|
||||
Dropping one token from the prefix leaves ')"', which the full prompt does not contain. A
|
||||
tail like '."}' stays ['.', '"}'] either way, which is the ordinary case."""
|
||||
|
||||
def apply_chat_template(self, messages, tokenize, add_generation_prompt, enable_thinking):
|
||||
assert tokenize is False and add_generation_prompt is True and enable_thinking is False
|
||||
return "<s>" + "|".join(m["content"] for m in messages) + "<a>"
|
||||
|
||||
def encode(self, text, add_special_tokens):
|
||||
assert add_special_tokens is False
|
||||
out, i = [], 0
|
||||
while i < len(text):
|
||||
if text.startswith(')"},', i):
|
||||
out += [')', '"},']; i += 4
|
||||
elif text.startswith(')"', i) or text.startswith('"}', i):
|
||||
out.append(text[i:i + 2]); i += 2
|
||||
else:
|
||||
out.append(text[i]); i += 1
|
||||
return out
|
||||
|
||||
|
||||
def messages(row): # the shape of semif_phase1.core.direct_messages: evidence first, then the criterion
|
||||
return [{"role": "user", "content": json.dumps(
|
||||
{"evidence": row["state"], "criterion": row["question"], "options": row["options"]}, ensure_ascii=False)}]
|
||||
|
||||
|
||||
def upstream_prefix(tokenizer, state): # what SemIf's _state_prefix does: through the state, minus one token
|
||||
return tokenizer.encode("<s>" + json.dumps({"evidence": state}, ensure_ascii=False)[:-1], add_special_tokens=False)[:-1]
|
||||
Reference in New Issue
Block a user