SemIf's shared scorer trims one token at the state boundary. When an object state's last value ends in ')', ';' or '}', the JSON that follows re-merges two tokens back, so score_shared refused the request with 422. The engine now wraps semif_phase1.shared._state_prefix to keep only the tokens the full prompts share. Each row scores the same token sequence; only the prefill/suffix split moves. Startup proves the fix is in effect, not just installed (heid bug hunt SKAL, folded). It checks that the hook is callable and is what score_shared resolves, that an ordinary state keeps upstream's whole prefix, and that a merge-prone state scores through the shared path. Real tokenizer: 154 states, 23 refused before and 0 after, with no ordinary or authored144 prefix changed. Acceptance: 144/144 parity. Shared vs direct 71/72; the miss is a bf16 tie that flipped across a plain restart (see README).
38 lines
1.9 KiB
Python
38 lines
1.9 KiB
Python
"""Test doubles for SemIf's tokenizer seam (INV-7), shared by test_prefix and test_engine_load.
|
|
MergeTokenizer reproduces the real failure's shape without the Qwen vocabulary; the real
|
|
tokenizer is checked on the card, in acceptance."""
|
|
import json
|
|
|
|
|
|
class MergeTokenizer:
|
|
"""Char-level with three merges, in the shape of the real failure. ')"' and '"}' are single
|
|
tokens, but ')"},' splits as [')', '"},']. So a state tail ')"}' encodes as [')"', '}'] when
|
|
nothing follows it (the prefix), and the full prompt's ')"},' encodes as [')', '"},'].
|
|
Dropping one token from the prefix leaves ')"', which the full prompt does not contain. A
|
|
tail like '."}' stays ['.', '"}'] either way, which is the ordinary case."""
|
|
|
|
def apply_chat_template(self, messages, tokenize, add_generation_prompt, enable_thinking):
|
|
assert tokenize is False and add_generation_prompt is True and enable_thinking is False
|
|
return "<s>" + "|".join(m["content"] for m in messages) + "<a>"
|
|
|
|
def encode(self, text, add_special_tokens):
|
|
assert add_special_tokens is False
|
|
out, i = [], 0
|
|
while i < len(text):
|
|
if text.startswith(')"},', i):
|
|
out += [')', '"},']; i += 4
|
|
elif text.startswith(')"', i) or text.startswith('"}', i):
|
|
out.append(text[i:i + 2]); i += 2
|
|
else:
|
|
out.append(text[i]); i += 1
|
|
return out
|
|
|
|
|
|
def messages(row): # the shape of semif_phase1.core.direct_messages: evidence first, then the criterion
|
|
return [{"role": "user", "content": json.dumps(
|
|
{"evidence": row["state"], "criterion": row["question"], "options": row["options"]}, ensure_ascii=False)}]
|
|
|
|
|
|
def upstream_prefix(tokenizer, state): # what SemIf's _state_prefix does: through the state, minus one token
|
|
return tokenizer.encode("<s>" + json.dumps({"evidence": state}, ensure_ascii=False)[:-1], add_special_tokens=False)[:-1]
|