Files
esh-pfi-infrastructure/services/semif-serve/tests/fake_tokenizer.py
T
vh 47cad33dd1 fix(semif): 0.1.4 — object states ending in ) ; } no longer 422 (INV-7)
SemIf's shared scorer trims one token at the state boundary. When an object
state's last value ends in ')', ';' or '}', the JSON that follows re-merges two
tokens back, so score_shared refused the request with 422. The engine now wraps
semif_phase1.shared._state_prefix to keep only the tokens the full prompts
share. Each row scores the same token sequence; only the prefill/suffix split
moves.

Startup proves the fix is in effect, not just installed (heid bug hunt SKAL,
folded). It checks that the hook is callable and is what score_shared resolves,
that an ordinary state keeps upstream's whole prefix, and that a merge-prone
state scores through the shared path.

Real tokenizer: 154 states, 23 refused before and 0 after, with no ordinary or
authored144 prefix changed. Acceptance: 144/144 parity. Shared vs direct 71/72;
the miss is a bf16 tie that flipped across a plain restart (see README).
2026-09-27 10:23:14 -07:00

38 lines
1.9 KiB
Python

"""Test doubles for SemIf's tokenizer seam (INV-7), shared by test_prefix and test_engine_load.
MergeTokenizer reproduces the real failure's shape without the Qwen vocabulary; the real
tokenizer is checked on the card, in acceptance."""
import json
class MergeTokenizer:
"""Char-level with three merges, in the shape of the real failure. ')"' and '"}' are single
tokens, but ')"},' splits as [')', '"},']. So a state tail ')"}' encodes as [')"', '}'] when
nothing follows it (the prefix), and the full prompt's ')"},' encodes as [')', '"},'].
Dropping one token from the prefix leaves ')"', which the full prompt does not contain. A
tail like '."}' stays ['.', '"}'] either way, which is the ordinary case."""
def apply_chat_template(self, messages, tokenize, add_generation_prompt, enable_thinking):
assert tokenize is False and add_generation_prompt is True and enable_thinking is False
return "<s>" + "|".join(m["content"] for m in messages) + "<a>"
def encode(self, text, add_special_tokens):
assert add_special_tokens is False
out, i = [], 0
while i < len(text):
if text.startswith(')"},', i):
out += [')', '"},']; i += 4
elif text.startswith(')"', i) or text.startswith('"}', i):
out.append(text[i:i + 2]); i += 2
else:
out.append(text[i]); i += 1
return out
def messages(row): # the shape of semif_phase1.core.direct_messages: evidence first, then the criterion
return [{"role": "user", "content": json.dumps(
{"evidence": row["state"], "criterion": row["question"], "options": row["options"]}, ensure_ascii=False)}]
def upstream_prefix(tokenizer, state): # what SemIf's _state_prefix does: through the state, minus one token
return tokenizer.encode("<s>" + json.dumps({"evidence": state}, ensure_ascii=False)[:-1], add_special_tokens=False)[:-1]