fix(semif): 0.1.4 — object states ending in ) ; } no longer 422 (INV-7)
SemIf's shared scorer trims one token at the state boundary. When an object state's last value ends in ')', ';' or '}', the JSON that follows re-merges two tokens back, so score_shared refused the request with 422. The engine now wraps semif_phase1.shared._state_prefix to keep only the tokens the full prompts share. Each row scores the same token sequence; only the prefill/suffix split moves. Startup proves the fix is in effect, not just installed (heid bug hunt SKAL, folded). It checks that the hook is callable and is what score_shared resolves, that an ordinary state keeps upstream's whole prefix, and that a merge-prone state scores through the shared path. Real tokenizer: 154 states, 23 refused before and 0 after, with no ordinary or authored144 prefix changed. Acceptance: 144/144 parity. Shared vs direct 71/72; the miss is a bf16 tie that flipped across a plain restart (see README).
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
# Persistent memory — eshpfi-management
|
||||
|
||||
_Last updated: 2026-09-27 ~0947 PT (SemIf spikes done and ruled build-nothing; SemIf-as-Cicada-mood measured slower and worse; semif-serve 0.1.3 live; overnight backups all green.)_
|
||||
_Last updated: 2026-09-27 ~1025 PT (semif-serve 0.1.4 live: the ")" 422 fixed; SemIf spikes done and ruled build-nothing; SemIf-as-Cicada-mood measured slower and worse, idea 88 dropped; semif-serve 0.1.3 live; overnight backups all green.)_
|
||||
|
||||
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
|
||||
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
|
||||
@@ -144,7 +144,7 @@ _As of 2026-09-27 ~0900 PT._
|
||||
|
||||
### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime)
|
||||
|
||||
- **LIVE: `semif-serve` 0.1.3** at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging
|
||||
- **LIVE: `semif-serve` 0.1.4** (was 0.1.3 until 1014 on 2026-09-27) at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging
|
||||
and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code +
|
||||
contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.**
|
||||
- 0.1.3 acceptance: 144/144 parity with upstream; averaging through the service 78.6% → 88.1%
|
||||
@@ -167,6 +167,12 @@ _As of 2026-09-27 ~0900 PT._
|
||||
and +94 ms sequential vs today's 246 ms, because the pose header costs only ~31 ms and SemIf
|
||||
shares GPU 1 with the LLM. Acceptable pose 67% vs 92%; the mood carried 7/15 vs 14/15. Upside:
|
||||
gestures at 13% vs 58%.
|
||||
- **Prime 0948: idea 88 dropped; fix the 422 → DONE, 0.1.4 live (the `fix(semif): 0.1.4` commit).** INV-7 wraps SemIf's
|
||||
`shared._state_prefix` so the prefix is only the tokens the full prompts share. Startup proves the fix
|
||||
is in effect (the hook must be what score_shared resolves, the prefix unchanged on an ordinary state,
|
||||
and a merge-prone state scored through the shared path). Folded from heid bug hunt SKAL (Hulda, thread
|
||||
`01M3HXMXN27F3K534Q6QS45AHV`). Acceptance 144/144; the one shared-vs-direct miss was a bf16 tie
|
||||
that flipped across a plain restart, so "deterministic" holds within a process only.
|
||||
|
||||
### restic: credential leak fixed (2026-09-27, Prime)
|
||||
|
||||
@@ -222,6 +228,7 @@ _As of 2026-09-27 ~0900 PT._
|
||||
|
||||
- `[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** (SVOS v2.1.12: `propose_decision` gained `seat_up`, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling `zellij-fleet@Claude` at boot remains Prime's call.
|
||||
- `[2026-09-27]` **SemIf LIVE on fv-ml1 GPU 1 (semif-serve 0.1.2, Prime):** wrapper + contract + 39 tests, 142/144 upstream parity, two card-only memory defects fixed. → `persistent-memory.d/2026-09-27-semif-live-on-fv-ml1-gpu1.md`
|
||||
- `[2026-09-27]` **semif-serve 0.1.4: object states ending in `)`, `;` or `}` no longer 422 (INV-7, a prefix wrapper proven at startup); numerics are deterministic within a process but a bf16 near-tie can flip across a restart.** Prime ruled; heid bug hunt folded. → `stacks/semif/README.md`
|
||||
- `[2026-09-27]` **SemIf as Cicada's mood source: slower (+32 ms async, +94 ms sequential) and worse (67% vs 92% apt; carry 7/15 vs 14/15); only the gesture restraint is a win.** Build nothing (Prime). Henge 88 carries it. → `persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md`
|
||||
- `[2026-09-27]` **SemIf consumer-fit spikes (Prime): Cicada affect gate 30/31 with descriptive wording and 19/31 terse; Wyrd "left this place?" 21/21 on the second wording, exit choice 18/21.** Recommendations await Prime. → `persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md`
|
||||
- `[2026-09-27]` **SemIf order-averaging spiked (+9.1 pts accuracy, agreement = strong confidence signal); Prime ruled: build it in as 0.1.3 and trial the fast kernels.** Tracked in the in-flight SemIf section. → `persistent-memory.d/2026-09-27-semif-order-averaging.md` — **DONE:** 0.1.3 live, with fast kernels adopted (`77b8cb4`).
|
||||
|
||||
@@ -0,0 +1,63 @@
|
||||
{
|
||||
"url": "http://10.251.50.54:8032",
|
||||
"health": {
|
||||
"status": "ok",
|
||||
"semif_commit": "23cf1f39fc9534fe81437200959b6dfc7106e45a",
|
||||
"model": {
|
||||
"source": "Qwen/Qwen3.5-4B",
|
||||
"revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
|
||||
"dtype": "bfloat16",
|
||||
"device": "cuda:0",
|
||||
"torch_version": "2.10.0+cu128",
|
||||
"transformers_version": "5.17.0",
|
||||
"device_name": "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition",
|
||||
"allocated_gib": 7.84,
|
||||
"reserved_gib": 8.12
|
||||
},
|
||||
"vram_cap_gib": 12.0,
|
||||
"max_tokens": 4096,
|
||||
"max_decisions": 64,
|
||||
"workloads": []
|
||||
},
|
||||
"1_parity_vs_upstream": {
|
||||
"rows": 144,
|
||||
"top_choice_agree": 144,
|
||||
"max_abs_prob_gap": 0.0595381559570336
|
||||
},
|
||||
"1_prompt_sha256_equal": 144,
|
||||
"2_noise_floor_a_vs_b": {
|
||||
"rows": 144,
|
||||
"top_choice_agree": 144,
|
||||
"max_abs_prob_gap": 0.0
|
||||
},
|
||||
"3_negative_rotated_options": {
|
||||
"rows": 144,
|
||||
"top_choice_agree": 14,
|
||||
"max_abs_prob_gap": 0.9987391090270772
|
||||
},
|
||||
"4_shared_vs_direct": {
|
||||
"groups": 36,
|
||||
"rows": 72,
|
||||
"top_choice_agree": 71,
|
||||
"max_abs_prob_gap": 0.05927145076874801
|
||||
},
|
||||
"5_speed_21_binary": {
|
||||
"prefix_tokens": 62,
|
||||
"shared_s": {
|
||||
"runs": [
|
||||
0.13865972400526516,
|
||||
0.1382300010009203,
|
||||
0.13749008599552326
|
||||
],
|
||||
"median": 0.1382300010009203
|
||||
},
|
||||
"sequential_decide_s": {
|
||||
"runs": [
|
||||
0.9489945390087087,
|
||||
0.9463060130074155,
|
||||
0.9458249929884914
|
||||
],
|
||||
"median": 0.9463060130074155
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,6 +1,6 @@
|
||||
[project]
|
||||
name = "semif-serve"
|
||||
version = "0.1.3"
|
||||
version = "0.1.4"
|
||||
description = "HTTP wrapper around SemIf's direct and shared option-logit scorers"
|
||||
requires-python = ">=3.12"
|
||||
dependencies = [
|
||||
|
||||
@@ -104,6 +104,26 @@ Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from
|
||||
- **INV-5 no network at runtime.** Weights come from the mounted HF cache at the
|
||||
pinned revision. The entry point sets `HF_HUB_OFFLINE=1` itself before torch or
|
||||
transformers load, so this holds outside the image too (S10).
|
||||
- **INV-7 boundary-safe shared prefix (0.1.4).** At load, the engine wraps SemIf's
|
||||
module-global `semif_phase1.shared._state_prefix`, which `score_shared` calls. The wrapper
|
||||
keeps only the leading tokens that upstream's prefix shares with a real full prompt for
|
||||
the same state (a probe row: evidence, then a placeholder criterion).
|
||||
- **Why:** upstream drops just one token at the state boundary. An object state whose
|
||||
last value ends in `)`, `;` or `}` re-tokenises two tokens back once `, "criterion"`
|
||||
follows, so it was refused with 422 "The fixed state prefix does not match every full
|
||||
prompt" (found 2026-09-27).
|
||||
- **Effect:** each row still scores the same token sequence; only the prefill/suffix split
|
||||
moves, costing a few tokens of sharing. An ordinary state keeps upstream's prefix
|
||||
unchanged. `score_shared` still checks every real row and fails closed.
|
||||
- **Startup proves the fix is in effect, not just installed** (heid bug hunt SKAL,
|
||||
2026-09-27). Before the weights load, it refuses to start if `_state_prefix` is
|
||||
missing or not callable, or if `score_shared` does not resolve it from
|
||||
`semif_phase1.shared`'s globals (a SemIf bump that moves or re-exports it). After the
|
||||
warm-up, it checks two things. The wrapper must return upstream's exact prefix for
|
||||
the warm-up state, since a wrapper rendering the wrong prompt would silently drop all
|
||||
sharing. And a state upstream alone refuses, `{"person_said": "ok :)"}`, must score
|
||||
through `score_shared` itself; a hook bound before the patch would fail here. A
|
||||
reload wraps the original again rather than stacking wrappers.
|
||||
- **INV-6 constant-time auth.** Token comparison uses `hmac.compare_digest`. The
|
||||
token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything
|
||||
else, because a CR, LF or NUL in the token can never arrive in a header (S2).
|
||||
@@ -169,7 +189,13 @@ cancels a position bias exactly; `all` sends n! rows and is 422 above 4 options;
|
||||
agreement and spread are computed from the orderings; a mixed shared request
|
||||
(averaged + plain) is one engine call, with results in request order and plain
|
||||
results unchanged; expanded rows count toward the cap; `workload` + `orderings`
|
||||
→ 422.
|
||||
→ 422. **Prefix (INV-7):** against a tokenizer whose merge reaches two tokens back, the
|
||||
wrapped prefix is a token-prefix of every row's full prompt and gives up exactly the
|
||||
mismatched token; an ordinary state keeps the whole upstream prefix; load() installs
|
||||
one wrapper however many times it runs, and refuses to start when `_state_prefix` is
|
||||
gone or not callable, when `score_shared` binds it early or resolves its globals
|
||||
elsewhere, or when the wrapper renders the wrong prompt. The fake `score_shared` is
|
||||
compiled into the fake module, so that it resolves its globals as the real one does.
|
||||
|
||||
## Acceptance (on fv-ml1, real model; not unit tests)
|
||||
|
||||
@@ -180,3 +206,6 @@ results unchanged; expanded rows count toward the cap; `workload` + `orderings`
|
||||
4. **Shared vs direct:** the same rows agree within the A-vs-A floor.
|
||||
5. **Speed:** 21 binary criteria over one state, N ≥ 3, p50 + spread.
|
||||
6. **VRAM:** the peak at a 4096-token input sets `SEMIF_VRAM_CAP_GIB`.
|
||||
7. **Boundary (0.1.4):** states whose last value ends in `)`, `;` and `}`, as objects,
|
||||
are all answered by `/decide/shared`, with the same top choice as a string state
|
||||
holding the same text; parity (1) still holds.
|
||||
|
||||
@@ -28,6 +28,44 @@ WARMUP_ROW = {
|
||||
}
|
||||
|
||||
|
||||
# INV-7 startup proof: a state SemIf's own prefix refuses (measured with the real tokenizer,
|
||||
# 2026-09-27), sent through score_shared itself, so the fix is shown to be IN EFFECT, not just
|
||||
# installed (bug hunt SKAL, R2/H1/H2).
|
||||
BOUNDARY_ROW = {**WARMUP_ROW, "id": "semif-serve-boundary-check", "state": {"person_said": "ok :)"}}
|
||||
PREFIX_PROBE_ROW = {
|
||||
"id": "semif-serve-prefix-probe",
|
||||
"question": "prefix boundary placeholder",
|
||||
"options": [{"id": "yes", "description": "Yes"}, {"id": "no", "description": "No"}],
|
||||
}
|
||||
|
||||
|
||||
def boundary_safe_prefix(state_prefix: Callable, messages: Callable) -> Callable:
|
||||
"""INV-7: wrap SemIf's `shared._state_prefix` so the prefill never covers a token that the full
|
||||
prompts do not share.
|
||||
|
||||
Upstream encodes the prompt up to the end of the state and drops ONE token, because the JSON
|
||||
punctuation that follows the state can merge with it. One is not always enough: an object
|
||||
state whose last value ends in ")", ";" or "}" re-tokenises TWO tokens back once `, "criterion"`
|
||||
follows, and score_shared then refused the request with 422 (found 2026-09-27). This keeps only
|
||||
the leading tokens that upstream's prefix shares with a real full prompt for the same state.
|
||||
The suffix starts that much earlier and scores the same token sequence; the cost is a few
|
||||
tokens of lost sharing. Only the punctuation run at the boundary can merge, and it is the same
|
||||
in every row, so a probe row stands for all of them. score_shared still checks every real row
|
||||
and fails closed if that ever stops holding."""
|
||||
def prefix(tokenizer, state):
|
||||
ids = state_prefix(tokenizer, state)
|
||||
full = tokenizer.encode(tokenizer.apply_chat_template(
|
||||
messages({**PREFIX_PROBE_ROW, "state": state}), tokenize=False, add_generation_prompt=True,
|
||||
enable_thinking=False), add_special_tokens=False)
|
||||
shared = 0
|
||||
while shared < min(len(ids), len(full)) and ids[shared] == full[shared]:
|
||||
shared += 1
|
||||
return ids[:shared]
|
||||
|
||||
prefix.semif_serve_wraps = state_prefix
|
||||
return prefix
|
||||
|
||||
|
||||
def _first_line(exc: BaseException) -> str:
|
||||
lines = str(exc).splitlines()
|
||||
return lines[0] if lines else ""
|
||||
@@ -44,10 +82,24 @@ class TorchEngine:
|
||||
@classmethod
|
||||
def load(cls, settings: Settings) -> "TorchEngine":
|
||||
import torch
|
||||
from semif_phase1.core import load_causal_model
|
||||
import semif_phase1.shared as upstream_shared
|
||||
from semif_phase1.core import direct_messages, load_causal_model
|
||||
from semif_phase1.direct import score
|
||||
from semif_phase1.shared import score_shared
|
||||
|
||||
# INV-7: score_shared looks `_state_prefix` up as a module global, so the wrapper goes there.
|
||||
# A SemIf bump that renames it, or moves score_shared so it resolves its globals elsewhere,
|
||||
# must stop startup, not silently bring the 422 back.
|
||||
original = getattr(upstream_shared, "_state_prefix", None)
|
||||
if not callable(original):
|
||||
raise RuntimeError("semif_phase1.shared._state_prefix is gone or not callable at this SemIf "
|
||||
"commit: re-check the boundary-safe prefix (INV-7) before serving")
|
||||
original = getattr(original, "semif_serve_wraps", original) # one wrapper, however many loads
|
||||
upstream_shared._state_prefix = boundary_safe_prefix(original, direct_messages)
|
||||
if getattr(score_shared, "__globals__", {}).get("_state_prefix") is not upstream_shared._state_prefix:
|
||||
raise RuntimeError("INV-7: score_shared does not resolve semif_phase1.shared._state_prefix, "
|
||||
"so the boundary-safe prefix would be inert")
|
||||
|
||||
if settings.device == "cuda":
|
||||
if not torch.cuda.is_available():
|
||||
raise RuntimeError("SEMIF_DEVICE=cuda but torch sees no CUDA device")
|
||||
@@ -70,10 +122,23 @@ class TorchEngine:
|
||||
raise RuntimeError(f"model landed on {placed}, expected {settings.device}")
|
||||
engine = cls(torch, model, tokenizer, metadata, settings, direct_fn=score, shared_fn=score_shared)
|
||||
engine.direct(WARMUP_ROW) # INV-3: one decision must score
|
||||
engine._prove_prefix_hook(upstream_shared._state_prefix, original)
|
||||
if settings.device == "cuda": # INV-4: the resting footprint
|
||||
engine._release_above = torch.cuda.memory_reserved(0) + RELEASE_SLACK_BYTES
|
||||
return engine
|
||||
|
||||
def _prove_prefix_hook(self, hook: Callable, original: Callable) -> None:
|
||||
"""INV-7, at startup: the wrapper keeps upstream's whole prefix on an ordinary state (a wrapper
|
||||
rendering the wrong prompt would silently give up all sharing), and a state upstream alone
|
||||
refuses scores through score_shared itself (a hook score_shared never calls would not)."""
|
||||
state = WARMUP_ROW["state"]
|
||||
if hook(self._tokenizer, state) != original(self._tokenizer, state):
|
||||
raise RuntimeError("INV-7: the boundary-safe prefix does not keep upstream's prefix on an ordinary state")
|
||||
try:
|
||||
self.shared([BOUNDARY_ROW])
|
||||
except ValueError as exc:
|
||||
raise RuntimeError(f"INV-7: shared scoring still refuses a merge-prone state: {exc}") from None
|
||||
|
||||
def health(self) -> dict:
|
||||
info = dict(self._metadata)
|
||||
if self._settings.device == "cuda":
|
||||
|
||||
@@ -0,0 +1,37 @@
|
||||
"""Test doubles for SemIf's tokenizer seam (INV-7), shared by test_prefix and test_engine_load.
|
||||
MergeTokenizer reproduces the real failure's shape without the Qwen vocabulary; the real
|
||||
tokenizer is checked on the card, in acceptance."""
|
||||
import json
|
||||
|
||||
|
||||
class MergeTokenizer:
|
||||
"""Char-level with three merges, in the shape of the real failure. ')"' and '"}' are single
|
||||
tokens, but ')"},' splits as [')', '"},']. So a state tail ')"}' encodes as [')"', '}'] when
|
||||
nothing follows it (the prefix), and the full prompt's ')"},' encodes as [')', '"},'].
|
||||
Dropping one token from the prefix leaves ')"', which the full prompt does not contain. A
|
||||
tail like '."}' stays ['.', '"}'] either way, which is the ordinary case."""
|
||||
|
||||
def apply_chat_template(self, messages, tokenize, add_generation_prompt, enable_thinking):
|
||||
assert tokenize is False and add_generation_prompt is True and enable_thinking is False
|
||||
return "<s>" + "|".join(m["content"] for m in messages) + "<a>"
|
||||
|
||||
def encode(self, text, add_special_tokens):
|
||||
assert add_special_tokens is False
|
||||
out, i = [], 0
|
||||
while i < len(text):
|
||||
if text.startswith(')"},', i):
|
||||
out += [')', '"},']; i += 4
|
||||
elif text.startswith(')"', i) or text.startswith('"}', i):
|
||||
out.append(text[i:i + 2]); i += 2
|
||||
else:
|
||||
out.append(text[i]); i += 1
|
||||
return out
|
||||
|
||||
|
||||
def messages(row): # the shape of semif_phase1.core.direct_messages: evidence first, then the criterion
|
||||
return [{"role": "user", "content": json.dumps(
|
||||
{"evidence": row["state"], "criterion": row["question"], "options": row["options"]}, ensure_ascii=False)}]
|
||||
|
||||
|
||||
def upstream_prefix(tokenizer, state): # what SemIf's _state_prefix does: through the state, minus one token
|
||||
return tokenizer.encode("<s>" + json.dumps({"evidence": state}, ensure_ascii=False)[:-1], add_special_tokens=False)[:-1]
|
||||
@@ -6,6 +6,8 @@ import types
|
||||
|
||||
import pytest
|
||||
|
||||
from fake_tokenizer import MergeTokenizer, messages, upstream_prefix
|
||||
|
||||
from semif_serve.config import Settings
|
||||
|
||||
TOKEN = "t" * 40
|
||||
@@ -24,6 +26,25 @@ class Model:
|
||||
yield Param(self._device)
|
||||
|
||||
|
||||
# The shape of SemIf's score_shared, compiled INTO the fake module so that, like the real one, it
|
||||
# resolves _state_prefix and direct_messages as module globals at call time and refuses (ValueError)
|
||||
# a prefix that is not a token-prefix of every row.
|
||||
SHARED_SRC = """
|
||||
def score_shared(model, tokenizer, rows, metadata, max_tokens=4096):
|
||||
prefix = _state_prefix(tokenizer, rows[0]["state"])
|
||||
for row in rows:
|
||||
ids = tokenizer.encode(tokenizer.apply_chat_template(direct_messages(row), tokenize=False,
|
||||
add_generation_prompt=True, enable_thinking=False), add_special_tokens=False)
|
||||
if not prefix or ids[:len(prefix)] != prefix:
|
||||
raise ValueError("The fixed state prefix does not match every full prompt")
|
||||
CALLS.append(("shared", rows[0]["state"]))
|
||||
return [{"id": row["id"]} for row in rows], {}
|
||||
"""
|
||||
# A SemIf revision that binds the helper when score_shared is defined: the module global can be
|
||||
# replaced all day and scoring never sees it.
|
||||
SHARED_SRC_EARLY_BOUND = SHARED_SRC.replace("max_tokens=4096):", "max_tokens=4096, _state_prefix=_state_prefix):")
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def fakes(monkeypatch):
|
||||
calls = []
|
||||
@@ -42,7 +63,7 @@ def fakes(monkeypatch):
|
||||
|
||||
def load_causal_model(model, revision, device, dtype):
|
||||
calls.append(("load", model, revision, device, dtype))
|
||||
return Model(state["device"]), object(), {"source": model}
|
||||
return Model(state["device"]), MergeTokenizer(), {"source": model}
|
||||
|
||||
def score(model, tok, row, meta, max_tokens):
|
||||
calls.append(("score", row["id"]))
|
||||
@@ -52,10 +73,12 @@ def fakes(monkeypatch):
|
||||
|
||||
core = types.ModuleType("semif_phase1.core")
|
||||
core.load_causal_model = load_causal_model
|
||||
core.direct_messages = messages
|
||||
direct = types.ModuleType("semif_phase1.direct")
|
||||
direct.score = score
|
||||
shared = types.ModuleType("semif_phase1.shared")
|
||||
shared.score_shared = lambda *a: ([], {})
|
||||
shared._state_prefix, shared.direct_messages, shared.CALLS = upstream_prefix, messages, calls
|
||||
exec(SHARED_SRC, shared.__dict__)
|
||||
pkg = types.ModuleType("semif_phase1")
|
||||
for name, mod in {"torch": torch, "semif_phase1": pkg, "semif_phase1.core": core,
|
||||
"semif_phase1.direct": direct, "semif_phase1.shared": shared}.items():
|
||||
@@ -67,7 +90,7 @@ def test_load_caps_before_the_weights_land_then_warms_up(fakes):
|
||||
from semif_serve.engine import TorchEngine
|
||||
_torch, calls, _ = fakes
|
||||
TorchEngine.load(Settings(api_token=TOKEN, vram_cap_gib=12.0))
|
||||
assert [c[0] for c in calls] == ["cap", "load", "score"]
|
||||
assert [c[0] for c in calls] == ["cap", "load", "score", "shared"]
|
||||
assert calls[0] == ("cap", round(12 / 96, 4))
|
||||
assert calls[2] == ("score", "semif-serve-warmup")
|
||||
|
||||
@@ -108,3 +131,77 @@ def test_the_entry_point_forces_offline_mode_before_the_engine_loads(fakes, monk
|
||||
classmethod(lambda cls, s: seen.update(offline=os.environ.get("HF_HUB_OFFLINE")) or object()))
|
||||
main.app_from_env()
|
||||
assert seen["offline"] == "1"
|
||||
|
||||
|
||||
def test_load_wraps_the_upstream_state_prefix_once_even_across_reloads(fakes):
|
||||
"""INV-7: score_shared looks _state_prefix up as a module global, so the wrapper must replace
|
||||
it there. A second load must wrap the ORIGINAL again, not stack wrapper on wrapper."""
|
||||
import semif_phase1.shared as upstream
|
||||
from semif_serve.engine import TorchEngine
|
||||
for _ in range(2):
|
||||
TorchEngine.load(Settings(api_token=TOKEN))
|
||||
assert upstream._state_prefix is not upstream_prefix
|
||||
assert upstream._state_prefix.semif_serve_wraps is upstream_prefix
|
||||
|
||||
|
||||
def test_load_fails_closed_when_the_pinned_semif_no_longer_has_the_prefix_hook(fakes, monkeypatch):
|
||||
import semif_phase1.shared as upstream
|
||||
from semif_serve.engine import TorchEngine
|
||||
_torch, calls, _ = fakes
|
||||
monkeypatch.delattr(upstream, "_state_prefix")
|
||||
with pytest.raises(RuntimeError, match="_state_prefix"):
|
||||
TorchEngine.load(Settings(api_token=TOKEN))
|
||||
assert not any(c[0] == "load" for c in calls)
|
||||
|
||||
|
||||
def test_load_proves_the_hook_on_a_merge_prone_state_through_the_real_shared_path(fakes):
|
||||
"""Bug hunt SKAL (R2/H1, H2): existence is not effect. Startup scores one state that upstream
|
||||
alone refuses, through score_shared itself."""
|
||||
import semif_phase1.shared as upstream
|
||||
from semif_serve.engine import BOUNDARY_ROW, TorchEngine
|
||||
_torch, calls, _ = fakes
|
||||
with pytest.raises(ValueError): # the probe really is merge-prone here
|
||||
upstream.score_shared(None, MergeTokenizer(), [BOUNDARY_ROW], {})
|
||||
TorchEngine.load(Settings(api_token=TOKEN))
|
||||
assert ("shared", BOUNDARY_ROW["state"]) in calls
|
||||
|
||||
|
||||
def test_load_fails_closed_when_score_shared_binds_the_prefix_before_the_patch(fakes):
|
||||
import semif_phase1.shared as upstream
|
||||
from semif_serve.engine import TorchEngine
|
||||
exec(SHARED_SRC_EARLY_BOUND, upstream.__dict__)
|
||||
with pytest.raises(RuntimeError, match="INV-7"):
|
||||
TorchEngine.load(Settings(api_token=TOKEN))
|
||||
|
||||
|
||||
def test_load_fails_closed_when_score_shared_resolves_its_globals_elsewhere(fakes):
|
||||
import semif_phase1.shared as upstream
|
||||
from semif_serve.engine import TorchEngine
|
||||
_torch, calls, _ = fakes
|
||||
elsewhere = {"_state_prefix": upstream_prefix, "direct_messages": messages, "CALLS": calls}
|
||||
exec(SHARED_SRC, elsewhere)
|
||||
upstream.score_shared = elsewhere["score_shared"] # re-exported from another module
|
||||
with pytest.raises(RuntimeError, match="INV-7"):
|
||||
TorchEngine.load(Settings(api_token=TOKEN))
|
||||
assert not any(c[0] == "load" for c in calls)
|
||||
|
||||
|
||||
def test_load_fails_closed_when_the_hook_is_not_callable(fakes, monkeypatch):
|
||||
import semif_phase1.shared as upstream
|
||||
from semif_serve.engine import TorchEngine
|
||||
_torch, calls, _ = fakes
|
||||
monkeypatch.setattr(upstream, "_state_prefix", "not a function")
|
||||
with pytest.raises(RuntimeError, match="_state_prefix"):
|
||||
TorchEngine.load(Settings(api_token=TOKEN))
|
||||
assert not any(c[0] == "load" for c in calls)
|
||||
|
||||
|
||||
def test_load_fails_closed_when_the_wrapper_renders_the_wrong_prompt(fakes, monkeypatch):
|
||||
"""The SURVIVED row of bug hunt SKAL: a wrapper that renders nothing still yields a valid
|
||||
(degenerate) prefix, so scoring stays correct and silently loses all sharing. Startup
|
||||
requires the wrapper to keep upstream's whole prefix on an ordinary state."""
|
||||
import semif_phase1.core as core
|
||||
from semif_serve.engine import TorchEngine
|
||||
monkeypatch.setattr(core, "direct_messages", lambda row: [])
|
||||
with pytest.raises(RuntimeError, match="INV-7"):
|
||||
TorchEngine.load(Settings(api_token=TOKEN))
|
||||
|
||||
@@ -0,0 +1,39 @@
|
||||
"""The shared prefix must never cover a token that the full prompts do not share (0.1.4, INV-7).
|
||||
|
||||
Found 2026-09-27 by the Cicada mood spike. An object state whose last value ends in ")", ";" or
|
||||
"}" was refused with 422 "The fixed state prefix does not match every full prompt". SemIf's
|
||||
_state_prefix encodes the prompt up to the end of the state and drops ONE token, because the
|
||||
JSON punctuation that follows can merge with it. Here the merge reaches two tokens back.
|
||||
MergeTokenizer reproduces that shape without the real Qwen vocabulary. The real tokenizer is
|
||||
checked on the card, in acceptance.
|
||||
"""
|
||||
from fake_tokenizer import MergeTokenizer, messages, upstream_prefix
|
||||
|
||||
from semif_serve.engine import boundary_safe_prefix
|
||||
|
||||
ROW = {"question": "Which pose?", "options": [{"id": "a", "description": "A"}, {"id": "b", "description": "B"}]}
|
||||
|
||||
|
||||
def full(tok, state, row=ROW):
|
||||
return tok.encode(tok.apply_chat_template(messages({**row, "state": state}), tokenize=False,
|
||||
add_generation_prompt=True, enable_thinking=False), add_special_tokens=False)
|
||||
|
||||
|
||||
def test_the_merge_that_broke_upstream_is_reproduced():
|
||||
tok, state = MergeTokenizer(), {"person_said": "ok :)"}
|
||||
ids = upstream_prefix(tok, state)
|
||||
assert full(tok, state)[:len(ids)] != ids # the 422: upstream's prefix is not a prefix
|
||||
|
||||
|
||||
def test_a_state_ending_in_a_merging_character_gets_a_prefix_every_row_shares():
|
||||
tok, state = MergeTokenizer(), {"person_said": "ok :)"}
|
||||
ids = boundary_safe_prefix(upstream_prefix, messages)(tok, state)
|
||||
for row in (ROW, {"question": "A different criterion entirely?", "options": ROW["options"][::-1]}):
|
||||
assert full(tok, state, row)[:len(ids)] == ids
|
||||
assert ids == upstream_prefix(tok, state)[:-1] # it gave up exactly the one mismatched token
|
||||
|
||||
|
||||
def test_an_ordinary_state_keeps_the_whole_upstream_prefix():
|
||||
tok = MergeTokenizer()
|
||||
for state in ({"person_said": "Turn off the lights."}, "ok :)", "Nothing was said."):
|
||||
assert boundary_safe_prefix(upstream_prefix, messages)(tok, state) == upstream_prefix(tok, state)
|
||||
Generated
+1
-1
@@ -991,7 +991,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "semif-serve"
|
||||
version = "0.1.3"
|
||||
version = "0.1.4"
|
||||
source = { editable = "." }
|
||||
dependencies = [
|
||||
{ name = "fastapi" },
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# semif — copy to /opt/docker/compose/semif/.env on fv-ml1 (mode 0600).
|
||||
# Built on fv-ml1 from services/semif-serve (see README "Building").
|
||||
IMAGE=semif-serve:0.1.3
|
||||
IMAGE=semif-serve:0.1.4
|
||||
PORT=8032
|
||||
HOST_IP=10.251.50.54
|
||||
# fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero, scriberr).
|
||||
|
||||
+37
-10
@@ -53,12 +53,15 @@ curl -s -H "Authorization: Bearer $T" http://10.251.50.54:8032/decide -d '{
|
||||
644 suffix tokens, 5 calls after warm-up). One cold call at that size returned 503. The
|
||||
VRAM table below covers binary decisions only. For many options, use one ordering or
|
||||
shorter option text.
|
||||
- ⚠ **`/decide/shared` refuses some object states.** If the state is an object whose
|
||||
LAST value ends in `)`, `;` or `}`, the service returns 422 "The fixed state prefix
|
||||
does not match every full prompt". The closing `"}` merges with that character into one
|
||||
token. The same text as a plain string state works, and `.`, `!`, `?`, `]`, `…` and
|
||||
`—` endings work. Not fixed yet. Callers that pass user text last should
|
||||
append a full stop or send a string state.
|
||||
- **Object states ending in `)`, `;` or `}` work as of 0.1.4.** Until then, if the state was
|
||||
an object whose last value ended in one of those characters, the service returned 422 "The
|
||||
fixed state prefix does not match every full prompt". SemIf trims one token at the state
|
||||
boundary, and the JSON that follows re-merged two tokens back. The fix (contract INV-7)
|
||||
moves only where the shared prefix ends; every row still scores the same tokens. Measured
|
||||
with the real tokenizer: of 154 states (22 endings × 7 shapes) SemIf alone refused 23, and
|
||||
with the fix none. None of the 131 ordinary states, nor any of the authored144 states,
|
||||
changed its prefix. Startup proves the fix is live by scoring `{"person_said": "ok :)"}`
|
||||
through the shared path.
|
||||
|
||||
## ⚠ Probabilities are uncalibrated
|
||||
|
||||
@@ -123,6 +126,21 @@ warm-up fails, and startup fails closed. Build without the kernels:
|
||||
| 6 orderings, short | 118 ms | 81 ms |
|
||||
| 3 rotations, ~2,000-token state | 200 ms | 158 ms |
|
||||
|
||||
## Acceptance (2026-09-27, v0.1.4)
|
||||
|
||||
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.4.json`. Parity with upstream
|
||||
144/144 (identical prompt hashes, max prob gap 0.060). Deterministic within the process
|
||||
(A-vs-A gap 0.0). The negative control fails as it should (14/144). Shared vs direct 71/72.
|
||||
The one miss is an exact bf16 tie in the shared result (0.444/0.444). INV-7 did not move that
|
||||
row's prefix. After a plain `docker restart` of the same image, the row read 0.369/0.537 and
|
||||
agreed.
|
||||
|
||||
⚠ **So "deterministic" holds within one process, not across restarts.** Logits come in bf16
|
||||
steps (0.125 here), and a near-tie can land differently after a restart, most likely because
|
||||
the fast kernels autotune at startup. That is n=1 row across one restart, and the cross-restart
|
||||
floor is otherwise unmeasured. When comparing two versions, compare them against that floor
|
||||
and not against zero.
|
||||
|
||||
## Acceptance (2026-09-27, v0.1.3)
|
||||
|
||||
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json`,
|
||||
@@ -149,16 +167,25 @@ now matches the label. So the misses come from the numeric path, not the wrapper
|
||||
speed figure includes one network round trip (~27 ms). The burst release costs ~16 ms
|
||||
on it (0.1.1 measured 143 ms without the release).
|
||||
|
||||
## Building
|
||||
## Building and deploying
|
||||
|
||||
```bash
|
||||
# from nh3-dev
|
||||
tar -C services/semif-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ --exclude=acceptance . \
|
||||
| ssh infra-ops@10.251.50.54 'mkdir -p /opt/docker/src/semif-serve-X.Y.Z && tar -x -C /opt/docker/src/semif-serve-X.Y.Z'
|
||||
# from nh3-dev. /opt/docker/src is root-owned, so create the version dir with sudo first.
|
||||
ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g infra-ops /opt/docker/src/semif-serve-X.Y.Z'
|
||||
tar -C services/semif-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ \
|
||||
--exclude=acceptance --exclude=spike --exclude='*.egg-info' . \
|
||||
| ssh infra-ops@10.251.50.54 'tar -x -C /opt/docker/src/semif-serve-X.Y.Z'
|
||||
# on fv-ml1
|
||||
cd /opt/docker/src/semif-serve-X.Y.Z && docker build -t semif-serve:X.Y.Z .
|
||||
# ⚠ /opt/docker/compose/semif is root-owned, so `sed -i` cannot write its temp file there. .env
|
||||
# itself is infra-ops's: rewrite it IN PLACE, which keeps its owner and mode (0600).
|
||||
cd /opt/docker/compose/semif && new=$(sed 's/^IMAGE=.*/IMAGE=semif-serve:X.Y.Z/' .env) \
|
||||
&& printf '%s\n' "$new" > .env && docker compose up -d
|
||||
```
|
||||
|
||||
Startup fails closed (INV-3, INV-7), so a container that does not reach healthy did not pass
|
||||
its own checks: read `docker logs semif`. Record the deploy with `scripts/ops-log`.
|
||||
|
||||
To move SemIf forward: bump the commit in `services/semif-serve/pyproject.toml`
|
||||
(and `SEMIF_COMMIT` in `config.py`), run `uv lock`, rebuild, and **re-run the
|
||||
acceptance**. Upstream is research code that changes weekly, which is why it is
|
||||
|
||||
Reference in New Issue
Block a user