diff --git a/persistent-memory.md b/persistent-memory.md index b47e495..c13acf5 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -1,6 +1,6 @@ # Persistent memory — eshpfi-management -_Last updated: 2026-09-27 ~0947 PT (SemIf spikes done and ruled build-nothing; SemIf-as-Cicada-mood measured slower and worse; semif-serve 0.1.3 live; overnight backups all green.)_ +_Last updated: 2026-09-27 ~1025 PT (semif-serve 0.1.4 live: the ")" 422 fixed; SemIf spikes done and ruled build-nothing; SemIf-as-Cicada-mood measured slower and worse, idea 88 dropped; semif-serve 0.1.3 live; overnight backups all green.)_ > **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its > `Written:` stamp is under **8 hours** old, read it (it carries the in-flight @@ -144,7 +144,7 @@ _As of 2026-09-27 ~0900 PT._ ### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime) -- **LIVE: `semif-serve` 0.1.3** at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging +- **LIVE: `semif-serve` 0.1.4** (was 0.1.3 until 1014 on 2026-09-27) at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code + contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.** - 0.1.3 acceptance: 144/144 parity with upstream; averaging through the service 78.6% → 88.1% @@ -167,6 +167,12 @@ _As of 2026-09-27 ~0900 PT._ and +94 ms sequential vs today's 246 ms, because the pose header costs only ~31 ms and SemIf shares GPU 1 with the LLM. Acceptable pose 67% vs 92%; the mood carried 7/15 vs 14/15. Upside: gestures at 13% vs 58%. +- **Prime 0948: idea 88 dropped; fix the 422 → DONE, 0.1.4 live (the `fix(semif): 0.1.4` commit).** INV-7 wraps SemIf's + `shared._state_prefix` so the prefix is only the tokens the full prompts share. Startup proves the fix + is in effect (the hook must be what score_shared resolves, the prefix unchanged on an ordinary state, + and a merge-prone state scored through the shared path). Folded from heid bug hunt SKAL (Hulda, thread + `01M3HXMXN27F3K534Q6QS45AHV`). Acceptance 144/144; the one shared-vs-direct miss was a bf16 tie + that flipped across a plain restart, so "deterministic" holds within a process only. ### restic: credential leak fixed (2026-09-27, Prime) @@ -222,6 +228,7 @@ _As of 2026-09-27 ~0900 PT._ - `[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** (SVOS v2.1.12: `propose_decision` gained `seat_up`, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling `zellij-fleet@Claude` at boot remains Prime's call. - `[2026-09-27]` **SemIf LIVE on fv-ml1 GPU 1 (semif-serve 0.1.2, Prime):** wrapper + contract + 39 tests, 142/144 upstream parity, two card-only memory defects fixed. → `persistent-memory.d/2026-09-27-semif-live-on-fv-ml1-gpu1.md` +- `[2026-09-27]` **semif-serve 0.1.4: object states ending in `)`, `;` or `}` no longer 422 (INV-7, a prefix wrapper proven at startup); numerics are deterministic within a process but a bf16 near-tie can flip across a restart.** Prime ruled; heid bug hunt folded. → `stacks/semif/README.md` - `[2026-09-27]` **SemIf as Cicada's mood source: slower (+32 ms async, +94 ms sequential) and worse (67% vs 92% apt; carry 7/15 vs 14/15); only the gesture restraint is a win.** Build nothing (Prime). Henge 88 carries it. → `persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md` - `[2026-09-27]` **SemIf consumer-fit spikes (Prime): Cicada affect gate 30/31 with descriptive wording and 19/31 terse; Wyrd "left this place?" 21/21 on the second wording, exit choice 18/21.** Recommendations await Prime. → `persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md` - `[2026-09-27]` **SemIf order-averaging spiked (+9.1 pts accuracy, agreement = strong confidence signal); Prime ruled: build it in as 0.1.3 and trial the fast kernels.** Tracked in the in-flight SemIf section. → `persistent-memory.d/2026-09-27-semif-order-averaging.md` — **DONE:** 0.1.3 live, with fast kernels adopted (`77b8cb4`). diff --git a/services/semif-serve/acceptance/result-2026-09-27-v0.1.4.json b/services/semif-serve/acceptance/result-2026-09-27-v0.1.4.json new file mode 100644 index 0000000..e0479f4 --- /dev/null +++ b/services/semif-serve/acceptance/result-2026-09-27-v0.1.4.json @@ -0,0 +1,63 @@ +{ + "url": "http://10.251.50.54:8032", + "health": { + "status": "ok", + "semif_commit": "23cf1f39fc9534fe81437200959b6dfc7106e45a", + "model": { + "source": "Qwen/Qwen3.5-4B", + "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", + "dtype": "bfloat16", + "device": "cuda:0", + "torch_version": "2.10.0+cu128", + "transformers_version": "5.17.0", + "device_name": "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition", + "allocated_gib": 7.84, + "reserved_gib": 8.12 + }, + "vram_cap_gib": 12.0, + "max_tokens": 4096, + "max_decisions": 64, + "workloads": [] + }, + "1_parity_vs_upstream": { + "rows": 144, + "top_choice_agree": 144, + "max_abs_prob_gap": 0.0595381559570336 + }, + "1_prompt_sha256_equal": 144, + "2_noise_floor_a_vs_b": { + "rows": 144, + "top_choice_agree": 144, + "max_abs_prob_gap": 0.0 + }, + "3_negative_rotated_options": { + "rows": 144, + "top_choice_agree": 14, + "max_abs_prob_gap": 0.9987391090270772 + }, + "4_shared_vs_direct": { + "groups": 36, + "rows": 72, + "top_choice_agree": 71, + "max_abs_prob_gap": 0.05927145076874801 + }, + "5_speed_21_binary": { + "prefix_tokens": 62, + "shared_s": { + "runs": [ + 0.13865972400526516, + 0.1382300010009203, + 0.13749008599552326 + ], + "median": 0.1382300010009203 + }, + "sequential_decide_s": { + "runs": [ + 0.9489945390087087, + 0.9463060130074155, + 0.9458249929884914 + ], + "median": 0.9463060130074155 + } + } +} \ No newline at end of file diff --git a/services/semif-serve/pyproject.toml b/services/semif-serve/pyproject.toml index 7bab3e5..4472b70 100644 --- a/services/semif-serve/pyproject.toml +++ b/services/semif-serve/pyproject.toml @@ -1,6 +1,6 @@ [project] name = "semif-serve" -version = "0.1.3" +version = "0.1.4" description = "HTTP wrapper around SemIf's direct and shared option-logit scorers" requires-python = ">=3.12" dependencies = [ diff --git a/services/semif-serve/semif-serve.contract.md b/services/semif-serve/semif-serve.contract.md index f4b365f..8142e23 100644 --- a/services/semif-serve/semif-serve.contract.md +++ b/services/semif-serve/semif-serve.contract.md @@ -104,6 +104,26 @@ Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from - **INV-5 no network at runtime.** Weights come from the mounted HF cache at the pinned revision. The entry point sets `HF_HUB_OFFLINE=1` itself before torch or transformers load, so this holds outside the image too (S10). +- **INV-7 boundary-safe shared prefix (0.1.4).** At load, the engine wraps SemIf's + module-global `semif_phase1.shared._state_prefix`, which `score_shared` calls. The wrapper + keeps only the leading tokens that upstream's prefix shares with a real full prompt for + the same state (a probe row: evidence, then a placeholder criterion). + - **Why:** upstream drops just one token at the state boundary. An object state whose + last value ends in `)`, `;` or `}` re-tokenises two tokens back once `, "criterion"` + follows, so it was refused with 422 "The fixed state prefix does not match every full + prompt" (found 2026-09-27). + - **Effect:** each row still scores the same token sequence; only the prefill/suffix split + moves, costing a few tokens of sharing. An ordinary state keeps upstream's prefix + unchanged. `score_shared` still checks every real row and fails closed. + - **Startup proves the fix is in effect, not just installed** (heid bug hunt SKAL, + 2026-09-27). Before the weights load, it refuses to start if `_state_prefix` is + missing or not callable, or if `score_shared` does not resolve it from + `semif_phase1.shared`'s globals (a SemIf bump that moves or re-exports it). After the + warm-up, it checks two things. The wrapper must return upstream's exact prefix for + the warm-up state, since a wrapper rendering the wrong prompt would silently drop all + sharing. And a state upstream alone refuses, `{"person_said": "ok :)"}`, must score + through `score_shared` itself; a hook bound before the patch would fail here. A + reload wraps the original again rather than stacking wrappers. - **INV-6 constant-time auth.** Token comparison uses `hmac.compare_digest`. The token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything else, because a CR, LF or NUL in the token can never arrive in a header (S2). @@ -169,7 +189,13 @@ cancels a position bias exactly; `all` sends n! rows and is 422 above 4 options; agreement and spread are computed from the orderings; a mixed shared request (averaged + plain) is one engine call, with results in request order and plain results unchanged; expanded rows count toward the cap; `workload` + `orderings` -→ 422. +→ 422. **Prefix (INV-7):** against a tokenizer whose merge reaches two tokens back, the +wrapped prefix is a token-prefix of every row's full prompt and gives up exactly the +mismatched token; an ordinary state keeps the whole upstream prefix; load() installs +one wrapper however many times it runs, and refuses to start when `_state_prefix` is +gone or not callable, when `score_shared` binds it early or resolves its globals +elsewhere, or when the wrapper renders the wrong prompt. The fake `score_shared` is +compiled into the fake module, so that it resolves its globals as the real one does. ## Acceptance (on fv-ml1, real model; not unit tests) @@ -180,3 +206,6 @@ results unchanged; expanded rows count toward the cap; `workload` + `orderings` 4. **Shared vs direct:** the same rows agree within the A-vs-A floor. 5. **Speed:** 21 binary criteria over one state, N ≥ 3, p50 + spread. 6. **VRAM:** the peak at a 4096-token input sets `SEMIF_VRAM_CAP_GIB`. +7. **Boundary (0.1.4):** states whose last value ends in `)`, `;` and `}`, as objects, + are all answered by `/decide/shared`, with the same top choice as a string state + holding the same text; parity (1) still holds. diff --git a/services/semif-serve/src/semif_serve/engine.py b/services/semif-serve/src/semif_serve/engine.py index c75eda6..a5f741f 100644 --- a/services/semif-serve/src/semif_serve/engine.py +++ b/services/semif-serve/src/semif_serve/engine.py @@ -28,6 +28,44 @@ WARMUP_ROW = { } +# INV-7 startup proof: a state SemIf's own prefix refuses (measured with the real tokenizer, +# 2026-09-27), sent through score_shared itself, so the fix is shown to be IN EFFECT, not just +# installed (bug hunt SKAL, R2/H1/H2). +BOUNDARY_ROW = {**WARMUP_ROW, "id": "semif-serve-boundary-check", "state": {"person_said": "ok :)"}} +PREFIX_PROBE_ROW = { + "id": "semif-serve-prefix-probe", + "question": "prefix boundary placeholder", + "options": [{"id": "yes", "description": "Yes"}, {"id": "no", "description": "No"}], +} + + +def boundary_safe_prefix(state_prefix: Callable, messages: Callable) -> Callable: + """INV-7: wrap SemIf's `shared._state_prefix` so the prefill never covers a token that the full + prompts do not share. + + Upstream encodes the prompt up to the end of the state and drops ONE token, because the JSON + punctuation that follows the state can merge with it. One is not always enough: an object + state whose last value ends in ")", ";" or "}" re-tokenises TWO tokens back once `, "criterion"` + follows, and score_shared then refused the request with 422 (found 2026-09-27). This keeps only + the leading tokens that upstream's prefix shares with a real full prompt for the same state. + The suffix starts that much earlier and scores the same token sequence; the cost is a few + tokens of lost sharing. Only the punctuation run at the boundary can merge, and it is the same + in every row, so a probe row stands for all of them. score_shared still checks every real row + and fails closed if that ever stops holding.""" + def prefix(tokenizer, state): + ids = state_prefix(tokenizer, state) + full = tokenizer.encode(tokenizer.apply_chat_template( + messages({**PREFIX_PROBE_ROW, "state": state}), tokenize=False, add_generation_prompt=True, + enable_thinking=False), add_special_tokens=False) + shared = 0 + while shared < min(len(ids), len(full)) and ids[shared] == full[shared]: + shared += 1 + return ids[:shared] + + prefix.semif_serve_wraps = state_prefix + return prefix + + def _first_line(exc: BaseException) -> str: lines = str(exc).splitlines() return lines[0] if lines else "" @@ -44,10 +82,24 @@ class TorchEngine: @classmethod def load(cls, settings: Settings) -> "TorchEngine": import torch - from semif_phase1.core import load_causal_model + import semif_phase1.shared as upstream_shared + from semif_phase1.core import direct_messages, load_causal_model from semif_phase1.direct import score from semif_phase1.shared import score_shared + # INV-7: score_shared looks `_state_prefix` up as a module global, so the wrapper goes there. + # A SemIf bump that renames it, or moves score_shared so it resolves its globals elsewhere, + # must stop startup, not silently bring the 422 back. + original = getattr(upstream_shared, "_state_prefix", None) + if not callable(original): + raise RuntimeError("semif_phase1.shared._state_prefix is gone or not callable at this SemIf " + "commit: re-check the boundary-safe prefix (INV-7) before serving") + original = getattr(original, "semif_serve_wraps", original) # one wrapper, however many loads + upstream_shared._state_prefix = boundary_safe_prefix(original, direct_messages) + if getattr(score_shared, "__globals__", {}).get("_state_prefix") is not upstream_shared._state_prefix: + raise RuntimeError("INV-7: score_shared does not resolve semif_phase1.shared._state_prefix, " + "so the boundary-safe prefix would be inert") + if settings.device == "cuda": if not torch.cuda.is_available(): raise RuntimeError("SEMIF_DEVICE=cuda but torch sees no CUDA device") @@ -70,10 +122,23 @@ class TorchEngine: raise RuntimeError(f"model landed on {placed}, expected {settings.device}") engine = cls(torch, model, tokenizer, metadata, settings, direct_fn=score, shared_fn=score_shared) engine.direct(WARMUP_ROW) # INV-3: one decision must score + engine._prove_prefix_hook(upstream_shared._state_prefix, original) if settings.device == "cuda": # INV-4: the resting footprint engine._release_above = torch.cuda.memory_reserved(0) + RELEASE_SLACK_BYTES return engine + def _prove_prefix_hook(self, hook: Callable, original: Callable) -> None: + """INV-7, at startup: the wrapper keeps upstream's whole prefix on an ordinary state (a wrapper + rendering the wrong prompt would silently give up all sharing), and a state upstream alone + refuses scores through score_shared itself (a hook score_shared never calls would not).""" + state = WARMUP_ROW["state"] + if hook(self._tokenizer, state) != original(self._tokenizer, state): + raise RuntimeError("INV-7: the boundary-safe prefix does not keep upstream's prefix on an ordinary state") + try: + self.shared([BOUNDARY_ROW]) + except ValueError as exc: + raise RuntimeError(f"INV-7: shared scoring still refuses a merge-prone state: {exc}") from None + def health(self) -> dict: info = dict(self._metadata) if self._settings.device == "cuda": diff --git a/services/semif-serve/tests/fake_tokenizer.py b/services/semif-serve/tests/fake_tokenizer.py new file mode 100644 index 0000000..a84dc0a --- /dev/null +++ b/services/semif-serve/tests/fake_tokenizer.py @@ -0,0 +1,37 @@ +"""Test doubles for SemIf's tokenizer seam (INV-7), shared by test_prefix and test_engine_load. +MergeTokenizer reproduces the real failure's shape without the Qwen vocabulary; the real +tokenizer is checked on the card, in acceptance.""" +import json + + +class MergeTokenizer: + """Char-level with three merges, in the shape of the real failure. ')"' and '"}' are single + tokens, but ')"},' splits as [')', '"},']. So a state tail ')"}' encodes as [')"', '}'] when + nothing follows it (the prefix), and the full prompt's ')"},' encodes as [')', '"},']. + Dropping one token from the prefix leaves ')"', which the full prompt does not contain. A + tail like '."}' stays ['.', '"}'] either way, which is the ordinary case.""" + + def apply_chat_template(self, messages, tokenize, add_generation_prompt, enable_thinking): + assert tokenize is False and add_generation_prompt is True and enable_thinking is False + return "" + "|".join(m["content"] for m in messages) + "" + + def encode(self, text, add_special_tokens): + assert add_special_tokens is False + out, i = [], 0 + while i < len(text): + if text.startswith(')"},', i): + out += [')', '"},']; i += 4 + elif text.startswith(')"', i) or text.startswith('"}', i): + out.append(text[i:i + 2]); i += 2 + else: + out.append(text[i]); i += 1 + return out + + +def messages(row): # the shape of semif_phase1.core.direct_messages: evidence first, then the criterion + return [{"role": "user", "content": json.dumps( + {"evidence": row["state"], "criterion": row["question"], "options": row["options"]}, ensure_ascii=False)}] + + +def upstream_prefix(tokenizer, state): # what SemIf's _state_prefix does: through the state, minus one token + return tokenizer.encode("" + json.dumps({"evidence": state}, ensure_ascii=False)[:-1], add_special_tokens=False)[:-1] diff --git a/services/semif-serve/tests/test_engine_load.py b/services/semif-serve/tests/test_engine_load.py index 9143270..faee356 100644 --- a/services/semif-serve/tests/test_engine_load.py +++ b/services/semif-serve/tests/test_engine_load.py @@ -6,6 +6,8 @@ import types import pytest +from fake_tokenizer import MergeTokenizer, messages, upstream_prefix + from semif_serve.config import Settings TOKEN = "t" * 40 @@ -24,6 +26,25 @@ class Model: yield Param(self._device) +# The shape of SemIf's score_shared, compiled INTO the fake module so that, like the real one, it +# resolves _state_prefix and direct_messages as module globals at call time and refuses (ValueError) +# a prefix that is not a token-prefix of every row. +SHARED_SRC = """ +def score_shared(model, tokenizer, rows, metadata, max_tokens=4096): + prefix = _state_prefix(tokenizer, rows[0]["state"]) + for row in rows: + ids = tokenizer.encode(tokenizer.apply_chat_template(direct_messages(row), tokenize=False, + add_generation_prompt=True, enable_thinking=False), add_special_tokens=False) + if not prefix or ids[:len(prefix)] != prefix: + raise ValueError("The fixed state prefix does not match every full prompt") + CALLS.append(("shared", rows[0]["state"])) + return [{"id": row["id"]} for row in rows], {} +""" +# A SemIf revision that binds the helper when score_shared is defined: the module global can be +# replaced all day and scoring never sees it. +SHARED_SRC_EARLY_BOUND = SHARED_SRC.replace("max_tokens=4096):", "max_tokens=4096, _state_prefix=_state_prefix):") + + @pytest.fixture def fakes(monkeypatch): calls = [] @@ -42,7 +63,7 @@ def fakes(monkeypatch): def load_causal_model(model, revision, device, dtype): calls.append(("load", model, revision, device, dtype)) - return Model(state["device"]), object(), {"source": model} + return Model(state["device"]), MergeTokenizer(), {"source": model} def score(model, tok, row, meta, max_tokens): calls.append(("score", row["id"])) @@ -52,10 +73,12 @@ def fakes(monkeypatch): core = types.ModuleType("semif_phase1.core") core.load_causal_model = load_causal_model + core.direct_messages = messages direct = types.ModuleType("semif_phase1.direct") direct.score = score shared = types.ModuleType("semif_phase1.shared") - shared.score_shared = lambda *a: ([], {}) + shared._state_prefix, shared.direct_messages, shared.CALLS = upstream_prefix, messages, calls + exec(SHARED_SRC, shared.__dict__) pkg = types.ModuleType("semif_phase1") for name, mod in {"torch": torch, "semif_phase1": pkg, "semif_phase1.core": core, "semif_phase1.direct": direct, "semif_phase1.shared": shared}.items(): @@ -67,7 +90,7 @@ def test_load_caps_before_the_weights_land_then_warms_up(fakes): from semif_serve.engine import TorchEngine _torch, calls, _ = fakes TorchEngine.load(Settings(api_token=TOKEN, vram_cap_gib=12.0)) - assert [c[0] for c in calls] == ["cap", "load", "score"] + assert [c[0] for c in calls] == ["cap", "load", "score", "shared"] assert calls[0] == ("cap", round(12 / 96, 4)) assert calls[2] == ("score", "semif-serve-warmup") @@ -108,3 +131,77 @@ def test_the_entry_point_forces_offline_mode_before_the_engine_loads(fakes, monk classmethod(lambda cls, s: seen.update(offline=os.environ.get("HF_HUB_OFFLINE")) or object())) main.app_from_env() assert seen["offline"] == "1" + + +def test_load_wraps_the_upstream_state_prefix_once_even_across_reloads(fakes): + """INV-7: score_shared looks _state_prefix up as a module global, so the wrapper must replace + it there. A second load must wrap the ORIGINAL again, not stack wrapper on wrapper.""" + import semif_phase1.shared as upstream + from semif_serve.engine import TorchEngine + for _ in range(2): + TorchEngine.load(Settings(api_token=TOKEN)) + assert upstream._state_prefix is not upstream_prefix + assert upstream._state_prefix.semif_serve_wraps is upstream_prefix + + +def test_load_fails_closed_when_the_pinned_semif_no_longer_has_the_prefix_hook(fakes, monkeypatch): + import semif_phase1.shared as upstream + from semif_serve.engine import TorchEngine + _torch, calls, _ = fakes + monkeypatch.delattr(upstream, "_state_prefix") + with pytest.raises(RuntimeError, match="_state_prefix"): + TorchEngine.load(Settings(api_token=TOKEN)) + assert not any(c[0] == "load" for c in calls) + + +def test_load_proves_the_hook_on_a_merge_prone_state_through_the_real_shared_path(fakes): + """Bug hunt SKAL (R2/H1, H2): existence is not effect. Startup scores one state that upstream + alone refuses, through score_shared itself.""" + import semif_phase1.shared as upstream + from semif_serve.engine import BOUNDARY_ROW, TorchEngine + _torch, calls, _ = fakes + with pytest.raises(ValueError): # the probe really is merge-prone here + upstream.score_shared(None, MergeTokenizer(), [BOUNDARY_ROW], {}) + TorchEngine.load(Settings(api_token=TOKEN)) + assert ("shared", BOUNDARY_ROW["state"]) in calls + + +def test_load_fails_closed_when_score_shared_binds_the_prefix_before_the_patch(fakes): + import semif_phase1.shared as upstream + from semif_serve.engine import TorchEngine + exec(SHARED_SRC_EARLY_BOUND, upstream.__dict__) + with pytest.raises(RuntimeError, match="INV-7"): + TorchEngine.load(Settings(api_token=TOKEN)) + + +def test_load_fails_closed_when_score_shared_resolves_its_globals_elsewhere(fakes): + import semif_phase1.shared as upstream + from semif_serve.engine import TorchEngine + _torch, calls, _ = fakes + elsewhere = {"_state_prefix": upstream_prefix, "direct_messages": messages, "CALLS": calls} + exec(SHARED_SRC, elsewhere) + upstream.score_shared = elsewhere["score_shared"] # re-exported from another module + with pytest.raises(RuntimeError, match="INV-7"): + TorchEngine.load(Settings(api_token=TOKEN)) + assert not any(c[0] == "load" for c in calls) + + +def test_load_fails_closed_when_the_hook_is_not_callable(fakes, monkeypatch): + import semif_phase1.shared as upstream + from semif_serve.engine import TorchEngine + _torch, calls, _ = fakes + monkeypatch.setattr(upstream, "_state_prefix", "not a function") + with pytest.raises(RuntimeError, match="_state_prefix"): + TorchEngine.load(Settings(api_token=TOKEN)) + assert not any(c[0] == "load" for c in calls) + + +def test_load_fails_closed_when_the_wrapper_renders_the_wrong_prompt(fakes, monkeypatch): + """The SURVIVED row of bug hunt SKAL: a wrapper that renders nothing still yields a valid + (degenerate) prefix, so scoring stays correct and silently loses all sharing. Startup + requires the wrapper to keep upstream's whole prefix on an ordinary state.""" + import semif_phase1.core as core + from semif_serve.engine import TorchEngine + monkeypatch.setattr(core, "direct_messages", lambda row: []) + with pytest.raises(RuntimeError, match="INV-7"): + TorchEngine.load(Settings(api_token=TOKEN)) diff --git a/services/semif-serve/tests/test_prefix.py b/services/semif-serve/tests/test_prefix.py new file mode 100644 index 0000000..1744f3c --- /dev/null +++ b/services/semif-serve/tests/test_prefix.py @@ -0,0 +1,39 @@ +"""The shared prefix must never cover a token that the full prompts do not share (0.1.4, INV-7). + +Found 2026-09-27 by the Cicada mood spike. An object state whose last value ends in ")", ";" or +"}" was refused with 422 "The fixed state prefix does not match every full prompt". SemIf's +_state_prefix encodes the prompt up to the end of the state and drops ONE token, because the +JSON punctuation that follows can merge with it. Here the merge reaches two tokens back. +MergeTokenizer reproduces that shape without the real Qwen vocabulary. The real tokenizer is +checked on the card, in acceptance. +""" +from fake_tokenizer import MergeTokenizer, messages, upstream_prefix + +from semif_serve.engine import boundary_safe_prefix + +ROW = {"question": "Which pose?", "options": [{"id": "a", "description": "A"}, {"id": "b", "description": "B"}]} + + +def full(tok, state, row=ROW): + return tok.encode(tok.apply_chat_template(messages({**row, "state": state}), tokenize=False, + add_generation_prompt=True, enable_thinking=False), add_special_tokens=False) + + +def test_the_merge_that_broke_upstream_is_reproduced(): + tok, state = MergeTokenizer(), {"person_said": "ok :)"} + ids = upstream_prefix(tok, state) + assert full(tok, state)[:len(ids)] != ids # the 422: upstream's prefix is not a prefix + + +def test_a_state_ending_in_a_merging_character_gets_a_prefix_every_row_shares(): + tok, state = MergeTokenizer(), {"person_said": "ok :)"} + ids = boundary_safe_prefix(upstream_prefix, messages)(tok, state) + for row in (ROW, {"question": "A different criterion entirely?", "options": ROW["options"][::-1]}): + assert full(tok, state, row)[:len(ids)] == ids + assert ids == upstream_prefix(tok, state)[:-1] # it gave up exactly the one mismatched token + + +def test_an_ordinary_state_keeps_the_whole_upstream_prefix(): + tok = MergeTokenizer() + for state in ({"person_said": "Turn off the lights."}, "ok :)", "Nothing was said."): + assert boundary_safe_prefix(upstream_prefix, messages)(tok, state) == upstream_prefix(tok, state) diff --git a/services/semif-serve/uv.lock b/services/semif-serve/uv.lock index c68f554..9d319df 100644 --- a/services/semif-serve/uv.lock +++ b/services/semif-serve/uv.lock @@ -991,7 +991,7 @@ dependencies = [ [[package]] name = "semif-serve" -version = "0.1.3" +version = "0.1.4" source = { editable = "." } dependencies = [ { name = "fastapi" }, diff --git a/stacks/semif/.env.example b/stacks/semif/.env.example index 08899fa..8b489e1 100644 --- a/stacks/semif/.env.example +++ b/stacks/semif/.env.example @@ -1,6 +1,6 @@ # semif — copy to /opt/docker/compose/semif/.env on fv-ml1 (mode 0600). # Built on fv-ml1 from services/semif-serve (see README "Building"). -IMAGE=semif-serve:0.1.3 +IMAGE=semif-serve:0.1.4 PORT=8032 HOST_IP=10.251.50.54 # fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero, scriberr). diff --git a/stacks/semif/README.md b/stacks/semif/README.md index 8e010bb..1271674 100644 --- a/stacks/semif/README.md +++ b/stacks/semif/README.md @@ -53,12 +53,15 @@ curl -s -H "Authorization: Bearer $T" http://10.251.50.54:8032/decide -d '{ 644 suffix tokens, 5 calls after warm-up). One cold call at that size returned 503. The VRAM table below covers binary decisions only. For many options, use one ordering or shorter option text. -- ⚠ **`/decide/shared` refuses some object states.** If the state is an object whose - LAST value ends in `)`, `;` or `}`, the service returns 422 "The fixed state prefix - does not match every full prompt". The closing `"}` merges with that character into one - token. The same text as a plain string state works, and `.`, `!`, `?`, `]`, `…` and - `—` endings work. Not fixed yet. Callers that pass user text last should - append a full stop or send a string state. +- **Object states ending in `)`, `;` or `}` work as of 0.1.4.** Until then, if the state was + an object whose last value ended in one of those characters, the service returned 422 "The + fixed state prefix does not match every full prompt". SemIf trims one token at the state + boundary, and the JSON that follows re-merged two tokens back. The fix (contract INV-7) + moves only where the shared prefix ends; every row still scores the same tokens. Measured + with the real tokenizer: of 154 states (22 endings × 7 shapes) SemIf alone refused 23, and + with the fix none. None of the 131 ordinary states, nor any of the authored144 states, + changed its prefix. Startup proves the fix is live by scoring `{"person_said": "ok :)"}` + through the shared path. ## ⚠ Probabilities are uncalibrated @@ -123,6 +126,21 @@ warm-up fails, and startup fails closed. Build without the kernels: | 6 orderings, short | 118 ms | 81 ms | | 3 rotations, ~2,000-token state | 200 ms | 158 ms | +## Acceptance (2026-09-27, v0.1.4) + +Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.4.json`. Parity with upstream +144/144 (identical prompt hashes, max prob gap 0.060). Deterministic within the process +(A-vs-A gap 0.0). The negative control fails as it should (14/144). Shared vs direct 71/72. +The one miss is an exact bf16 tie in the shared result (0.444/0.444). INV-7 did not move that +row's prefix. After a plain `docker restart` of the same image, the row read 0.369/0.537 and +agreed. + +⚠ **So "deterministic" holds within one process, not across restarts.** Logits come in bf16 +steps (0.125 here), and a near-tie can land differently after a restart, most likely because +the fast kernels autotune at startup. That is n=1 row across one restart, and the cross-restart +floor is otherwise unmeasured. When comparing two versions, compare them against that floor +and not against zero. + ## Acceptance (2026-09-27, v0.1.3) Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json`, @@ -149,16 +167,25 @@ now matches the label. So the misses come from the numeric path, not the wrapper speed figure includes one network round trip (~27 ms). The burst release costs ~16 ms on it (0.1.1 measured 143 ms without the release). -## Building +## Building and deploying ```bash -# from nh3-dev -tar -C services/semif-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ --exclude=acceptance . \ - | ssh infra-ops@10.251.50.54 'mkdir -p /opt/docker/src/semif-serve-X.Y.Z && tar -x -C /opt/docker/src/semif-serve-X.Y.Z' +# from nh3-dev. /opt/docker/src is root-owned, so create the version dir with sudo first. +ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g infra-ops /opt/docker/src/semif-serve-X.Y.Z' +tar -C services/semif-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ \ + --exclude=acceptance --exclude=spike --exclude='*.egg-info' . \ + | ssh infra-ops@10.251.50.54 'tar -x -C /opt/docker/src/semif-serve-X.Y.Z' # on fv-ml1 cd /opt/docker/src/semif-serve-X.Y.Z && docker build -t semif-serve:X.Y.Z . +# ⚠ /opt/docker/compose/semif is root-owned, so `sed -i` cannot write its temp file there. .env +# itself is infra-ops's: rewrite it IN PLACE, which keeps its owner and mode (0600). +cd /opt/docker/compose/semif && new=$(sed 's/^IMAGE=.*/IMAGE=semif-serve:X.Y.Z/' .env) \ + && printf '%s\n' "$new" > .env && docker compose up -d ``` +Startup fails closed (INV-3, INV-7), so a container that does not reach healthy did not pass +its own checks: read `docker logs semif`. Record the deploy with `scripts/ops-log`. + To move SemIf forward: bump the commit in `services/semif-serve/pyproject.toml` (and `SEMIF_COMMIT` in `config.py`), run `uv lock`, rebuild, and **re-run the acceptance**. Upstream is research code that changes weekly, which is why it is