Commit Graph
4 Commits
Author SHA1 Message Date
vh e268ff7c99 spike(semif): consumer fit for Wyrd scene change and Cicada affect gate (no service change)
Harness consumer_fit.py runs a consumer's per-turn decisions over hand-labelled
cases, with rotations, a content-free null control and tagged positive controls.

Cicada: an input-only 'does this earn a visible reaction?' gate scored 30/31
with descriptive options and 19/31 with terse yes/no options. Scoped by
Cicada's 2026-09-20 ruling (affect is emitted once, no mood-ring classifier).

Wyrd: on 3 real seed graphs, the first place-change wording failed its
positive controls (1/6 moves). A location-anchored rewording scored 21/21, and
exit selection scored 18/21. Semif fits the choice, not writing the node.
2026-09-27 09:20:35 -07:00
vh 77b8cb449c feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)
Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
  ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
  top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
  (group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.

Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.

Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
  >= 1 (S1); the token must be visible ASCII (S2); the calibration file must
  exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
  Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
  gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
  (C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
  the body read, a shared-route lock, calibration pass-through, the gc cycle,
  the exact caps, TorchEngine.load's arch and device checks, and the offline
  entry point.
86 tests.

Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
2026-09-27 03:27:15 -07:00
vh 739aa03123 spike(semif): latency profile and order-averaging measurement (no service change)
Latency, measured from nh3-dev (3 runs x 20 per condition; network floor 31 ms):
- /decide short: 71 ms end to end, 38 ms server-side;
- /decide with a ~2,000-token state: 210 / 169 ms;
- shared, 3 rotations: 113 / 79 ms;
- shared, 6 orderings: 137 / 99 ms.
Qwen3.5's fast kernels (causal_conv1d, flash-linear-attention) are not installed,
so transformers falls back to its reference PyTorch paths. That is a speed lever,
and using it needs a parity re-check.

Averaging over option orderings, on SemIf authored144 + perturbations108 (252 rows,
72 groups):
- a single ordering scores 78.6%;
- log-mean over the 3 rotations scores 87.7% (+9.1 pts, group-bootstrap 95% CI +4.7
  to +13.8);
- all 6 permutations score 88.1%.
Rotations capture nearly all of the gain. Rows where the rotations agree unanimously
(161) are 94.4% accurate; split rows (91) are 75.8%.
2026-09-27 02:50:14 -07:00
vh 069725c4b3 feat(semif): SemIf option-logit decisions on fv-ml1 GPU 1 (Prime)
services/semif-serve is a FastAPI wrapper around SemIf's direct and shared torch
scorers (SemIf-OpenJev @ 23cf1f39, MIT). Upstream ships only a batch CLI. The
wrapper loads the pinned Qwen3.5-4B (851bf6e8, BF16) once from the offline HF
cache and returns SemIf's result dicts unchanged, with an optional per-workload
temperature-calibrated view. Contract: semif-serve.contract.md. Built with a
short contract, TDD (39 tests, fake engine and fake torch, no GPU) and a heid
bug-hunt panel (pending).

On the card:
- torch 2.10.0+cu128 with sm_120 kernels, which is SemIf's own stack;
- a hard 12 GiB VRAM cap.
Two defects surfaced only on the card, and each fix is covered by a test:
- 0.1.1: an OOM raised as a chained exception kept the failed request's tensors
  alive (11.9 GiB after the 503). It is now raised unchained, after gc.
- 0.1.2: a large request left 12.6 GB reserved on the shared card. After each
  call, reserved memory over the baseline + 512 MiB is now released.

Acceptance against SemIf's committed torch predictions (authored144):
- 142/144 same top choice; both misses are exact bf16 ties;
- 144/144 identical prompt hashes;
- deterministic A-vs-A;
- negative control 14/144;
- shared vs direct 72/72.
21 binary criteria over one state take 159 ms. The shared-mode capacity table
under the cap is in stacks/semif/README.md.

The Dockerfile installs dependencies from a manifest with the project version
blanked, so a version bump reuses the ~4 GB torch layer. Verified: 41 s rebuild,
dependency layer CACHED.

DNS: semif.fv.internal. Token: vault semif/api-token.
2026-09-27 02:36:56 -07:00