spike(semif): consumer fit for Wyrd scene change and Cicada affect gate (no service change)

Harness consumer_fit.py runs a consumer's per-turn decisions over hand-labelled
cases, with rotations, a content-free null control and tagged positive controls.

Cicada: an input-only 'does this earn a visible reaction?' gate scored 30/31
with descriptive options and 19/31 with terse yes/no options. Scoped by
Cicada's 2026-09-20 ruling (affect is emitted once, no mood-ring classifier).

Wyrd: on 3 real seed graphs, the first place-change wording failed its
positive controls (1/6 moves). A location-anchored rewording scored 21/21, and
exit selection scored 18/21. Semif fits the choice, not writing the node.
This commit is contained in:
vh
2026-09-27 09:20:35 -07:00
parent 3048e4194f
commit e268ff7c99
19 changed files with 8114 additions and 6 deletions
@@ -0,0 +1,99 @@
# SemIf consumer-fit spikes: Wyrd scene change, Cicada affect gate (2026-09-27, Prime)
Prime, 0904: "Two semif spikes -- Wyrd has a couple flags per turn on whether the scene should
change, check the codebase, see if semif is a clean fit and then spike a couple of sample
scenarios. likewise, cicada has a speech module and some reactions-- spike whether or not she
should show emotion." No service change. Code and raw output: `services/semif-serve/spike/`
(`consumer_fit.py` harness; `wyrd_cases.py`, `cicada_cases.py` build the scenario files;
`consumer-fit-2026-09-27/` holds scenarios, `.result.json` and `.log` per run).
## Method (applies to both)
- One `/decide/shared` per turn, every decision with `"orderings": "rotations"`. Target:
semif-serve 0.1.3 on fv-ml1 GPU 1, client on nh3-dev.
- **Noise floor:** the service is deterministic (acceptance: 144 rows twice, gap 0.0), and the
harness asserts the same top on all 3 runs of every case. Repeats measure latency only.
- **Null control:** every case is also asked over a content-free state. It always picks the
"nothing happened" option, so its accuracy is that label's base rate (Cicada 16/31, Wyrd
15/21). Evidence conditions are read against that, not against 0.
- **Positive controls:** cases tagged `control`, unambiguous by construction.
- **Limits, stated:** one labeller (me), authored inputs, 31 Cicada and 21 Wyrd cases per
decision. Wilson 95% intervals are wide (30/31 → 0.84–0.99; 21/21 → 0.85–1.0), so nothing
here resolves differences under ~10–15 points. The winning wordings were chosen on this same
set, so their scores are optimistic until checked on held-out real play.
## Wyrd: scene change (fit: yes for the CHOICE, no for the WRITING)
How it works today (from the code): no dedicated flag. The GM's `gen` structured-output calls
carry it two ways. **R8** (reflect): `plot_direction.target_node`, accepted only if it is an
advanceable exit; `plot_direction.intrude` is "advisory" and never read anywhere. **Path B**:
raw `AddStoryNode` + `AdvanceToNode` events, gated only by prompt prose ("Add + advance to a NEW
node ONLY when the scene moves to a genuinely different place — never every turn"). Wyrd's
archival memory already says "if it churns, tighten the prompt (or gate node-adds)".
**Three of the real campaigns in `.wyrd-data/canon` have ZERO story edges, so R8 can never fire
there and every scene change went through ungated path B.**
Fit: both are typed choices over one state, exactly semif's shape. Semif's exit options can only
be "stay" plus the real advanceable exits, so it cannot name an illegal target. It cannot write the
new node (title, summary, mood, image prompt) or the steering directive; the GM keeps those. So
it would be a gate/selector in front of the GM, not a replacement.
Spike: 3 real seed graphs (The Fletcher, The Toll Road, The Last Vigil + one added edge), 21
authored turns in the campaigns' register (the transcripts are gone from conv-api:
`session_not_found`). State = current scene + player action + narration.
| decision | wording | correct | real moves found | controls |
|---|---|---|---|---|
| `place`: left this place? | r1 "has the scene moved … or does it still hold?" | 16/21 (null 15/21) | **1/6** | 4/7 |
| `place2` | r2 location-anchored: "began at the {loc}; by the end, where are they?" | **21/21**, all unanimous, p ≥ 0.91 | 6/6 | 7/7 |
| `exit`: which exit? | r1 "which one does this turn move the story into?" | **18/21** | 6/6 | 6/7 |
| `exit2` | r2 "only once the narration actually gets there" | 16/21 | 1/6 | 4/7 |
- Round 1 `place` failed its positive controls (no better than the blind null), so the
instrument, not the model, was broken. `place2` was written after seeing that.
- `exit` r1 misses: two false advances on Toll-Road conversation turns (p 0.59 split, p 0.67
unanimous), and `v-to-pub`, where my label said stay and semif chose "Outside the Sanctum". On
re-reading, semif's pick is defensible (they do run outside first). All its correct advances
had p ≥ 0.87 except one at 0.67. A threshold would be tuned on the test set, so none is proposed.
- Latency: 4 decisions per turn, 133–179 ms end to end (101–141 ms server). The first request
of two files took ~700–800 ms (n=2, cause not isolated). Either way it is in dead time and is
far below the GM's reflect.
- Side findings from the code read, not fixed (they are Wyrd's): `story_node.created_turn` is
never written, so "scenes so far (in order)" is alphabetical and includes unvisited seeds;
and a late "ready" for an abandoned node's backdrop can still replace the one on screen.
## Cicada: should she show emotion (a gate, not a mood ring)
Constraint found first: Cicada persistent-memory 2026-09-20, operator: **"Affect is emitted once
and fanned out … nothing derives affect a second time … A classifier inferring emotion from
output text is the mood-ring failure."** So semif does not pick her pose and never reads her
reply. What is open: talk's `/face` loop measured the model gesturing on 8–11 of 18 mundane
turns (target under 1 in 3), and its README names "the page dropping gestures" as the next lever,
"a design call, not a tuning pass". The spike measures that lever: a gate on what the PERSON
said, run beside the LLM.
| wording | correct | mundane → "ordinary" | earned → "react" | controls |
|---|---|---|---|---|
| w1: descriptive options (news, joke, surprise, distress, a turn / command, plain question, small talk) | **30/31** | 16/16 | 14/15 | 5/5 |
| w2: "Should the assistant's face visibly react to this?" Yes/No | 19/31 | 16/16 | **3/15** | 4/5 |
- w2 said "ordinary" to "My dog died this morning" at p 0.93. **The wording carries the whole
definition.** Terse yes/no options fail; descriptive ones work.
- Adding the earlier lines of the conversation changed no scored answer.
- The one w1 miss: "Order four hundred pounds of cheese" → ordinary. Open (unscored) cases:
"Goodnight" → react (p 0.58), "Thanks, you're the best" → react (0.89), "STOP THE TIMER" →
ordinary.
- Latency ~105 ms end to end (76 ms server), under talk's ~0.4–0.6 s pose decode, so it adds
nothing to the turn if it starts at end of utterance.
## Recommendations (surfaced to Prime; nothing built)
- **Cicada:** the cleanest fit of the two. Gate **gestures only**: on "ordinary", drop the model's
gesture and leave its pose alone. A miss then costs one missing gesture; gating poses would
leave her neutral on bad news. It reads input, never her output, and never picks WHICH feeling,
but it does override the model's affect. That is Prime's ruling to make, and the code would
live in tts-stack talk `/face`, not in cicada (affect is not on cicada's v1 list).
- **Wyrd:** `place2` is the node-add gate the archival note asked for. R8 exit choice works but
its wording is fragile. Wyrd's v1 is met and no churn has been observed (2–6 nodes per real
campaign), so by the roadmap gate this is a parking-lot item with a trigger: build it if live
play shows churn, after labelling ~50 real turns as a held-out set.