spike(semif): consumer fit for Wyrd scene change and Cicada affect gate (no service change)
Harness consumer_fit.py runs a consumer's per-turn decisions over hand-labelled cases, with rotations, a content-free null control and tagged positive controls. Cicada: an input-only 'does this earn a visible reaction?' gate scored 30/31 with descriptive options and 19/31 with terse yes/no options. Scoped by Cicada's 2026-09-20 ruling (affect is emitted once, no mood-ring classifier). Wyrd: on 3 real seed graphs, the first place-change wording failed its positive controls (1/6 moves). A location-anchored rewording scored 21/21, and exit selection scored 18/21. Semif fits the choice, not writing the node.
This commit is contained in:
@@ -0,0 +1,99 @@
|
||||
# SemIf consumer-fit spikes: Wyrd scene change, Cicada affect gate (2026-09-27, Prime)
|
||||
|
||||
Prime, 0904: "Two semif spikes -- Wyrd has a couple flags per turn on whether the scene should
|
||||
change, check the codebase, see if semif is a clean fit and then spike a couple of sample
|
||||
scenarios. likewise, cicada has a speech module and some reactions-- spike whether or not she
|
||||
should show emotion." No service change. Code and raw output: `services/semif-serve/spike/`
|
||||
(`consumer_fit.py` harness; `wyrd_cases.py`, `cicada_cases.py` build the scenario files;
|
||||
`consumer-fit-2026-09-27/` holds scenarios, `.result.json` and `.log` per run).
|
||||
|
||||
## Method (applies to both)
|
||||
|
||||
- One `/decide/shared` per turn, every decision with `"orderings": "rotations"`. Target:
|
||||
semif-serve 0.1.3 on fv-ml1 GPU 1, client on nh3-dev.
|
||||
- **Noise floor:** the service is deterministic (acceptance: 144 rows twice, gap 0.0), and the
|
||||
harness asserts the same top on all 3 runs of every case. Repeats measure latency only.
|
||||
- **Null control:** every case is also asked over a content-free state. It always picks the
|
||||
"nothing happened" option, so its accuracy is that label's base rate (Cicada 16/31, Wyrd
|
||||
15/21). Evidence conditions are read against that, not against 0.
|
||||
- **Positive controls:** cases tagged `control`, unambiguous by construction.
|
||||
- **Limits, stated:** one labeller (me), authored inputs, 31 Cicada and 21 Wyrd cases per
|
||||
decision. Wilson 95% intervals are wide (30/31 → 0.84–0.99; 21/21 → 0.85–1.0), so nothing
|
||||
here resolves differences under ~10–15 points. The winning wordings were chosen on this same
|
||||
set, so their scores are optimistic until checked on held-out real play.
|
||||
|
||||
## Wyrd: scene change (fit: yes for the CHOICE, no for the WRITING)
|
||||
|
||||
How it works today (from the code): no dedicated flag. The GM's `gen` structured-output calls
|
||||
carry it two ways. **R8** (reflect): `plot_direction.target_node`, accepted only if it is an
|
||||
advanceable exit; `plot_direction.intrude` is "advisory" and never read anywhere. **Path B**:
|
||||
raw `AddStoryNode` + `AdvanceToNode` events, gated only by prompt prose ("Add + advance to a NEW
|
||||
node ONLY when the scene moves to a genuinely different place — never every turn"). Wyrd's
|
||||
archival memory already says "if it churns, tighten the prompt (or gate node-adds)".
|
||||
**Three of the real campaigns in `.wyrd-data/canon` have ZERO story edges, so R8 can never fire
|
||||
there and every scene change went through ungated path B.**
|
||||
|
||||
Fit: both are typed choices over one state, exactly semif's shape. Semif's exit options can only
|
||||
be "stay" plus the real advanceable exits, so it cannot name an illegal target. It cannot write the
|
||||
new node (title, summary, mood, image prompt) or the steering directive; the GM keeps those. So
|
||||
it would be a gate/selector in front of the GM, not a replacement.
|
||||
|
||||
Spike: 3 real seed graphs (The Fletcher, The Toll Road, The Last Vigil + one added edge), 21
|
||||
authored turns in the campaigns' register (the transcripts are gone from conv-api:
|
||||
`session_not_found`). State = current scene + player action + narration.
|
||||
|
||||
| decision | wording | correct | real moves found | controls |
|
||||
|---|---|---|---|---|
|
||||
| `place`: left this place? | r1 "has the scene moved … or does it still hold?" | 16/21 (null 15/21) | **1/6** | 4/7 |
|
||||
| `place2` | r2 location-anchored: "began at the {loc}; by the end, where are they?" | **21/21**, all unanimous, p ≥ 0.91 | 6/6 | 7/7 |
|
||||
| `exit`: which exit? | r1 "which one does this turn move the story into?" | **18/21** | 6/6 | 6/7 |
|
||||
| `exit2` | r2 "only once the narration actually gets there" | 16/21 | 1/6 | 4/7 |
|
||||
|
||||
- Round 1 `place` failed its positive controls (no better than the blind null), so the
|
||||
instrument, not the model, was broken. `place2` was written after seeing that.
|
||||
- `exit` r1 misses: two false advances on Toll-Road conversation turns (p 0.59 split, p 0.67
|
||||
unanimous), and `v-to-pub`, where my label said stay and semif chose "Outside the Sanctum". On
|
||||
re-reading, semif's pick is defensible (they do run outside first). All its correct advances
|
||||
had p ≥ 0.87 except one at 0.67. A threshold would be tuned on the test set, so none is proposed.
|
||||
- Latency: 4 decisions per turn, 133–179 ms end to end (101–141 ms server). The first request
|
||||
of two files took ~700–800 ms (n=2, cause not isolated). Either way it is in dead time and is
|
||||
far below the GM's reflect.
|
||||
- Side findings from the code read, not fixed (they are Wyrd's): `story_node.created_turn` is
|
||||
never written, so "scenes so far (in order)" is alphabetical and includes unvisited seeds;
|
||||
and a late "ready" for an abandoned node's backdrop can still replace the one on screen.
|
||||
|
||||
## Cicada: should she show emotion (a gate, not a mood ring)
|
||||
|
||||
Constraint found first: Cicada persistent-memory 2026-09-20, operator: **"Affect is emitted once
|
||||
and fanned out … nothing derives affect a second time … A classifier inferring emotion from
|
||||
output text is the mood-ring failure."** So semif does not pick her pose and never reads her
|
||||
reply. What is open: talk's `/face` loop measured the model gesturing on 8–11 of 18 mundane
|
||||
turns (target under 1 in 3), and its README names "the page dropping gestures" as the next lever,
|
||||
"a design call, not a tuning pass". The spike measures that lever: a gate on what the PERSON
|
||||
said, run beside the LLM.
|
||||
|
||||
| wording | correct | mundane → "ordinary" | earned → "react" | controls |
|
||||
|---|---|---|---|---|
|
||||
| w1: descriptive options (news, joke, surprise, distress, a turn / command, plain question, small talk) | **30/31** | 16/16 | 14/15 | 5/5 |
|
||||
| w2: "Should the assistant's face visibly react to this?" Yes/No | 19/31 | 16/16 | **3/15** | 4/5 |
|
||||
|
||||
- w2 said "ordinary" to "My dog died this morning" at p 0.93. **The wording carries the whole
|
||||
definition.** Terse yes/no options fail; descriptive ones work.
|
||||
- Adding the earlier lines of the conversation changed no scored answer.
|
||||
- The one w1 miss: "Order four hundred pounds of cheese" → ordinary. Open (unscored) cases:
|
||||
"Goodnight" → react (p 0.58), "Thanks, you're the best" → react (0.89), "STOP THE TIMER" →
|
||||
ordinary.
|
||||
- Latency ~105 ms end to end (76 ms server), under talk's ~0.4–0.6 s pose decode, so it adds
|
||||
nothing to the turn if it starts at end of utterance.
|
||||
|
||||
## Recommendations (surfaced to Prime; nothing built)
|
||||
|
||||
- **Cicada:** the cleanest fit of the two. Gate **gestures only**: on "ordinary", drop the model's
|
||||
gesture and leave its pose alone. A miss then costs one missing gesture; gating poses would
|
||||
leave her neutral on bad news. It reads input, never her output, and never picks WHICH feeling,
|
||||
but it does override the model's affect. That is Prime's ruling to make, and the code would
|
||||
live in tts-stack talk `/face`, not in cicada (affect is not on cicada's v1 list).
|
||||
- **Wyrd:** `place2` is the node-add gate the archival note asked for. R8 exit choice works but
|
||||
its wording is fragile. Wyrd's v1 is met and no churn has been observed (2–6 nodes per real
|
||||
campaign), so by the roadmap gate this is a parking-lot item with a trigger: build it if live
|
||||
play shows churn, after labelling ~50 real turns as a held-out set.
|
||||
Reference in New Issue
Block a user