Against talk /face's guided pose (first paragraph 246 ms median), SemIf in parallel adds 32 ms and SemIf-first adds 94 ms (n=72 each, noise floor 16.5 ms). Removing the pose header saves only ~31 ms, and SemIf shares GPU 1 with the LLM. Acceptable pose 67% vs 92% on clear-emotion lines, and the mood carried through mundane follow-ups 7/15 vs 14/15. SemIf gestures far less (13% vs 58%). README: rotations cost options^2 in suffix tokens, and /decide/shared returns 422 when an object state's last value ends in ) ; or }.
161 lines
11 KiB
Markdown
161 lines
11 KiB
Markdown
# SemIf consumer-fit spikes: Wyrd scene change, Cicada affect gate (2026-09-27, Prime)
|
||
|
||
Prime, 0904: "Two semif spikes -- Wyrd has a couple flags per turn on whether the scene should
|
||
change, check the codebase, see if semif is a clean fit and then spike a couple of sample
|
||
scenarios. likewise, cicada has a speech module and some reactions-- spike whether or not she
|
||
should show emotion." No service change. Code and raw output: `services/semif-serve/spike/`
|
||
(`consumer_fit.py` harness; `wyrd_cases.py`, `cicada_cases.py` build the scenario files;
|
||
`consumer-fit-2026-09-27/` holds scenarios, `.result.json` and `.log` per run).
|
||
|
||
## Method (applies to both)
|
||
|
||
- One `/decide/shared` per turn, every decision with `"orderings": "rotations"`. Target:
|
||
semif-serve 0.1.3 on fv-ml1 GPU 1, client on nh3-dev.
|
||
- **Noise floor:** the service is deterministic (acceptance: 144 rows twice, gap 0.0), and the
|
||
harness asserts the same top on all 3 runs of every case. Repeats measure latency only.
|
||
- **Null control:** every case is also asked over a content-free state. It always picks the
|
||
"nothing happened" option, so its accuracy is that label's base rate (Cicada 16/31, Wyrd
|
||
15/21). Evidence conditions are read against that, not against 0.
|
||
- **Positive controls:** cases tagged `control`, unambiguous by construction.
|
||
- **Limits, stated:** one labeller (me), authored inputs, 31 Cicada and 21 Wyrd cases per
|
||
decision. Wilson 95% intervals are wide (30/31 → 0.84–0.99; 21/21 → 0.85–1.0), so nothing
|
||
here resolves differences under ~10–15 points. The winning wordings were chosen on this same
|
||
set, so their scores are optimistic until checked on held-out real play.
|
||
|
||
## Wyrd: scene change (fit: yes for the CHOICE, no for the WRITING)
|
||
|
||
How it works today (from the code): no dedicated flag. The GM's `gen` structured-output calls
|
||
carry it two ways. **R8** (reflect): `plot_direction.target_node`, accepted only if it is an
|
||
advanceable exit; `plot_direction.intrude` is "advisory" and never read anywhere. **Path B**:
|
||
raw `AddStoryNode` + `AdvanceToNode` events, gated only by prompt prose ("Add + advance to a NEW
|
||
node ONLY when the scene moves to a genuinely different place — never every turn"). Wyrd's
|
||
archival memory already says "if it churns, tighten the prompt (or gate node-adds)".
|
||
**Three of the real campaigns in `.wyrd-data/canon` have ZERO story edges, so R8 can never fire
|
||
there and every scene change went through ungated path B.**
|
||
|
||
Fit: both are typed choices over one state, exactly semif's shape. Semif's exit options can only
|
||
be "stay" plus the real advanceable exits, so it cannot name an illegal target. It cannot write the
|
||
new node (title, summary, mood, image prompt) or the steering directive; the GM keeps those. So
|
||
it would be a gate/selector in front of the GM, not a replacement.
|
||
|
||
Spike: 3 real seed graphs (The Fletcher, The Toll Road, The Last Vigil + one added edge), 21
|
||
authored turns in the campaigns' register (the transcripts are gone from conv-api:
|
||
`session_not_found`). State = current scene + player action + narration.
|
||
|
||
| decision | wording | correct | real moves found | controls |
|
||
|---|---|---|---|---|
|
||
| `place`: left this place? | r1 "has the scene moved … or does it still hold?" | 16/21 (null 15/21) | **1/6** | 4/7 |
|
||
| `place2` | r2 location-anchored: "began at the {loc}; by the end, where are they?" | **21/21**, all unanimous, p ≥ 0.91 | 6/6 | 7/7 |
|
||
| `exit`: which exit? | r1 "which one does this turn move the story into?" | **18/21** | 6/6 | 6/7 |
|
||
| `exit2` | r2 "only once the narration actually gets there" | 16/21 | 1/6 | 4/7 |
|
||
|
||
- Round 1 `place` failed its positive controls (no better than the blind null), so the
|
||
instrument, not the model, was broken. `place2` was written after seeing that.
|
||
- `exit` r1 misses: two false advances on Toll-Road conversation turns (p 0.59 split, p 0.67
|
||
unanimous), and `v-to-pub`, where my label said stay and semif chose "Outside the Sanctum". On
|
||
re-reading, semif's pick is defensible (they do run outside first). All its correct advances
|
||
had p ≥ 0.87 except one at 0.67. A threshold would be tuned on the test set, so none is proposed.
|
||
- Latency: 4 decisions per turn, 133–179 ms end to end (101–141 ms server). The first request
|
||
of two files took ~700–800 ms (n=2, cause not isolated). Either way it is in dead time and is
|
||
far below the GM's reflect.
|
||
- Side findings from the code read, not fixed (they are Wyrd's): `story_node.created_turn` is
|
||
never written, so "scenes so far (in order)" is alphabetical and includes unvisited seeds;
|
||
and a late "ready" for an abandoned node's backdrop can still replace the one on screen.
|
||
|
||
## Cicada: should she show emotion (a gate, not a mood ring)
|
||
|
||
Constraint found first: Cicada persistent-memory 2026-09-20, operator: **"Affect is emitted once
|
||
and fanned out … nothing derives affect a second time … A classifier inferring emotion from
|
||
output text is the mood-ring failure."** So semif does not pick her pose and never reads her
|
||
reply. What is open: talk's `/face` loop measured the model gesturing on 8–11 of 18 mundane
|
||
turns (target under 1 in 3), and its README names "the page dropping gestures" as the next lever,
|
||
"a design call, not a tuning pass". The spike measures that lever: a gate on what the PERSON
|
||
said, run beside the LLM.
|
||
|
||
| wording | correct | mundane → "ordinary" | earned → "react" | controls |
|
||
|---|---|---|---|---|
|
||
| w1: descriptive options (news, joke, surprise, distress, a turn / command, plain question, small talk) | **30/31** | 16/16 | 14/15 | 5/5 |
|
||
| w2: "Should the assistant's face visibly react to this?" Yes/No | 19/31 | 16/16 | **3/15** | 4/5 |
|
||
|
||
- w2 said "ordinary" to "My dog died this morning" at p 0.93. **The wording carries the whole
|
||
definition.** Terse yes/no options fail; descriptive ones work.
|
||
- Adding the earlier lines of the conversation changed no scored answer.
|
||
- The one w1 miss: "Order four hundred pounds of cheese" → ordinary. Open (unscored) cases:
|
||
"Goodnight" → react (p 0.58), "Thanks, you're the best" → react (0.89), "STOP THE TIMER" →
|
||
ordinary.
|
||
- Latency ~105 ms end to end (76 ms server), under talk's ~0.4–0.6 s pose decode, so it adds
|
||
nothing to the turn if it starts at end of utterance.
|
||
|
||
## Recommendations (surfaced to Prime; nothing built)
|
||
|
||
- **Cicada:** the cleanest fit of the two. Gate **gestures only**: on "ordinary", drop the model's
|
||
gesture and leave its pose alone. A miss then costs one missing gesture; gating poses would
|
||
leave her neutral on bad news. It reads input, never her output, and never picks WHICH feeling,
|
||
but it does override the model's affect. That is Prime's ruling to make, and the code would
|
||
live in tts-stack talk `/face`, not in cicada (affect is not on cicada's v1 list).
|
||
- **Wyrd:** `place2` is the node-add gate the archival note asked for. R8 exit choice works but
|
||
its wording is fragile. Wyrd's v1 is met and no churn has been observed (2–6 nodes per real
|
||
campaign), so by the roadmap gate this is a parking-lot item with a trigger: build it if live
|
||
play shows churn, after labelling ~50 real turns as a held-out set.
|
||
|
||
## Follow-up: SemIf as Cicada's mood source — FASTER? WORKS? (Prime, 0937: "build nothing … test cicada on latency grounds with clear emotional context")
|
||
|
||
Parked design, henge id 88: SemIf replaces the chat model's guided pose. Each call sees the earlier
|
||
turns, each stamped with the pose SemIf chose, and the stamp drives the carry-over. Harness:
|
||
`services/semif-serve/spike/cicada_mood_latency.py` + `cicada_mood_analyze.py`, raw data in
|
||
`cicada-mood-2026-09-27/`. 24 emotionally clear single lines and three 5-turn arcs, 3 runs,
|
||
interleaved. The prompt and schema are imported from tts-stack talk `app.py`, so A matches /face
|
||
byte for byte. The chat model is char-rp-fast = G4-MeroMero-26B-A4B NVFP4A16 on fv-ml1:8021,
|
||
**on GPU 1, the same GPU as semif**.
|
||
|
||
**Answer: no on both counts.**
|
||
|
||
| (singles, n=72 each) | first paragraph ready, direction known | vs today |
|
||
|---|---|---|
|
||
| A: today, guided `{pose, gesture, text}` | **246 ms** (IQR 230–263); the pose closes at 86 ms | — |
|
||
| BA: SemIf and a text-only LLM call at once | 285 ms | **+32 ms** (95% CI +27..+41), faster in 10/72 |
|
||
| BP: SemIf first, mood in the prompt | 352 ms | **+94 ms** (CI +85..+117), faster in 0/72 |
|
||
|
||
Noise floor A vs A: median |diff| 16.5 ms. The arcs agree: BP is +98 ms.
|
||
- **Why it can't win on speed.** Dropping the pose+gesture header saves only ~31 ms (the LLM
|
||
alone, text-only, reaches first paragraph in 215 ms vs 246 ms). The prompt is prefix-cached, so
|
||
its 1,885-vs-909 tokens barely matter. SemIf's fast path costs 136 ms on its own. Run
|
||
concurrently, it slows the LLM's decode on the shared GPU: SemIf goes to 196 ms and the LLM's
|
||
first paragraph to 285 ms, and SemIf was the bottleneck in only 3/72 cases. DERIVED, not
|
||
measured: with SemIf on a different GPU, the async version would be about max(215, 136) =
|
||
~31 ms faster than today. That is the ceiling.
|
||
- **Rotations are out.** 16 options with rotations is 550–1,200 ms, so every live arm used one
|
||
ordering (pose + gesture ~110–136 ms).
|
||
|
||
| works (clear cases) | today (A) | SemIf (live config) |
|
||
|---|---|---|
|
||
| acceptable pose, singles | **66/72** (92%) | 48/72 (67%), 16/24 per case, deterministic |
|
||
| controls | 12/12 | 12/12 |
|
||
| egregious | 0 | 0 |
|
||
| arcs acceptable | **40/45** | 25/45 |
|
||
| mood carries through a mundane follow-up | **14/15** | **7/15** |
|
||
| distinct poses used | 10 | 9 (no collapse to neutral) |
|
||
| gesture rate, all / calm commands | 42/72 / 2/12 | **9/72** / 0/12 |
|
||
|
||
- SemIf's misses are systematic: hostility → annoyed, "someone's in the hallway" → curious,
|
||
"someone's trying the back door" → suspicious, jokes → curious or neutral.
|
||
- **The stamped previous mood did NOT carry.** In the vet arc, SemIf went concerned → neutral →
|
||
neutral in all 3 runs ("I don't want to talk about it", "play something quiet") while the model
|
||
held sad, sad, sad. The previous pose was in the state and it lost to the neutral-sounding line.
|
||
- SemIf's pick agrees with the model 43% of the time, against the model's own run-to-run
|
||
agreement of 83%, so this is a different judgment and not noise.
|
||
- Wording grid (deterministic, singles): authored 16/24, short 14/24, plain **19/24** but it
|
||
called "turn off the lights" sleepy. Rotations add about one case at 5–8× the cost. Null
|
||
control (content-free state) → "content" 24/24, i.e. 5/24 by base rate, so SemIf is reading
|
||
the evidence.
|
||
- No vocal-sound vs pose clashes in any arm (0/117, 0/72, 0/117; crude detector).
|
||
- The one bright spot matches the gate spike: **SemIf is far more sparing with gestures (13% vs
|
||
58%)**, and the model is well over its own "fewer than 1 in 3" target on emotionally loaded
|
||
lines.
|
||
- Limits: one labeller, 24 + 15 cases, one model size (4B vs 26B-A4B), wordings not tuned.
|
||
Tuning could lift accuracy some (plain 19/24), but it cannot create latency headroom that
|
||
isn't there.
|
||
- Found on the way (service, not fixed): `/decide/shared` 422s an object state whose last value
|
||
ends in `)`, `;` or `}`. Also, rotations over long option lists cost options², and hit one cold
|
||
503. Both are now in `stacks/semif/README.md`.
|