Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md
T
vh 7e11cf247b spike(semif): SemIf as Cicada's mood source is slower and less apt (no service change)
Against talk /face's guided pose (first paragraph 246 ms median), SemIf in
parallel adds 32 ms and SemIf-first adds 94 ms (n=72 each, noise floor 16.5 ms).
Removing the pose header saves only ~31 ms, and SemIf shares GPU 1 with the LLM.
Acceptable pose 67% vs 92% on clear-emotion lines, and the mood carried through
mundane follow-ups 7/15 vs 14/15. SemIf gestures far less (13% vs 58%).

README: rotations cost options^2 in suffix tokens, and /decide/shared returns
422 when an object state's last value ends in ) ; or }.
2026-09-27 09:47:28 -07:00

11 KiB
Raw Blame History

SemIf consumer-fit spikes: Wyrd scene change, Cicada affect gate (2026-09-27, Prime)

Prime, 0904: "Two semif spikes -- Wyrd has a couple flags per turn on whether the scene should change, check the codebase, see if semif is a clean fit and then spike a couple of sample scenarios. likewise, cicada has a speech module and some reactions-- spike whether or not she should show emotion." No service change. Code and raw output: services/semif-serve/spike/ (consumer_fit.py harness; wyrd_cases.py, cicada_cases.py build the scenario files; consumer-fit-2026-09-27/ holds scenarios, .result.json and .log per run).

Method (applies to both)

  • One /decide/shared per turn, every decision with "orderings": "rotations". Target: semif-serve 0.1.3 on fv-ml1 GPU 1, client on nh3-dev.
  • Noise floor: the service is deterministic (acceptance: 144 rows twice, gap 0.0), and the harness asserts the same top on all 3 runs of every case. Repeats measure latency only.
  • Null control: every case is also asked over a content-free state. It always picks the "nothing happened" option, so its accuracy is that label's base rate (Cicada 16/31, Wyrd 15/21). Evidence conditions are read against that, not against 0.
  • Positive controls: cases tagged control, unambiguous by construction.
  • Limits, stated: one labeller (me), authored inputs, 31 Cicada and 21 Wyrd cases per decision. Wilson 95% intervals are wide (30/31 → 0.84–0.99; 21/21 → 0.85–1.0), so nothing here resolves differences under ~10–15 points. The winning wordings were chosen on this same set, so their scores are optimistic until checked on held-out real play.

Wyrd: scene change (fit: yes for the CHOICE, no for the WRITING)

How it works today (from the code): no dedicated flag. The GM's gen structured-output calls carry it two ways. R8 (reflect): plot_direction.target_node, accepted only if it is an advanceable exit; plot_direction.intrude is "advisory" and never read anywhere. Path B: raw AddStoryNode + AdvanceToNode events, gated only by prompt prose ("Add + advance to a NEW node ONLY when the scene moves to a genuinely different place — never every turn"). Wyrd's archival memory already says "if it churns, tighten the prompt (or gate node-adds)". Three of the real campaigns in .wyrd-data/canon have ZERO story edges, so R8 can never fire there and every scene change went through ungated path B.

Fit: both are typed choices over one state, exactly semif's shape. Semif's exit options can only be "stay" plus the real advanceable exits, so it cannot name an illegal target. It cannot write the new node (title, summary, mood, image prompt) or the steering directive; the GM keeps those. So it would be a gate/selector in front of the GM, not a replacement.

Spike: 3 real seed graphs (The Fletcher, The Toll Road, The Last Vigil + one added edge), 21 authored turns in the campaigns' register (the transcripts are gone from conv-api: session_not_found). State = current scene + player action + narration.

decision wording correct real moves found controls
place: left this place? r1 "has the scene moved … or does it still hold?" 16/21 (null 15/21) 1/6 4/7
place2 r2 location-anchored: "began at the {loc}; by the end, where are they?" 21/21, all unanimous, p ≥ 0.91 6/6 7/7
exit: which exit? r1 "which one does this turn move the story into?" 18/21 6/6 6/7
exit2 r2 "only once the narration actually gets there" 16/21 1/6 4/7
  • Round 1 place failed its positive controls (no better than the blind null), so the instrument, not the model, was broken. place2 was written after seeing that.
  • exit r1 misses: two false advances on Toll-Road conversation turns (p 0.59 split, p 0.67 unanimous), and v-to-pub, where my label said stay and semif chose "Outside the Sanctum". On re-reading, semif's pick is defensible (they do run outside first). All its correct advances had p ≥ 0.87 except one at 0.67. A threshold would be tuned on the test set, so none is proposed.
  • Latency: 4 decisions per turn, 133–179 ms end to end (101–141 ms server). The first request of two files took ~700–800 ms (n=2, cause not isolated). Either way it is in dead time and is far below the GM's reflect.
  • Side findings from the code read, not fixed (they are Wyrd's): story_node.created_turn is never written, so "scenes so far (in order)" is alphabetical and includes unvisited seeds; and a late "ready" for an abandoned node's backdrop can still replace the one on screen.

Cicada: should she show emotion (a gate, not a mood ring)

Constraint found first: Cicada persistent-memory 2026-09-20, operator: "Affect is emitted once and fanned out … nothing derives affect a second time … A classifier inferring emotion from output text is the mood-ring failure." So semif does not pick her pose and never reads her reply. What is open: talk's /face loop measured the model gesturing on 8–11 of 18 mundane turns (target under 1 in 3), and its README names "the page dropping gestures" as the next lever, "a design call, not a tuning pass". The spike measures that lever: a gate on what the PERSON said, run beside the LLM.

wording correct mundane → "ordinary" earned → "react" controls
w1: descriptive options (news, joke, surprise, distress, a turn / command, plain question, small talk) 30/31 16/16 14/15 5/5
w2: "Should the assistant's face visibly react to this?" Yes/No 19/31 16/16 3/15 4/5
  • w2 said "ordinary" to "My dog died this morning" at p 0.93. The wording carries the whole definition. Terse yes/no options fail; descriptive ones work.
  • Adding the earlier lines of the conversation changed no scored answer.
  • The one w1 miss: "Order four hundred pounds of cheese" → ordinary. Open (unscored) cases: "Goodnight" → react (p 0.58), "Thanks, you're the best" → react (0.89), "STOP THE TIMER" → ordinary.
  • Latency ~105 ms end to end (76 ms server), under talk's ~0.4–0.6 s pose decode, so it adds nothing to the turn if it starts at end of utterance.

Recommendations (surfaced to Prime; nothing built)

  • Cicada: the cleanest fit of the two. Gate gestures only: on "ordinary", drop the model's gesture and leave its pose alone. A miss then costs one missing gesture; gating poses would leave her neutral on bad news. It reads input, never her output, and never picks WHICH feeling, but it does override the model's affect. That is Prime's ruling to make, and the code would live in tts-stack talk /face, not in cicada (affect is not on cicada's v1 list).
  • Wyrd: place2 is the node-add gate the archival note asked for. R8 exit choice works but its wording is fragile. Wyrd's v1 is met and no churn has been observed (2–6 nodes per real campaign), so by the roadmap gate this is a parking-lot item with a trigger: build it if live play shows churn, after labelling ~50 real turns as a held-out set.

Follow-up: SemIf as Cicada's mood source — FASTER? WORKS? (Prime, 0937: "build nothing … test cicada on latency grounds with clear emotional context")

Parked design, henge id 88: SemIf replaces the chat model's guided pose. Each call sees the earlier turns, each stamped with the pose SemIf chose, and the stamp drives the carry-over. Harness: services/semif-serve/spike/cicada_mood_latency.py + cicada_mood_analyze.py, raw data in cicada-mood-2026-09-27/. 24 emotionally clear single lines and three 5-turn arcs, 3 runs, interleaved. The prompt and schema are imported from tts-stack talk app.py, so A matches /face byte for byte. The chat model is char-rp-fast = G4-MeroMero-26B-A4B NVFP4A16 on fv-ml1:8021, on GPU 1, the same GPU as semif.

Answer: no on both counts.

(singles, n=72 each) first paragraph ready, direction known vs today
A: today, guided {pose, gesture, text} 246 ms (IQR 230–263); the pose closes at 86 ms —
BA: SemIf and a text-only LLM call at once 285 ms +32 ms (95% CI +27..+41), faster in 10/72
BP: SemIf first, mood in the prompt 352 ms +94 ms (CI +85..+117), faster in 0/72

Noise floor A vs A: median |diff| 16.5 ms. The arcs agree: BP is +98 ms.

  • Why it can't win on speed. Dropping the pose+gesture header saves only ~31 ms (the LLM alone, text-only, reaches first paragraph in 215 ms vs 246 ms). The prompt is prefix-cached, so its 1,885-vs-909 tokens barely matter. SemIf's fast path costs 136 ms on its own. Run concurrently, it slows the LLM's decode on the shared GPU: SemIf goes to 196 ms and the LLM's first paragraph to 285 ms, and SemIf was the bottleneck in only 3/72 cases. DERIVED, not measured: with SemIf on a different GPU, the async version would be about max(215, 136) = ~31 ms faster than today. That is the ceiling.
  • Rotations are out. 16 options with rotations is 550–1,200 ms, so every live arm used one ordering (pose + gesture ~110–136 ms).
works (clear cases) today (A) SemIf (live config)
acceptable pose, singles 66/72 (92%) 48/72 (67%), 16/24 per case, deterministic
controls 12/12 12/12
egregious 0 0
arcs acceptable 40/45 25/45
mood carries through a mundane follow-up 14/15 7/15
distinct poses used 10 9 (no collapse to neutral)
gesture rate, all / calm commands 42/72 / 2/12 9/72 / 0/12
  • SemIf's misses are systematic: hostility → annoyed, "someone's in the hallway" → curious, "someone's trying the back door" → suspicious, jokes → curious or neutral.
  • The stamped previous mood did NOT carry. In the vet arc, SemIf went concerned → neutral → neutral in all 3 runs ("I don't want to talk about it", "play something quiet") while the model held sad, sad, sad. The previous pose was in the state and it lost to the neutral-sounding line.
  • SemIf's pick agrees with the model 43% of the time, against the model's own run-to-run agreement of 83%, so this is a different judgment and not noise.
  • Wording grid (deterministic, singles): authored 16/24, short 14/24, plain 19/24 but it called "turn off the lights" sleepy. Rotations add about one case at 5–8× the cost. Null control (content-free state) → "content" 24/24, i.e. 5/24 by base rate, so SemIf is reading the evidence.
  • No vocal-sound vs pose clashes in any arm (0/117, 0/72, 0/117; crude detector).
  • The one bright spot matches the gate spike: SemIf is far more sparing with gestures (13% vs 58%), and the model is well over its own "fewer than 1 in 3" target on emotionally loaded lines.
  • Limits: one labeller, 24 + 15 cases, one model size (4B vs 26B-A4B), wordings not tuned. Tuning could lift accuracy some (plain 19/24), but it cannot create latency headroom that isn't there.
  • Found on the way (service, not fixed): /decide/shared 422s an object state whose last value ends in ), ; or }. Also, rotations over long option lists cost options², and hit one cold 503. Both are now in stacks/semif/README.md.