spike(semif): SemIf as Cicada's mood source is slower and less apt (no service change)

Against talk /face's guided pose (first paragraph 246 ms median), SemIf in
parallel adds 32 ms and SemIf-first adds 94 ms (n=72 each, noise floor 16.5 ms).
Removing the pose header saves only ~31 ms, and SemIf shares GPU 1 with the LLM.
Acceptable pose 67% vs 92% on clear-emotion lines, and the mood carried through
mundane follow-ups 7/15 vs 14/15. SemIf gestures far less (13% vs 58%).

README: rotations cost options^2 in suffix tokens, and /decide/shared returns
422 when an object state's last value ends in ) ; or }.
This commit is contained in:
vh
2026-09-27 09:47:28 -07:00
parent 89301a89fe
commit 7e11cf247b
7 changed files with 13163 additions and 1 deletions
@@ -97,3 +97,64 @@ said, run beside the LLM.
its wording is fragile. Wyrd's v1 is met and no churn has been observed (2–6 nodes per real
campaign), so by the roadmap gate this is a parking-lot item with a trigger: build it if live
play shows churn, after labelling ~50 real turns as a held-out set.
## Follow-up: SemIf as Cicada's mood source — FASTER? WORKS? (Prime, 0937: "build nothing … test cicada on latency grounds with clear emotional context")
Parked design, henge id 88: SemIf replaces the chat model's guided pose. Each call sees the earlier
turns, each stamped with the pose SemIf chose, and the stamp drives the carry-over. Harness:
`services/semif-serve/spike/cicada_mood_latency.py` + `cicada_mood_analyze.py`, raw data in
`cicada-mood-2026-09-27/`. 24 emotionally clear single lines and three 5-turn arcs, 3 runs,
interleaved. The prompt and schema are imported from tts-stack talk `app.py`, so A matches /face
byte for byte. The chat model is char-rp-fast = G4-MeroMero-26B-A4B NVFP4A16 on fv-ml1:8021,
**on GPU 1, the same GPU as semif**.
**Answer: no on both counts.**
| (singles, n=72 each) | first paragraph ready, direction known | vs today |
|---|---|---|
| A: today, guided `{pose, gesture, text}` | **246 ms** (IQR 230–263); the pose closes at 86 ms | — |
| BA: SemIf and a text-only LLM call at once | 285 ms | **+32 ms** (95% CI +27..+41), faster in 10/72 |
| BP: SemIf first, mood in the prompt | 352 ms | **+94 ms** (CI +85..+117), faster in 0/72 |
Noise floor A vs A: median |diff| 16.5 ms. The arcs agree: BP is +98 ms.
- **Why it can't win on speed.** Dropping the pose+gesture header saves only ~31 ms (the LLM
alone, text-only, reaches first paragraph in 215 ms vs 246 ms). The prompt is prefix-cached, so
its 1,885-vs-909 tokens barely matter. SemIf's fast path costs 136 ms on its own. Run
concurrently, it slows the LLM's decode on the shared GPU: SemIf goes to 196 ms and the LLM's
first paragraph to 285 ms, and SemIf was the bottleneck in only 3/72 cases. DERIVED, not
measured: with SemIf on a different GPU, the async version would be about max(215, 136) =
~31 ms faster than today. That is the ceiling.
- **Rotations are out.** 16 options with rotations is 550–1,200 ms, so every live arm used one
ordering (pose + gesture ~110–136 ms).
| works (clear cases) | today (A) | SemIf (live config) |
|---|---|---|
| acceptable pose, singles | **66/72** (92%) | 48/72 (67%), 16/24 per case, deterministic |
| controls | 12/12 | 12/12 |
| egregious | 0 | 0 |
| arcs acceptable | **40/45** | 25/45 |
| mood carries through a mundane follow-up | **14/15** | **7/15** |
| distinct poses used | 10 | 9 (no collapse to neutral) |
| gesture rate, all / calm commands | 42/72 / 2/12 | **9/72** / 0/12 |
- SemIf's misses are systematic: hostility → annoyed, "someone's in the hallway" → curious,
"someone's trying the back door" → suspicious, jokes → curious or neutral.
- **The stamped previous mood did NOT carry.** In the vet arc, SemIf went concerned → neutral →
neutral in all 3 runs ("I don't want to talk about it", "play something quiet") while the model
held sad, sad, sad. The previous pose was in the state and it lost to the neutral-sounding line.
- SemIf's pick agrees with the model 43% of the time, against the model's own run-to-run
agreement of 83%, so this is a different judgment and not noise.
- Wording grid (deterministic, singles): authored 16/24, short 14/24, plain **19/24** but it
called "turn off the lights" sleepy. Rotations add about one case at 5–8× the cost. Null
control (content-free state) → "content" 24/24, i.e. 5/24 by base rate, so SemIf is reading
the evidence.
- No vocal-sound vs pose clashes in any arm (0/117, 0/72, 0/117; crude detector).
- The one bright spot matches the gate spike: **SemIf is far more sparing with gestures (13% vs
58%)**, and the model is well over its own "fewer than 1 in 3" target on emotionally loaded
lines.
- Limits: one labeller, 24 + 15 cases, one model size (4B vs 26B-A4B), wordings not tuned.
Tuning could lift accuracy some (plain 19/24), but it cannot create latency headroom that
isn't there.
- Found on the way (service, not fixed): `/decide/shared` 422s an object state whose last value
ends in `)`, `;` or `}`. Also, rotations over long option lists cost options², and hit one cold
503. Both are now in `stacks/semif/README.md`.