spike(semif): SemIf as Cicada's mood source is slower and less apt (no service change)
Against talk /face's guided pose (first paragraph 246 ms median), SemIf in parallel adds 32 ms and SemIf-first adds 94 ms (n=72 each, noise floor 16.5 ms). Removing the pose header saves only ~31 ms, and SemIf shares GPU 1 with the LLM. Acceptable pose 67% vs 92% on clear-emotion lines, and the mood carried through mundane follow-ups 7/15 vs 14/15. SemIf gestures far less (13% vs 58%). README: rotations cost options^2 in suffix tokens, and /decide/shared returns 422 when an object state's last value ends in ) ; or }.
This commit is contained in:
@@ -97,3 +97,64 @@ said, run beside the LLM.
|
||||
its wording is fragile. Wyrd's v1 is met and no churn has been observed (2–6 nodes per real
|
||||
campaign), so by the roadmap gate this is a parking-lot item with a trigger: build it if live
|
||||
play shows churn, after labelling ~50 real turns as a held-out set.
|
||||
|
||||
## Follow-up: SemIf as Cicada's mood source — FASTER? WORKS? (Prime, 0937: "build nothing … test cicada on latency grounds with clear emotional context")
|
||||
|
||||
Parked design, henge id 88: SemIf replaces the chat model's guided pose. Each call sees the earlier
|
||||
turns, each stamped with the pose SemIf chose, and the stamp drives the carry-over. Harness:
|
||||
`services/semif-serve/spike/cicada_mood_latency.py` + `cicada_mood_analyze.py`, raw data in
|
||||
`cicada-mood-2026-09-27/`. 24 emotionally clear single lines and three 5-turn arcs, 3 runs,
|
||||
interleaved. The prompt and schema are imported from tts-stack talk `app.py`, so A matches /face
|
||||
byte for byte. The chat model is char-rp-fast = G4-MeroMero-26B-A4B NVFP4A16 on fv-ml1:8021,
|
||||
**on GPU 1, the same GPU as semif**.
|
||||
|
||||
**Answer: no on both counts.**
|
||||
|
||||
| (singles, n=72 each) | first paragraph ready, direction known | vs today |
|
||||
|---|---|---|
|
||||
| A: today, guided `{pose, gesture, text}` | **246 ms** (IQR 230–263); the pose closes at 86 ms | — |
|
||||
| BA: SemIf and a text-only LLM call at once | 285 ms | **+32 ms** (95% CI +27..+41), faster in 10/72 |
|
||||
| BP: SemIf first, mood in the prompt | 352 ms | **+94 ms** (CI +85..+117), faster in 0/72 |
|
||||
|
||||
Noise floor A vs A: median |diff| 16.5 ms. The arcs agree: BP is +98 ms.
|
||||
- **Why it can't win on speed.** Dropping the pose+gesture header saves only ~31 ms (the LLM
|
||||
alone, text-only, reaches first paragraph in 215 ms vs 246 ms). The prompt is prefix-cached, so
|
||||
its 1,885-vs-909 tokens barely matter. SemIf's fast path costs 136 ms on its own. Run
|
||||
concurrently, it slows the LLM's decode on the shared GPU: SemIf goes to 196 ms and the LLM's
|
||||
first paragraph to 285 ms, and SemIf was the bottleneck in only 3/72 cases. DERIVED, not
|
||||
measured: with SemIf on a different GPU, the async version would be about max(215, 136) =
|
||||
~31 ms faster than today. That is the ceiling.
|
||||
- **Rotations are out.** 16 options with rotations is 550–1,200 ms, so every live arm used one
|
||||
ordering (pose + gesture ~110–136 ms).
|
||||
|
||||
| works (clear cases) | today (A) | SemIf (live config) |
|
||||
|---|---|---|
|
||||
| acceptable pose, singles | **66/72** (92%) | 48/72 (67%), 16/24 per case, deterministic |
|
||||
| controls | 12/12 | 12/12 |
|
||||
| egregious | 0 | 0 |
|
||||
| arcs acceptable | **40/45** | 25/45 |
|
||||
| mood carries through a mundane follow-up | **14/15** | **7/15** |
|
||||
| distinct poses used | 10 | 9 (no collapse to neutral) |
|
||||
| gesture rate, all / calm commands | 42/72 / 2/12 | **9/72** / 0/12 |
|
||||
|
||||
- SemIf's misses are systematic: hostility → annoyed, "someone's in the hallway" → curious,
|
||||
"someone's trying the back door" → suspicious, jokes → curious or neutral.
|
||||
- **The stamped previous mood did NOT carry.** In the vet arc, SemIf went concerned → neutral →
|
||||
neutral in all 3 runs ("I don't want to talk about it", "play something quiet") while the model
|
||||
held sad, sad, sad. The previous pose was in the state and it lost to the neutral-sounding line.
|
||||
- SemIf's pick agrees with the model 43% of the time, against the model's own run-to-run
|
||||
agreement of 83%, so this is a different judgment and not noise.
|
||||
- Wording grid (deterministic, singles): authored 16/24, short 14/24, plain **19/24** but it
|
||||
called "turn off the lights" sleepy. Rotations add about one case at 5–8× the cost. Null
|
||||
control (content-free state) → "content" 24/24, i.e. 5/24 by base rate, so SemIf is reading
|
||||
the evidence.
|
||||
- No vocal-sound vs pose clashes in any arm (0/117, 0/72, 0/117; crude detector).
|
||||
- The one bright spot matches the gate spike: **SemIf is far more sparing with gestures (13% vs
|
||||
58%)**, and the model is well over its own "fewer than 1 in 3" target on emotionally loaded
|
||||
lines.
|
||||
- Limits: one labeller, 24 + 15 cases, one model size (4B vs 26B-A4B), wordings not tuned.
|
||||
Tuning could lift accuracy some (plain 19/24), but it cannot create latency headroom that
|
||||
isn't there.
|
||||
- Found on the way (service, not fixed): `/decide/shared` 422s an object state whose last value
|
||||
ends in `)`, `;` or `}`. Also, rotations over long option lists cost options², and hit one cold
|
||||
503. Both are now in `stacks/semif/README.md`.
|
||||
|
||||
Reference in New Issue
Block a user