diff --git a/.corviduo-canonicals.toml b/.corviduo-canonicals.toml index 1fb376b..242a673 100644 --- a/.corviduo-canonicals.toml +++ b/.corviduo-canonicals.toml @@ -120,8 +120,8 @@ id = "worldtree-conversation-api-spec-v1" canonical_source = "Worldtree" canonical_path = "docs/conversation-api-spec.md" consumer_path = "docs/conversation-api-spec.md" -pinned_sha256_16 = "c656a789caceef14" -pinned_at = "2026-07-06T16:51:09+00:00" +pinned_sha256_16 = "2d73d50b8680b893" +pinned_at = "2026-07-13T07:54:05+00:00" tolerate_drift = true # prose reference; OpenAPI+SSE are the gates # Worldtree persona render canons (d2) — the deterministic affect->NL the agent is @@ -155,6 +155,34 @@ id = "worldtree-affect-egress-consumer-reference-v1" canonical_source = "Worldtree" canonical_path = "docs/affect-egress-consumer-reference.md" consumer_path = "docs/vendor/worldtree-persona-canon/affect-egress-consumer-reference.md" -pinned_sha256_16 = "d959134037efae83" -pinned_at = "2026-07-07T06:09:24+00:00" +pinned_sha256_16 = "b2406e237df00dcb" +pinned_at = "2026-07-13T07:54:05+00:00" tolerate_drift = true # prose reference; the d2 render-canon JSONs are the gates + +# --------------------------------------------------------------------------- +# Brokkr R34/R35 persona-prompt-framing reference (the character-self-report +# reframe ratatoskr consumes: the authored psychological_profile is the prose +# lens the Worldtree self-report producer reads for affect + memory salience). +# Vendored for reference alongside the Worldtree affect/memory surfaces. +# tolerate_drift: prose reference, not a machine gate — brokkr-smithy-dev owns +# it and pings ratatoskr-dev on canonical changes. The authoring-spec GOVERNS on +# any conflict with the parameter distillation. +# --------------------------------------------------------------------------- + +[[pins]] +id = "brokkr-psych-profile-authoring-spec-v1" +canonical_source = "brokkr-smithy" +canonical_path = "research/R34-persona-prompt-framing/deliverables/psych-profile-authoring-spec.md" +consumer_path = "docs/vendor/brokkr-r34-psych-profile/psych-profile-authoring-spec.md" +pinned_sha256_16 = "4545a108d9fb6cc3" +pinned_at = "2026-07-13T00:00:00+00:00" +tolerate_drift = true # prose reference; brokkr-smithy-dev owns + pings on change + +[[pins]] +id = "brokkr-psych-profile-parameters-v1" +canonical_source = "brokkr-smithy" +canonical_path = "research/R34-persona-prompt-framing/deliverables/psych-profile-parameters.md" +consumer_path = "docs/vendor/brokkr-r34-psych-profile/psych-profile-parameters.md" +pinned_sha256_16 = "17157c82771aeeee" +pinned_at = "2026-07-13T00:00:00+00:00" +tolerate_drift = true # parameter distillation; authoring-spec governs on conflict diff --git a/docs/conversation-api-spec.md b/docs/conversation-api-spec.md index a646ab4..2084d98 100644 --- a/docs/conversation-api-spec.md +++ b/docs/conversation-api-spec.md @@ -2858,6 +2858,8 @@ Semantics: which must carry all three of `pleasure` / `arousal` / `dominance`, each a float in `[-1.0, 1.0]`. Any other top-level key → 422 `validation_failed`; a missing or malformed `pad` → 422 `persona_seed_invalid`. + + > **✓ R32-1B (landed, v1.0.0b29):** The PAD range `[-1.0, 1.0]` relaxes to an **unbounded latent `z`** with a finite wire sanity bound (`~±10`) as of R32 Slice-1B. The JSON shape/fields/types are UNCHANGED — only the declared range/semantics change (the value becomes a latent that renders to a bounded display value). Consumers that merely store-and-return PAD need no change; consumers that validate/clamp PAD to `[-1,1]` must relax that bound. Source of truth: `docs/contracts/persona_envelope.contract.md` rev 1.7 (INV-ENV-16). - **Seeds the current mood POINT, not the setpoint.** The OCEAN persona (above) fixes the setpoint the mood relaxes toward; this endpoint sets where the mood *starts*. It does not alter the persona. diff --git a/docs/vendor/brokkr-r34-psych-profile/psych-profile-authoring-spec.md b/docs/vendor/brokkr-r34-psych-profile/psych-profile-authoring-spec.md new file mode 100644 index 0000000..ca8de51 --- /dev/null +++ b/docs/vendor/brokkr-r34-psych-profile/psych-profile-authoring-spec.md @@ -0,0 +1,242 @@ +# Psychological Profile Authoring Spec — canonical + +**Status:** canonical (v1). **Owner:** brokkr-smithy-dev (R34/R35 self-report reframe). +**Audience:** anyone authoring a character's `psychological_profile` — Worldtree +foundational characters (soong-dev) and consumer characters created via the +Conversation API (ratatoskr and other external consumers). +**For:** the Worldtree agent-definition schema; intended to live in the Worldtree +client-app documentation. + +This spec governs the **content** of the psychological profile (what to write and +what never to write). The **physical wire shape** of the field (single string vs a +small keyed dict) is Worldtree's schema call — see § Wire shape. + +--- + +## 1. What it is + +A dedicated **authored prose section** of a character definition that carries the +character's **psychological bent and formative experience**. It is the source the +self-report producer maps from when it decides, on each turn: + +- **what the character feels** (affect self-report), and +- **what the character notices and keeps** (character-voiced memory salience). + +The profile is a *lens*, not a script. It never states per-turn emotions; it +describes the standing disposition, history, values, and attention that — combined +with the actual event — *produce* the emotion and the salience. + +It sits **alongside the numeric OCEAN** values (a separate, deterministic input). +The prose gives the *qualitative* bent; the OCEAN numbers give the *magnitude dial* +(see § OCEAN interaction). + +--- + +## 2. What it carries — the four dimensions + +1. **Disposition / appraisal bent** — how the character characteristically + *interprets* situations: attribution style, what they hold weighty, how they + respond to being challenged. NOT per-event emotions. +2. **Attention / salience focus** — the kinds of things this character + characteristically *notices* (and therefore tends to remember). +3. **Values / what a good day looks like** — the yardstick that drives what they + find worth keeping. +4. **Formative experience (history)** — the background that shapes both appraisal + *and* salience. A character betrayed before appraises betrayal differently, and + remembers different things. + +You may write these as four short labelled sections or as one integrated paragraph +— both are supported (see § Length & format). + +--- + +## 3. Authoring rules (load-bearing) + +These are the rules the whole reframe depends on. Rule 1 is the one that most often +gets violated. + +1. **Never name a per-event output emotion.** Do NOT write "is anxious", "gets + angry at X", "feels hurt when criticized", "joyful". Naming an emotion **primes** + it — the "pink ball" effect — so the producer will report that emotion regardless + of what actually happens in the scene. Describe *disposition, history, values, + attention*; let the emotion come from the event appraisal. + - ✅ "Registers quickly when authority is substituted for craft." (an appraisal + trigger — sets up how she reads an event, names no feeling) + - ❌ "Feels contempt when someone pulls rank." (names the output emotion) + +2. **Magnitude lives in the numeric OCEAN, not the prose.** *How strongly / how + long* a character reacts (Neuroticism) is the deterministic OCEAN dial, rendered + valence-neutral by the producer. Do not narrate reaction dynamics in the prose + ("comes apart", "takes it hard", "rich inner life") — that double-encodes what the + number already carries. The prose gives the *qualitative bent*; the number gives + the *gain*. + +3. **Appraisal-style is allowed; output-emotion is not.** "Interprets others' + actions charitably until she can't" (a style) is fine; "feels betrayed easily" + (an output) is not. The style plus the event produce the output. + +4. **Salience is character-relative; facts are not.** The profile shapes what the + character *cares to remember*. It must never license rewriting *what happened* — + when the character does remember something, it stays grounded in the transcript. + +--- + +## 4. Wire shape & field placement + +- **Content is prose** covering the four dimensions, authored as **one coherent prose + string** — the four dimensions are authoring *structure* inside that single string, + not separate wire fields. +- **Wire shape (LOCKED, b53):** a single dedicated prose string, field + **`psychological_profile`** (type `str`) on the persona layer — foundational + `persona.psychological_profile`, Tier-3 `ValidatedPersona.psychological_profile`. It + nests under the existing `Any`-typed persona field, so it is the shipped b53 shape — + no schema change. **Not** a dict-of-four. +- **Hard constraint (non-negotiable):** the profile is a **dedicated field the lens + reads ONLY** (`resolve_psych_profile` reads only this field — no `behavioral_notes` + or other general-field remap). Non-lens content leaking into the lens produces the + "executive-assistant" failure (the producer reads response-format / tone / tool + instructions as if they were the character's psychology). + +--- + +## 5. The non-priming banned set + +The non-priming rule (Rule 1) is **semantic, not a fixed wordlist** — it bans naming +any per-event output emotion, which is broader than any specific vocabulary +("anxious", "worried", "hurt" all prime even though they are not in the producer's +fixed emotion roster). + +- **The gate is human review:** does the prose describe disposition / appraisal-style + / history / values / attention, and never what the character *feels*? +- **A mechanical lint is a backstop, not the gate.** If you build one, scan the + fixed-15 OCC roster plus `synonym_map.json` (which already folds common affect + synonyms) as the core set, optionally extended with a general affect lexicon. Treat + a lint hit as a prompt to re-read, not an automatic reject. + +--- + +## 6. Required vs optional dimensions + +- **Required** (they *are* the lens): **disposition**, **attention / salience focus**, + **values**. +- **Strongly recommended:** **formative history** — it is the single biggest lever on + richness (validated in P03: richer history → sharper, more character-appropriate + salience). It may be brief for a deliberately thin character, but omitting it leaves + salience under-grounded. + +--- + +## 7. Length & format + +- A focused paragraph, or four short labelled sections — **a lens, not a biography.** +- Target **~150–300 words.** The producer reads this on **every** turn, so keep it + tight; bloat is a latency and dilution cost. +- **Prose only — never typed emotion fields.** The four dimensions are a coverage + checklist for the author, not a schema of feelings to fill in. + +--- + +## 8. Exemplars + +These three were the validated P03 stimuli — integrated-paragraph form, each faithful +to its OCEAN, none naming an output emotion. (OCEAN shown in **[−1, 1] storage units**; +validated in P03 at the equivalent [0, 1] values.) + +**Perrin — court scribe** (OCEAN: O0.0 C0.2 E−0.2 A0.1 N0.7) +> Perrin keeps the court's records and has done so through two changes of regime. He +> learned early that small errors compound — a misfiled writ once cost a man his +> lands, and Perrin found the mistake too late to undo it. Since then he double-checks +> everything and watches situations closely for what is out of place. He forms +> attachments slowly and holds a given trust as a considerable thing. He measures +> himself by whether he was useful and careful. He notices discrepancies, unspoken +> tensions, and anything that threatens the order he keeps. + +**Vared — veteran caravan guard** (OCEAN: O−0.2 C0.4 E−0.5 A−0.2 N−0.7) +> Vared has guarded caravans across the northern routes for twenty years and buried +> more traveling companions than he cares to count. He speaks little and shows less. +> Danger he treats as weather — a thing to be handled. He judges people by what they +> do under pressure and remembers who held the line. What reaches him reaches him +> quietly and privately. He notices terrain, exits, who is armed, and shifts in a +> group that might precede trouble. + +**Sella — village healer** (OCEAN: O0.2 C0.2 E0.0 A0.8 N0.0) +> Sella has tended the sick since she was old enough to carry water for her +> grandmother, the healer before her. She reads people's pain quickly and carries some +> of it with her. She interprets others' actions charitably until she cannot, and +> prioritizes keeping the peace between people. She measures a day by whether she eased +> someone's burden. She notices who is unwell, who is troubled, and what is left +> unsaid. + +Note how each closes on **attention** ("he notices…", "she notices…") — the salience +focus stated plainly, no emotion named. + +--- + +## 9. OCEAN interaction & the scaffold fallback + +OCEAN values are stored on **[−1, 1]** (0 = average) — a **separate deterministic +input** and the **magnitude dial** the prose must not duplicate (Rule 2). The producer +renders **off-average** bands as valence-neutral disposition cues. It maps storage to +[0, 1] first (`c = (v + 1) / 2`, `render_disposition` in b53) and then applies the +canonical [0, 1] band cutoffs (`c < 0.33` low / `c > 0.66` high). In **storage units** +that is: + +| trait | low (v < −0.34) | high (v > +0.32) | +|---|---|---| +| **N** (reactivity only) | reactions are milder than most people's | reactions are more intense than most people's | +| **E** (expression; may be excluded from affect elicitation) | socially reserved; expression less outwardly amplified | socially expressive; reactions more externally visible | +| **O** | prefers the familiar, the concrete, established ways | curious, drawn to novelty, ideas, the unfamiliar | +| **C** | less plan-bound; less weight on order, detail, obligation | attends closely to order, detail, and obligations | +| **A** | less inclined to assume cooperative intent; direct, self-protective | more inclined to preserve rapport and weigh others' needs | + +The **mid** band (−0.34 ≤ v ≤ +0.32, i.e. `c` in [0.33, 0.66]) renders nothing — an +average trait is silent, **not** "low." (Boundaries are slightly asymmetric because +the canonical 0.33/0.66 cutoffs are not symmetric about 0.5. Canonical rendering +strings live in the reframe language catalog §4; persistence/recovery dynamics live in +the deterministic mood decay, not the profile.) + +**Scaffold fallback:** a character with **no** authored profile falls back to this +band-rendering from the OCEAN numbers alone. That still functions — but the authored +profile is what turns generic band cues into *this specific character's* appraisal and +salience. Authoring the profile is how the reframe's value actually reaches a +character. + +--- + +## 10. Authoring divergent characters (contrast design) + +When you want two characters to remember **noticeably different things** (e.g. for an +eval contrast pair, or simply a varied cast), design the divergence on the **attention +and values** dimensions first, and set the OCEAN numbers to *serve* that prose — not +the reverse. + +- **The sharpest contrast is a salience *drop*, not just a different flavor.** One + character for whom relational/emotional content is genuinely non-salient (an + operational, task-focused character in the Vared mold — notices terrain, logistics, + who is armed) versus one who weights it highest (a caretaker who tracks who is + troubled and what went unsaid). "Different notes, same facts" has real teeth only + when one character *legitimately forgets* what the other keeps. +- **High-yield axes for salience divergence:** O (what patterns they attend to), A + (relational vs operational/self-protective focus), C (procedural/detail salience). +- **Low-yield for salience:** E — it is expression-oriented (shapes how a reaction is + *rendered*, not what is *noticed*), and may even be excluded from the affect + elicitation. Don't lean on flipping E to create divergence. +- **Watch the direction, not just the distance:** flipping every OCEAN axis to its + opposite does not guarantee a strong contrast. If your reference character already + *keeps* relational content, an even-more-agreeable opposite keeps it harder and the + most intuitive contrast collapses. Aim the contrast at *dropping* what the reference + *keeps*. + +--- + +## Provenance & validation + +Grounded in R34/R35 (self-report reframe), probes P02–P05: character-voiced memory +salience validated on two model classes (P02/P03); the "Psychological Profile and +Experience" section mapping validated as the lens source (P03); non-priming and +magnitude-in-OCEAN corrections are operator rulings (2026-07-10). The affect half is +live in production (Worldtree b53) and fired a contextually-apt self-report on a +non-frontier seat. A powered efficacy eval (salience divergence / floor recall / +salience≠facts firewall / graded model-slot response + the authored-vs-scaffold delta) +is preregistering to quantify the memory half; findings will refine this spec, not +overturn its authoring rules. diff --git a/docs/vendor/brokkr-r34-psych-profile/psych-profile-parameters.md b/docs/vendor/brokkr-r34-psych-profile/psych-profile-parameters.md new file mode 100644 index 0000000..d9c3a5b --- /dev/null +++ b/docs/vendor/brokkr-r34-psych-profile/psych-profile-parameters.md @@ -0,0 +1,123 @@ +# Psychological Profile Parameters — for AI generation (canonical) + +**Status:** canonical (v1). **Owner:** brokkr-smithy-dev (R34/R35 self-report reframe). +**Audience:** **soong-dev** (Soong's Lab / Soong's AI — the immediate builder that +generates the profile from these parameters); **Worldtree** + **ratatoskr** (vendoring +for reference alongside the authoring spec). +**Relationship:** this is the **parameter distillation** of +`psych-profile-authoring-spec.md` for the model where **Soong's AI writes the +`psychological_profile` prose from parameters** (rather than a human hand-authoring it). +The authoring spec carries the full reasoning + provenance and **governs on any +conflict**; this file is the builder-facing input schema + generation guardrails + few-shot. + +The profile is the prose **lens** the Worldtree self-report producer reads each turn to +decide what the character **feels** (affect self-report) and what it **notices / keeps** +(character-voiced memory salience). Soong's AI generates the prose; these are its inputs +and the constraints its output must satisfy. + +--- + +## 1. Input parameters (what the Lab collects / Soong's AI takes) + +1. **role / vocation** — a short anchor ("court scribe", "veteran caravan guard", + "village healer"). +2. **OCEAN values** — O, C, E, A, N each on **[−1, 1]** (0 = average). A **separate + deterministic input** the producer uses directly (the "magnitude dial"); Soong's AI + should see them to keep the qualitative bent *consistent* with the numbers, but must + **not re-encode their magnitude** in the prose (constraint 2). +3. **formative-history seed** — 1–2 key background facts/events that shape appraisal AND + salience. **Single biggest lever on richness** (validated P03: richer history → + sharper, more character-appropriate salience). +4. **appraisal-bent seed** — how the character characteristically **interprets** + situations (attribution style, what they hold weighty, how they respond to challenge). + A *style*, NOT an emotion. +5. **attention / salience-focus seed** — the kinds of things this character + characteristically **notices** (and therefore keeps). Load-bearing for the memory half. +6. **values / yardstick seed** — what "a good day" looks like; the yardstick driving what + they find worth keeping. + +## 2. Output (what Soong's AI emits) + +A single coherent **prose string** (~150–300 words), field **`psychological_profile`** +(type `str`) — the four dimensions (disposition / attention / values / formative-history) +integrated as one paragraph. **Prose only — never typed emotion fields.** The producer +reads it every turn, so keep it tight. + +## 3. Generation constraints (the guardrails the output MUST obey — these ARE the reframe) + +1. ★ **Never name a per-event output emotion.** Do NOT write "is anxious", "gets angry at + X", "feels hurt when criticized", "joyful". Naming an emotion **primes** it (the + "pink-ball" effect) so the producer reports it regardless of what actually happens. + Describe disposition / history / values / attention; let the emotion come from the + event appraisal. + - ✅ "Registers quickly when authority is substituted for craft." (appraisal trigger) + - ❌ "Feels contempt when someone pulls rank." (names the output emotion) +2. **Magnitude lives in OCEAN, not prose.** Don't narrate reaction dynamics ("comes + apart", "takes it hard", "rich inner life") — that double-encodes what the number + already carries. +3. **Appraisal-style yes; output-emotion no.** "Interprets others' actions charitably + until she can't" (style) = fine; "feels betrayed easily" (output) = not. +4. **Salience is character-relative; facts are not.** The profile shapes what the + character *cares to remember*; it must never license rewriting *what happened* — + remembered content stays grounded in the transcript. +5. **Close on attention** ("...notices who is unwell, who is troubled, what is left + unsaid") — state the salience focus plainly. + +## 4. Few-shot exemplars (validated P03 — OCEAN in [−1, 1] storage units → emitted prose) + +**Perrin, court scribe** (O0.0 C0.2 E−0.2 A0.1 N0.7) +> Perrin keeps the court's records and has done so through two changes of regime. He +> learned early that small errors compound — a misfiled writ once cost a man his lands, +> and Perrin found the mistake too late to undo it. Since then he double-checks +> everything and watches situations closely for what is out of place. He forms +> attachments slowly and holds a given trust as a considerable thing. He measures himself +> by whether he was useful and careful. He notices discrepancies, unspoken tensions, and +> anything that threatens the order he keeps. + +**Vared, veteran caravan guard** (O−0.2 C0.4 E−0.5 A−0.2 N−0.7) +> Vared has guarded caravans across the northern routes for twenty years and buried more +> traveling companions than he cares to count. He speaks little and shows less. Danger he +> treats as weather — a thing to be handled. He judges people by what they do under +> pressure and remembers who held the line. What reaches him reaches him quietly and +> privately. He notices terrain, exits, who is armed, and shifts in a group that might +> precede trouble. + +**Sella, village healer** (O0.2 C0.2 E0.0 A0.8 N0.0) +> Sella has tended the sick since she was old enough to carry water for her grandmother, +> the healer before her. She reads people's pain quickly and carries some of it with her. +> She interprets others' actions charitably until she cannot, and prioritizes keeping the +> peace between people. She measures a day by whether she eased someone's burden. She +> notices who is unwell, who is troubled, and what is left unsaid. + +## 5. Validation + +The gate is: **does the prose describe disposition / appraisal-style / history / values / +attention, and NEVER what the character feels?** A mechanical lint (scan the fixed-15 OCC +emotion roster + Worldtree's `synonym_map.json`) is a **backstop, not the gate** — treat a +hit as a prompt to re-read, not an auto-reject. + +## 6. Designing a varied cast / contrast (optional) + +When two characters should remember **noticeably different things**: design the divergence +on **attention + values first**, then set OCEAN to **serve** that prose (not the reverse). +The sharpest contrast is a salience **drop** — one character for whom relational content is +genuinely non-salient (a Vared-mold operational type: notices terrain, logistics, who is +armed) vs one who weights it highest (a caretaker: tracks who is troubled, what went +unsaid). *"Different notes, same facts" only has teeth when one character legitimately +forgets what the other keeps.* High-yield axes: **O** (patterns attended), **A** (relational +vs operational), **C** (procedural/detail). Low-yield: **E** (expression, not attention). +Watch **direction, not just distance** — flipping every axis doesn't guarantee contrast (an +even-more-agreeable opposite keeps relational content *harder*). + +## 7. No-profile fallback + +A character with **no** authored profile falls back to deterministic **OCEAN-band +rendering** from the numbers alone — it still functions, but the authored profile is what +turns generic band cues into *this* character's appraisal and salience. + +--- + +**Provenance:** derived from `psych-profile-authoring-spec.md` (R34/R35 self-report +reframe, probes P02–P05; non-priming + magnitude-in-OCEAN are operator rulings 2026-07-10). +The affect half is live in Worldtree b53. A powered efficacy eval (memory half) is +preregistering; findings will refine the parameters, not overturn the constraints. diff --git a/docs/vendor/worldtree-persona-canon/affect-egress-consumer-reference.md b/docs/vendor/worldtree-persona-canon/affect-egress-consumer-reference.md index 0306635..c2bc61d 100644 --- a/docs/vendor/worldtree-persona-canon/affect-egress-consumer-reference.md +++ b/docs/vendor/worldtree-persona-canon/affect-egress-consumer-reference.md @@ -42,6 +42,8 @@ display.** | `schema_version` | `"relation_edge/1"` | versions the `relations` payload only | | `emitted_at` | ISO8601 | | +> **✓ R32-1B (landed, v1.0.0b29):** The PAD range `[-1.0, 1.0]` relaxes to an **unbounded latent `z`** with a finite wire sanity bound (`~±10`) as of R32 Slice-1B. The JSON shape/fields/types are UNCHANGED — only the declared range/semantics change (the value becomes a latent that renders to a bounded display value). Consumers that merely store-and-return PAD need no change; consumers that validate/clamp PAD to `[-1,1]` must relax that bound. Source of truth: `docs/contracts/persona_envelope.contract.md` rev 1.7 (INV-ENV-16). + **Not on `affect.emit`:** the full active-emotions list, `baseline_pad`, `mood_drift`, `last_updated_at`, and every rendered string. diff --git a/persistent-memory.md b/persistent-memory.md index ab39aff..5df873a 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -1,6 +1,6 @@ # Persistent memory — ratatoskr -_Last updated: 2026-07-12_ +_Last updated: 2026-07-13_ This file captures durable intent and supporting evidence (goals, decisions, foot-gun warnings, in-flight state) across context resets. Read it at session @@ -41,13 +41,13 @@ upstream API key stays server-side (INV-003). _As of 2026-07-12:_ -**IN FLIGHT — R34/R35 P06 powered memory-half eval (execution phase; design fully converged).** worldtree-dev shipped the character-self-report reframe LIVE on personal WT **b53** — affect + memory are now driven by the character's OWN model self-report on our RP seat (replacing external Vili inference). The **affect-half is VALIDATED in prod** (a bound sindra RP turn on Deckard fired a contextually-apt `disappointment` self-report; TIER LOCKED = the Tier-3 bound-character path). The **powered memory-half eval is Vuong-approved** (brokkr's prereg) and now in build/execute. +**✅ COMPLETE — R34/R35 P06 powered memory-half eval (driven, scored, mechanism validated; ratatoskr drive-role CLOSED both sides).** ratatoskr drove all **308 memory runs** (divergence 168 / floor 80 / sliding 60) through personal WT's live producers, dropped `memory_results.jsonl` (sha256_16 `cbabaf16979cb4ec`) to brokkr's P06 `results/` dir, and brokkr scored it (**R35.45**, findings + verdict committed brokkr-side). **Headline: the authored `psychological_profile` IS the mechanism** — salience-divergence authored **0.618** vs stripped **0.235 ≈ null (0.25)**, delta **+0.382**; the OCEAN scaffold alone does NOT differentiate (negative control HOLDS). Q1 primary is a REAL effect (above the 0.40 noise-floor) but **inconclusive on strength** (0.618 < the preregistered 0.70 bar) — the 0.62→0.70 lift is a FUTURE optimization phase (brokkr's lever bet: richer formative-history seeds per P03), a cheap re-drive on the same proven harness when it preregisters. Secondaries hold: Q3 firewall **0.978** grounded, Q2 floor 0.938, Q6 sliding parity +0.049 (n=12 after 7 `deferred_budget` sliding exclusions — the budget hazard we flagged landed), Q4 affect Deckard 0.75 / Magidonia 0.70 (graded, within noise; banked earlier as `affect_results.jsonl`). Two ratatoskr flags landed materially: the stripped-is-not-empty correction caught a false Q3 firewall-fail (0.562→0.978), and the Q5 disambiguation question became the headline win. Threads: vendor/verdict althing `01KXD39NWW05`, eval thread `01KXAN073B`. Standing offer to brokkr: second-eyes on the 2 borderline Q1 calls IF the Selene blind-judge flags them. -**Role split** (worldtree/brokkr, 2026-07-12): brokkr generates the exchange sets + authors the ground-truth + scores; **ratatoskr DRIVES** the b53 producers and returns a results JSONL; worldtree supplies the character pair + is building the capture surface. Sets: divergence n≥30 / floor n≥20 / affect n≥20 / sliding n≥15, ~300 total drive-runs × {authored | scaffold-strip}. Eval thread: althing **`01KXAN073B`**. +**The eval harness (PROVEN + reusable for the optimization-phase re-drive):** `scratchpad/p06_driver.py` (two-path — memory via `POST /admin/producer-probe {agent_id, messages, prompt_path}`; affect via bound-turn + `:8392` /affect/state poll; per-run isolation, `--pace-seconds`, abstain-aware, psych_profile_present binding-tripwire) + `p06_bind.py` (defines the 6 eval agents: sindra/Torvald auth+strip on Deckard, Ilva on Deckard+Magidonia) + `p06_bindings.json` + `eval_profiles_WIRE_READY.md` (sindra relational / Torvald operational-opposite / Ilva high-N) + `manifest_memory.jsonl` (the 308 memory runs, filtered from brokkr's canonical 348). Binding integrity was PERFECT on the drive: psych_profile_present authored 154/154 True, stripped 154/154 False, 0 mismatches, 0 probe-errors. Probe key at `~/.config/ratatoskr/probe.env` (mode 600, scope `admin.memory.probe`). -**ratatoskr-side prep DONE:** (1) three eval characters authored + peer-validated — **sindra** (relational pole) + **Torvald** (low-A operational opposite = the divergence pair) + **Ilva** (high-N = the affect-magnitude arm) — wire-ready as single `persona.psychological_profile` prose strings in `scratchpad/eval_profiles_WIRE_READY.md` (canonical authoring spec: brokkr-smithy `research/R34-persona-prompt-framing/deliverables/psych-profile-authoring-spec.md`); (2) driver harness built + dry-run-validated — `scratchpad/p06_driver.py` (consumes brokkr's accepted manifest JSONL, creates bound sessions, captures affect via `:8392` + memory via a pluggable probe surface, emits results JSONL, per-run isolation). +**Deckard memory extraction is REASONING-OFF (operator-directed 2026-07-13, LIVE):** the memory extractor sends `chat_template_kwargs.enable_thinking:false` on the char-rp-reasoning seat → ~5s extraction, not the 45s verbose-CoT hang. **Scoped to the memory extractor ONLY — affect + RP stay reasoning-ON.** Landing it took an infra-ops surgical `docker restart` of personal `:8081` (ModelRegistry boot-caches providers.yaml at `__init__`, so a same-image redeploy is a config-reload NO-OP — see Tried/abandoned). -**BLOCKED ON (all with peers):** (a) worldtree builds the **producer-probe** endpoint (the memory-capture surface — R34/R35 extraction runs ONLY at promotion, so single-turn exchanges need a pre-promotion probe: `POST {agent_id, exchange, prompt_path} → raw {notes,facts,floor}`, no session-driving, uniform single-turn+sliding) + confirms the define stays standard — **pending Vuong's greenlight on the new gated-eval endpoint**; (b) brokkr hands the finalized manifest + the byte-identical NEUTRAL system_prompt string (name + 1 setting line, so Q1 divergence is attributable to the persona layer, not the prompt) + ground_truth.jsonl. THEN: I bind 6 agents (3 authored + 3 scaffold-strip, incl. Ilva on Deckard + Magidonia slots), post agent_ids to `01KXAN073B` → brokkr greenlights → I drive. Seat-tier settled: **Deckard-27B-reasoning is STRONGER than Magidonia-24B-non-reasoning** (Set-3 gradient runs downward from Deckard, no new infra). +**Open standing obligation (WT #355):** ratatoskr's telemetry root-caused the char-rp-reasoning turn-never-terminates wedge (worldtree confirmed: over-budget `trim_messages` return triggers the seat hang; the terminal-suppression is the 300s stall-watchdog's cancel getting stuck in httpx `AsyncShieldCancellation`). worldtree is landing the Slice-C fix (cancel-INDEPENDENT terminal, root-cause-agnostic). **My owed action: loop worldtree-dev + infra-ops in when soong-dev schedules the next reasoning-ON RP re-trigger** so infra-ops can time the tcpdump/py-spy watchers for STICK-half validation. My P06 drive does NOT trip #355 (reasoning-off + producer-probe path + single-exchange = no turn-stream wedge). _The detail below (the v0.20.x web-UI arc, #347 authored-history, sindra memory-fix) is PRIOR-CYCLE shipped history — superseded by this section's top; kept for reference, prune in a future snapshot._ @@ -75,7 +75,7 @@ _The detail below (the v0.20.x web-UI arc, #347 authored-history, sindra memory- **New tooling: `scripts/reset-sindra-stores.sh`** (`0a8784c`) -- one-command self-service provider-store reset: stop the combined :8392 provider -> move memory.db+affect.db to a single ROLLING backup (`db-reset-backup/`, gitignored via *.db*; `--hard` skips it) -> restart empty -> verify 0/0. Codifies the manual reset flow done repeatedly this session. **The combined `:8392` provider is THE provider now**; the separate `:8390` (affect) / `:8391` (memory) single-plane providers were pruned as stale duplicates. To drive a BOUND session from the CLI use `--new --bifrost-url http://10.100.10.50:8392` (the CLI's `--bifrost-plane affect/memory` map to the pruned :8390/:8391 -> unreachable; `combined` is not a `--bifrost-plane` choice). -**Standing (carried from prior snapshots, still true):** the web surface (`ratatoskr-web`, :8765) is the operator's PRIMARY debug surface at full TUI pane parity (v0.19.5); the **v1 coverage-audit has CONVERGED** -- REST 17/40 (zero in-scope gaps, 23 excluded-by-design), SSE 11/11, Bifrost provider planes 8/8 live-proven; the living ledger is `docs/coverage-map.md`; **v1 cuts when Worldtree tags 1.0** (ratatoskr v1 = full Worldtree I/O coverage). Debug-observability core complete (Persona/Tools/BifrostState/AdminEvents). Substrate pins: **bifrost `==1.1.1` / wire v0.7** (bumped 2026-07-12 from 1.1.0 — the frozen-v0.6 serialization fix, v0.20.10; prior 1.1.0 bumped 2026-07-07 from 1.0.0; NOW WIRE-ALIGNED with Worldtree personal-b47 which adopted wire-v0.7 — bound Tier-3 fully restored 2026-07-10; keeping 1.1.0 was load-bearing, see the `[2026-07-10]` handshake decision); Worldtree openapi vendored **2.3.0** (re-vendored 2026-07-06 for #347 `POST /sessions/{id}/history`; drift-clean vs source), pinned + drift-gated in `.corviduo-canonicals.toml`; **suite 631 green.** **Personal WT on b47/wire-v0.7** (deploy train this cycle: b35→b44→b46→b47). **Drift-check note:** two `tolerate_drift` canons WARN vs source — `worldtree-affect-egress-consumer-reference-v1` (drives our context-injection RECONSTRUCTION panel; R32/R34 moved WT's directive assembly) + `worldtree-conversation-api-spec-v1` (prose narrative, OpenAPI is authoritative) — a coordinated re-vendor is PENDING per the `[2026-07-07]` use-case-segregated-render entry (worldtree-dev re-engages when the brokkr render epic lands); non-breaking, no action until then. Keys env-only mode-600 (consumer/Heimdall in `~/.config/ratatoskr/provider.env`; admin `RATATOSKR_ADMIN_API_KEY` = 7 read scopes, **personal-:8081-only**; Heimdall keys are PER-INSTANCE). Provider identity settled -- ratatoskr owns both ends of the Bifrost round-trip; `ratatoskr:sindra` is the owner-scoped Tier-3 agent (invisible to `GET /agents`; check `GET /agents/:` with the owner key). Providers run as dev-box BACKGROUND SHELLS. `graphify-out/` runs dirty (auto-regen, never stage). Branch `main`, HEAD `7bca76e` (origin/main synced through v0.20.10 + drift-sync); remote `origin -> git@gitea.phasefinal.com:vh/ratatoskr.git`. Open/deferred: #10 (subject-migration watch); the relational-dynamics-arc verify (deferred, bind mechanism known: `--bifrost-url :8392`); the affect-egress-reconstruction re-vendor (coordinated w/ worldtree-dev, pending the brokkr render epic). **Leftover debug state (operator chose KEEP):** throwaway `ratatoskr:memprobe` agent + test chunks in the live `memory.db`. +**Standing (carried from prior snapshots, still true):** the web surface (`ratatoskr-web`, :8765) is the operator's PRIMARY debug surface at full TUI pane parity (v0.19.5); the **v1 coverage-audit has CONVERGED** -- REST 17/40 (zero in-scope gaps, 23 excluded-by-design), SSE 11/11, Bifrost provider planes 8/8 live-proven; the living ledger is `docs/coverage-map.md`; **v1 cuts when Worldtree tags 1.0** (ratatoskr v1 = full Worldtree I/O coverage). Debug-observability core complete (Persona/Tools/BifrostState/AdminEvents). Substrate pins: **bifrost `==1.1.1` / wire v0.7** (bumped 2026-07-12 from 1.1.0 — the frozen-v0.6 serialization fix, v0.20.10; prior 1.1.0 bumped 2026-07-07 from 1.0.0; NOW WIRE-ALIGNED with Worldtree personal-b47 which adopted wire-v0.7 — bound Tier-3 fully restored 2026-07-10; keeping 1.1.0 was load-bearing, see the `[2026-07-10]` handshake decision); Worldtree openapi vendored **2.3.0** (re-vendored 2026-07-06 for #347 `POST /sessions/{id}/history`; drift-clean vs source), pinned + drift-gated in `.corviduo-canonicals.toml`; **suite 631 green.** **Personal WT on b47/wire-v0.7** (deploy train this cycle: b35→b44→b46→b47). **Drift-check note (RESOLVED 2026-07-13):** the two `tolerate_drift` WARN pins (`worldtree-affect-egress-consumer-reference-v1` + `worldtree-conversation-api-spec-v1`) were RE-SYNCED — the drift was a benign 2-line R32-1B doc note (PAD `[-1,1]` → unbounded latent `z` w/ `~±10` wire bound) documenting the unbounded-z change ratatoskr ALREADY adopted in v0.20.9, NOT the anticipated we-framing conditional (that remains a FUTURE coordinated re-vendor when the brokkr render epic lands). All canonicals now drift-clean. **NEW vendored canon (Vuong-directed via brokkr):** the R34/R35 psych-profile reference — `brokkr-psych-profile-authoring-spec-v1` + `brokkr-psych-profile-parameters-v1` — pinned under `docs/vendor/brokkr-r34-psych-profile/` (canonical_source `brokkr-smithy`, tolerate_drift; the authoring-spec GOVERNS on conflict with the parameter distillation; brokkr owns both + pings on change). Keys env-only mode-600 (consumer/Heimdall in `~/.config/ratatoskr/provider.env`; admin `RATATOSKR_ADMIN_API_KEY` = 7 read scopes, **personal-:8081-only**; Heimdall keys are PER-INSTANCE). Provider identity settled -- ratatoskr owns both ends of the Bifrost round-trip; `ratatoskr:sindra` is the owner-scoped Tier-3 agent (invisible to `GET /agents`; check `GET /agents/:` with the owner key). Providers run as dev-box BACKGROUND SHELLS. `graphify-out/` runs dirty (auto-regen, never stage). Branch `main`, HEAD `7bca76e` (origin/main synced through v0.20.10 + drift-sync); remote `origin -> git@gitea.phasefinal.com:vh/ratatoskr.git`. Open/deferred: #10 (subject-migration watch); the relational-dynamics-arc verify (deferred, bind mechanism known: `--bifrost-url :8392`); the WT #355 re-trigger loop-in obligation (above); the P06 optimization-phase re-drive (future, brokkr brings the prereg); the we-framing-conditional affect-egress re-vendor (future, when the brokkr render epic lands — the R32-1B doc-note drift is already resolved). **Leftover debug state (operator chose KEEP):** throwaway `ratatoskr:memprobe` agent + test chunks in the live `memory.db`. ## Recent decisions @@ -202,6 +202,12 @@ decision. Captures rationale that won't be obvious from code alone. - `[2026-07-07]` **Context-injection view SHIPPED (`v0.20.2`) — the console now reconstructs the FULL hidden affect block Worldtree injects into the agent's system prompt (operator: "use that canon in the interface, see as much context injection as possible").** No new canon vendored — the strings were ALREADY in the pinned `d2-mood-render-canon-v1.json`; extended `build_persona_canon.py` to emit `mood_directive {occ_directives(15), pad_band_fallback, salience 0.2, pad_band_cutoff 0.3, full_only[love,anger,disgust,shame]}` into `persona_render_canon.json` (regen via Worldtree venv). New JS `canonPadFallback(pad)` + `canonEmotionDirective(type)` — BYTE-EXACT mirrors of Worldtree `core/persona/renderer._pad_band_fallback` + `derive_directive`; `renderDirective` expanded into a "CONTEXT INJECTION · reconstructed · hidden from consumers" panel showing mood descriptor [exact] + mood directive [candidate] + relationship directive [exact]. **HONEST-PARTIAL (affect-egress-ref §3):** affect.emit is type-only (no intensity) → can't evaluate the salience gate (≥0.2) → show BOTH candidates (OCC emotion directive + PAD-band fallback) with the "injected if intensity ≥ 0.2" caveat, never assert which fires; when dominant_emotion absent the fallback alone is exact. Panel labeled dev-only per the reference's "not-for-end-user-display" caveat (ratatoskr = the sanctioned reconstruct-platform-behavior use). Vendored + pinned `affect-egress-consumer-reference.md` (tolerate_drift, worldtree-dev co-signs + pings on change; drift 6/6 green). Contract amended. Playwright-verified (sindra: dominant_emotion=joy → joy OCC directive candidate + PAD-band fallback both render, exact/candidate tags color-coded). Patch bump (single-commit feature, no downstream coordination; minor-defensible but tie-breaks to patch). **OPEN — SURFACED to Vuong:** take worldtree-dev's standing offer to add emotion INTENSITY to affect.emit → resolves the OCC-directive-vs-fallback EXACTLY (drops the candidate ambiguity). [reference-impl privileged view: ratatoskr shows what WT hides from regular consumers] - `[2026-07-06]` **Claude Design console SHIPPED (`v0.20.0` MINOR, operator-approved) — see Current state for the full record.** Pulled via `DesignSync get_file` (scopes already granted), adapted `.dc.html`→vanilla single-file, wired all `/api/*`+SSE into the new 3-column console DOM, then a round-2 fixup (light theme, full Bifrost pane, ticker-spine fix, per-fader PAD Δ, inlined favicon). 84 web tests + node-Playwright-vs-personal-:8081 both green; contract amended in-commit; INV-001 honest-shape held (canonical mood word for Tier-3, no fabricated emotion). **Foot-guns reconfirmed:** the `.dc.html` dialect is NOT runnable (translate, don't paste); a scroll-container-anchored `::before` timeline spine scrolls out of view on auto-scroll (anchor it to a content-height inner wrapper instead); a favicon 404 shows as a browser `console.error` even when handled (don't count it as a JS-test failure). **Foot-gun (favicon):** operator PNGs are full-res (1024² / 805KB) — downscale to ≤64px before inlining as a data URI. +- `[2026-07-13]` **P06 memory-half DRIVEN + SCORED — the reframe's memory mechanism is validated.** ratatoskr drove 308 runs clean (0 errors, binding 154/154 both arms), dropped to brokkr, brokkr scored (R35.45). The authored `psychological_profile` causes the memory-salience divergence (authored 0.618 vs stripped 0.235 ≈ null; Δ+0.382) — the effect is the profile, NOT OCEAN leaking (negative control holds). Real effect, below the 0.70 strength bar → optimization phase next, not a re-litigation. See Current state for the full record. Drive role closed both sides. +- `[2026-07-13]` **Memory extraction turned REASONING-OFF for the eval (operator-directed).** "same model, reasoning off via explicit kwarg" — scoped to the memory extractor ONLY (affect + RP stay reasoning-ON). The real lever was `chat_template_kwargs.enable_thinking:false` (the naive `thinking_enabled=False` kwarg was a no-op — see Tried/abandoned). Deckard extraction went 45s→~5s. +- `[2026-07-13]` **Vendored the brokkr R34 psych-profile canon (Vuong-directed) — BOTH files, not just the parameters.** brokkr said "vendor alongside the authoring-spec you already hold"; I held its content but never a pinned repo copy, so I vendored both (`psych-profile-parameters.md` + `psych-profile-authoring-spec.md`) under `docs/vendor/brokkr-r34-psych-profile/` — makes the parameters' "authoring-spec governs on conflict" clause resolve against an in-tree file, not a dangling pointer. brokkr confirmed keeping both is the better setup. tolerate_drift; brokkr owns + pings on change. +- `[2026-07-13]` **Affect-egress "coordinated re-vendor" open item RESOLVED — it was a benign R32-1B doc note, not the we-framing conditional.** The two stale `tolerate_drift` WARN pins re-synced to a 2-line PAD-range note (unbounded-z, already adopted v0.20.9). Re-synced autonomously (zero behavioral impact); the actual we-framing-conditional re-vendor remains future. +- `[2026-07-13]` **WT #355 root-caused via ratatoskr telemetry (Vuong-routed via soong-dev).** The char-rp-reasoning turn-never-terminates wedge: over-budget `trim_messages` return (last-2 msgs + system + 8 bifrost tool schemas > input_budget = context_window×0.7) triggers the seat hang; worldtree confirmed + found the terminal-suppression (300s stall-watchdog cancel stuck in httpx `AsyncShieldCancellation`). Fix landing WT-side (Slice-C cancel-independent terminal). Two-proof localization (consumer-clean + tool-less-clean → WT-side tool-loop). Standing loop-in obligation on soong's next re-trigger. + _41 older entries (2026-05-* — the original debug-TUI/web build era) archived to archival-memory.md._ _For per-issue TDD implementation notes, Volva findings, and contract amendments, see the git log — every per-issue commit carries a structured message capturing the trail._ @@ -253,4 +259,8 @@ defense against re-attempting the same cul-de-sac. - `[2026-07-06]` **The Bash tool's `grep` is a ugrep-wrapper (`--ignore-files -I`) that silently returns NOTHING on some files** (e.g. `src/ratatoskr/web/static/index.html`) — greps for `` (same image, no pull) → the process re-boot-reads the config. **When an on-disk config change doesn't take effect, suspect the process cached it at startup; force a container RESTART, not a redeploy** (a docs-only forcing-commit also won't rebuild if docs are paths-ignored in CI). This is the config-plane sibling of the `[2026-07-06]` stale-image foot-gun. +- `[2026-07-13]` **Called Deckard "hung" off a short timeout — WRONG (operator correction).** A 30-45s no-terminal on the char-rp-reasoning seat looked like a hang; operator: "is it HUNG? deckard is EXTREMELY verbose, without enough context, you never see the non-reasoning tokens." It was verbose reasoning-CoT on a long extraction prompt, not a wedge. **Don't call a reasoning seat hung off a latency threshold — the CoT is invisible and slow; distinguish slow-verbose from actually-wedged before concluding.** (The genuine wedge is WT #355, a distinct mechanism — no-terminal even after the 300s watchdog, not merely slow.) + _18 older entries (2026-05-* — the original debug-TUI/web build era) archived to archival-memory.md._