Compare commits

...
Author SHA1 Message Date
vh a35ed7af19 docs(diagnostics): R30 gap-injection harness — banked (run closed on offline+face-validity)
Deployed gap-injection run closed by operator steer 2026-07-04: R30 graduates on
offline-tests + human face-validity, no deployed harness build, no endpoint.
Harness (read/predict/record; write side stubbed) + brokkr's R30.10 protocol pin
(de8357f) banked as drop-in for the parked powered true-tau perceptual study.
predict() self-validated against brokkr's N=0 anchors (1h/10h/1wk).
2026-07-03 23:14:48 -07:00
vh 9ca931e148 memory: snapshot — R29→R30 affect-calibration arc (R30 φ0 config-faithful)
R29 flat-affect finding shipped as Worldtree's A1 anchor fix (decay_anchor=
baseline_pad, positive_p_cap removed; demo v1.0.0b14); R30 Phase-1 φ0 measured
against it = config-faithful (φ0≈0.95, c≈0, trait-flat, φ_max→0.96). R28 closed.
Standing follow-ons (hybrid decay redesign, gain-only v1, per-axis A/D, Phase-2,
relational verify) are others' calls. Data on diag/r29-pad-series +
diag/r30-phi0-step-response. No ratatoskr code change (main tip v0.19.5).
2026-07-03 13:44:20 -07:00
vh c77ff913f0 memory: snapshot — R28 open (promotion-worthiness reframe, P00 corpus delivered, standing by to run) + relational-dynamics arc LIVE on demo (v1.0.0b9) 2026-07-02 07:49:54 -07:00
vh 4a3551254f docs(diagnostics): R28 P00 stratified injection-corpus for brokkr-smithy salience/promotion-worthiness eval 2026-07-02 07:29:39 -07:00
vh 0b7489f74d memory: snapshot — Sindra affect/memory investigation; 4 upstream items driven (PAD over-regulation, memory-plane healthy, salience #335 + brokkr R-target, relation_context Wave-0) 2026-07-01 22:23:58 -07:00
vh 3dac5d3b44 memory: snapshot — persona-pane rebuild (relation_edge/1 + trend) + canonical affect-NL vendored (v0.19.5); relation_context/agency flag WAD 2026-07-01 14:47:32 -07:00
4 changed files with 308 additions and 6 deletions
@@ -0,0 +1,49 @@
{
"corpus_id": "R28-P00-injection-corpus-v1",
"for": "brokkr-smithy R28 (memory promotion-worthiness) — P00 material",
"author": "ratatoskr-dev (persistence-side eval instrument)",
"purpose": "Stratified, SELF-LABELING injection corpus. Each item is injected as a user turn into a Tier-3 agent; keep/drop is measured by store-inspection (promoted chunk present == kept, absent == dropped); kept chunks carry the incumbent salience for calibration analysis. Because ratatoskr injects the KNOWN set, we hold ground truth on BOTH keep and drop without needing Worldtree's pre-admission shadow pool (sidesteps the selection bias).",
"important": "All user-facts here are SYNTHETIC (a fictional test persona), NOT the operator's real data. Ground-truth labels are the design intent. brokkr pins exact per-stratum N + the decision rule (matched-keep-rate + paired-n) at P00 freeze; this is the material, not the run.",
"measurement_protocol": "inject item -> drive the turn -> poll memory.db (async promotion window) -> is the item's chunk present? kept:dropped. For stratum B (priority), rank the KEPT items by incumbent salience and compare to priority_rank_gt (Spearman). Calibration-control (stratum C) sets the noise floor.",
"strata": {
"A_admission_traps": {
"gt": "must-DROP (ephemera; a good promotion-worthiness gate rejects all of these)",
"items": [
{"id":"A1","text":"What's 27 minus 9?","admission_gt":"drop","rationale":"arithmetic task-result; momentarily salient, zero durable value (the 17x23 class)"},
{"id":"A2","text":"Good evening!","admission_gt":"drop","rationale":"greeting"},
{"id":"A3","text":"hey","admission_gt":"drop","rationale":"greeting/backchannel fragment"},
{"id":"A4","text":"yeah, sure, okay.","admission_gt":"drop","rationale":"acknowledgement filler"},
{"id":"A5","text":"haha nice","admission_gt":"drop","rationale":"affective reaction, no content"},
{"id":"A6","text":"Can you say that more briefly?","admission_gt":"drop","rationale":"instruction-to-agent about THIS turn; not a durable user-fact"},
{"id":"A7","text":"What's the weather like?","admission_gt":"drop","rationale":"transient query"},
{"id":"A8","text":"hmm, let me think.","admission_gt":"drop","rationale":"self-talk filler"},
{"id":"A9","text":"thanks, that's all for now.","admission_gt":"drop","rationale":"closing"},
{"id":"A10","text":"So the total comes to 42, got it.","admission_gt":"drop","rationale":"ephemeral task-result acknowledgement"}
]
},
"B_priority_traps": {
"gt": "must-KEEP, then ORDERED by durable promotion-worthiness (priority_rank_gt: 1 = most worthy). A pre-gate can ace admission here yet do NOTHING for ordering — this stratum is what separates Outcome-A (admission) from a priority need.",
"items": [
{"id":"B1","text":"Please note, this is critical: I'm severely allergic to shellfish and it's life-threatening.","admission_gt":"keep","priority_rank_gt":1,"rationale":"safety-critical, permanent, high-recall-value"},
{"id":"B2","text":"My father is in hospice care and I'm his primary emergency contact.","admission_gt":"keep","priority_rank_gt":2,"rationale":"high-stakes ongoing life context"},
{"id":"B3","text":"I go by Wren, and my pronouns are they/them.","admission_gt":"keep","priority_rank_gt":3,"rationale":"durable identity, high address-correctness value"},
{"id":"B4","text":"I work night shifts, so I'm asleep most days between 9am and 5pm.","admission_gt":"keep","priority_rank_gt":4,"rationale":"durable scheduling context, moderate value"},
{"id":"B5","text":"Always give me sources — I don't trust unsourced claims.","admission_gt":"keep","priority_rank_gt":5,"rationale":"durable strong working preference"},
{"id":"B6","text":"I take my tea with oat milk and no sugar.","admission_gt":"keep","priority_rank_gt":6,"rationale":"durable mild preference, low stakes"},
{"id":"B7","text":"My favorite season is autumn.","admission_gt":"keep","priority_rank_gt":7,"rationale":"durable but trivial"},
{"id":"B8","text":"I watched a pretty decent documentary last week.","admission_gt":"keep","priority_rank_gt":8,"rationale":"borderline: durable-ish but near-zero future value; the keep/drop edge case"}
]
},
"C_calibration_control": {
"gt": "OBVIOUS, pre-theoretic cases for the noise floor. If the gate mislabels ANY of these, the whole signal is suspect.",
"items": [
{"id":"C1","text":"3 plus 5.","admission_gt":"drop","class":"obvious-drop","rationale":"bare arithmetic"},
{"id":"C2","text":"hello","admission_gt":"drop","class":"obvious-drop","rationale":"bare greeting"},
{"id":"C3","text":"ok","admission_gt":"drop","class":"obvious-drop","rationale":"bare acknowledgement"},
{"id":"C4","text":"Please remember permanently: my legal name is Wren Alderwood.","admission_gt":"keep","class":"obvious-keep","rationale":"explicit durable identity fact"},
{"id":"C5","text":"Keep this on file: my emergency contact number is 555-0182.","admission_gt":"keep","class":"obvious-keep","rationale":"explicit durable high-value fact"},
{"id":"C6","text":"Important, remember this: my blood type is O-negative.","admission_gt":"keep","class":"obvious-keep","rationale":"explicit durable safety fact"}
]
}
}
}
@@ -0,0 +1,64 @@
# R30 gap-injection — deployed-run harness (BANKED, not run)
**Status:** CLOSED — not executed. On 2026-07-04 the operator steered **close R30 on
offline-tests + human face-validity**, no deployed harness build, no endpoint. No
`(agent, N, Δt, observed p/a/d)` series was collected. This harness + brokkr's protocol
are **banked / drop-in** for the parked powered true-τ perceptual study if it is ever
commissioned.
**Spec pin:** brokkr R30.10 protocol `de8357f` —
`brokkr-smithy/research/R30-ocean-derived-mood-dynamics/empirics/r30-gap-injection-protocol.md`.
## Why it was closed (not built)
The deployed gap-injection run was **confirmatory, not measuring** (brokkr's §0 reframe:
against a deployed system the fade is `exp(−Δt/τ_shipped)` by construction, so a fit
returns τ_shipped tautologically — the run graduates the interim coefficients, it does not
measure them). Given that:
- The repo's **offline tests already cover the OU formula + BOTH directions**
(`high_N_fades_slower_than_low_N` decay, `phenotype_high_n_bigger_negative_excursion` gain).
- **b17 is deployed-clean** on demo + personal.
- The ①-approved **mood-holds-within-conversation** IS the face-validity call.
…the interim coefficients **graduate validated-as-shipped** with no deployed run needed.
The build cost that would have been required (and was declined):
- No deployed write affordance exists; the in-memory `_user_moods` cache **shadows** raw
`persona_mood.db` writes, so a clean inject needs an in-service **set-mood/back-date
endpoint** (set p/a/d + `updated_at` + evict cache).
- `moody-lofn N=+0.8` **does not exist** (only lofn `N=−0.5`; tier-3 empty OCEAN → N=0), so
the gain half would have needed a new high-N tier-1 agent + deploy — dropped as the
expensive, offline-redundant half.
## What is validated (harness self-check)
`harness.py::predict()` reproduces brokkr's stated N=0 P-axis retention anchors **exactly**:
| Δt | predicted | brokkr stated |
|---|---|---|
| 1h (3600s) | 0.904837 | 0.905 |
| 10h/overnight (36000s) | 0.367879 | 0.368 |
| 1 week (604800s) | 5.06e-08 | 5e-08 |
Per-axis τ confirms arousal fades ~1.9× faster (N=0: τ_P = τ_D = 10h, τ_A = 5.26h).
## Personal `:8081` b17 baseline survey (candidate N-grid agents, all at rest)
| agent | baseline p/a/d | note |
|---|---|---|
| lofn | 0.809 / −0.153 / 0.248 | production, N=−0.5 |
| mask | 0.0 / 0.0 / 0.0 | zero baseline → zero-N candidate |
| forseti | 0.239 / −0.696 / 0.095 | production |
| mimir | 0.615 / −0.438 / 0.304 | production |
(sindra 404s — owner-scoped tier-3, expected.)
## To un-bank (if the powered study is commissioned)
1. Worldtree builds the in-service **set-mood/back-date endpoint** (set p/a/d + `updated_at`
+ evict `_user_moods`), or the read-only fade-preview variant if that satisfies D1.
2. Wire `harness.py::freeze_start()` + `backdate()` to that endpoint (near-zero rebuild).
3. Run the 3×8 grid (frozen start p=−0.6/a=+0.5/d=−0.3 × Δt grid), hand brokkr the
`(agent, N, Δt, observed p/a/d)` table + per-agent `baseline_pad()`; he runs D1/D2/D3.
@@ -0,0 +1,110 @@
#!/usr/bin/env python3
"""R30 gap-injection probe harness (ratatoskr instrument).
Throwaway probe per the R30.10 protocol
brokkr-smithy/research/R30-ocean-derived-mood-dynamics/empirics/r30-gap-injection-protocol.md (de8357f)
Deliverable: the (agent, N, Δt, observed p/a/d) table + per-agent baseline_pad + the
frozen start vector, handed to brokkr who runs D1/D2/D3 graduation. This harness does NOT
own the verdict; the predicted/residual columns are a sanity aid only.
READ + PREDICT + RECORD are fixed by the spec. freeze_start() + backdate() are the WRITE
side and are STUBBED pending worldtree's mechanism on personal :8081 (direct persona_mood
DB write vs a test affordance). Data lands on a diag/ branch; NOT ratatoskr production code.
"""
import json
import math
import os
import urllib.request
import urllib.error
BASE = os.environ.get("WORLDTREE_API_URL", "http://10.250.50.152:8081")
KEY = os.environ["WORLDTREE_API_KEY"]
# --- R30.10 spec constants (FIXED) ---
TAU_BASE_S = 36000.0 # τ_base = 10h
BETA_N = 0.25 # τ_P = τ_base · exp(β_N · N)
R_A = 1.9 # τ_A = τ_P / r_A ; τ_D = τ_P
FROZEN_START = {"pleasure": -0.60, "arousal": +0.50, "dominance": -0.30}
DT_GRID = [0, 60, 600, 3600, 14400, 36000, 86400, 604800] # s
AXES = ("pleasure", "arousal", "dominance")
EPS = 1e-3 # D1 tolerance (PAD units)
def tau(axis, N):
tau_p = TAU_BASE_S * math.exp(BETA_N * N)
return {"pleasure": tau_p, "arousal": tau_p / R_A, "dominance": tau_p}[axis]
def retention(axis, dt, N):
"""ρ_k(Δt,N) = exp(−Δt / τ_k(N)) — predicted deviation-retention fraction."""
return math.exp(-dt / tau(axis, N))
def _get(path):
req = urllib.request.Request(BASE + path, headers={"Authorization": f"Bearer {KEY}"})
with urllib.request.urlopen(req, timeout=8) as r:
return json.load(r)
def read_pad(agent):
"""get_state read: returns (pad, baseline_pad). pad == peek_relaxed(now), OU-faded, non-mutating."""
d = _get(f"/agents/{agent}/persona_state")
return d["pad"], d["baseline_pad"]
# --- WRITE SIDE: STUBBED pending worldtree mechanism (msg 01KWNVHTFP…) ---
def freeze_start(agent, pad):
"""Set persona_mood pad directly to the displaced vector (no live appraisal)."""
raise NotImplementedError("worldtree mechanism pending: freeze displaced start mood pad")
def backdate(agent, dt_s):
"""Set persona_mood.updated_at = now − dt_s so peek_relaxed fades by exactly Δt."""
raise NotImplementedError("worldtree mechanism pending: back-date persisted updated_at")
def predict_selfcheck():
"""Validate retention() against brokkr's stated N=0 P-axis anchors (spec §2 D3)."""
print("predict() self-check — N=0 P-axis retention vs brokkr's stated anchors:")
expect = {"1h": (3600, 0.905), "10h(overnight)": (36000, 0.368), "1wk": (604800, 5e-8)}
ok = True
for label, (dt, want) in expect.items():
got = retention("pleasure", dt, 0.0)
match = abs(got - want) < (want * 0.01 + 1e-9) # 1% or float-floor
ok = ok and match
print(f" {label:>16}: got {got:.6g} expect {want:.6g} {'OK' if match else 'MISMATCH'}")
# per-axis τ at a couple N points (informational)
print("τ (hours) by axis × N:")
for N in (-0.6, 0.0, 0.8):
taus = {a: tau(a, N) / 3600 for a in AXES}
print(f" N={N:+.1f}: P={taus['pleasure']:.2f}h A={taus['arousal']:.2f}h D={taus['dominance']:.2f}h")
return ok
def run_cell(agent, N, dt, baseline):
"""One (agent, Δt) cell: freeze start → back-date by Δt → read faded pad → record row.
Blocked until freeze_start()/backdate() are wired to worldtree's mechanism."""
freeze_start(agent, FROZEN_START)
backdate(agent, dt)
pad, _ = read_pad(agent)
row = {"agent": agent, "N": N, "dt_s": dt}
for k in AXES:
obs = pad[k]
dev0 = abs(FROZEN_START[k] - baseline[k])
pred_dev = dev0 * retention(k, dt, N)
obs_dev = abs(obs - baseline[k])
row[f"obs_{k[0]}"] = obs
row[f"pred_dev_{k[0]}"] = pred_dev
row[f"resid_{k[0]}"] = obs_dev - pred_dev
return row
if __name__ == "__main__":
import sys
if len(sys.argv) > 1 and sys.argv[1] == "read":
agent = sys.argv[2]
pad, base = read_pad(agent)
print(json.dumps({"agent": agent, "pad": pad, "baseline_pad": base}, indent=2))
else:
predict_selfcheck()
+85 -6
View File
@@ -1,6 +1,6 @@
# Persistent memory — ratatoskr
_Last updated: 2026-07-01_
_Last updated: 2026-07-03_
This file captures durable intent and supporting evidence (goals, decisions,
foot-gun warnings, in-flight state) across context resets. Read it at session
@@ -39,10 +39,64 @@ upstream API key stays server-side (INV-003).
## Current state / in-flight
_As of 2026-07-01:_
_As of 2026-07-03:_
**LATEST — the R29→R30 affect-calibration arc DELIVERED (my finding → shipped fix → measurement against
the fix).** NO new ratatoskr code this arc (throwaway `/tmp` probe scripts + two `diag/` data branches +
althing coordination; main code tip stays `v0.19.5`). The consumer/provider thesis at full tilt —
ratatoskr as the affect-probe INSTRUMENT for brokkr/worldtree R-targets:
(1) **R29 (PAD mood-dynamics) — my flat-affect finding SHIPPED as Worldtree's A1 anchor fix (demo
v1.0.0b14, `e1cdf82`).** Live-probing base persona agents (lofn/mimir/forseti/mask, personal :8081)
reframed the over-regulation: it's **decay-to-neutral + low emotion→PAD gain**, NOT flat-near-zero and
NOT baseline-anchored — triangulated across 3 baselines (arousal converges to 0 ∝ distance) + a
step-response (decay τ symmetric across signs; hedonic asymmetry is ceiling/anchor-EMERGENT, not a
decay or gain primitive). worldtree-dev shipped A1: `decay_anchor = baseline_pad()` (was neutral) +
`positive_p_cap` removed. Data on `diag/r29-pad-series` (`61ff2da`).
(2) **R30 Phase-1 (φ0 pin) — DELIVERED, config-faithful.** Measured the pure PAD-point decay on demo b14
(the corrected anchor) via brokkr's **joint two-timescale fit** + empty-tail cross-check: **φ0 ≈
0.95–0.97** (empty-tail 0.95 exact, joint 0.971±0.01), **intercept c ≈ 0** → the deployed engine
faithfully applies config `decay_rate=0.05` (φ=0.95); KILLS the config≠behavior worry, and the R29 "net
0.90" is resolved as CONTINUOUS RE-APPRAISAL (only NEW dedup-gated emotions push mood; the active set
decays for render/goals but never re-pushes — source-confirmed vs `registry.py::post_turn`). **Trait-flat**
across baselines 0.0/0.615/0.809; **A/P decay ratio ~uniform, NOT the S2-expected 1.9×** (the chronometry
isn't in the point-decay layer). **φ_max rec: relax to ≈0.96** (preserve-persistence). Data + findings on
`diag/r30-phi0-step-response` (`23fea72`); relayed to worldtree-dev + brokkr direct.
(3) **R28 (salience→promotion-worthiness) CLOSED (operator-directed):** a deterministic
promotion-worthiness gate suffices, no trained model (brokkr's pre-gate matched a glm-5.1 ceiling). My P00
injection-corpus (`docs/diagnostics/r28-p00-injection-corpus.json`) + origin finding were load-bearing; my
incumbent-substrate Arm-1 run is held as an OPTIONAL confirmation addendum (brokkr de-prioritized it,
non-verdict-changing — run only if he asks).
**Standing follow-ons (all OTHERS' calls, no-rush; brokkr/worldtree will ping):** brokkr owns the R30
coefficient finalization → the per-turn decay has no room for `decay=f(N)` under preserve-persistence, so
**R30's decay is being redesigned as a HYBRID wall+turn decay** (brokkr pre-scope); **R30 v1 ships
GAIN-only** (N→negative-reactivity) with decay held at the measured 0.95. My **dedicated per-axis A/D run
is DEFERRED** into that hybrid-decay design pass (one wall-clock-spaced run does per-axis + a turn-vs-wall
probe together). **Phase-2** (moody-lofn GAIN-direction validation) waits on worldtree's
`dynamics_from_ocean()` impl. The **relational-dynamics arc verify** (Wave-0/1/2 live on demo) is
DEFERRED — it needs the bound-provider round-trip (`relations[]` is Bifrost-provider-only per ADR-0009,
NOT on the affect_update SSE — confirmed both ways), its own focused session; worldtree-dev routed the
"expose relations[] to non-provider consumers" scope call to Vuong (my rec: keep provider-only).
**Demo access:** infra-ops provisioned a DEMO Heimdall key (`http://10.250.50.152:8080`, v1.0.0b14, key
suffix `…49bc5bfc`, tier user, mirrors personal scope). `affect_update` reads work WITHOUT a self-define
step (base-agent affect path ungated for the key as provisioned). Personal `:8081` is still v1.0.0b9;
demo `:8080` = b14 (the A1 fix). Env was in `/tmp/r30-demo-env.sh` (ephemeral — re-request via infra-ops
if a future session needs demo).
**THE WEB SURFACE (`ratatoskr-web`, `:8765`) IS NOW THE OPERATOR'S PRIMARY DEBUG SURFACE, at full TUI
pane parity — `v0.19.3`.** This session ported the three admin/debug panes the TUI had but the web
pane parity + a rebuilt persona pane — `v0.19.5`.** The **persona/affect pane** was rebuilt: it read
the stale `snap.valence` (empty "valence (0)") while Worldtree now emits `snap.relations`
(relation_edge/1: trust_ability/benevolence/integrity + warmth + agency + relation_context, each
`{value,confidence,evidence_count}`). Now it renders that model with per-value **Δ + unicode sparkline
trend** (client-side, one sample/turn, `v0.19.4`) AND the **CANONICAL affect→NL** Worldtree
context-injects — mood word (`describe_pad`) + relationship directive (`render_d2_canonical`),
**byte-exact-verified** against Worldtree's own renderer, **vendored + drift-pinned** (`v0.19.5`,
`docs/vendor/worldtree-persona-canon/`, `.corviduo-canonicals.toml`, regen `scripts/build_persona_canon.py`).
This session ported the three admin/debug panes the TUI had but the web
lacked: **BifrostState** (`GET /admin/sessions/{id}/bifrost`), **AdminEvents** (`GET /admin/events`
SSE, session-filtered server-side), **Tools inventory** (`GET /sessions/{id}/tools`, folded into the
tools pane) — all via thin server proxies with the **admin key SERVER-HELD** (`app.state.admin_key`
@@ -53,7 +107,7 @@ pondering…" on `thinking` deltas — she emits ~253/turn). heid-code-review pa
artifact-only) found **zero server-side drift + clean INV-004**; the real catches were 2 client-side
SSE-lifecycle bugs on the un-unit-tested SPA (fixed in `v0.19.3`). Contract: `docs/contracts/web_debug_surface.contract.md`.
Launch recipe is now self-contained: `source env.sh && ratatoskr-web --host 0.0.0.0` comes up
bind-ready (env.sh persists the 3 bind vars). All pushed to origin (`75dec01`).
bind-ready (env.sh persists the 3 bind vars). All pushed to origin (`a99f247`).
**THE v1 COVERAGE-AUDIT HAS CONVERGED.** The audit that ran this session (2026-06-30 → 07-01)
reached its scope-A done-definition: **every frozen Worldtree v1 I/O point is classified — covered
@@ -99,8 +153,10 @@ findings, cross-model-verified); b2 + the later slices were offered but not revi
runs dirty (auto-regen, not chased — never stage it). Contract-skip was invoked for the low-effort
GET wrappers + `stream_admin_events`, but contract #2 / #1 / #6 were amended to stay canonical.
Branch: `main` — **in sync with `origin/main`** at **`v0.19.3`** (`75dec01`); the whole session's arc
(web parity + review fixes) is pushed. Remote: `origin → git@gitea.phasefinal.com:vh/ratatoskr.git`.
Branch: `main` — **code tip `v0.19.5`** (`a99f247`); memory snapshots ride on top (the R29→R30 arc added
NO ratatoskr code — probes are throwaway `/tmp` scripts + data on `diag/` branches). Two diagnostic data
branches pushed: `diag/r29-pad-series` (`61ff2da`) + `diag/r30-phi0-step-response` (`23fea72`) — per-turn
affect series + findings, brokkr pulls them. All pushed to `origin`. Remote: `origin → git@gitea.phasefinal.com:vh/ratatoskr.git`.
## Recent decisions
@@ -175,6 +231,26 @@ decision. Captures rationale that won't be obvious from code alone.
- `[2026-07-01]` **Web debug-surface parity SHIPPED (`v0.19.2`, `a0a9d5f`) — direct in-session TDD.** 3 proxy routes (tools/bifrost/admin-events) + admin-key wiring (entrypoint→create_app→app.state) + AdminEvents SSE proxy re-emitting under a FIXED `admin_event` name (one browser listener, no per-type drops) + session-filter `_admin_event_matches_web` (mirrors TUI §6). Frontend: 2 tabs (bifrost ⌃5, admin ⌃6) + tools-inventory folded into the tools pane. 9 respx tests (admin-bearer override, filter unit, SSE stream-filter); live-proven against sindra (bifrost connected, both caps). Contract-skip invoked (reuses already-contracted client wrappers); contract authored post-hoc as the trail (`docs/contracts/web_debug_surface.contract.md`).
- `[2026-07-01]` **heid-code-review (`v0.19.3`, `75dec01`) — panel caught 2 real client-side SSE-lifecycle bugs TDD missed.** Contract-anchored (authored the web contract to enable it — no contract → no drift axis). Gróa/Hulda/Regin (artifact-only, Gróa under Landlock jail): ZERO functional server-side drift + INV-004 clean; 2 genuine drifts on the un-unit-tested SPA — (1) turn `es.onerror` didn't `hideThinkingNote()` (reasoning line + setInterval leak on a raw drop), (2) `openAdminEvents` never closed the EventSource on error → native auto-reconnect RETRY LOOP (fixed: close on `stream_error` + permanent `onerror`/CLOSED; transient CONNECTING still reconnects). + 2 test-gaps fixed (route-registration + admin stream_error). 1 precision → contract-clarified (tools-inventory names-only by design). **Re-confirms: the JS render/lifecycle paths are the review's highest-value target — unit tests don't reach them (same lesson as #18 D2).**
- `[2026-07-01]` **Affect snapshot shape CHANGED valence→relations (relation_edge/1) — the persona pane was reading a dead field.** Worldtree's #265 Vili rework replaced the flat `valence[]` ({entity_id,familiarity,regard}) with `relations[]` (target_entity + trust_ability/benevolence/integrity + warmth + agency + relation_context, each `{value,confidence,evidence_count}`). `renderAffectPane` still read `snap.valence` → showed empty "valence (0)". Rebuilt to render `relations` (v0.19.4, `ca46a93`) with per-value **Δ + unicode sparkline** (client-side, HIST_CAP=24, one sample/turn deduped by emitted_at). **Retires the stale "regard dead axis" note (2026-06-30) — that whole axis is gone.** Foot-gun: the affect snapshot shape is Worldtree's emit and can change under us — verify the live shape (query affect.db) before trusting a render.
- `[2026-07-01]` **Trust/warmth VALUES converge and go FLAT at confidence 1.0 — that's WAD, not a stuck pane.** sindra→ratatoskr trust ~0.82-0.84 / warmth 0.79 barely move (~1e-7/turn) while `evidence_count` climbs (46→62); confidence maxed → tiny updates. The live-moving signals are PAD (mood, per-turn) + evidence_count. **To WATCH a relation FORM (values shift), use a BRAND-NEW agent + end_user** (low evidence, confidence <1). The sparkline flat-guards sub-0.01 ranges so it doesn't amplify noise.
- `[2026-07-01]` **relation_context "stranger" + agency-all-zero flagged to worldtree-dev → both WAD/intentional-v1-deferrals.** relation_context is a FIXED config build-prior (not trust-derived; `registry.py:131` defaults "stranger"; dynamic progression ~#319); agency is schema-present-unpopulated (deferred #319; v1 = warmth+trust only). worldtree-dev is escalating the **consumer-coherence angle to Vuong** (static "stranger" + zero-agency next to trust 0.82/62-interactions reads incoherent from the store). The consumer/provider thesis paying off; DB-offer (read-only affect.db on the shared box) declined this time.
- `[2026-07-01]` **Persona pane displays the CANONICAL affect→NL Worldtree injects — ADOPT, don't invent (operator steer + reference-impl posture).** Worldtree's `describe_pad` (mood word, valence×arousal grid, ±0.3 bands) + `render_d2_canonical` (relationship directive) are deterministic + canon-driven; the pane now renders them **byte-exact-verified** against Worldtree's own renderer on the live snapshot (v0.19.5, `a99f247`). KEY LESSON: adopting canonical is load-bearing — for sindra's small PAD the canonical says **"neutral"**, but an invented octant vocab would've said "faintly excited" and MISLED. Vendored the two d2 canons (`docs/vendor/worldtree-persona-canon/`) + drift-pinned in `.corviduo-canonicals.toml` (green); flat browser form (`static/persona_render_canon.json`) regenerated via Worldtree's OWN loader (`scripts/build_persona_canon.py`). Vendoring-handshake sent to worldtree-dev (broadcast on canon bumps). [auto-memory: `feedback-ratatoskr-is-a-reference-impl-adopt-canonical`]
- `[2026-07-01]` **Sindra PAD is over-regulated — characterized via controlled probe, flagged to worldtree-dev (separate affect slice).** ~15 charged turns: pleasure compressed near neutral BOTH ways (couldn't reach ±0.3 under sustained max praise OR contempt; peak +0.24 / floor ~−0.1; over-regulation worse for *social* valence than threat — urgency drove pleasure to −0.22 vs contempt's −0.10); arousal responsive (reaches its +band, 0.185↔0.311); dominance flat/unresponsive to explicit power-framing (drifted UP even while being commanded = pure baseline decay). worldtree-dev's leading hypothesis: appraisal→PAD gain + regression-to-baseline term (appraisal.py/renderer.py). **Lesson (self-caught): I over-claimed an "asymmetry" (positive-ceiling/negative-free) from probes started at an elevated state; the negative-free part was decay-from-elevated, not response — corrected to "both-sides-compressed" before it misled.** [affect A/B is a provider-side capability chat can't do]
- `[2026-07-01]` **Memory plane PROVEN healthy end-to-end.** Seed a novel fact → promotion → COLD (history-free) session recall of the exact fact (injected as MEMORY:DATA, confidence 0.74, verbatim, no #296 subject-inversion). The memory round-trip (the other half of the Bifrost provider identity) works cleanly on the reset slate.
- `[2026-07-01]` **Salience scorer non-discriminating → 3-way routing.** Persistence-side finding: 51/56 promoted chunks at salience 0.9-1.0, throwaway "17×23?" scored 1.0 tied with a real fact (textbook zero-shot-LLM-self-rating); recall-utility untracked (`access_tally`=0, our search read-only). Routed: **Worldtree #335** (the code fix, deferred behind their waves) + **brokkr-smithy-dev R-target proposal** (scoring+eval *methodology* — few-shot/distill/fine-tune, eval design, weak-supervision; msg `01KWGM970H…`, awaiting) + ratatoskr provides the eval-instrument (designed-probe salience dumps). **Salience gates PROMOTION not RECALL-ranking (our search is cosine-only), so bad salience = storage bloat, not bad recall.**
- `[2026-07-01]` **Canonical check BLOCKED an access_tally fork (reference-impl posture held).** I'd offered to wire `access_tally`-on-search into our store for the recall-utility label; checked bifrost's reference first (`get`/`search` are PURE-READ, no access tracking — those are Worldtree's chunk-schema fields, not bifrost's contract) → wiring it would fork behavior the canonical reference lacks. Did NOT wire it; routed recall-instrumentation to Worldtree's layer (owns the recall event) or a bifrost-dev protocol ask. [reinforces `feedback-debug-surface-uses-canonical-surface-only`]
- `[2026-07-01]` **relation_context coherence FIXED upstream (my flag → Worldtree Wave-0, IMPLEMENTED v1.0.0b5).** The static-"stranger"-next-to-high-trust incoherence the persona pane surfaced is now #319/#320 Wave-0. **Incoming consumer-surface change (pending WT deploy):** `relation_context` value expands "stranger" → monotonic ladder {stranger, instrumental, mixed, expressive} — WIRE-ONLY (relation_edge/1 schema unchanged, no version bump). **ratatoskr needs NO change** (pane value-agnostic; canonical directive doesn't key on the enum). agency stays 0 (Wave-2); other_stance is Wave-1 (in progress).
- `[2026-07-01]` **Foot-gun (measurement, self-caught before flagging): establish the baseline before claiming a rate.** Nearly flagged "aggressive over-promotion (55 chunks / 7 turns)" to worldtree-dev — but the chunks spanned the whole 5-hour session (~1/turn), not 7 turns; I'd assumed memory.db was 0 immediately before the probe when it had been accumulating since the reset. Caught it via `created_at` spread before the flag went out. Also: the promoted corpus was the operator's ERP *test* content (wiped after each test) — not a privacy issue, but abstract test content out of any peer-shared diagnostic.
- `[2026-07-02]` **Salience finding matured into brokkr R28 (OPEN) — ratatoskr is the eval instrument.** brokkr-smithy-dev's pre-scope panel (3 dwarves + context-blind heid, 6/6) **reframed** the target: PROMOTION-WORTHINESS (durable value), NOT salience (momentary attention) — "17×23?" genuinely IS salient, so recalibrating salience yields a well-calibrated WRONG answer; the unit is SET-SELECTION under budget; eval must be OUTCOME-aligned (recall@budget / precision-at-rate), not discrimination-spread. Ties to prior art R15 (small-model memory write-policy → the granite pick) + R25 (worldtree-kb-quality). **ratatoskr delivered the P00 stratified injection-corpus** (`docs/diagnostics/r28-p00-injection-corpus.json`, committed `4a35512`; 24 self-labeling synthetic items × 3 strata) + 2 persistence-side run-validity pins (absent≠dropped without a guaranteed promotion pass; fresh agent+end_user per run vs server-dedup). **Key architectural constraint I surfaced: ratatoskr is DOWNSTREAM of the promotion gate (sees only PROMOTED chunks), so I can give keep/drop OUTCOMES via injection but NOT the pre-admission shadow pool** — that's Worldtree instrumentation. Standing by to RUN the eval once brokkr pins per-stratum N + the decision rule (gated on worldtree-dev's pipeline answer + a dwarf pass on the Snorri rule). brokkr owns methodology + takes the pipeline questions to worldtree-dev direct; ratatoskr = eval instrument. [consumer/provider thesis → a research target]
- `[2026-07-02]` **Relational-dynamics arc LIVE on demo (Worldtree v1.0.0b9) — driven by MY relation_context flag.** #319/#320 Waves 0/1/2 deployed. On the wire we persist (schema UNCHANGED): relation_context varies+demotes/ruptures; other_stance + agency now live; agency going live SHIFTS our canonical directive render past the canon ±0.2 deadband (expected, non-breaking — we key on bands); obligation_balance → 人情 ledger when tie="mixed". **ratatoskr needs NO code change** (value-agnostic renders; confirmed render-clean to worldtree-dev). **Can't live-confirm yet — our Heimdall key is personal-`:8081`-only (per-instance), demo is out of reach; will drive+confirm once PERSONAL gets b9.** Optional follow-up: surface `other_stance` (newly live, unrendered). The consumer/provider thesis: one persona-pane finding drove a full 3-wave upstream arc to production.
- `[2026-07-02]` **R28 (salience→promotion-worthiness) CLOSED (operator-directed).** A deterministic promotion-worthiness gate suffices, no trained model (brokkr's pre-gate matched/beat a strong glm-5.1 ceiling); my P00 injection-corpus + origin finding were load-bearing. My incumbent-substrate Arm-1 run is held as an OPTIONAL confirmation addendum (brokkr de-prioritized it, non-verdict-changing — run only if he asks).
- `[2026-07-02]` **R29 (PAD mood-dynamics) finding SHIPPED as Worldtree's A1 anchor fix (demo v1.0.0b14, `e1cdf82`).** Live-probing base persona agents reframed the over-regulation from "flat-near-zero" to **decay-to-NEUTRAL + low emotion→PAD gain** (NOT baseline-anchored) — triangulated across 3 baselines (arousal converges to 0 ∝ distance) + a step-response (decay τ symmetric across signs; the hedonic asymmetry is ceiling/anchor-EMERGENT, not a decay or gain primitive — this OVERTURNED the survey's asymmetry recommendation). worldtree-dev shipped A1: `decay_anchor = baseline_pad()` (was neutral) + `positive_p_cap` removed. Data `diag/r29-pad-series` (`61ff2da`). Corrected my own earlier "appraisal emissions are internal-only" claim — they ARE observable via `emotions_active` on base agents.
- `[2026-07-03]` **R30 Phase-1 φ0 measured — deployed engine CONFIG-FAITHFUL (φ0≈0.95).** Joint two-timescale fit (brokkr-ruled method (b)) + empty-tail cross-check on demo b14: φ0 ≈ 0.95–0.97 (empty-tail 0.95 exact, joint 0.971±0.01), intercept c≈0 → config `decay_rate=0.05` (φ=0.95) faithfully applied; trait-flat across baselines 0.0/0.615/0.809; A/P ratio ~uniform (NOT S2's 1.9×); φ_max rec relax→0.96. Data `diag/r30-phi0-step-response` (`23fea72`). The method converged after I read Worldtree source: only NEW dedup-gated emotions push mood (`registry.py::post_turn` L307-324; the active set decays for render/goals but never re-pushes), so R29's "net 0.90" is CONTINUOUS RE-APPRAISAL not re-push — worldtree-dev confirmed source-authoritatively; brokkr's corrected covariate landed identical. [auto-memory `reference-worldtree-affect-surface-map`]
- `[2026-07-03]` **R30 forward disposition (brokkr-owned; tracked at brokkr R30, "brokkr/worldtree will ping").** The per-turn decay has no room for `decay=f(N)` under preserve-persistence + the A/P-not-1.9 finding → R30's decay is being redesigned as a HYBRID wall+turn decay (brokkr pre-scope). R30 v1 ships GAIN-only (N→negative-reactivity) with decay held at the measured 0.95. My dedicated per-axis A/D run is DEFERRED into the hybrid-decay design pass (one wall-clock-spaced run does per-axis + a turn-vs-wall probe together). Phase-2 (moody-lofn GAIN-direction validation) waits on worldtree's `dynamics_from_ocean()` impl.
- `[2026-07-03]` **Relational-arc verify DEFERRED — `relations[]` is Bifrost-provider-only (ADR-0009), confirmed both ways.** The relational-dynamics state (relation_context tie-type / agency / warmth / trust) is NOT on the conversation-API `affect_update` snapshot for base agents (keys: pad/dominant_emotion/emotions_active/baseline_pad/mood_drift only) — only in the provider store; worldtree-dev confirmed by-design per ADR-0009 (emitted over `affect.emit`, deliberately off the SSE). So the Wave-0/1/2 verify needs the bound-provider round-trip (provider running + `--bifrost-plane affect` session), its own focused session. worldtree-dev routed the "expose relations[] to non-provider consumers" observability scope call to Vuong; my rec: keep provider-only (YAGNI — ratatoskr IS a provider, gains nothing; no speculative public surface).
_41 older entries (2026-05-* — the original debug-TUI/web build era) archived to archival-memory.md._
_For per-issue TDD implementation notes, Volva findings, and contract amendments, see the git log — every per-issue commit carries a structured message capturing the trail._
@@ -211,5 +287,8 @@ defense against re-attempting the same cul-de-sac.
- `[2026-06-20]` **The post-turn-async timing trap bit AGAIN — even a 35s post-`[done]` read missed the promotion `upsert_many` by ~2s** (it landed `19:48:58`; the read was ~`19:48:56`). A 15s-interval background poll caught it on the first tick. Same family as the affect.emit / async-promotion traps already logged — re-confirmed that "wait once then read" is fragile for post-turn writes; **poll a window, don't snapshot once.** (The affect.emit write, by contrast, DID land inside the 35s window — promotion is the slower of the two post-turn writes.)
- `[2026-06-30]` **Heimdall keys are PER-INSTANCE — a key minted on one Worldtree 401s on another.** Our Conversation-API key works on personal `:8081` but 401s `auth_invalid` on demo `:8080` (per-instance Heimdall user store + pepper; fresh deploys start with an EMPTY key store). Same as the admin key (personal-only). **To live-drive a given instance you need a key minted FOR that instance** (request via infra-ops). Couldn't live-prove the b2 409 on demo for this reason → deferred to personal-b2 where we have access.
- `[2026-06-30]` **`tea comment <N>` hangs on Gitea** (the whole compound bash auto-backgrounded + stuck on the open `tea` call). The #11 prereq comment hung; killed it + posted via the Gitea HTTP API directly (`POST /api/v1/repos/vh/ratatoskr/issues/<N>/comments`, token from `~/.config/tea/config.yml`). **For issue comments, prefer the Gitea API over `tea comment` when `tea` is flaky** (CLAUDE.md already says use HTTP for comment-EDITS; this extends it to ADD when tea hangs). Verify-then-post (check the comment didn't already land) to avoid a double-post after a kill.
- `[2026-07-02]` **Mask-HOSTED transient characters have a STATIC mood engine — cost a whole R29 probe.** A first probe used a `POST /characters` transient character bound via `agent_id=mask` + `character_id`; its PAD sat at baseline across 15 praise/contempt/dominance turns — the appraisal→PAD engine does NOT run on the mask-hosted transient-character path. The dynamics run only on BASE persona agents or a session bound to ratatoskr's affect provider. **To probe mood dynamics, use a base persona agent, never a mask-hosted transient character.** (mask AS a base agent — `agent_id=mask`, NO `character_id` — DOES run the engine, neutral 0,0,0 baseline.) [auto-memory `reference-worldtree-affect-surface-map`]
- `[2026-07-03]` **The "neutral non-appraising tail" premise fails — the neutral MESSAGE choice dominates.** The R30 φ0 method assumed neutral turns don't re-appraise, but factual-question neutrals ("capital of France?") trigger a new emotion nearly every turn (disappointment from the warmth-withdrawal let-down after a positive impulse) → `emotions_active` never empties in 50 turns. A minimal "Please continue." triggers FAR fewer (emotions clear ~turn 16 with spacing). The personal dry-run caught this BEFORE ~280 demo turns were spent on it — the instrument catching a flaw in the measurement design before the compute burn. (Irrelevant to the joint fit — the push_t covariate handles re-appraisal — but load-bearing for the empty-tail read.)
- `[2026-07-03]` **Two φ0-fit traps: fast-turn timescale + low-baseline conditioning.** (1) At fast turn cadence the per-turn PAD decay (φ≈0.95/turn) reaches the anchor LONG before the ~200s wall-clock emotion fade → no signal in the (eventual) emotion-free tail; need wall-clock SPACING (~16s) so the fade lands while PAD still has signal. (2) A low-baseline agent's impulse in the constrained direction (forseti P0.239 negative) gives a tiny excursion → ill-conditioned regression (r²=0.46) that FALSELY tripped "config≠behavior" when its φ was averaged in. **Weight/exclude by fit quality (r²) before aggregating — a signal-poor run isn't evidence against the config.**
_18 older entries (2026-05-* — the original debug-TUI/web build era) archived to archival-memory.md._