|
|
|
@@ -1,6 +1,6 @@
|
|
|
|
|
# Persistent memory — eshpfi-management
|
|
|
|
|
|
|
|
|
|
_Last updated: 2026-09-28 ~0850 PT (Zigbee2MQTT live on esh-docker-vm + HA-MQTT macvlan route fixed; Blender on fv-ml1 GPU 3 (MCP per working session, blender-run batch for draupnir); semif-serve 0.1.4; Worldtree reward config on demo+personal and instance-configs pushed; restic repository-file move verified all-green 2026-09-28.)_
|
|
|
|
|
_Last updated: 2026-09-29 ~2345 PT (Blender extensions live on both paths + blender-run --cpu; Bonsai spike + MMVQ follow-up + weights acquired; phasefinal.com cleanup live; Worldtree U10 backfill done on demo+personal, U11a prepped (staged, not flipped); NEXT: build the U8 gate-batch wrapper for infra-hermes, Prime ruled 2340.)_
|
|
|
|
|
|
|
|
|
|
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
|
|
|
|
|
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
|
|
|
|
@@ -115,7 +115,18 @@ no longer deployed sidecars here. See Recent decisions.)
|
|
|
|
|
|
|
|
|
|
## Current state / in-flight
|
|
|
|
|
|
|
|
|
|
_As of 2026-09-27 ~0900 PT._
|
|
|
|
|
_As of 2026-09-29 ~2345 PT._
|
|
|
|
|
|
|
|
|
|
### Worldtree memory-split: U10 done, U11a prepped, U8 window wrapper IN PROGRESS (2026-09-29)
|
|
|
|
|
|
|
|
|
|
- U10 legacy backfill DONE: demo 5 filed, personal 797 filed (mimir's 377 `dropped:too_large` are ONE interests record at its ceiling, re-fileable after consolidation; that is the U11 checkpoint's call). Pinned: Prime ruled "leave pinned".
|
|
|
|
|
- U11a (`memory.legacy.mode` read_only): demo runs 78a509b90267 (U11a code) and boots `live`, proven by the gauge `worldtree_memory_legacy_mode{kind="live"} 1.0` on `127.0.0.1:8080/metrics`. The INFO boot line is not emitted in this deployment's logs. The flip config is STAGED and inert at `/opt/worldtree/config/defaults.yaml.staged-u11a-read_only` on corviduo-dev. **The flip waits on Prime's go**, relayed by worldtree-dev (thread `01M3RGGYR5EHKZVGM7ABM5RMTF`). Apply = back up, copy over `defaults.yaml`, `docker restart worldtree-worldtree-api-1`, then check that the gauge reads read_only=1. Open Prime ruling: skaldsong/wizard-v2, personal's only Tier-3 client, has unknown record-profile adoption.
|
|
|
|
|
- **IN PROGRESS, Prime's ruling of 2026-09-29 ~2340: "infra-hermes runs the batches, go ahead".** Pending: build a wrapper (e.g. `scripts/wt-memory-gate-batch`) around worldtree-dev's U8 harness, then brief infra-hermes to run one batch per day through the read_only window, with infra-ops on escalation. Harness facts (worldtree-dev thread `01M3RG8RDFSFGDK6BDFTE6ZCJE`):
|
|
|
|
|
- nh3-dev `~/development/Worldtree`, pinned to demo's deployed SHA. The main checkout was at 78a509b9, the same as demo, at 2340. It is worldtree-dev's working tree: NEVER checkout or reset it; use a separate `git worktree` if the SHAs diverge.
|
|
|
|
|
- Command: `source .venv/bin/activate && source env.sh && python -m core.memory_acceptance.live --runs 3 --gate --legacy-mode read_only --label window`.
|
|
|
|
|
- Output lands in `docs/eval/memory_acceptance_runs/<UTC>Z/`. The verdict is in `trace1.md`: `Aggregate label: **gate**, verdict: **PASS|FAIL**`. Report the Semantic axis section when the judge was not controlled, but do not fail on it.
|
|
|
|
|
- About 10-15 min per batch; needs the LiteLLM gateway; never two at once (flock); FAIL goes to worldtree-dev the same day with the run stamp, never a retry; the wrapper must NOT git-commit (worldtree-dev commits the run dirs).
|
|
|
|
|
- Ask worldtree-dev whether daily batches start now or at the flip.
|
|
|
|
|
|
|
|
|
|
### nh3-pve + nh3-ml1: post-visit, all live (2026-09-25/26)
|
|
|
|
|
|
|
|
|
@@ -270,17 +281,17 @@ _As of 2026-09-27 ~0900 PT._
|
|
|
|
|
|
|
|
|
|
### Live threads
|
|
|
|
|
|
|
|
|
|
- git: `main` is ahead of origin by 19 at `6f0c480`, plus this snapshot. Pushing
|
|
|
|
|
is Prime's call.
|
|
|
|
|
- git: `main` == origin at `274b817` (pushed 2026-09-29 on Prime's word), plus this snapshot. `graphify-out/GRAPH_REPORT.md` stays modified
|
|
|
|
|
and uncommitted on purpose: it is auto-regenerated.
|
|
|
|
|
- Booth submit-all fix (Prime's report) is LIVE since 2026-09-27 1705, via booth-dev (booth `50bfc7b`).
|
|
|
|
|
Pushing it is booth's call, per Prime; it is not ours.
|
|
|
|
|
- ESH has a single outside route (esh-scale on esh-pve). Noted, untracked.
|
|
|
|
|
|
|
|
|
|
## Recent decisions
|
|
|
|
|
|
|
|
|
|
- `[2026-09-29]` **Worldtree U11a (legacy memory → read_only) is PREPPED, not flipped.** A staged, inert `/opt/worldtree/config/defaults.yaml.staged-u11a-read_only` sits on corviduo-dev. Applying it needs Prime's go AND demo on 48bdf235 or later (demo ran 1b8746e6, which predates U11a). Apply = back up, copy over `defaults.yaml`, `docker restart`, then read the boot line `memory cutover: legacy plane read_only (writer on, reader on)`. **Once the window opens, infra-hermes runs one U8 batch per day and I own the wrapper and the escalations (proposed, worldtree-dev agreed):** on nh3-dev in `~/development/Worldtree`, at demo's deployed SHA, `python -m core.memory_acceptance.live --runs 3 --gate --legacy-mode read_only --label window`. The verdict is in `trace1.md`; FAIL goes to worldtree-dev the same day, never a retry; no git commit; never two at once. The wrapper is NOT built yet, deliberately: it waits for the go.
|
|
|
|
|
- `[2026-09-28]` **Worldtree U10 legacy-memory backfill: demo COMMITTED (Prime, direct go 2350).** Dry run then commit, both as `-u worldtree`: 5 filed (lofn 4, mimir 1), 16 `unslotted_pre_364` skipped, 0 aborted. **The first commit exited 1 because demo's hand-managed `model_roles.yaml` lacked the `memory_tagger` role**, so I synced it verbatim from the image's config-defaults (backup `.bak-pre-memory-tagger`). **2026-09-29 0854 presync for personal (Prime: "complete u10"):** personal got `memory_tagger` and the U9 `admin.memory.forget` policy delta. **Demo had also been missing that U9 policy delta since 09-24**, which left readonly-admin's `admin.*` rule with no exclusion for the destructive forget. Fixed, and demo's api was restarted to load it. **Personal COMMITTED 2026-09-29 1007** (b191 e2f74706, a snapshot-copy build; Prime's standing "complete u10" via worldtree-dev's go): 797 filed (mimir 795, lofn 2), mimir `dropped:too_large` 377, 1 unroutable held. The first personal dry run (b190) had crashed on 45 mimir rows with no vector (HNSW frozen since 08-06, #417, accepted by Prime until U11). Pinned (446e5807; Prime: "leave pinned") keeps its legacy chroma in the container's WRITABLE LAYER (`/app/agents/*/memory/.chroma`, 7 agents, 196K each): a recreate erases them.
|
|
|
|
|
- `[2026-09-28]` **Bonsai ternary spike on fv-ml1 GPU 3 (Prime via brokkr): at the 275 W cap, PQ2_0 is 1.93x Q4_K_XL at N=1 but only 1.06x at N=8 (PTQ1_0 1.24x), recovering to ~1.37x at N=16.** Positive control passed on tg128 (+1.1%). nvidia-smi's sw_power_cap flag never fires on this card, so "at cap" is judged from board draw. Follow-up (brokkr): moving PQ2_0 onto MMQ from batch 6, like Q4_K (mmvq.cu, `ne11 <= 5`), lifts N=8 to 1.21x, so the dip was partly a kernel threshold. The residual gap is not power. Runs are in `fv-ml1:/tank/spikes/bonsai-2026-09-28/runs{,-mmvq5}/` and on the Booth. The build image was removed; `Dockerfile.build` recreates it. **Acquired for keeps (Prime, 1540):** GGUF PQ2_0/PTQ1_0/mmproj-Q8_0 + the fork source pin in `/tank/aimodels/llm/prism-ml_Ternary-Bonsai-2-27B-gguf/` (runtime/), the MLX 2-bit pack in `/tank/aimodels/mlx/`, 34/34 hash-verified. Weights single copy (no snapshots, /tank not in restic); the fork pin also at `/mnt/smithy/runtime-pins/prism-llama.cpp-87268f77/` (brokkr).
|
|
|
|
|
- `[2026-09-29]` **Worldtree U11a prepped, not flipped: a staged demo config plus the agreed U8 window plan (infra-hermes runs the batches).** → `persistent-memory.d/2026-09-29-worldtree-u11a-prepped.md`
|
|
|
|
|
- `[2026-09-28]` **Worldtree U10 backfill done on demo (5) and personal (797). model_roles drift needed a memory_tagger sync first; mimir had missing vectors.** → `persistent-memory.d/2026-09-28-worldtree-u10-backfill.md`
|
|
|
|
|
- `[2026-09-28]` **Bonsai ternary vs Q4_K_XL at concurrency on the 275 W card (1.93x at N=1 falls to 1.06x at N=8, 1.21x with the MMVQ fix); the weights are acquired.** → `persistent-memory.d/2026-09-28-bonsai-ternary-spike.md`
|
|
|
|
|
- `[2026-09-28]` **blender-run gained `--cpu` (no GPU attached) and a fixed hostname `fv-ml1-blender` (draupnir).** The design stage never renders, so it stays off GPU 3.
|
|
|
|
|
- `[2026-09-28]` **Blender extensions live in a read-only System repo built from a sha256 lock, enabled by a hook, opt-in for blender-run (`--extensions`).** SurfacePsycho's eval() is patched to literal_eval (a proven safe-mode escape). → `stacks/blender/README.md` § Extensions
|
|
|
|
|
- `[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** (SVOS v2.1.12: `propose_decision` gained `seat_up`, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling `zellij-fleet@Claude` at boot remains Prime's call.
|
|
|
|
@@ -481,14 +492,8 @@ _As of 2026-09-27 ~0900 PT._
|
|
|
|
|
|
|
|
|
|
- `[2026-09-15]` **Hermes bearer rotation hold released — svos-dev split their HS256 signing key off the shared value (svos `7165272`)** → `persistent-memory.d/2026-09-15-hermes-bearer-rotation-hold-released-svos-dev-split-their.md`
|
|
|
|
|
|
|
|
|
|
- `[2026-09-14]` **fv-ml1 rebalance: cyberprev→`sec` (mog-sec retired), NEW gen-small A3B seat, all sec/gen/char at native 262K in-band, coder reclaimed, seat catalog + bench shipped.** cyberprev = hotdogs cyber-SFT (name-repaired past a tripled-prefix unsloth export bug, house NVFP4 quant); gen-small = llmfan46 Qwen3.6-35B-A3B Heretic (already on disk), MTP 69.6%. Serial depth-tested all seats clean (0 OOM); warm tok/s 62.7-337.3. Commits 1418edb→dfa91a8. → `persistent-memory.d/2026-09-14-fv-seat-rebalance-gen-small.md`
|
|
|
|
|
|
|
|
|
|
- `[2026-09-14]` fv-ml1 all-night seat reorg — MTP k=3 on gen-large (+52%@conc1), gen consolidated onto flash-next (27B dense retired, 38 GB freed), char-rp restored to MeroMero-v2-31B, Sentinel-R3 served + dflash cutover (beat MTP 2.40 vs 2.18). ✅ gen-large RESOLVED 2026-09-14 — orcarouter serving: PLE bf16→FP8 convert + `ple_embedding_dtype` + `layer_types` rename; NO source build needed. → `persistent-memory.d/2026-09-14-fv-seat-reorg-and-orca-blocker.md`
|
|
|
|
|
|
|
|
|
|
- `[2026-09-13]` **STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint.** → `persistent-memory.d/2026-09-13-standing-policy-operator-cap-gpu-power.md`
|
|
|
|
|
|
|
|
|
|
- `[2026-09-13]` ⭐⭐ **Qwen3.8-Flash-Next serving on ONE card with its 51B n-gram table in host RAM — the first seat whose weights do not fit its GPU.** `stacks/flash-next-seat/`, fv-ml1 GPU 2 `:8022`, plus a `gen-large` LiteLLM alias. Measured: 74.36 GiB weights resident, 14.00 GiB KV = 560,654 tokens at the full 262,144 context, 67 GiB host RSS, 75.5/212.3/387.8 tok/s at conc 1/4/8 (⚠ n=1). ⭐ The offload is vLLM **#54371 (UVA, merged 2026-09-09)** which **supersedes the paused #53899** — it has no worker process, so #53899's whole bug family (TP=1 deadlock #53960, `pidfd_getfd`/ptrace gate, stale-output-under-graphs) is designed out; in `v0.29.1rc0`, **not** `v0.29.0`. ⚠ **`text_config.ple_embedding_dtype` is the load-or-fail discriminator** for any community build. ⚠⚠ **`--kv-cache-memory` makes vLLM SKIP MEMORY PROFILING and ignore `--gpu-memory-utilization`** — 16 GiB nearly OOM'd on a 155K prefill with no visible failure; 14 GiB is the measured-safe value and vLLM's own "17.46 GiB to fully utilize" is 3.5 GiB too high. ⚠ MTP is off **pending measurement here, not written off** — the recipe's number is cross-harness and tested k=3 only, while the head is ONE layer run autoregressively, so k=1 is unpublished and may win (`services/flash-next-mtp-bench/`, one `off_A` rep banked before the outage). ⚠ A container once ran `(healthy)` with `PORTS=[]` — verify `docker port`, not the healthcheck. → `persistent-memory.d/2026-09-13-flash-next-seat-and-fv-outage.md`
|
|
|
|
|
|
|
|
|
|
- `[2026-09-11]` **Worldtree memory-split (U6) — PROTOCOL AGREED with worldtree-dev: nobody flips memory.reader.enabled or m** → `persistent-memory.d/2026-09-11-worldtree-memory-split-u6-protocol-agreed.md`
|
|
|
|
|
|
|
|
|
|
- `[2026-09-03]` **Run 3c is STAGED on pfi-gx10 and deliberately NOT launched** — the launch is a 13.3 h commitment and the operator stood this port down once already. Base shards AND the encoded corpus sha256-verified identical to ana-ml2's, so the transformers 5.15.1→5.16.1 / x86-64→aarch64 delta is *measured* inert. ⚠ The encode-cache FILENAME differs by design (`base_model_path` is in the key) — input hash, not output. ⚠ Tripped the `pkill -f` ssh self-match again; the launcher guards on a pidfile because of it. → `persistent-memory.d/2026-09-03-gx10-run3c-staged.md`
|
|
|
|
@@ -499,7 +504,7 @@ _As of 2026-09-27 ~0900 PT._
|
|
|
|
|
|
|
|
|
|
- `[2026-08-19]` **AI-tab Dormant regrouping BELAYED by the operator** — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong… → `persistent-memory.d/2026-08-19-ai-tab-dormant-regrouping-belayed-by-the-operator.md`
|
|
|
|
|
|
|
|
|
|
_109 older entries archived to archival-memory.md._
|
|
|
|
|
_112 older entries archived to archival-memory.md._
|
|
|
|
|
|
|
|
|
|
## Tried and abandoned
|
|
|
|
|
|
|
|
|
|