# Persistent memory — eshpfi-management _Last updated: 2026-10-01 ~0420 PT (Parakeet seat → unified-en under NeMo LIVE + audited; gen-small util 0.36 → .env 0.33; leftover bench weights + spike dirs deleted; Scriberr no-upstream; eshpfi + worldtree-instance-configs pushed. Prior: U11a off, SemIf → intern-decision, Scriberr GPU 3 + patches.)_ > **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its > `Written:` stamp is under **8 hours** old, read it (it carries the in-flight > handoff from the previous session), then delete it. Older than 8 hours: > stale — delete it unread. > > _(Raised from 1 h to 8 h by operator 2026-09-13 — a one-hour window deleted the > handoff unread across any overnight gap, which is the exact case it exists for. > 8 h also matches the global CLAUDE.md and the `/snapshot` skill default.)_ ## Repo purpose - **2026-09-10 Beszel fleet wiring:** all seven requested hosts plus existing corviduo-dev report up. `/tank` and other data filesystems now have real usage metrics; NVIDIA telemetry covers ana-ml2 and irv-ml1. Thirty alerts deliver to **infra-ops**, explicitly chosen by operator; Miranda routing is deferred. A real low-threshold disk alert reached althing, then the threshold was restored to 85%/5 min. Homepage has one native overview widget (reachability counts, not degraded health). Dedicated superuser approved and stored in Vaultwarden. See `persistent-memory.d/2026-09-10-beszel-fleet-wiring.md` and `stacks/beszel/README.md`. Reference workspace for PFI infrastructure: server inventory, canonical Docker Compose stacks, ops playbooks, and conventions. Authoritative copies of compose files live on the servers under `/opt/docker/compose//`; this repo mirrors them for version control, editing, planning, and CI-driven deploys. **It was originally spun up to handle the fleet backups** — keep that lens when triaging backup/storage issues. ## Tools and conventions Sister repos (separate gitea repos, deployed by playbooks here): | Repo | Role | CI status | |---|---|---| | `vh/task-board` | MCP + web dashboard for assistant task state (port 7878) — **[2026-09-24] MOTHBALLED** by Prime (superseded by the High Seat + ledger); container removed on ana-docker, data/image/compose kept (`e6da607`) | push-to-main → CI deploys (2026-04-29) | | `vh/vor` | Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) | | `vh/nevermore` | Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) | | `vh/asset-engine` | Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) | | `vh/althing` | Lean trusted inter-agent message bus — **v3.0.0 "the post office" as of 2026-08-28 (U9b flag day, one-way, no rollback)**: ONE container on nh3-dev at `http://10.100.50.40:8390` is the only stateful component; `althing-po-herald` one per box; `althing-listen` one per session; `postbox` is the client. **Every v2 command was DELETED, not deprecated** — `althing-cli`→`postbox`, `althing-wake-listener`→`althing-listen`, `althing-light-monitor`/`althing-receiver` gone. Sessions need BOTH `ALTHING_POST_OFFICE` and `ALTHING_HANDLE`; there is no default address. ⚠ An unreachable post office is an OUTAGE, never an empty inbox. → `persistent-memory.d/2026-08-28-althing-v3-cutover.md` | per-box install (NOT CI-deploy); **nh3-dev** = container host + repo; **nh3-extdev** = system WHEEL at `/opt/uv-tools`, needs its own wheel install (`playbooks/nh3-extdev-althing-v3.yaml`) | | `vh/mead-hall` | Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) | | `vh/skaldsong` | Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) | | `vh/Worldtree` | Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. **gitea-runner builds on ana-docker**; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. | push-to-main → CI build-and-deploy (runner on ana-docker) | | `vh/yt-voice-clipper` | YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → **gitea-webhook auto-deploy** to irv-ml1 (2026-06-03) — see `docs/runbooks/ytvc-autodeploy.md` | | `vh/arbo` | Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) | push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook | | `vh/zonos-gateway` | OpenAI-compatible TTS gateway over stock ZONOS2 (`:8890` irv-ml1); emotion **dials-first** + voice mapping; reached via LiteLLM `ext-tts` alias. **v0.2.1 (2026-07-18): voice-resolved emotion presets** (`resolve_preset(name,voice)`; angry/happy/startled_happy per-voice). 8 voices incl. 4 clones | pushed to gitea (main `8f1885b`/`v0.2.1`); **deployed irv-ml1 tree still NON-git** (hand-updated build context — CI-wire = open follow-up). Spec `docs/EMOTION-DIALS-SPEC.md`; host-managed voices bind-mount (`./voices:/app/voices`, drop wav + restart, no rebuild) | | `vh/soong-lab` | Noonien Soong character-design studio (SPA + /api + WT `/bifrost/tool-call`); **containerized 2026-07-18**, LIVE on corviduo-dev `:8443` (image `vh/soong-lab:latest`). soong-dev owns Dockerfile/compose/workflow; infra-ops owns the host | CI = Gitea Actions build+push+**DEPLOY** on tag/dispatch (fleet recipe: docker:cli + raw buildx, pushes AS vh; **auto-redeploy LIVE 2026-07-18** — runner SSHes corviduo-dev as `deploy`, `compose pull && up -d` from **/opt/soong-lab**, health-gated on /api/version). Manual redeploy `sudo -u deploy bash -c 'cd /opt/soong-lab && docker compose pull && docker compose up -d'`. → `archival-memory.md` (archived 2026-08-16) | | `model-training-forge` (mtf-dev) | Fine-tuning recipe forge; **T1 = E-RP writing LoRA, retargeted qwopus-122B→AEON-27B (2026-07-06)** (SFT→DPO, LitBench-RM reward) | training runs, not a deployed sidecar | (`vh/volva` + Heid were re-architected from systemd daemons to Claude Code session orchestrators 2026-06-08; their nh3-dev `.service` units were removed — no longer deployed sidecars here. See Recent decisions.) - **Two-layer backups** — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see `docs/runbooks/disaster-recovery.md` for the blast-radius matrix. **⚠️ The restic file+DB layer routes through TWO rest-servers** (`rest-server-ana` @ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas; `rest-server-nh3` @ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of `/mnt/backup`. (rest-server-ana recovered 2026-06-20.) - **`pull-hf-repo.yaml`** is the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at `/tank/aimodels/huggingface/`" playbook. Supports `--var repo_type=model|dataset|space`. Replaces ad-hoc `huggingface_hub.snapshot_download` patterns. - **Worldtree admin auth — per-instance.** Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (`key_id 61419c92`) at `ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin` auths against **demo only**. Personal-instance admin (the `~/.config/worldtree/personal-admin-token`, mode 600) POSTs `/admin/keys` (mints per-project keys; takes `user_id`+`label`, **no scope param** — scopes are tier-derived). **On-instance mint recipe (cleaner than DB-manip):** `docker exec worldtree-worldtree-api-1` POST `/admin/keys` with the in-container `WORLDTREE_BOOTSTRAP_ADMIN_KEY`; cleartext once in `.key`=`wt_live_+16hex`. auto-memory `reference_worldtree_demo_key_mint`. - **Per-project user keys against personal Worldtree** (issued 2026-05-19): `skaldsong:79744637`, `skaldsong:7c1dbbbe`, `althing:50d85460`, `mead-hall:a360822d`. Mint via `/admin/keys`, drop value to `/tmp/wt-personal-.key` mode 600, dev collects + shreds (DO NOT cat to chat transcript). - **Skaldsong CD pattern (registry-pull).** vh/skaldsong's CI builds and pushes `gitea.phasefinal.com/vh/skaldsong:` + `:latest`; `playbooks/deploy-skaldsong.yaml` on ana-docker pulls + recreates. SHA-pin only. Prereq: host needs `docker login gitea.phasefinal.com` once. - **gitea internal route for fleet hosts.** gitea is a container on **ana-docker** — git-SSH `10.250.50.70:222`, HTTP `:3000`. Fleet/colo hosts must use this internal route, NOT public `gitea.phasefinal.com` (`38.120.12.44`) — the public path fail2bans the host egress IP. Full gotcha in `docs/orientation.md` → Git/gitea. - **docker-as-root pattern** (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo): `docker run --rm -v :/wt docker:cli sh -c "..."`. docker-group membership is effectively root via bind-mount. **Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass `-e VAR=/abs/path` for any relative-default config dir.** - **`scripts/elway` sudo handling** — elway prompts for the sudo password ONCE via `getpass` before the first `sudo: true` step → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH. - **Per-host SSH identity matters for sudo.** infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: **default `ssh ana-docker` = `lkraven`** (docker-group, NO passwordless sudo); **`ssh infra-ops@ana-docker` HAS NOPASSWD root**. **→ For any sudo op on ana-docker, use `ssh infra-ops@ana-docker`.** `ssh infra-ops@10.100.10.50` (nh3-dev) ALSO NOPASSWD sudo; on **nh3-extdev** infra-ops is sudo-LESS by design (`ssh lkraven@10.100.50.42` is the NOPASSWD path) — **[2026-09-23] measured: infra-ops HAS NOPASSWD sudo on nh3-extdev** (fleet ownership audit; also CLAUDE.md 2026-09-05). **irv-ml1: `ssh irv-ml1` = lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD)** — stage model pulls to `/home`, not root-owned `/worktank`. ## Current state / in-flight _As of 2026-10-01 ~0420 PT._ ### Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01) - **⚠ INCIDENT 04:21 PT 2026-10-01, MITIGATED, root fix in flight:** - `vllm-gen-small`'s EngineCore CUDA-OOM'd when it needed a 394 MiB runtime workspace and GPU 0 had 388 MiB Free. The parakeet seat was parked at its 3,582 MiB window cache. - vLLM grows ~0.8 GB at runtime beyond its preallocation; my audit checked gen-small's boot margin, not its runtime growth. gen-small auto-restarted, healthy at 04:23. - I restarted parakeet-nemo to drop its cache (rest 2,084 MiB). gen-small then answered 3/3 via LiteLLM, and GPU 0 Free is ~1,075 MiB. - **infra-hermes is tasked with nemo-0.1.1:** `empty_cache` after windowed requests, plus a hard memory ceiling so the seat 503s instead of starving gen-small. It also measures gen-small's runtime growth. - **Awaiting Prime:** trim gen-small's KV pin (8 → 7 GiB frees ~1 GiB; 670k → ~586k tokens), or move the seat to GPU 3. - Until fixed, a long transcription can re-grow the seat's cache and starve gen-small. - **LIVE since ~0126 PT 2026-10-01** as `parakeet-nemo` (`stacks/parakeet-nemo`, `local/parakeet-nemo:nemo-0.1.0`, built by infra-hermes) on fv-ml1 GPU 0 :8300, with LiteLLM `ext-stt`/`whisper-1` unchanged. **infra-ops audit PASSED 0137.** - p50 on GPU 0 for 1–3 / 3–8 / 8–20 / 20–60 s: 33 / 36 / 42 / 71 ms, against 187 / 308 / 626 ms for the old seat. - WER: LibriSpeech clean 1.965, other 3.026. - Files longer than 6 min run in 360 s windows. That avoids NeMo's T×T attention mask; a seam can lose a space or a word. - **Rollback:** `docker stop parakeet-nemo && docker start parakeet`. The old container and image are kept. - **GPU 0 is FULL:** - The seat's steady state is **3,582 MiB** (its cached window peak); Free is **385 MiB**. - `vllm-gen-small` runs at util 0.36, and its `.env` holds **0.33** for the next restart (~3 GiB of boot-check margin). Its KV is byte-pinned: 670,142 tokens / 2.56×. - Before restarting any vLLM seat on this card, check that util × 95.6 GiB ≤ measured Free + the seat's own resident memory. Do not trial-boot. A trial-boot sequence took gen-small down for 34 min on 2026-10-01. - Seat invariants (in its README): cast to bf16 AFTER change_attention_model; uvicorn pinned with `--http h11` (httptools 0.8.0 emits `HTTP/1.1 200\x00OK`, which LiteLLM/httpx rejects). - NVIDIA Open Model License accepted for internal use. → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md` ### Worldtree U11 memory cutover (demo + personal) - **Legacy plane OFF since 0115/0120 PT 2026-09-30** (config repo 63cf268; personal /metrics b6fdd81; the repo is pushed). → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md` - **Daily gate batches:** infra-hermes runs `scripts/wt-memory-gate-batch` from 2026-10-01 and copies me on every verdict. The count is **1 of 3** consecutive PASS at off (20260930T090608Z); a FAIL restarts it. - **⚠ U11b STEP 5 IS MINE, triggered by the 3rd consecutive PASS:** 1. Run `docker exec -i python - < scripts/wt-h2-count.py` VERBATIM, right before each instance's deletion. Exit 2 means STOP and send worldtree-dev the output. 2. Delete LIVE, with the api running, using literal paths only: `agents/{forseti,lofn,mimir}/memory/.chroma` and `memory/context_promotion`, on BOTH instances. 3. Send worldtree-dev the stamp; b193 ships after it. - A b192 restart re-creates an empty schema-only `ledger.db`. That is residue: say so in the stamp and remove it after b193. - **Legacy archive:** DESTROY it whole by **2026-10-30**, or at retirement-done, or on a subject-erasure request, whichever comes first. The runbook is in the detail file. - **TODO:** re-sweep both api logs after real traffic. After the b193 push, remove the retired config keys. ### fv-ml1 GPU layout (as of 2026-10-01) - **GPU 0:** cyberprev (47.1 GB), gen-small (35.3 GB), voices (10.8 GB), parakeet-nemo (3.6 GB steady). Free 385 MiB, FULL. - **GPU 1:** vllm-coder, erp-seat, meromero-rp, plus intern-decision (cap 14.4 GiB, 32k tokens, peak 15,220 of a 15,437 MiB budget). FULL. - **GPU 3:** the full-size-seat reserve (Flash-Next is parked). On-demand tenants: Blender, and Scriberr (0 idle, ~5.5 GB per job). When a full-size seat claims GPU 3, Scriberr steps aside to **irv-ml1's A6000**, not back to GPU 1. ### intern-decision (replaced SemIf on 2026-09-30) - **LIVE 0.1.3** at `intern-decision.fv.internal:8033`: semif-compatible `/decide` plus Jev `/v1/systemone`, 32k tokens, a Triton cache volume. Run `scripts/intern-decision-warmup` after an IMAGE change. → `persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md` - **Open:** label ~50 real Wyrd/Cicada turns before trusting it in production; its card makes no contamination claim. ### Scriberr (fv-ml1 GPU 3) - **LIVE `scriberr:local-blackwell-a353078-dropout2`:** upstream a353078 plus patch 0001 (overlap slicer) and patch 0002 (gap retry, `PARAKEET_MODEL_PATH`), carried LOCALLY ONLY (Prime 2026-10-01: no upstream). v3 stays. `scripts/scriberr-rebuild` re-applies both. → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md` ### nh3-pve + nh3-ml1: post-visit, all live (2026-09-25/26) - **nh3-pve:** Secure Boot OFF; IGFX restored; NVIDIA 580.178.04 DKMS; kernel 6.8.12-43 (benign btmtk oops every boot). **AMT live**: static `10.100.250.61` on nh3-mgmt (UDM port 6), KVM on, Opt-in None, password in the vault as `nh3-pve/amt-admin`. → `servers/nh3-pve/README.md` - **nh3-ml1** (CT 109 @ `10.100.50.80`): - TEI embed/rerank is **load-shared with esh-ml1** through the gateway, with `router_settings.enable_weighted_failover`. - brokkr's foundry seats: - `lfm-vl` `:8030` (gateway `lfm25-vl-3b`); - `lfm-vl-uncensored` `:8032` (direct; passed brokkr's eval); - `vibevoice-asr` `:8031` (audio.cpp, **Q8_0**). - GPU ~11.3/16 GB. → `servers/nh3-ml1/README.md` - Not yet proven: nh3-ml1 coming back by itself after an nh3-pve reboot (only the config was checked). ### esh-matter: Matter server for Home Assistant (2026-09-26) - CT 111 @ `10.0.90.20`, on VLAN 90 (esh-iot) only. matter.js 1.4.0; `:5580` is firewalled to HA `10.0.50.46`. HA's matter integration is loaded (ha-dev). - **Next is Prime's:** share the Aqara W200 into HA via Matter multi-admin. If commissioning misbehaves, the UniFi mDNS reflector on esh-iot is the first knob (a Prime/infra-ops change). → `servers/esh-matter/README.md` ### Worldtree reward path: code + config live (2026-09-27, Prime) - `d8726a13` (Domari's Skywork/Selene fix) was ALREADY on origin/main when Prime said to push it: it was pushed as part of `79cce92f` at 1221. Demo runs image `79cce92f`. **Personal still runs `51168a2e`**, which is older and lacks the fix; its image follows worldtree CI, never a manual up. - `deploy-wt-config` deployed demo and personal (1437): `providers.yaml` `reward_models.scalar-judge` (6d4ac44), plus a77639d's corrected selene description, which had never reached the hosts. Both are healthy, with 0 drift. It is inert on personal until that image picks up d8726a13. - worldtree-instance-configs: its 6 long-unpushed commits (53349f8…6d4ac44, Aug 2 to Sep 27) were **pushed on Prime's go-ahead** (a9d091e..6d4ac44). Origin, the repo and both hosts now agree. ### Blender on fv-ml1 GPU 3, agent-driven (2026-09-27, Prime) - Prime: "go ahead with gpu 3, both", and he does not use Blender, so **agents drive it through MCP**. Blender 5.2.2 LTS (linuxserver/selkies image, digest-pinned) with a web desktop at `https://10.251.50.54:3001` (vault `fv-ml1/blender-web-password`). **On demand only:** `scripts/blender-mcp up|down|status`. It was left DOWN (GPU 3 back to 2 MiB). - MCP: `mcp-for-blender` 2.1.1 (Dvalin's research pick). The add-on is vendored at `41a18432`. The server runs INSIDE the container, and agents reach it as ssh+docker-exec stdio via `scripts/blender-mcp`. No port is published, and screenshots work because server and Blender share a filesystem. Telemetry is off and safe mode is on. Verified end to end (render, screenshot, safe-mode refusal). **Batch path: `scripts/blender-run`** (one-shot `docker run --rm`, `--job` staging; first user is draupnir). **Access: the shared fleet `infra-ops` login, with no render-only key (Prime, 2026-09-28).** **Registration + up/down: PER WORKING SESSION (Prime, 2026-09-28, relayed by draupnir; superseded per-task of 09-27)**: `blender-mcp up` + `claude mcp add` at session start, `claude mcp remove` + `down` at its end. Never always-on, never user- or project-wide config. → `stacks/blender/README.md` - **Extensions (2026-09-28, draupnir; Prime ruled Blender a MANDATORY pipeline stage):** 8 pinned add-ons (`stacks/blender/extensions.lock`) built by `scripts/blender-extensions sync` into `fv-ml1:/tank/blender-extensions/5.2/system` (LIVE), mounted read-only as the System repo; `fleet_extensions.py` enables them (GUI startup timer; `blender-run --extensions`). Headless acceptance 8/9 (CAD Sketcher sketching is GUI-only); **MCP 9/9 after Prime ran the deploy himself (1505; the classifier had refused mine)**. SurfacePsycho's eval() is patched to literal_eval (held under MCP). Open: agent-drawn CAD Sketcher geometry (its stateful ops want point picks), and MeasureIt overlays not seen in MCP screenshots. GUI left DOWN. ### Zigbee2MQTT on esh-docker-vm (2026-09-27, Prime go-ahead; ha-dev request) - **LIVE since 1240:** Z2M 2.14.1 (digest-pinned) on `http://10.0.50.45:8099` (auth token), radio SLZB-MR1U chip 0 `tcp://10.0.90.10:6638` (ember), channel 25, **PAN 0xCFF4**. State and the network key are in `/opt/docker/data/zigbee2mqtt` (root 0700, restic via /opt/docker, never in git). Vault: `esh-docker-vm/zigbee2mqtt-{network-key,frontend-token,mqtt-password}`. Broker user `zigbee2mqtt` added to mosquitto (passwd backup `.bak-20260927-z2m`). - **ha-dev deleted the ZHA entry**; pairing the Aqara T1 is theirs. ⚠ The "bridge in HA" acceptance first FAILED because HA had had no MQTT since 2026-09-25 06:13. Cause: the esh-docker-vm macvlan shim had no host route to HA (two equal /24s, ens18 wins). Fixed at 1249 with a /32 via the shim (if-up.d hook, `playbooks/esh-docker-vm-macvlan-shim-route.yaml`), and HA reconnected. My first report called HA "connected" from a client count that included my own probe; `$SYS` counts are not identities. - ⚠ The first credential-mint attempt was blocked by the auto-mode classifier on a peer-relayed approval. It went ahead only on Prime's own go-ahead in this session. Treat that as correct. → `stacks/zigbee2mqtt/README.md` ### restic: credential leak fixed (2026-09-27, Prime) - Seven hosts were moved from `env-file` to `repository-file` (`playbooks/restic-repository-file.yaml`), so the systemd units no longer carry the rest-server password. Secrets are vaulted as `/etc/restic/{repository,password}`. `restic.env` is KEPT for manual snippets, so **a rotation must update the vault, `restic.env` and `repository`.** - **vm-esh-nas done too** (Prime bootstrapped infra-ops there at uid 850, NOPASSWD). **All 8 hosts are clean and vaulted.** - The passwords were readable until today, so the rotation (Prime's) remains the real fix. - 2026-09-27 0859 freshness check: all restic and PBS repos ✅. Those 0100 runs PRE-date the move (done ~0140–0200). - **VERIFIED 2026-09-28 (infra-hermes, read-only):** the first nightly on `repository-file` (0100 PT) succeeded on all 8 hosts (unit Result=success AND a today-dated snapshot in each repo), and the 0800 freshness check was all-green, PBS included. nh3-dev also passed a content restore (`af5580c3`). ### augaman: face recognition for Cicada - **v0.1.3 LIVE on esh-ml1:8040** (v0.1.1 first deployed 2026-09-26 2347 PT), `stacks/augaman`, healthy on CUDA, in `nvidia-smi`. Built on-box from a `git archive` of the tag. `pytest -m gpu tests/vision` 3/3 PASS on v0.1.3. - **Gallery backup WIRED + RESTORE-VERIFIED 2026-09-27** (Prime OK'd off-site): restic daily 0100 PT → rest-server-ana `esh-ml1/`, fail-closed hook, secrets vaulted under `esh-ml1/etc/restic/`. Identity-level restore matched the live gallery (canary, snapshot `fd3061a1`). **The backup gate for real enrollments is MET.** - **Deploy CLOSED by augaman-dev 2026-09-27 0009 PT:** HTTP checks passed; the canary survived the recreate and was then deleted, so the gallery is empty and ready for real enrollments. `/recognize` p50 186 ms (1080p, one face). The detector-latency follow-up is theirs. Their next release pins the container gid to 10001; nothing is owed by infra-ops. - **The fv-ml1 instance was REMOVED by Prime on 2026-09-27 (~0130 PT)** after the bench. esh-ml1 is the only instance: in the house, backed up, 48 ms per face on v0.1.3. - **v0.1.3 went live on both hosts (2026-09-27 ~0105 PT); the fv-ml1 instance was removed later.** Speed bench (`docs/pfi/augaman-speed-bench/`), server-side one face, v0.1.2 → v0.1.3: esh GPU 144 → 48 ms, fv GPU 75 → 27 ms. CPU mode REGRESSED (fv CPU6 152 → 205 ms; suspected ORT thread-pool spinning) and is reported to augaman-dev; the deployment does not use CPU mode. ### esh-ml1 - The reward seat's consumer is FIXED: Worldtree d8726a13 speaks vLLM /classify through the gateway `/scalar-judge` passthrough, which is key-gated per key via `allowed_passthrough_routes`, granted by infra-hermes on 2026-09-27. It is live on demo; personal waits on a newer image (see "Worldtree reward path"). - The nh3-docker Dozzle agent has been stopped by hand since ~2026-04. Revive or drop. - The Beszel superuser password was echoed into a session transcript (local only, not in memory). Rotation offered to Prime. ### Live threads - git: **eshpfi-management pushed to `6b66207` and worldtree-instance-configs to `b6fdd81` (2026-10-01 ~0418, Prime's go).** Anything after that is unpushed. ⚠ The working tree AND index are shared with infra-hermes and subagents: commit with `git commit -- ` (auto-memory `feedback_shared_git_index_commit_pathspecs`). `graphify-out/GRAPH_REPORT.md` stays modified and uncommitted on purpose: it is auto-regenerated. - nh3-dev root disk was cleaned 2026-09-30 1704 (uv prune, dangling images, old build cache): 86% → 82%. The Beszel 85% alert flaps near the line. - Booth submit-all fix (Prime's report) is LIVE since 2026-09-27 1705, via booth-dev (booth `50bfc7b`). Pushing it is booth's call, per Prime; it is not ours. - ESH has a single outside route (esh-scale on esh-pve). Noted, untracked. ## Recent decisions - `[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`.** It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner `irv-ml1-arbo` was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it. - `[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push.** DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; **unified-en KEPT, the live seat mounts it**), plus `/tank/spikes/scriberr-slicer` (including the private copies of Prime's recordings) and `/tank/spikes/parakeet-ab`. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed. - `[2026-10-01]` **irv-ml1 /storetank reclaim done:** Prime ruled through comfy-dev, which deleted 84 files of its own (260 → 299 GB free). Tiers B/C/D got no ruling (infra-hermes thread `01M3TCSYRSFNAPA9BPQTMFQ6KJ`). - `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. DONE 2026-10-01 0126 by infra-hermes; infra-ops audit passed 0137.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md` - `[2026-09-30]` **Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30.** → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md` - `[2026-09-30]` **SemIf replaced by intern-decision (Intern-Decision-4B, the Jev bench pick): semif-compatible plus Jev `/v1/systemone` at 32k tokens on GPU 1, with a Triton warm-up cache volume.** → `persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md` - `[2026-09-30]` **Scriberr moved to GPU 3 (on demand); our build carries the overlap slicer (0001) and the Parakeet gap retry (0002); v3 kept.** → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md` - `[2026-09-29]` **Worldtree U11a prepped, not flipped: a staged demo config plus the agreed U8 window plan (infra-hermes runs the batches).** → `persistent-memory.d/2026-09-29-worldtree-u11a-prepped.md` - `[2026-09-28]` **Worldtree U10 backfill done on demo (5) and personal (797). model_roles drift needed a memory_tagger sync first; mimir had missing vectors.** → `persistent-memory.d/2026-09-28-worldtree-u10-backfill.md` - `[2026-09-28]` **Bonsai ternary vs Q4_K_XL at concurrency on the 275 W card (1.93x at N=1 falls to 1.06x at N=8, 1.21x with the MMVQ fix); the weights are acquired.** → `persistent-memory.d/2026-09-28-bonsai-ternary-spike.md` - `[2026-09-28]` **blender-run gained `--cpu` (no GPU attached) and a fixed hostname `fv-ml1-blender` (draupnir).** The design stage never renders, so it stays off GPU 3. - `[2026-09-28]` **Blender extensions live in a read-only System repo built from a sha256 lock, enabled by a hook, opt-in for blender-run (`--extensions`).** SurfacePsycho's eval() is patched to literal_eval (a proven safe-mode escape). → `stacks/blender/README.md` § Extensions - `[2026-09-27]` **hermes-gateway restarted 0401 for highseat-dev** (SVOS v2.1.12: `propose_decision` gained `seat_up`, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enabling `zellij-fleet@Claude` at boot remains Prime's call. - `[2026-09-27]` **SemIf LIVE on fv-ml1 GPU 1 (semif-serve 0.1.2, Prime):** wrapper + contract + 39 tests, 142/144 upstream parity, two card-only memory defects fixed. → `persistent-memory.d/2026-09-27-semif-live-on-fv-ml1-gpu1.md` - `[2026-09-27]` **Blender 5.2 on fv-ml1 GPU 3 (on demand), agent-driven via mcp-for-blender running in-container over ssh stdio; safe mode on, no published port.** → `stacks/blender/README.md` - `[2026-09-27]` **esh-docker-vm host→HA traffic needs a /32 over macvlan-shim (if-up.d hook); without it HA silently loses MQTT (3× since August).** → `servers/esh-docker-vm/README.md` - `[2026-09-27]` **Zigbee2MQTT live on esh-docker-vm :8099 (PAN 0xCFF4, ch 25) replacing ZHA; network key vaulted, data host-only.** Prime go-ahead in this session; a peer-relayed approval was blocked by the permission gate. → `stacks/zigbee2mqtt/README.md` - `[2026-09-27]` **semif-serve 0.1.4: object states ending in `)`, `;` or `}` no longer 422 (INV-7, a prefix wrapper proven at startup); numerics are deterministic within a process but a bf16 near-tie can flip across a restart.** Prime ruled; heid bug hunt folded. → `stacks/semif/README.md` - `[2026-09-27]` **SemIf as Cicada's mood source: slower (+32 ms async, +94 ms sequential) and worse (67% vs 92% apt; carry 7/15 vs 14/15); only the gesture restraint is a win.** Build nothing (Prime). Henge 88 carries it. → `persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md` - `[2026-09-27]` **SemIf consumer-fit spikes (Prime): Cicada affect gate 30/31 with descriptive wording and 19/31 terse; Wyrd "left this place?" 21/21 on the second wording, exit choice 18/21.** Recommendations await Prime. → `persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md` - `[2026-09-27]` **SemIf order-averaging spiked (+9.1 pts accuracy, agreement = strong confidence signal); Prime ruled: build it in as 0.1.3 and trial the fast kernels.** Tracked in the in-flight SemIf section. → `persistent-memory.d/2026-09-27-semif-order-averaging.md` — **DONE:** 0.1.3 live, with fast kernels adopted (`77b8cb4`). - `[2026-09-27]` **restic repository URL moved out of world-readable systemd units on all 8 hosts, and vaulted (Prime).** Rotation stays Prime's. → `persistent-memory.d/2026-09-27-restic-repository-file-fleetwide.md` - `[2026-09-27]` **infra-ops bootstrapped on vm-esh-nas by Prime** (the first account at the fleet-pinned uid/gid 850; `bootstrap-infra-ops-user.yaml` now pins 850 when free, `d775a01`). - `[2026-09-27]` **augaman: second instance on fv-ml1 benched, then REMOVED (Prime).** esh-ml1 on v0.1.3 does 48 ms/face in the house with the verified backup (`ad484c3`, bench `docs/pfi/augaman-speed-bench/`). - `[2026-09-27]` Created empty private repo `corviduo/svoperatingsystem` (SVOS, the operating system) as claude-bot, at brokkr-smithy-dev's relay of Prime's ruling. `svos-dev` → the OS session; The High Seat's session is `highseat-dev` again. - `[2026-09-26]` **augaman: infra-ops owes the gallery backup + first deploy — DEFERRED until augaman-dev reaches deploy.** → `persistent-memory.d/2026-09-26-augaman-infra-ops-owes-the-gallery-backup-first-deploy.md` — `[2026-09-27]` **DONE:** deployed (v0.1.3 on esh-ml1), backup wired, identity-level restore verified (`d8f59a1`). - `[2026-09-26]` **`zellij-fleet@.service` installed, NOT enabled** (for svos-dev's seat_up; owns the fleet zellij server in its own cgroup). Enabling `@Claude` at boot is Prime's call when seat_up ships. Tracked: `0ad7799`, `services/zellij-fleet/README.md`. - `[2026-09-26]` **btusb blacklist on nh3-pve + esh-pve — proposed, awaiting Prime (low priority).** Every 6.8.12-4x boot oopses in btmtk (benign). Also noted for the next NH3 visit: nh3-pve runs ~20 °C warmer than esh-pve (workload-confounded; check airflow). Tracked: `servers/nh3-pve/README.md`. - `[2026-09-26]` **GPU-LXC Temperature alerts watch the GPU, not the host CPU** (`SENSORS=-coretemp_*,acpitz`); hypervisor CPU alerts at >95 °C on nh3-pve/esh-pve. The first nh3-ml1 alert was nh3-pve's CPU during vzdump. `6d901c0`. - `[2026-09-26]` **Abliterated LFM2.5-VL-3B parallel seat (:8032) passed brokkr's 6-image eval and stays** (direct only; stock seat kept for SFW A/B). `9cc3824`, `944bb36`. - `[2026-09-25]` **AMT follow-ups PARKED until the two MS-03s arrive and esh-pve's MS-01 is on AMT (Prime 2326).** — Parked: MeshCommander test; dummy HDMI plug, then NanoKVM → gx10; an OOB path not through nh3-scale… → `persistent-memory.d/2026-09-25-amt-follow-ups-parked-until-the-two-ms-03s-arrive-and-esh.md` - `[2026-09-25]` **nh3-pve AMT LIVE: static `10.100.250.61` on nh3-mgmt (UDM port 6), KVM on, Opt-in None** — (Prime). → `persistent-memory.d/2026-09-25-nh3-pve-amt-live-static-10-100-250-61-on-nh3-mgmt-udm-port.md` - `[2026-09-25]` **pfi-gx10 AC-restore VALIDATED by Prime's AC pull; Homepage stays on esh-docker-vm (Prime) after the mmap_lock wedge reboot.** `servers/pfi-gx10/README.md`, `servers/esh-docker-vm/README.md`. - `[2026-09-26]` **VibeVoice ASR → Q8_0 (Prime).** WER on the 4 bundled LibriSpeech clips 3/69 → 2/69 (only the I'm/I am artifact left), ×3 identical; RTF 0.09–0.17; +1.1 GB VRAM (nh3-ml1 ~11.3/16 GB). Q4_K file removed. - `[2026-09-26]` **esh-matter LIVE: a Matter server (matter.js 1.4.0) on CT 111 @ 10.0.90.20, VLAN 90 only** — , for ha-dev (operator-approved, relayed). → `persistent-memory.d/2026-09-26-esh-matter-live-a-matter-server-matter-js-1-4-0-on-ct-111.md` - `[2026-09-26]` **Embed/rerank LOAD-SHARED across esh-ml1 + nh3-ml1 (Prime).** — Second deployments were added for qwen3-embedding and reranker (config) and for reranker-a3-bge-v2-m3 (DB… → `persistent-memory.d/2026-09-26-embed-rerank-load-shared-across-esh-ml1-nh3-ml1-prime.md` - `[2026-09-26]` **Two dataset-foundry utility seats LIVE on nh3-ml1 (brokkr; operator approval relayed):** — LFM2.5-VL-3B on llama.cpp `:8030` (gateway `lfm25-vl-3b`, LiteLLM restarted 36 s at 0039) and… → `persistent-memory.d/2026-09-26-two-dataset-foundry-utility-seats-live-on-nh3-ml1-brokkr.md` - `[2026-09-26]` **Coder seat STAYS on fv-ml1 (Prime).** — The nh3-ml1 copy gave the same quality (teacher-forced true-code logprob diff +0.008 ± 0.019) but ran ~5×… → `persistent-memory.d/2026-09-26-coder-seat-stays-on-fv-ml1-prime.md` - `[2026-09-25]` **nh3-ml1 LIVE after the NH3 visit (SB off, IGFX restored): driver + CT 109 + TEI, parity vs esh-ml1 indistinguishable (cos min 0.999993 = own floor; overlap@10 1.000 vs MRL-256 control 0.684), same speed; monitoring/DNS wired. Autonomous: lxc-pve 6.0.0-2 upgrade on nh3-pve (Docker-in-LXC fix #7006). Gateway routing + AMT follow-up are Prime's.** → `persistent-memory.d/2026-09-25-nh3-ml1-live.md` - `[2026-09-25]` **nh3-ml1 = CT 109 @ 10.100.50.80, second TEI embed/rerank backend on nh3-pve's RTX 2000E; GPU playbooks made host-generic — BUILD DEFERRED: blocked on nh3-pve Secure Boot until the NH3 visit; SB handling and gateway routing are Prime's calls.** Tracked: `7ddd116`, Current state. → `persistent-memory.d/2026-09-25-nh3-ml1-standup.md` - `[2026-09-25]` **nh3-pve OOB: HOLD blind until the next NH3 visit (Prime via Miranda, 1333).** Target: NanoKVM → pfi-gx10, and the MS-01 on its own vPro. The IGFX BIOS fix rides the same visit, since AMT KVM captures only the iGPU. Tracked: `servers/nh3-pve/README.md` (`2b49be4`, `cdd7605`). - `[2026-09-25]` **MS-01 foot-gun: a GPU in the PCIe slot renames every NIC** — (the slot's root port takes bus 01, so the X710 goes `enp2s0f0np0`→`enp3s0f0np0`). → `persistent-memory.d/2026-09-25-ms-01-foot-gun-a-gpu-in-the-pcie-slot-renames-every-nic.md` - `[2026-09-25]` **Direct ESH→esh-ml1 consumer path — PARKED** (no ESH-side callers in 7 days; consumers go through the gateway). Tracked at `persistent-memory.d/2026-09-24-esh-ml1-embed-rerank.md`. - `[2026-09-25]` Created empty private repo `corviduo/norn` (git@gitea.phasefinal.com:corviduo/norn.git) as claude-bot, at brokkr-smithy-dev's relay of the operator's Norn ruling; the `norn-dev` handle is the operator's to declare. - `[2026-09-25]` **Reward seat audited (Skywork-Reward-V2-Llama-3.1-8B still #1 of 188 on RewardBench 2; our AWQ ≈ bf16 within noise; double BOS costs ~2.7 pts) and moved fv-ml1 → esh-ml1; esh-ml1 monitoring wired (Beszel+GPU, Kuma, Homepage, Dozzle); Dozzle hub's stale agent IPs fixed; nh3-dev Beszel agent revived.** → `persistent-memory.d/2026-09-25-reward-move-and-esh-ml1-monitoring.md` - `[2026-09-25]` **TEI adopted as the fleet embed/rerank engine (Prime); esh-ml1 is the sole gateway backend for `qwen3-embedding` + `reranker`; fv-ml1's vLLM embed/rerank seats retired (~6.1 GB freed); parakeet stays on fv-ml1.** → `persistent-memory.d/2026-09-25-tei-fleet-embed-rerank.md` - `[2026-09-24]` **esh-ml1 built — RTX 2000E Ada as CT 110 on esh-pve (host driver 580.178.04 DKMS, loaded live, no reboot), vLLM embed+rerank at measured parity with fv-ml1, wired as LiteLLM `order: 2` failover; dead DB alias `reranker-a3-bge-v2-m3` (pre-relocation IP) repaired.** → `persistent-memory.d/2026-09-24-esh-ml1-embed-rerank.md` - `[2026-09-24]` **esh-pve: VM 102 retired, a 14-day hung VFIO process found and cleared, T400 → RTX 2000 Ada; Prime chose LXC + host driver for embed/rerank** (implementation deferred to next session, tracked here + `servers/esh-pve/README.md`). → `persistent-memory.d/2026-09-24-esh-pve-vm102-retired-gpu-swap.md` - `[2026-09-24]` **pfi-gx10 AC-restore UEFI patch applied but NOT validated; box OFF until Prime's AC pull 2026-09-25** (tracking: `servers/pfi-gx10/README.md`, `e5197a3`). — `[2026-09-25]` ✅ **VALIDATED**: Prime pulled and replugged AC, and it came up by itself at 1108:54 (`2b49be4`). → `persistent-memory.d/2026-09-24-gx10-ac-restore-patch.md` - `[2026-09-24]` **NH3 power outage recovered** — pbs-nh3 had no `onboot` (set), NFS boot race fixed with automount (`1cbde50`), every other Claude session on nh3-dev died. → `persistent-memory.d/2026-09-24-nh3-power-outage-recovery.md` - `[2026-09-24]` **Miranda standing order is a repo CLAUDE.md operating parameter** — (`4b29492`, aligned to the global send protocol in `bcf3342`): high-urgency matters go to her, fixed or not… → `persistent-memory.d/2026-09-24-miranda-standing-order-is-a-repo-claude-md-operating.md` - `[2026-09-24]` **task-board mothballed** (Prime): container removed on ana-docker, data/image/compose kept, Kuma monitor deleted, task_* instructions removed from CLAUDE.md and the fork template (`e6da607`). Its hooks had sent no traffic in 30 days. - `[2026-09-24]` **Military 24-hour Pacific clock times** carried into the Codex/Grok shared bootstrap `docs/fleettools/AGENT-BOOTSTRAP.md` (`ad2b4d9`); Claude seats get it from the global CLAUDE.md. - `[2026-09-24]` **Worldtree `admin.memory.forget` stays OFF on demo/personal** until an instance needs it — a destructive erase; enabling it is a per-change operator yes via deploy-wt-config. - `[2026-09-24]` **A git checkout under `root:docker` needs `safe.directory` for its deploy user** — the 09-14 normalization (`826a63b`) silently broke yt-voice-clipper's webhook deploy until v0.3.13; fixed on irv-ml1 and recorded in fleet conventions (`eea9eb2`). Sweep found no other case. - `[2026-09-23]` **elway sudo uploads land root:root, validated and staged; fleet ownership audit built; 33 mis-owned root files fixed on 10 hosts.** → `persistent-memory.d/2026-09-23-elway-ownership-fix-fleet-audit.md` - `[2026-09-23]` **headscale-ddns hardened** — (`fedd4b6`): Cloudflare calls retry and validate every body (`pick()`), no write without both IDs, the run… → `persistent-memory.d/2026-09-23-headscale-ddns-hardened.md` - `[2026-09-23]` **esh-docker-vm restic was skipped 09-22..23 by my own Kuma move** — a dead uptime-kuma lookup aborted `pre-backup.sh` under `set -e` (`25e41d2`). → `persistent-memory.d/2026-09-23-esh-docker-vm-restic-was-skipped-09-22-23-by-my-own-kuma.md` - `[2026-09-23]` **hermes-gateway restart exit-1 is a Hermes race, not a crash** — the planned-stop watcher consumes the marker before SIGTERM re-runs the handler; drop-in… → `persistent-memory.d/2026-09-23-hermes-gateway-restart-exit-1-is-a-hermes-race-not-a-crash.md` - `[2026-09-23]` **Booth link board cleared to 14 durable links** (208 removed: drops, release posts, research links, token-bearing URLs; Scriberr reposted at `scriberr.fv.internal:8080`). Prime's rule: durable debug links only. - `[2026-09-22]` **Both carried calls approved** — build the `NRestarts` flap sampler (`163bb97`); the restic content-assertion ruling is ratified and stays (`ba60fda`). - `[2026-09-22]` **safe-rm installed on nh3-dev and delegated fleet-wide to infra-hermes.** — ⚠ The package installs INERT and looks fine — Debian's `/etc/zsh/zprofile` has 0 non-comment lines so the… → `persistent-memory.d/2026-09-22-safe-rm-installed-on-nh3-dev-and-delegated-fleet-wide-to.md` - `[2026-09-22]` **The acceptance probe for a guard must not be able to destroy what it tests** — (infra-hermes). → `persistent-memory.d/2026-09-22-the-acceptance-probe-for-a-guard-must-not-be-able-to.md` - `[2026-09-22]` ⚠ **`{"sent": true}` is a claim about transmission, never about effect.** `pane_send` structurally cannot deliver a harness command (every relay is prefixed with the card id), so D-0010 promised an unachievable `/clear` and its receipt reported success. Consumed an operator approval. → `persistent-memory.d/2026-09-22-instrument-errors.md` - `[2026-09-22]` **D-0010/D-0011 were misrouted to this seat by `pane_find` matching a ROLLING PANE TITLE.** — Genuine and operator-approved, wrong seat; `fleet_telemetry` held the right mapping and carries the warning… → `persistent-memory.d/2026-09-22-d-0010-d-0011-were-misrouted-to-this-seat-by-panefind.md` - `[2026-09-22]` **`headscale-ddns` exited 1 silently and the alarm carried no cause** — both failure paths were `|| exit 1` with no message. Now every exit names its reason and the WAN lookup retries 3×. The script was also untracked; a fix to what every mesh client resolves through lived on one disk. `30517fd`. - `[2026-09-22]` ⭐⭐ **Five instrument errors in one day, and the shape is one thing: a tool that enumerates "things that are fine" has selected against its own subject.** `--state=running` skipped the units most needing hooks; `awk '{print $1}'` dropped systemd's `●`-decorated FAILED rows; `grep -ic restic` on the wrapper missed the check script; `restic ls`'s header line made an absent path read as "blobs gone"; and a cooldown test that invoked before failing the unit. **Every one reported cleanly while looking at the wrong thing.** The rule is not "verify" — it is *verify, then ask what the verification could not have seen*. → `persistent-memory.d/2026-09-22-instrument-errors.md` - `[2026-09-22]` **Fleet alert bridge generalized** — `beszel-althing` → `althing-alert-bridge`, route registry (`/beszel` + `/kuma`), each with its own… → `persistent-memory.d/2026-09-22-fleet-alert-bridge-generalized.md` - `[2026-09-22]` **Uptime Kuma rebuilt from scratch on 2.5.5, moved esh-docker-vm → ana-docker.** — ⚠ `:latest` is a TRAP — it tracks 1.x, so an Aug-2026 pull gave a Dec-2024 build. → `persistent-memory.d/2026-09-22-uptime-kuma-rebuilt-from-scratch-on-2-5-5-moved-esh-docker.md` - `[2026-09-22]` **Beszel and Uptime Kuma are DISJOINT, not redundant** — Beszel's alerts bind to a *system* with a threshold; there is no URL column, so it is structurally incapable… → `persistent-memory.d/2026-09-22-beszel-and-uptime-kuma-are-disjoint-not-redundant.md` - `[2026-09-22]` **Failed-START alarms on 23 nh3-dev units** — (`services/althing-notify-failure/`). → `persistent-memory.d/2026-09-22-failed-start-alarms-on-23-nh3-dev-units.md` - `[2026-09-22]` **Backup coverage is a property of the SYSTEM, never one job's scope** — establish it by querying the repo for the path in a real snapshot, never by reading a job's `SRC=`. → `persistent-memory.d/2026-09-22-backup-coverage-is-a-property-of-the-system-never-one-job-s.md` - `[2026-09-22]` **restic checks now assert CONTENT and are DISCOVERED not enumerated** — conjunctive (recent AND present AND restores non-zero bytes); repos discovered from each NAS, hand list… → `persistent-memory.d/2026-09-22-restic-checks-now-assert-content-and-are-discovered-not.md` - `[2026-09-22]` **irv-ml1: Irvine is a TENANCY behind a Fortinet PFI does not control** — its TLS inspection breaks Tailscale relay/control intermittently (41 cert warnings/week). → `persistent-memory.d/2026-09-22-irv-ml1-irvine-is-a-tenancy-behind-a-fortinet-pfi-does-not.md` - `[2026-09-21]` ⭐⭐⭐ **lv-mccarthy GATED and NOT SHIPPED — voice passes, memorisation is the cleanest in the line, and the length discipline is gone.** 720 generations, 3 arms. Voice +0.152 at **2.9×** floor for ckpt450 (ckpt900 +0.172 at only 1.2×, its spread one outlier seed), and **~3/4 of the gain survives stripping every punctuation mark**, so it is not the cheap win. ⭐ Memorisation: **ckpt450 at 0.12 against the author's own held-out 0.12 — identical, longest match 11 words against the author's coincidental 12**, all 96 matches READ and every one stock grammar (`he looked at the wolf and he looked at him`); the name-shaped hits are the RENAMED inventions. Axis C is the blocker: **20% / 28% of generations overshoot the 90–140 band against base's 1%**, worst case a degenerate loop at 279 words. ⚠ My own prereg's axis C transcribed `score_beats.py`'s **v1** criteria including "in-band up on base", which the operator RETIRED 2026-09-15 because base maxes it — under the v2 (ran-on only) ckpt450 passes by **0.01 against a 0.200 floor**, but that reading was found AFTER the numbers and was not used. Fix the prereg prospectively. → `persistent-memory.d/2026-09-21-lv-mccarthy-gate.md` - `[2026-09-21]` **The two-epoch recipe is now 0 for 3, and this time the loss curve was CONFIDENTLY wrong.** — On Brontë and Hemingway the epoch-1/epoch-2 checkpoints were TIED on eval loss, so preferring the earlier one… → `persistent-memory.d/2026-09-21-the-two-epoch-recipe-is-now-0-for-3-and-this-time-the-loss.md` - `[2026-09-21]` **The memorisation control that shipped broken is now committed, and it turned a 12× red flag into a clean pass.** → `persistent-memory.d/2026-09-21-the-memorisation-control-that-shipped-broken-is-now.md` - `[2026-09-21]` **My confound detector was counting apostrophes as quote marks, and I saw it FIRE before I saw the bug.** → `persistent-memory.d/2026-09-21-my-confound-detector-was-counting-apostrophes-as-quote.md` - `[2026-09-21]` ⭐⭐⭐ **The ops log shipped — and then failed FOUR ways in its first hours, every one recording something unfindable.** Claim dropped by a sub-tool, hook behind graphify's eight `exit 0`s, no handle in the env, ssh-target written as a hostname. The general shape is **configured ≠ effective**; twelve instruments reported confidently and wrongly across three days, five of them mine. → `persistent-memory.d/2026-09-21-ops-log-and-the-instruments-that-lied.md` - `[2026-09-21]` ⭐⭐ **The Booth gained blur + a closed keep round trip, after shipping TWO controls that did nothing** — a reveal handler Jinja discarded for sitting after `{% endblock %}`, and a `×` a sibling form covered by 30×22 px. Both operator-found, both from reading templates instead of rendering them. `scripts/layout-probe.py` took four iterations to become trustworthy. ⚠ Blur is COSMETIC and a test asserts the 200 on purpose. → `persistent-memory.d/2026-09-21-booth-two-dead-controls.md` - `[2026-09-21]` ⭐⭐ **claude-bot is an Owner in corviduo/pfi/vastblue; cicada + draupnir moved to `pfi`; vh-token use is now standing-authorized from the vault.** ⚠ `vh` is a USER not an org, so no namespace-scoped admin exists — the only realization of the original ask was site-admin, surfaced rather than executed. ⚠⚠ `~/.config/claude-bot/gitea-token` is DEAD and had been misreporting permissions; the working one is `gitea-token-repo-create`. → `persistent-memory.d/2026-09-21-gitea-orgs-and-the-dead-token.md` - `[2026-09-21]` ⭐⭐ **nh3-dev 84%→72%: 27 GB reclaimed, and a 252 MB LoRA adapter rescued from `/tmp`, which this box sweeps at 3 days.** `babyyarros` existed nowhere else; hash-verified to smithy before any deletion. ⚠⚠ My 7-day session prune then deleted an ACTIVE session's dir — directory mtime does not reflect subdirectory writes. Do not re-run that predicate. → `persistent-memory.d/2026-09-21-disk-triage-and-the-rescued-adapter.md` - `[2026-09-21]` **Draupnir geometry engine provisioned on irv-ml1 and acceptance-tested** — (34c4179, e574b91, 2e08edc). → `persistent-memory.d/2026-09-21-draupnir-geometry-engine-provisioned-on-irv-ml1-and.md` - `[2026-09-21]` **Backup alarm verdict split: `STALE` (exit 1) vs `ERRORED-JOBS` (exit 3)** — (7fe4102). → `persistent-memory.d/2026-09-21-backup-alarm-verdict-split-stale-exit-1-vs-errored-jobs.md` - `[2026-09-21]` **`vh/forgefirm` mirrored** — from `github.com/openglow-org/forgefirm`, following the house convention read off the existing 17: `vh/`… → `persistent-memory.d/2026-09-21-vh-forgefirm-mirrored.md` - `[2026-09-20]` **`ravenpen.com` REGISTERED by the operator — the 09-18 hold is discharged and hamr-dev is answered.** — infra-ops deliberately did NOT execute this twice, because it was a non-refundable purchase reaching us as a… → `persistent-memory.d/2026-09-20-ravenpen-com-registered-by-the-operator-the-09-18-hold-is.md` - `[2026-09-19]` **FV is a dedicated 20 A circuit carrying fv-ml1 AND the R420 running OPNsense, and nothing else (operator-confirmed).** → `persistent-memory.d/2026-09-19-fv-is-a-dedicated-20-a-circuit-carrying-fv-ml1-and-the-r420.md` - `[2026-09-19]` **althing 3.7.0 rolled out to both infra-ops surfaces — and the rollout broke the claim tooling built that morning.** → `persistent-memory.d/2026-09-19-althing-3-7-0-rolled-out-to-both-infra-ops-surfaces-and-the.md` - `[2026-09-19]` **Three agents commit as one git author, and closing that gap took three instruments to get right.** — An unattributable commit (`e43e262`) appeared in the push set between two of mine — unidentifiable from git… → `persistent-memory.d/2026-09-19-three-agents-commit-as-one-git-author-and-closing-that-gap.md` - `[2026-09-19]` **The backup alarm had no wire for three weeks, and fixing it exposed two more instruments that pass without looking.** → `persistent-memory.d/2026-09-19-the-backup-alarm-had-no-wire-for-three-weeks-and-fixing-it.md` - `[2026-09-19]` **elway evaluated `when:` / `creates:` / `removes:` / `changed_when:` WITHOUT the step's sudo, and it fails silently in the dangerous direction.** → `persistent-memory.d/2026-09-19-elway-evaluated-when-creates-removes-changedwhen-without.md` - `[2026-09-19]` **ESH VM 102 (`esh-vm-workstation`) excluded from the nightly backup job — operator ruling.** — It is a Windows 11 Parsec/RDP **sandbox** (no password, no state to recover), and its vzdump had failed… → `persistent-memory.d/2026-09-19-esh-vm-102-esh-vm-workstation-excluded-from-the-nightly.md` - `[2026-09-19]` **The ops log is BUILT — `scripts/ops-log`, automatic writers, and a detector for the path they cannot cover.** → `persistent-memory.d/2026-09-19-the-ops-log-is-built-scripts-ops-log-automatic-writers-and.md` - `[2026-09-19]` **`infra-hermes` is this session's ASSISTANT, and the division of labour is now standing policy.** — infra-ops keeps **improving infrastructure tooling** plus the hard calls; infra-hermes does **day-to-day… → `persistent-memory.d/2026-09-19-infra-hermes-is-this-session-s-assistant-and-the-division.md` - `[2026-09-18]` ⭐⭐⭐ **NH3↔Anaheim had been running over a throttled DERP relay, not a direct path — 78 GB of fleet traffic on someone else's free infrastructure.** Four additive objects on ana-gw gave ana-scale a stable inbound UDP 41641 endpoint; `tailscale ping` 373–522 ms → **6 ms direct**, cross-site HTTP 1.2 s → 0.015 s, STT via the ANA gateway 1.4 s → 0.25 s. ⚠ That box runs `central-nat`, so a policy `dstaddr` is the REAL internal address, not the VIP. No OOB access — back up with `show` to a local file and make additive changes ONLY. irv-ml1 still relayed. → `persistent-memory.d/2026-09-18-nh3-ana-derp-relay.md` - `[2026-09-18]` ⭐⭐⭐ **`.internal` DNS was failing ~10% of lookups fleet-wide, from two independent causes.** A PUBLIC resolver was the fallback for a PRIVATE zone (Cloudflare answers NXDOMAIN authoritatively, so a transient miss became a hard failure) — now a cross-site ring, each site local-first with a different site as backup. Then the root cause: all three AdGuards shipped `ratelimit: 20` shared across an entire **/24**, silently dropping queries at 5 s each. Set to 0. Hard failures 3/40 → 0/40; burst timeouts 40/60 → 0/60. ⚠ `resolv.conf` is DHCP-managed — change it at the UDM/FortiGate, not the file. → `persistent-memory.d/2026-09-18-fleet-dns-ring-and-ratelimit.md` - `[2026-09-18]` ⭐⭐ **SearXNG had ONE working general web engine and every health check said fine.** 7 of 55 were enabled-by-default and six of those are dictionary/translation engines — `inactive: false` only makes an engine SELECTABLE, `disabled: false` puts it in the DEFAULT set. Now seven. ⭐ This stack tracks `:latest` ON PURPOSE (upstream ships engine-handler fixes continuously; a pin freezes breakage). ⭐ The Brave key is committed in plaintext by explicit operator decision — scoped to one low-value credential, NOT a change to the no-secrets rule. → `persistent-memory.d/2026-09-18-searxng-one-engine-to-seven.md` - `[2026-09-18]` ⭐⭐ **althing search returned ZERO for every hyphenated query — silently, for every handle, and on this fleet that is most of our hostnames.** FTS5 read the hyphen as a column filter; `search()` caught the error, its probe passed, and it returned `[]`. Routed to forseti (they own the code, I own rollout) → **v3.6.3** deployed, 17 s bus outage. The reported symptom, a message that never arrived, was NOT a defect: the reporter's evidence file was truncated at 8 KiB. ⚠ No CI builds the post-office image. → `persistent-memory.d/2026-09-18-althing-363-hyphen-search.md` - `[2026-09-18]` ⭐ **FleetTools: one capability index, autoloaded by Claude, Codex and Grok from a single symlinked file.** `docs/fleettools/` + `~/FLEETTOOLS.md`; absolute detail paths because a non-Claude agent cats them. Rule zero is query-live-inventories-never-a-written-list. The global CLAUDE.md tools section went 231 lines → 39, keeping only the three rules that govern behaviour rather than lookup. → `persistent-memory.d/2026-09-18-fleettools-agent-index.md` - `[2026-09-18]` **Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** — (operator-approved). → `persistent-memory.d/2026-09-18-miranda-s-hermes-plugin-install-is-now-a-symlink-to-the.md` - `[2026-09-18]` **Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal… → `persistent-memory.d/2026-09-18-worldtree-s-env-sh-secrets-are-vaulted.md` - `[2026-09-17]` **`dragonfireacoustics.com` expires 2026-10-30 — six weeks — at eNom with NO transfer lock, and its sibling domain was already lost exactly this way.** → `persistent-memory.d/2026-09-17-dragonfireacoustics-com-expires-2026-10-30-six-weeks-at.md` ⚠ **If it is ever transferred, DNS does NOT come with the registration** — the nameservers are eNom's `name-services.com` and the zone must be recreated first or mail dies. The whole zone is two facts plus a landmine: `*` (WILDCARD) → 199.250.192.76 which is **dead** (no HTTP at all, and it is what the apex/mail/webmail/admin/ftp all answer with), `www` → 38.120.12.45 (us), and **7 Google Workspace MX records that must not be lost**. No DNSSEC (`delegationSigned: false`), so no transfer complication. ⚠ Also found: **no SPF and no DMARC** at the apex on a Google Workspace domain — a live deliverability problem independent of everything else. Transfer gate is the **TAC/EPP code from the eNom account**, not the lock; the missing lock is not authorization. 60-day rule is satisfied (last changed 2025-10-24). - `[2026-09-17]` **`dragonfireacoustics.com` IS configured on `pfi-ana-webhost`, and the whole thing is dead — a forgotten public-facing VM.** → `persistent-memory.d/2026-09-17-dragonfireacoustics-com-is-configured-on-pfi-ana-webhost.md` - `[2026-09-17]` **PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the…** → `persistent-memory.d/2026-09-17-parked-lv-krakauer-and-the-reason-is-a-selection-criterion.md` - `[2026-09-17]` **The althing route-declaring SessionStart hook is documented but NOT installed on nh3-dev** — `dev_launch.py` has zero occurrences of "route", no hook declares one, and every live route was hand-declared… → `persistent-memory.d/2026-09-17-the-althing-route-declaring-sessionstart-hook-is-documented.md` - `[2026-09-17]` **PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the…** → `persistent-memory.d/2026-09-17-parked-lv-krakauer-and-the-reason-is-a-selection-criterion-2.md` - `[2026-09-15]` ⚠⚠ **`--gpu-memory-utilization` DOES NOT PREDICT RESIDENT VRAM — measure it, never compute it.** Wrong in **both** directions on fv-ml1: `vllm-cyberprev` util 0.40 (expect ~39,155 MiB) holds **47,124** (+8 GB over); `vllm-gen-small` util 0.48 (expect ~46,986) holds **36,942** (−10 GB under). Planning a placement off the fractions would have been 8 GB wrong. Read `nvidia-smi --query-compute-apps`. Full per-seat residency table + the breeze shuffle arithmetic → `persistent-memory.d/2026-09-15-breeze-placement-sizing.md` - `[2026-09-15]` **breeze-tts stays on irv-ml1; the TTS-stack move to fv-ml1 is PARKED (park id 75, `move-the-tts-stack-breeze-tts-bragi-tts-gateway`), triggered on evacuating embed/rerank/reward.** ⚠ Trigger as stated says "gpu0" but those three are on **GPU 1** (~0.16 util, ~15.7 GB; GPU 1 is the tight card at 0.975 / 4,336 MiB free) — confirm which he meant before executing. All three services move together because only `breeze-tts` is GPU-resident (~10.3 GiB, **growing**) while `bragi` and `tts-gateway` are CPU proxies, and co-location is what avoids a cross-site hop per TTS call. **breeze-tts sizing — original recommendation NOT to move it.** ~**10.3 GiB** measured under load at 53 min uptime, **up from 9.2 GiB** shortly after warm-up (it grows; n=2, plateau unmeasured) — so GPU 0's 11,982 MiB free is a **1.7 GB margin and shrinking**, on the live chat serving path. ⚠ Two measurement traps: it reports **nothing at idle on the wrong card** (`BREEZE_GPU_DEVICES=0` = the **3090**, not the A6000), and an early reading understates it. ⭐ The real objection is **topology**: `tts-gateway` is on irv-ml1 and reaches it same-box, so moving breeze alone adds a cross-site hop to every TTS call against a 478 ms first-sample budget. GPU 3 would fit it but spends the reserve. → `persistent-memory.d/2026-09-15-breeze-placement-sizing.md` - `[2026-09-15]` **Hermes bearer rotation hold released — svos-dev split their HS256 signing key off the shared value (svos `7165272`)** → `persistent-memory.d/2026-09-15-hermes-bearer-rotation-hold-released-svos-dev-split-their.md` - `[2026-09-13]` **STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint.** → `persistent-memory.d/2026-09-13-standing-policy-operator-cap-gpu-power.md` - `[2026-08-25]` **Fused MoE kernel path — DEFERRED, tracked at park `fused-moe-kernel-path-for-gemma-4-moe-training` (id 47).** → `persistent-memory.d/2026-08-25-fused-moe-kernel-path-deferred-tracked-at-park-fused-moe.md` - `[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`. _167 older entries archived to archival-memory.md._ ## Tried and abandoned - `[2026-10-01]` **Size + mtime scans to pick reclaimable "staging" on a model store** (infra-hermes, irv-ml1). `_inbound/retro-diffusion` looked like staging leftovers but is the TARGET of 47 model-store symlinks, holding a purchased bundle used by a live commission. Resolve symlinks INTO any staging dir before proposing to delete it. - `[2026-09-30]` **Shorter Parakeet slices as Scriberr's memory fix.** I predicted ~6× less memory from the attention-matrix arithmetic. Measured, 300 → 120 s only went 9,384 → 6,510 MiB: a ~5.6 GB fixed floor dominates. `expandable_segments:True` was the real lever (5,496). Measure the process peak; never extrapolate it from one tensor. - `[2026-09-30]` **Whole-file local-attention Parakeet in Scriberr (context 255/255).** OOM past 16 GB on a 35-min file. Local attention inside chunks is also non-deterministic run to run. - `[2026-09-30]` **Start-time midpoint stitching of overlapped Parakeet chunks.** It duplicated a word at 26 of 108 stitches, because Parakeet timestamps a post-pause word anywhere inside the pause. Hand over at a word both chunks agree on instead. - `[2026-09-30]` **int8 ONNX (sherpa-onnx) as the low-latency Parakeet runtime.** The int8 graph runs on ONE CPU thread with the GPU at 2–9%. Unified-en int8 was slower than the seat; fp32 ONNX was 4–12× faster, and NeMo was fastest. - `[2026-09-30]` **GPU budgets computed as total − used.** nvidia-smi `Free` is ~640 MiB lower per card (driver reserve). Budget from `Free`. - `[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel.** Every tool worked except `get_viewport_screenshot`, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. → `stacks/blender/README.md` - `[2026-09-27]` **A TCP connect as the "is Blender ready" probe.** docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on to `ping`. The same trap applies to any service behind a published port. - `[2026-09-27]` **`log.exception()` in a GPU failure path** — the record keeps `exc_info`, so any retaining handler (pytest's capture does) pins the traceback's frames and tensors. Log `traceback.format_exc()` text instead (semif-serve `engine._guard`). - `[2026-09-27]` **fla/triton in a slim image without gcc** — triton compiles its CUDA driver shim at runtime ("Failed to find C compiler"). The warm-up died and startup failed closed. The semif image now installs gcc + libc6-dev with the fast extra. - `[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** — the export drops uv's per-package index routing; with `--extra-index-url … --index-strategy unsafe-best-match`, triton came from the pytorch index and failed the lock's hash. Keep `uv sync` and blank the project version in a deps-only stage instead (`services/semif-serve/Dockerfile`). - `[2026-09-27]` **Translating a CUDA OOM by raising inside `except` (or `from exc`)** — the chained exception's traceback pins the failed call's frames, and their GPU tensors (11.9 GiB after the 503). Raise after the block, unchained, `gc.collect()` before `empty_cache()`. - `[2026-09-27]` **`deploy-stack.sh --yes` / `dns-sync.py` fed a blind `y`** — the harness refuses a blind apply. Review with `echo n |` first, then apply with `echo y |`. - `[2026-09-27]` **Relying on the build cache surviving on esh-ml1** — the augaman v0.1.3 build missed its dependency layer there (fv-ml1 hit it; reason unexplained, suspect the `docker builder prune` runs) and the rootfs peaked at 90%. Check for ≥15 GB free, or remove the old image, before building there. - `[2026-09-26]` **Greedy exact-match as parity for a small generative seat (coder)** — same-host repeats matched only 30–38% (run B hit the prefix cache, so a different numeric path; near-tie… → `persistent-memory.d/2026-09-26-greedy-exact-match-as-parity-for-a-small-generative-seat.md` - `[2026-09-26]` **zsh traps in ad-hoc test loops:** `set -- $a` does NOT word-split, so alert POSTs went out with empty fields; and `local path=…` inside a function CLOBBERS `$PATH` (zsh's tied array), so the failover test ran nothing while TEI sat stopped. Use other names and `${=var}`. - `[2026-09-25]` **Concluding "not on the UDM" from a port table read 15 s after link-up** — UniFi polls (~60 s); AMT was on port 6 as Prime said. Wait a poll, and prefer a which-port-lists-the-MAC check. auto-memory `feedback_polled_stats_lag_the_event`. - `[2026-09-25]` **The esh-pve NVIDIA DKMS recipe on nh3-pve** — it built, then the kernel refused to load it: nh3-pve has SECURE BOOT ON (esh-pve, the same MS-01, has it off). The installer rolled back. The playbook now pre-flights `mokutil --sb-state` (`7ddd116`). - `[2026-09-25]` **elway against an unpinned host-key name** — every `when:` hit ssh rc 255, so all 9 steps showed SKIPPED, not failed. Only the verify phase went red. Pin the name (the IP key matched) or use the IP. Untracked elway defect: rc 255 in `when:` should fail. - `[2026-09-24]` **Testing "Restore AC Power Loss" with an OS shutdown** — a shutdown stays off BY DESIGN; only pulling and restoring AC tests it. A community README claimed the two are indistinguishable; trusting it cost Prime two trips to the pfi-gx10 power button. - `[2026-09-23]` **`booth link --help`** — there is no help flag; it posts `--help` to the operator's link board as a link. Read `booth` with no args for usage. - `[2026-09-21]` **Using directory mtime as a liveness test when pruning session scratchpads** — `find /tmp/claude-1000 -mindepth 2 -maxdepth 2 -type d -mtime +7` deleted an ACTIVE session's working… → `persistent-memory.d/2026-09-21-using-directory-mtime-as-a-liveness-test-when-pruning.md` - `[2026-09-18]` **Routing SearXNG's egress through a SOCKS5 proxy on esh-scale** — one day live, reverted. It fixed nothing, and the reason I gave for reverting it was itself wrong: the rollback was the counterfactual and it falsified my own published claim. Kept reverted on its own merits (no measurable gain, added a hard ESH dependency for all fleet search). → `persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md` - `[2026-09-18]` **`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten… → `persistent-memory.d/2026-09-18-apikey-env-searxngbraveapikey-in-searxng-settings.md` _118 older entries archived to archival-memory.md._