66 KiB
Persistent memory — eshpfi-management
Last updated: 2026-10-01 ~0420 PT (Parakeet seat → unified-en under NeMo LIVE + audited; gen-small util 0.36 → .env 0.33; leftover bench weights + spike dirs deleted; Scriberr no-upstream; eshpfi + worldtree-instance-configs pushed. Prior: U11a off, SemIf → intern-decision, Scriberr GPU 3 + patches.)
Always check for
/tmp/infra-ops-handoff.md— if it exists and itsWritten:stamp is under 8 hours old, read it (it carries the in-flight handoff from the previous session), then delete it. Older than 8 hours: stale — delete it unread.(Raised from 1 h to 8 h by operator 2026-09-13 — a one-hour window deleted the handoff unread across any overnight gap, which is the exact case it exists for. 8 h also matches the global CLAUDE.md and the
/snapshotskill default.)
Repo purpose
- 2026-09-10 Beszel fleet wiring: all seven requested hosts plus existing corviduo-dev report up.
/tankand other data filesystems now have real usage metrics; NVIDIA telemetry covers ana-ml2 and irv-ml1. Thirty alerts deliver to infra-ops, explicitly chosen by operator; Miranda routing is deferred. A real low-threshold disk alert reached althing, then the threshold was restored to 85%/5 min. Homepage has one native overview widget (reachability counts, not degraded health). Dedicated superuser approved and stored in Vaultwarden. Seepersistent-memory.d/2026-09-10-beszel-fleet-wiring.mdandstacks/beszel/README.md.
Reference workspace for PFI infrastructure: server inventory, canonical
Docker Compose stacks, ops playbooks, and conventions. Authoritative
copies of compose files live on the servers under
/opt/docker/compose/<stack>/; this repo mirrors them for version
control, editing, planning, and CI-driven deploys. It was originally
spun up to handle the fleet backups — keep that lens when triaging
backup/storage issues.
Tools and conventions
Sister repos (separate gitea repos, deployed by playbooks here):
| Repo | Role | CI status |
|---|---|---|
vh/task-board |
MCP + web dashboard for assistant task state (port 7878) — [2026-09-24] MOTHBALLED by Prime (superseded by the High Seat + ledger); container removed on ana-docker, data/image/compose kept (e6da607) |
push-to-main → CI deploys (2026-04-29) |
vh/vor |
Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
vh/nevermore |
Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
vh/asset-engine |
Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
vh/althing |
Lean trusted inter-agent message bus — v3.0.0 "the post office" as of 2026-08-28 (U9b flag day, one-way, no rollback): ONE container on nh3-dev at http://10.100.50.40:8390 is the only stateful component; althing-po-herald one per box; althing-listen one per session; postbox is the client. Every v2 command was DELETED, not deprecated — althing-cli→postbox, althing-wake-listener→althing-listen, althing-light-monitor/althing-receiver gone. Sessions need BOTH ALTHING_POST_OFFICE and ALTHING_HANDLE; there is no default address. ⚠ An unreachable post office is an OUTAGE, never an empty inbox. → persistent-memory.d/2026-08-28-althing-v3-cutover.md |
per-box install (NOT CI-deploy); nh3-dev = container host + repo; nh3-extdev = system WHEEL at /opt/uv-tools, needs its own wheel install (playbooks/nh3-extdev-althing-v3.yaml) |
vh/mead-hall |
Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
vh/skaldsong |
Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
vh/Worldtree |
Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. gitea-runner builds on ana-docker; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. | push-to-main → CI build-and-deploy (runner on ana-docker) |
vh/yt-voice-clipper |
YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md |
vh/arbo |
Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) | push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook |
vh/zonos-gateway |
OpenAI-compatible TTS gateway over stock ZONOS2 (:8890 irv-ml1); emotion dials-first + voice mapping; reached via LiteLLM ext-tts alias. v0.2.1 (2026-07-18): voice-resolved emotion presets (resolve_preset(name,voice); angry/happy/startled_happy per-voice). 8 voices incl. 4 clones |
pushed to gitea (main 8f1885b/v0.2.1); deployed irv-ml1 tree still NON-git (hand-updated build context — CI-wire = open follow-up). Spec docs/EMOTION-DIALS-SPEC.md; host-managed voices bind-mount (./voices:/app/voices, drop wav + restart, no rebuild) |
vh/soong-lab |
Noonien Soong character-design studio (SPA + /api + WT /bifrost/tool-call); containerized 2026-07-18, LIVE on corviduo-dev :8443 (image vh/soong-lab:latest). soong-dev owns Dockerfile/compose/workflow; infra-ops owns the host |
CI = Gitea Actions build+push+DEPLOY on tag/dispatch (fleet recipe: docker:cli + raw buildx, pushes AS vh; auto-redeploy LIVE 2026-07-18 — runner SSHes corviduo-dev as deploy, compose pull && up -d from /opt/soong-lab, health-gated on /api/version). Manual redeploy sudo -u deploy bash -c 'cd /opt/soong-lab && docker compose pull && docker compose up -d'. → archival-memory.md (archived 2026-08-16) |
model-training-forge (mtf-dev) |
Fine-tuning recipe forge; T1 = E-RP writing LoRA, retargeted qwopus-122B→AEON-27B (2026-07-06) (SFT→DPO, LitBench-RM reward) | training runs, not a deployed sidecar |
(vh/volva + Heid were re-architected from systemd daemons to Claude Code
session orchestrators 2026-06-08; their nh3-dev .service units were removed —
no longer deployed sidecars here. See Recent decisions.)
-
Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see
docs/runbooks/disaster-recovery.mdfor the blast-radius matrix. ⚠️ The restic file+DB layer routes through TWO rest-servers (rest-server-ana@ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas;rest-server-nh3@ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of/mnt/backup. (rest-server-ana recovered 2026-06-20.) -
pull-hf-repo.yamlis the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at/tank/aimodels/huggingface/" playbook. Supports--var repo_type=model|dataset|space. Replaces ad-hochuggingface_hub.snapshot_downloadpatterns. -
Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (
key_id 61419c92) atana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-adminauths against demo only. Personal-instance admin (the~/.config/worldtree/personal-admin-token, mode 600) POSTs/admin/keys(mints per-project keys; takesuser_id+label, no scope param — scopes are tier-derived). On-instance mint recipe (cleaner than DB-manip):docker exec worldtree-worldtree-api-1POST/admin/keyswith the in-containerWORLDTREE_BOOTSTRAP_ADMIN_KEY; cleartext once in.key=wt_live_+16hex. auto-memoryreference_worldtree_demo_key_mint. -
Per-project user keys against personal Worldtree (issued 2026-05-19):
skaldsong:79744637,skaldsong:7c1dbbbe,althing:50d85460,mead-hall:a360822d. Mint via/admin/keys, drop value to/tmp/wt-personal-<name>.keymode 600, dev collects + shreds (DO NOT cat to chat transcript). -
Skaldsong CD pattern (registry-pull). vh/skaldsong's CI builds and pushes
gitea.phasefinal.com/vh/skaldsong:<sha>+:latest;playbooks/deploy-skaldsong.yamlon ana-docker pulls + recreates. SHA-pin only. Prereq: host needsdocker login gitea.phasefinal.comonce. -
gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH
10.250.50.70:222, HTTP:3000. Fleet/colo hosts must use this internal route, NOT publicgitea.phasefinal.com(38.120.12.44) — the public path fail2bans the host egress IP. Full gotcha indocs/orientation.md→ Git/gitea. -
docker-as-root pattern (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo):
docker run --rm -v <target-dir>:/wt docker:cli sh -c "...". docker-group membership is effectively root via bind-mount. Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass-e VAR=/abs/pathfor any relative-default config dir. -
scripts/elwaysudo handling — elway prompts for the sudo password ONCE viagetpassbefore the firstsudo: truestep → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH. -
Per-host SSH identity matters for sudo. infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: default
ssh ana-docker=lkraven(docker-group, NO passwordless sudo);ssh infra-ops@ana-dockerHAS NOPASSWD root. → For any sudo op on ana-docker, usessh infra-ops@ana-docker.ssh infra-ops@10.100.10.50(nh3-dev) ALSO NOPASSWD sudo; on nh3-extdev infra-ops is sudo-LESS by design (ssh lkraven@10.100.50.42is the NOPASSWD path) — [2026-09-23] measured: infra-ops HAS NOPASSWD sudo on nh3-extdev (fleet ownership audit; also CLAUDE.md 2026-09-05). irv-ml1:ssh irv-ml1= lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD) — stage model pulls to/home, not root-owned/worktank.
Current state / in-flight
As of 2026-10-01 ~0420 PT.
Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01)
- ⚠ INCIDENT 04:21 PT 2026-10-01, MITIGATED, root fix in flight:
vllm-gen-small's EngineCore CUDA-OOM'd when it needed a 394 MiB runtime workspace and GPU 0 had 388 MiB Free. The parakeet seat was parked at its 3,582 MiB window cache.- vLLM grows ~0.8 GB at runtime beyond its preallocation; my audit checked gen-small's boot margin, not its runtime growth. gen-small auto-restarted, healthy at 04:23.
- I restarted parakeet-nemo to drop its cache (rest 2,084 MiB). gen-small then answered 3/3 via LiteLLM, and GPU 0 Free is ~1,075 MiB.
- infra-hermes is tasked with nemo-0.1.1:
empty_cacheafter windowed requests, plus a hard memory ceiling so the seat 503s instead of starving gen-small. It also measures gen-small's runtime growth. - Awaiting Prime: trim gen-small's KV pin (8 → 7 GiB frees ~1 GiB; 670k → ~586k tokens), or move the seat to GPU 3.
- Until fixed, a long transcription can re-grow the seat's cache and starve gen-small.
- LIVE since ~0126 PT 2026-10-01 as
parakeet-nemo(stacks/parakeet-nemo,local/parakeet-nemo:nemo-0.1.0, built by infra-hermes) on fv-ml1 GPU 0 :8300, with LiteLLMext-stt/whisper-1unchanged. infra-ops audit PASSED 0137.- p50 on GPU 0 for 1–3 / 3–8 / 8–20 / 20–60 s: 33 / 36 / 42 / 71 ms, against 187 / 308 / 626 ms for the old seat.
- WER: LibriSpeech clean 1.965, other 3.026.
- Files longer than 6 min run in 360 s windows. That avoids NeMo's T×T attention mask; a seam can lose a space or a word.
- Rollback:
docker stop parakeet-nemo && docker start parakeet. The old container and image are kept. - GPU 0 is FULL:
- The seat's steady state is 3,582 MiB (its cached window peak); Free is 385 MiB.
vllm-gen-smallruns at util 0.36, and its.envholds 0.33 for the next restart (~3 GiB of boot-check margin). Its KV is byte-pinned: 670,142 tokens / 2.56×.- Before restarting any vLLM seat on this card, check that util × 95.6 GiB ≤ measured Free + the seat's own resident memory. Do not trial-boot. A trial-boot sequence took gen-small down for 34 min on 2026-10-01.
- Seat invariants (in its README): cast to bf16 AFTER change_attention_model; uvicorn pinned with
--http h11(httptools 0.8.0 emitsHTTP/1.1 200\x00OK, which LiteLLM/httpx rejects). - NVIDIA Open Model License accepted for internal use. →
persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md
Worldtree U11 memory cutover (demo + personal)
- Legacy plane OFF since 0115/0120 PT 2026-09-30 (config repo 63cf268; personal /metrics b6fdd81; the repo is pushed). →
persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md - Daily gate batches: infra-hermes runs
scripts/wt-memory-gate-batchfrom 2026-10-01 and copies me on every verdict. The count is 1 of 3 consecutive PASS at off (20260930T090608Z); a FAIL restarts it. - ⚠ U11b STEP 5 IS MINE, triggered by the 3rd consecutive PASS:
- Run
docker exec -i <api> python - < scripts/wt-h2-count.pyVERBATIM, right before each instance's deletion. Exit 2 means STOP and send worldtree-dev the output. - Delete LIVE, with the api running, using literal paths only:
agents/{forseti,lofn,mimir}/memory/<agent>.chromaandmemory/context_promotion, on BOTH instances. - Send worldtree-dev the stamp; b193 ships after it.
- A b192 restart re-creates an empty schema-only
ledger.db. That is residue: say so in the stamp and remove it after b193.
- Run
- Legacy archive: DESTROY it whole by 2026-10-30, or at retirement-done, or on a subject-erasure request, whichever comes first. The runbook is in the detail file.
- TODO: re-sweep both api logs after real traffic. After the b193 push, remove the retired config keys.
fv-ml1 GPU layout (as of 2026-10-01)
- GPU 0: cyberprev (47.1 GB), gen-small (35.3 GB), voices (10.8 GB), parakeet-nemo (3.6 GB steady). Free 385 MiB, FULL.
- GPU 1: vllm-coder, erp-seat, meromero-rp, plus intern-decision (cap 14.4 GiB, 32k tokens, peak 15,220 of a 15,437 MiB budget). FULL.
- GPU 3: the full-size-seat reserve (Flash-Next is parked). On-demand tenants: Blender, and Scriberr (0 idle, ~5.5 GB per job). When a full-size seat claims GPU 3, Scriberr steps aside to irv-ml1's A6000, not back to GPU 1.
intern-decision (replaced SemIf on 2026-09-30)
- LIVE 0.1.3 at
intern-decision.fv.internal:8033: semif-compatible/decideplus Jev/v1/systemone, 32k tokens, a Triton cache volume. Runscripts/intern-decision-warmupafter an IMAGE change. →persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md - Open: label ~50 real Wyrd/Cicada turns before trusting it in production; its card makes no contamination claim.
Scriberr (fv-ml1 GPU 3)
- LIVE
scriberr:local-blackwell-a353078-dropout2: upstream a353078 plus patch 0001 (overlap slicer) and patch 0002 (gap retry,PARAKEET_MODEL_PATH), carried LOCALLY ONLY (Prime 2026-10-01: no upstream). v3 stays.scripts/scriberr-rebuildre-applies both. →persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md
nh3-pve + nh3-ml1: post-visit, all live (2026-09-25/26)
- nh3-pve: Secure Boot OFF; IGFX restored; NVIDIA 580.178.04 DKMS; kernel
6.8.12-43 (benign btmtk oops every boot). AMT live: static
10.100.250.61on nh3-mgmt (UDM port 6), KVM on, Opt-in None, password in the vault asnh3-pve/amt-admin. →servers/nh3-pve/README.md - nh3-ml1 (CT 109 @
10.100.50.80):- TEI embed/rerank is load-shared with esh-ml1 through the gateway, with
router_settings.enable_weighted_failover. - brokkr's foundry seats:
lfm-vl:8030(gatewaylfm25-vl-3b);lfm-vl-uncensored:8032(direct; passed brokkr's eval);vibevoice-asr:8031(audio.cpp, Q8_0).
- GPU ~11.3/16 GB. →
servers/nh3-ml1/README.md
- TEI embed/rerank is load-shared with esh-ml1 through the gateway, with
- Not yet proven: nh3-ml1 coming back by itself after an nh3-pve reboot (only the config was checked).
esh-matter: Matter server for Home Assistant (2026-09-26)
- CT 111 @
10.0.90.20, on VLAN 90 (esh-iot) only. matter.js 1.4.0;:5580is firewalled to HA10.0.50.46. HA's matter integration is loaded (ha-dev). - Next is Prime's: share the Aqara W200 into HA via Matter multi-admin. If
commissioning misbehaves, the UniFi mDNS reflector on esh-iot is the first knob
(a Prime/infra-ops change). →
servers/esh-matter/README.md
Worldtree reward path: code + config live (2026-09-27, Prime)
d8726a13(Domari's Skywork/Selene fix) was ALREADY on origin/main when Prime said to push it: it was pushed as part of79cce92fat 1221. Demo runs image79cce92f. Personal still runs51168a2e, which is older and lacks the fix; its image follows worldtree CI, never a manual up.deploy-wt-configdeployed demo and personal (1437):providers.yamlreward_models.scalar-judge(6d4ac44), plus a77639d's corrected selene description, which had never reached the hosts. Both are healthy, with 0 drift. It is inert on personal until that image picks up d8726a13.- worldtree-instance-configs: its 6 long-unpushed commits (53349f8…6d4ac44, Aug 2 to Sep 27) were pushed on Prime's go-ahead (a9d091e..6d4ac44). Origin, the repo and both hosts now agree.
Blender on fv-ml1 GPU 3, agent-driven (2026-09-27, Prime)
- Prime: "go ahead with gpu 3, both", and he does not use Blender, so agents drive it through MCP.
Blender 5.2.2 LTS (linuxserver/selkies image, digest-pinned) with a web desktop at
https://10.251.50.54:3001(vaultfv-ml1/blender-web-password). On demand only:scripts/blender-mcp up|down|status. It was left DOWN (GPU 3 back to 2 MiB). - MCP:
mcp-for-blender2.1.1 (Dvalin's research pick). The add-on is vendored at41a18432. The server runs INSIDE the container, and agents reach it as ssh+docker-exec stdio viascripts/blender-mcp. No port is published, and screenshots work because server and Blender share a filesystem. Telemetry is off and safe mode is on. Verified end to end (render, screenshot, safe-mode refusal). Batch path:scripts/blender-run(one-shotdocker run --rm,--jobstaging; first user is draupnir). Access: the shared fleetinfra-opslogin, with no render-only key (Prime, 2026-09-28). Registration + up/down: PER WORKING SESSION (Prime, 2026-09-28, relayed by draupnir; superseded per-task of 09-27):blender-mcp up+claude mcp addat session start,claude mcp remove+downat its end. Never always-on, never user- or project-wide config. →stacks/blender/README.md - Extensions (2026-09-28, draupnir; Prime ruled Blender a MANDATORY pipeline stage): 8 pinned
add-ons (
stacks/blender/extensions.lock) built byscripts/blender-extensions syncintofv-ml1:/tank/blender-extensions/5.2/system(LIVE), mounted read-only as the System repo;fleet_extensions.pyenables them (GUI startup timer;blender-run --extensions). Headless acceptance 8/9 (CAD Sketcher sketching is GUI-only); MCP 9/9 after Prime ran the deploy himself (1505; the classifier had refused mine). SurfacePsycho's eval() is patched to literal_eval (held under MCP). Open: agent-drawn CAD Sketcher geometry (its stateful ops want point picks), and MeasureIt overlays not seen in MCP screenshots. GUI left DOWN.
Zigbee2MQTT on esh-docker-vm (2026-09-27, Prime go-ahead; ha-dev request)
- LIVE since 1240: Z2M 2.14.1 (digest-pinned) on
http://10.0.50.45:8099(auth token), radio SLZB-MR1U chip 0tcp://10.0.90.10:6638(ember), channel 25, PAN 0xCFF4. State and the network key are in/opt/docker/data/zigbee2mqtt(root 0700, restic via /opt/docker, never in git). Vault:esh-docker-vm/zigbee2mqtt-{network-key,frontend-token,mqtt-password}. Broker userzigbee2mqttadded to mosquitto (passwd backup.bak-20260927-z2m). - ha-dev deleted the ZHA entry; pairing the Aqara T1 is theirs. ⚠ The "bridge in HA" acceptance first
FAILED because HA had had no MQTT since 2026-09-25 06:13. Cause: the esh-docker-vm macvlan shim had no
host route to HA (two equal /24s, ens18 wins). Fixed at 1249 with a /32 via the shim (if-up.d hook,
playbooks/esh-docker-vm-macvlan-shim-route.yaml), and HA reconnected. My first report called HA "connected" from a client count that included my own probe;$SYScounts are not identities. - ⚠ The first credential-mint attempt was blocked by the auto-mode classifier on a peer-relayed
approval. It went ahead only on Prime's own go-ahead in this session. Treat that as correct.
→
stacks/zigbee2mqtt/README.md
restic: credential leak fixed (2026-09-27, Prime)
- Seven hosts were moved from
env-filetorepository-file(playbooks/restic-repository-file.yaml), so the systemd units no longer carry the rest-server password. Secrets are vaulted as<host>/etc/restic/{repository,password}.restic.envis KEPT for manual snippets, so a rotation must update the vault,restic.envandrepository. - vm-esh-nas done too (Prime bootstrapped infra-ops there at uid 850, NOPASSWD). All 8 hosts are clean and vaulted.
- The passwords were readable until today, so the rotation (Prime's) remains the real fix.
- 2026-09-27 0859 freshness check: all restic and PBS repos ✅. Those 0100 runs PRE-date the move (done ~0140–0200).
- VERIFIED 2026-09-28 (infra-hermes, read-only): the first nightly on
repository-file(0100 PT) succeeded on all 8 hosts (unit Result=success AND a today-dated snapshot in each repo), and the 0800 freshness check was all-green, PBS included. nh3-dev also passed a content restore (af5580c3).
augaman: face recognition for Cicada
- v0.1.3 LIVE on esh-ml1:8040 (v0.1.1 first deployed 2026-09-26 2347 PT),
stacks/augaman, healthy on CUDA, innvidia-smi. Built on-box from agit archiveof the tag.pytest -m gpu tests/vision3/3 PASS on v0.1.3. - Gallery backup WIRED + RESTORE-VERIFIED 2026-09-27 (Prime OK'd off-site): restic
daily 0100 PT → rest-server-ana
esh-ml1/, fail-closed hook, secrets vaulted underesh-ml1/etc/restic/. Identity-level restore matched the live gallery (canary, snapshotfd3061a1). The backup gate for real enrollments is MET. - Deploy CLOSED by augaman-dev 2026-09-27 0009 PT: HTTP checks passed; the canary
survived the recreate and was then deleted, so the gallery is empty and ready for real
enrollments.
/recognizep50 186 ms (1080p, one face). The detector-latency follow-up is theirs. Their next release pins the container gid to 10001; nothing is owed by infra-ops. - The fv-ml1 instance was REMOVED by Prime on 2026-09-27 (~0130 PT) after the bench. esh-ml1 is the only instance: in the house, backed up, 48 ms per face on v0.1.3.
- v0.1.3 went live on both hosts (2026-09-27 ~0105 PT); the fv-ml1 instance was removed later. Speed bench (
docs/pfi/augaman-speed-bench/), server-side one face, v0.1.2 → v0.1.3: esh GPU 144 → 48 ms, fv GPU 75 → 27 ms. CPU mode REGRESSED (fv CPU6 152 → 205 ms; suspected ORT thread-pool spinning) and is reported to augaman-dev; the deployment does not use CPU mode.
esh-ml1
- The reward seat's consumer is FIXED: Worldtree d8726a13 speaks vLLM /classify through the
gateway
/scalar-judgepassthrough, which is key-gated per key viaallowed_passthrough_routes, granted by infra-hermes on 2026-09-27. It is live on demo; personal waits on a newer image (see "Worldtree reward path"). - The nh3-docker Dozzle agent has been stopped by hand since ~2026-04. Revive or drop.
- The Beszel superuser password was echoed into a session transcript (local only, not in memory). Rotation offered to Prime.
Live threads
- git: eshpfi-management pushed to
6b66207and worldtree-instance-configs tob6fdd81(2026-10-01 ~0418, Prime's go). Anything after that is unpushed. ⚠ The working tree AND index are shared with infra-hermes and subagents: commit withgit commit -- <paths>(auto-memoryfeedback_shared_git_index_commit_pathspecs).graphify-out/GRAPH_REPORT.mdstays modified and uncommitted on purpose: it is auto-regenerated. - nh3-dev root disk was cleaned 2026-09-30 1704 (uv prune, dangling images, old build cache): 86% → 82%. The Beszel 85% alert flaps near the line.
- Booth submit-all fix (Prime's report) is LIVE since 2026-09-27 1705, via booth-dev (booth
50bfc7b). Pushing it is booth's call, per Prime; it is not ours. - ESH has a single outside route (esh-scale on esh-pve). Noted, untracked.
Recent decisions
-
[2026-10-01]Prime: delete the bench leftovers, no upstream for Scriberr, push. DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; unified-en KEPT, the live seat mounts it), plus/tank/spikes/scriberr-slicer(including the private copies of Prime's recordings) and/tank/spikes/parakeet-ab. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed. -
[2026-10-01]irv-ml1 /storetank reclaim done: Prime ruled through comfy-dev, which deleted 84 files of its own (260 → 299 GB free). Tiers B/C/D got no ruling (infra-hermes thread01M3TCSYRSFNAPA9BPQTMFQ6KJ). -
[2026-09-30]Parakeet speech seat →parakeet-unified-en-0.6bunder NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. DONE 2026-10-01 0126 by infra-hermes; infra-ops audit passed 0137. →persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md -
[2026-09-30]Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30. →persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md -
[2026-09-30]SemIf replaced by intern-decision (Intern-Decision-4B, the Jev bench pick): semif-compatible plus Jev/v1/systemoneat 32k tokens on GPU 1, with a Triton warm-up cache volume. →persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md -
[2026-09-30]Scriberr moved to GPU 3 (on demand); our build carries the overlap slicer (0001) and the Parakeet gap retry (0002); v3 kept. →persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md -
[2026-09-29]Worldtree U11a prepped, not flipped: a staged demo config plus the agreed U8 window plan (infra-hermes runs the batches). →persistent-memory.d/2026-09-29-worldtree-u11a-prepped.md -
[2026-09-28]Worldtree U10 backfill done on demo (5) and personal (797). model_roles drift needed a memory_tagger sync first; mimir had missing vectors. →persistent-memory.d/2026-09-28-worldtree-u10-backfill.md -
[2026-09-28]Bonsai ternary vs Q4_K_XL at concurrency on the 275 W card (1.93x at N=1 falls to 1.06x at N=8, 1.21x with the MMVQ fix); the weights are acquired. →persistent-memory.d/2026-09-28-bonsai-ternary-spike.md -
[2026-09-28]blender-run gained--cpu(no GPU attached) and a fixed hostnamefv-ml1-blender(draupnir). The design stage never renders, so it stays off GPU 3. -
[2026-09-28]Blender extensions live in a read-only System repo built from a sha256 lock, enabled by a hook, opt-in for blender-run (--extensions). SurfacePsycho's eval() is patched to literal_eval (a proven safe-mode escape). →stacks/blender/README.md§ Extensions -
[2026-09-27]hermes-gateway restarted 0401 for highseat-dev (SVOS v2.1.12:propose_decisiongainedseat_up, and Hermes reads the plugin only at start). The plugin load was verified at file level; the end-to-end proof is Miranda's first seat_up card. Enablingzellij-fleet@Claudeat boot remains Prime's call. -
[2026-09-27]SemIf LIVE on fv-ml1 GPU 1 (semif-serve 0.1.2, Prime): wrapper + contract + 39 tests, 142/144 upstream parity, two card-only memory defects fixed. →persistent-memory.d/2026-09-27-semif-live-on-fv-ml1-gpu1.md -
[2026-09-27]Blender 5.2 on fv-ml1 GPU 3 (on demand), agent-driven via mcp-for-blender running in-container over ssh stdio; safe mode on, no published port. →stacks/blender/README.md -
[2026-09-27]esh-docker-vm host→HA traffic needs a /32 over macvlan-shim (if-up.d hook); without it HA silently loses MQTT (3× since August). →servers/esh-docker-vm/README.md -
[2026-09-27]Zigbee2MQTT live on esh-docker-vm :8099 (PAN 0xCFF4, ch 25) replacing ZHA; network key vaulted, data host-only. Prime go-ahead in this session; a peer-relayed approval was blocked by the permission gate. →stacks/zigbee2mqtt/README.md -
[2026-09-27]semif-serve 0.1.4: object states ending in),;or}no longer 422 (INV-7, a prefix wrapper proven at startup); numerics are deterministic within a process but a bf16 near-tie can flip across a restart. Prime ruled; heid bug hunt folded. →stacks/semif/README.md -
[2026-09-27]SemIf as Cicada's mood source: slower (+32 ms async, +94 ms sequential) and worse (67% vs 92% apt; carry 7/15 vs 14/15); only the gesture restraint is a win. Build nothing (Prime). Henge 88 carries it. →persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md -
[2026-09-27]SemIf consumer-fit spikes (Prime): Cicada affect gate 30/31 with descriptive wording and 19/31 terse; Wyrd "left this place?" 21/21 on the second wording, exit choice 18/21. Recommendations await Prime. →persistent-memory.d/2026-09-27-semif-consumer-fit-spikes.md -
[2026-09-27]SemIf order-averaging spiked (+9.1 pts accuracy, agreement = strong confidence signal); Prime ruled: build it in as 0.1.3 and trial the fast kernels. Tracked in the in-flight SemIf section. →persistent-memory.d/2026-09-27-semif-order-averaging.md— DONE: 0.1.3 live, with fast kernels adopted (77b8cb4). -
[2026-09-27]restic repository URL moved out of world-readable systemd units on all 8 hosts, and vaulted (Prime). Rotation stays Prime's. →persistent-memory.d/2026-09-27-restic-repository-file-fleetwide.md -
[2026-09-27]infra-ops bootstrapped on vm-esh-nas by Prime (the first account at the fleet-pinned uid/gid 850;bootstrap-infra-ops-user.yamlnow pins 850 when free,d775a01). -
[2026-09-27]augaman: second instance on fv-ml1 benched, then REMOVED (Prime). esh-ml1 on v0.1.3 does 48 ms/face in the house with the verified backup (ad484c3, benchdocs/pfi/augaman-speed-bench/). -
[2026-09-27]Created empty private repocorviduo/svoperatingsystem(SVOS, the operating system) as claude-bot, at brokkr-smithy-dev's relay of Prime's ruling.svos-dev→ the OS session; The High Seat's session ishighseat-devagain. -
[2026-09-26]augaman: infra-ops owes the gallery backup + first deploy — DEFERRED until augaman-dev reaches deploy. →persistent-memory.d/2026-09-26-augaman-infra-ops-owes-the-gallery-backup-first-deploy.md—[2026-09-27]DONE: deployed (v0.1.3 on esh-ml1), backup wired, identity-level restore verified (d8f59a1). -
[2026-09-26]zellij-fleet@.serviceinstalled, NOT enabled (for svos-dev's seat_up; owns the fleet zellij server in its own cgroup). Enabling@Claudeat boot is Prime's call when seat_up ships. Tracked:0ad7799,services/zellij-fleet/README.md. -
[2026-09-26]btusb blacklist on nh3-pve + esh-pve — proposed, awaiting Prime (low priority). Every 6.8.12-4x boot oopses in btmtk (benign). Also noted for the next NH3 visit: nh3-pve runs ~20 °C warmer than esh-pve (workload-confounded; check airflow). Tracked:servers/nh3-pve/README.md. -
[2026-09-26]GPU-LXC Temperature alerts watch the GPU, not the host CPU (SENSORS=-coretemp_*,acpitz); hypervisor CPU alerts at >95 °C on nh3-pve/esh-pve. The first nh3-ml1 alert was nh3-pve's CPU during vzdump.6d901c0. -
[2026-09-26]Abliterated LFM2.5-VL-3B parallel seat (:8032) passed brokkr's 6-image eval and stays (direct only; stock seat kept for SFW A/B).9cc3824,944bb36. -
[2026-09-25]AMT follow-ups PARKED until the two MS-03s arrive and esh-pve's MS-01 is on AMT (Prime 2326). — Parked: MeshCommander test; dummy HDMI plug, then NanoKVM → gx10; an OOB path not through nh3-scale… →persistent-memory.d/2026-09-25-amt-follow-ups-parked-until-the-two-ms-03s-arrive-and-esh.md -
[2026-09-25]nh3-pve AMT LIVE: static10.100.250.61on nh3-mgmt (UDM port 6), KVM on, Opt-in None — (Prime). →persistent-memory.d/2026-09-25-nh3-pve-amt-live-static-10-100-250-61-on-nh3-mgmt-udm-port.md -
[2026-09-25]pfi-gx10 AC-restore VALIDATED by Prime's AC pull; Homepage stays on esh-docker-vm (Prime) after the mmap_lock wedge reboot.servers/pfi-gx10/README.md,servers/esh-docker-vm/README.md. -
[2026-09-26]VibeVoice ASR → Q8_0 (Prime). WER on the 4 bundled LibriSpeech clips 3/69 → 2/69 (only the I'm/I am artifact left), ×3 identical; RTF 0.09–0.17; +1.1 GB VRAM (nh3-ml1 ~11.3/16 GB). Q4_K file removed. -
[2026-09-26]esh-matter LIVE: a Matter server (matter.js 1.4.0) on CT 111 @ 10.0.90.20, VLAN 90 only — , for ha-dev (operator-approved, relayed). →persistent-memory.d/2026-09-26-esh-matter-live-a-matter-server-matter-js-1-4-0-on-ct-111.md -
[2026-09-26]Embed/rerank LOAD-SHARED across esh-ml1 + nh3-ml1 (Prime). — Second deployments were added for qwen3-embedding and reranker (config) and for reranker-a3-bge-v2-m3 (DB… →persistent-memory.d/2026-09-26-embed-rerank-load-shared-across-esh-ml1-nh3-ml1-prime.md -
[2026-09-26]Two dataset-foundry utility seats LIVE on nh3-ml1 (brokkr; operator approval relayed): — LFM2.5-VL-3B on llama.cpp:8030(gatewaylfm25-vl-3b, LiteLLM restarted 36 s at 0039) and… →persistent-memory.d/2026-09-26-two-dataset-foundry-utility-seats-live-on-nh3-ml1-brokkr.md -
[2026-09-26]Coder seat STAYS on fv-ml1 (Prime). — The nh3-ml1 copy gave the same quality (teacher-forced true-code logprob diff +0.008 ± 0.019) but ran ~5×… →persistent-memory.d/2026-09-26-coder-seat-stays-on-fv-ml1-prime.md -
[2026-09-25]nh3-ml1 LIVE after the NH3 visit (SB off, IGFX restored): driver + CT 109 + TEI, parity vs esh-ml1 indistinguishable (cos min 0.999993 = own floor; overlap@10 1.000 vs MRL-256 control 0.684), same speed; monitoring/DNS wired. Autonomous: lxc-pve 6.0.0-2 upgrade on nh3-pve (Docker-in-LXC fix #7006). Gateway routing + AMT follow-up are Prime's. →persistent-memory.d/2026-09-25-nh3-ml1-live.md -
[2026-09-25]nh3-ml1 = CT 109 @ 10.100.50.80, second TEI embed/rerank backend on nh3-pve's RTX 2000E; GPU playbooks made host-generic — BUILD DEFERRED: blocked on nh3-pve Secure Boot until the NH3 visit; SB handling and gateway routing are Prime's calls. Tracked:7ddd116, Current state. →persistent-memory.d/2026-09-25-nh3-ml1-standup.md -
[2026-09-25]nh3-pve OOB: HOLD blind until the next NH3 visit (Prime via Miranda, 1333). Target: NanoKVM → pfi-gx10, and the MS-01 on its own vPro. The IGFX BIOS fix rides the same visit, since AMT KVM captures only the iGPU. Tracked:servers/nh3-pve/README.md(2b49be4,cdd7605). -
[2026-09-25]MS-01 foot-gun: a GPU in the PCIe slot renames every NIC — (the slot's root port takes bus 01, so the X710 goesenp2s0f0np0→enp3s0f0np0). →persistent-memory.d/2026-09-25-ms-01-foot-gun-a-gpu-in-the-pcie-slot-renames-every-nic.md -
[2026-09-25]Direct ESH→esh-ml1 consumer path — PARKED (no ESH-side callers in 7 days; consumers go through the gateway). Tracked atpersistent-memory.d/2026-09-24-esh-ml1-embed-rerank.md. -
[2026-09-25]Created empty private repocorviduo/norn(git@gitea.phasefinal.com:corviduo/norn.git) as claude-bot, at brokkr-smithy-dev's relay of the operator's Norn ruling; thenorn-devhandle is the operator's to declare. -
[2026-09-25]Reward seat audited (Skywork-Reward-V2-Llama-3.1-8B still #1 of 188 on RewardBench 2; our AWQ ≈ bf16 within noise; double BOS costs ~2.7 pts) and moved fv-ml1 → esh-ml1; esh-ml1 monitoring wired (Beszel+GPU, Kuma, Homepage, Dozzle); Dozzle hub's stale agent IPs fixed; nh3-dev Beszel agent revived. →persistent-memory.d/2026-09-25-reward-move-and-esh-ml1-monitoring.md -
[2026-09-25]TEI adopted as the fleet embed/rerank engine (Prime); esh-ml1 is the sole gateway backend forqwen3-embedding+reranker; fv-ml1's vLLM embed/rerank seats retired (~6.1 GB freed); parakeet stays on fv-ml1. →persistent-memory.d/2026-09-25-tei-fleet-embed-rerank.md -
[2026-09-24]esh-ml1 built — RTX 2000E Ada as CT 110 on esh-pve (host driver 580.178.04 DKMS, loaded live, no reboot), vLLM embed+rerank at measured parity with fv-ml1, wired as LiteLLMorder: 2failover; dead DB aliasreranker-a3-bge-v2-m3(pre-relocation IP) repaired. →persistent-memory.d/2026-09-24-esh-ml1-embed-rerank.md -
[2026-09-24]esh-pve: VM 102 retired, a 14-day hung VFIO process found and cleared, T400 → RTX 2000 Ada; Prime chose LXC + host driver for embed/rerank (implementation deferred to next session, tracked here +servers/esh-pve/README.md). →persistent-memory.d/2026-09-24-esh-pve-vm102-retired-gpu-swap.md -
[2026-09-24]pfi-gx10 AC-restore UEFI patch applied but NOT validated; box OFF until Prime's AC pull 2026-09-25 (tracking:servers/pfi-gx10/README.md,e5197a3). —[2026-09-25]✅ VALIDATED: Prime pulled and replugged AC, and it came up by itself at 1108:54 (2b49be4). →persistent-memory.d/2026-09-24-gx10-ac-restore-patch.md -
[2026-09-24]NH3 power outage recovered — pbs-nh3 had noonboot(set), NFS boot race fixed with automount (1cbde50), every other Claude session on nh3-dev died. →persistent-memory.d/2026-09-24-nh3-power-outage-recovery.md -
[2026-09-24]Miranda standing order is a repo CLAUDE.md operating parameter — (4b29492, aligned to the global send protocol inbcf3342): high-urgency matters go to her, fixed or not… →persistent-memory.d/2026-09-24-miranda-standing-order-is-a-repo-claude-md-operating.md -
[2026-09-24]task-board mothballed (Prime): container removed on ana-docker, data/image/compose kept, Kuma monitor deleted, task_* instructions removed from CLAUDE.md and the fork template (e6da607). Its hooks had sent no traffic in 30 days. -
[2026-09-24]Military 24-hour Pacific clock times carried into the Codex/Grok shared bootstrapdocs/fleettools/AGENT-BOOTSTRAP.md(ad2b4d9); Claude seats get it from the global CLAUDE.md. -
[2026-09-24]Worldtreeadmin.memory.forgetstays OFF on demo/personal until an instance needs it — a destructive erase; enabling it is a per-change operator yes via deploy-wt-config. -
[2026-09-24]A git checkout underroot:dockerneedssafe.directoryfor its deploy user — the 09-14 normalization (826a63b) silently broke yt-voice-clipper's webhook deploy until v0.3.13; fixed on irv-ml1 and recorded in fleet conventions (eea9eb2). Sweep found no other case. -
[2026-09-23]elway sudo uploads land root:root, validated and staged; fleet ownership audit built; 33 mis-owned root files fixed on 10 hosts. →persistent-memory.d/2026-09-23-elway-ownership-fix-fleet-audit.md -
[2026-09-23]headscale-ddns hardened — (fedd4b6): Cloudflare calls retry and validate every body (pick()), no write without both IDs, the run… →persistent-memory.d/2026-09-23-headscale-ddns-hardened.md -
[2026-09-23]esh-docker-vm restic was skipped 09-22..23 by my own Kuma move — a dead uptime-kuma lookup abortedpre-backup.shunderset -e(25e41d2). →persistent-memory.d/2026-09-23-esh-docker-vm-restic-was-skipped-09-22-23-by-my-own-kuma.md -
[2026-09-23]hermes-gateway restart exit-1 is a Hermes race, not a crash — the planned-stop watcher consumes the marker before SIGTERM re-runs the handler; drop-in… →persistent-memory.d/2026-09-23-hermes-gateway-restart-exit-1-is-a-hermes-race-not-a-crash.md -
[2026-09-23]Booth link board cleared to 14 durable links (208 removed: drops, release posts, research links, token-bearing URLs; Scriberr reposted atscriberr.fv.internal:8080). Prime's rule: durable debug links only. -
[2026-09-22]Both carried calls approved — build theNRestartsflap sampler (163bb97); the restic content-assertion ruling is ratified and stays (ba60fda). -
[2026-09-22]safe-rm installed on nh3-dev and delegated fleet-wide to infra-hermes. — ⚠ The package installs INERT and looks fine — Debian's/etc/zsh/zprofilehas 0 non-comment lines so the… →persistent-memory.d/2026-09-22-safe-rm-installed-on-nh3-dev-and-delegated-fleet-wide-to.md -
[2026-09-22]The acceptance probe for a guard must not be able to destroy what it tests — (infra-hermes). →persistent-memory.d/2026-09-22-the-acceptance-probe-for-a-guard-must-not-be-able-to.md -
[2026-09-22]⚠{"sent": true}is a claim about transmission, never about effect.pane_sendstructurally cannot deliver a harness command (every relay is prefixed with the card id), so D-0010 promised an unachievable/clearand its receipt reported success. Consumed an operator approval. →persistent-memory.d/2026-09-22-instrument-errors.md -
[2026-09-22]D-0010/D-0011 were misrouted to this seat bypane_findmatching a ROLLING PANE TITLE. — Genuine and operator-approved, wrong seat;fleet_telemetryheld the right mapping and carries the warning… →persistent-memory.d/2026-09-22-d-0010-d-0011-were-misrouted-to-this-seat-by-panefind.md -
[2026-09-22]headscale-ddnsexited 1 silently and the alarm carried no cause — both failure paths were|| exit 1with no message. Now every exit names its reason and the WAN lookup retries 3×. The script was also untracked; a fix to what every mesh client resolves through lived on one disk.30517fd. -
[2026-09-22]⭐⭐ Five instrument errors in one day, and the shape is one thing: a tool that enumerates "things that are fine" has selected against its own subject.--state=runningskipped the units most needing hooks;awk '{print $1}'dropped systemd's●-decorated FAILED rows;grep -ic resticon the wrapper missed the check script;restic ls's header line made an absent path read as "blobs gone"; and a cooldown test that invoked before failing the unit. Every one reported cleanly while looking at the wrong thing. The rule is not "verify" — it is verify, then ask what the verification could not have seen. →persistent-memory.d/2026-09-22-instrument-errors.md -
[2026-09-22]Fleet alert bridge generalized —beszel-althing→althing-alert-bridge, route registry (/beszel+/kuma), each with its own… →persistent-memory.d/2026-09-22-fleet-alert-bridge-generalized.md -
[2026-09-22]Uptime Kuma rebuilt from scratch on 2.5.5, moved esh-docker-vm → ana-docker. — ⚠:latestis a TRAP — it tracks 1.x, so an Aug-2026 pull gave a Dec-2024 build. →persistent-memory.d/2026-09-22-uptime-kuma-rebuilt-from-scratch-on-2-5-5-moved-esh-docker.md -
[2026-09-22]Beszel and Uptime Kuma are DISJOINT, not redundant — Beszel's alerts bind to a system with a threshold; there is no URL column, so it is structurally incapable… →persistent-memory.d/2026-09-22-beszel-and-uptime-kuma-are-disjoint-not-redundant.md -
[2026-09-22]Failed-START alarms on 23 nh3-dev units — (services/althing-notify-failure/). →persistent-memory.d/2026-09-22-failed-start-alarms-on-23-nh3-dev-units.md -
[2026-09-22]Backup coverage is a property of the SYSTEM, never one job's scope — establish it by querying the repo for the path in a real snapshot, never by reading a job'sSRC=. →persistent-memory.d/2026-09-22-backup-coverage-is-a-property-of-the-system-never-one-job-s.md -
[2026-09-22]restic checks now assert CONTENT and are DISCOVERED not enumerated — conjunctive (recent AND present AND restores non-zero bytes); repos discovered from each NAS, hand list… →persistent-memory.d/2026-09-22-restic-checks-now-assert-content-and-are-discovered-not.md -
[2026-09-22]irv-ml1: Irvine is a TENANCY behind a Fortinet PFI does not control — its TLS inspection breaks Tailscale relay/control intermittently (41 cert warnings/week). →persistent-memory.d/2026-09-22-irv-ml1-irvine-is-a-tenancy-behind-a-fortinet-pfi-does-not.md -
[2026-09-21]⭐⭐⭐ lv-mccarthy GATED and NOT SHIPPED — voice passes, memorisation is the cleanest in the line, and the length discipline is gone. 720 generations, 3 arms. Voice +0.152 at 2.9× floor for ckpt450 (ckpt900 +0.172 at only 1.2×, its spread one outlier seed), and ~3/4 of the gain survives stripping every punctuation mark, so it is not the cheap win. ⭐ Memorisation: ckpt450 at 0.12 against the author's own held-out 0.12 — identical, longest match 11 words against the author's coincidental 12, all 96 matches READ and every one stock grammar (he looked at the wolf and he looked at him); the name-shaped hits are the RENAMED inventions. Axis C is the blocker: 20% / 28% of generations overshoot the 90–140 band against base's 1%, worst case a degenerate loop at 279 words. ⚠ My own prereg's axis C transcribedscore_beats.py's v1 criteria including "in-band up on base", which the operator RETIRED 2026-09-15 because base maxes it — under the v2 (ran-on only) ckpt450 passes by 0.01 against a 0.200 floor, but that reading was found AFTER the numbers and was not used. Fix the prereg prospectively. →persistent-memory.d/2026-09-21-lv-mccarthy-gate.md -
[2026-09-21]The two-epoch recipe is now 0 for 3, and this time the loss curve was CONFIDENTLY wrong. — On Brontë and Hemingway the epoch-1/epoch-2 checkpoints were TIED on eval loss, so preferring the earlier one… →persistent-memory.d/2026-09-21-the-two-epoch-recipe-is-now-0-for-3-and-this-time-the-loss.md -
[2026-09-21]The memorisation control that shipped broken is now committed, and it turned a 12× red flag into a clean pass. →persistent-memory.d/2026-09-21-the-memorisation-control-that-shipped-broken-is-now.md -
[2026-09-21]My confound detector was counting apostrophes as quote marks, and I saw it FIRE before I saw the bug. →persistent-memory.d/2026-09-21-my-confound-detector-was-counting-apostrophes-as-quote.md -
[2026-09-21]⭐⭐⭐ The ops log shipped — and then failed FOUR ways in its first hours, every one recording something unfindable. Claim dropped by a sub-tool, hook behind graphify's eightexit 0s, no handle in the env, ssh-target written as a hostname. The general shape is configured ≠ effective; twelve instruments reported confidently and wrongly across three days, five of them mine. →persistent-memory.d/2026-09-21-ops-log-and-the-instruments-that-lied.md -
[2026-09-21]⭐⭐ The Booth gained blur + a closed keep round trip, after shipping TWO controls that did nothing — a reveal handler Jinja discarded for sitting after{% endblock %}, and a×a sibling form covered by 30×22 px. Both operator-found, both from reading templates instead of rendering them.scripts/layout-probe.pytook four iterations to become trustworthy. ⚠ Blur is COSMETIC and a test asserts the 200 on purpose. →persistent-memory.d/2026-09-21-booth-two-dead-controls.md -
[2026-09-21]⭐⭐ claude-bot is an Owner in corviduo/pfi/vastblue; cicada + draupnir moved topfi; vh-token use is now standing-authorized from the vault. ⚠vhis a USER not an org, so no namespace-scoped admin exists — the only realization of the original ask was site-admin, surfaced rather than executed. ⚠⚠~/.config/claude-bot/gitea-tokenis DEAD and had been misreporting permissions; the working one isgitea-token-repo-create. →persistent-memory.d/2026-09-21-gitea-orgs-and-the-dead-token.md -
[2026-09-21]⭐⭐ nh3-dev 84%→72%: 27 GB reclaimed, and a 252 MB LoRA adapter rescued from/tmp, which this box sweeps at 3 days.babyyarrosexisted nowhere else; hash-verified to smithy before any deletion. ⚠⚠ My 7-day session prune then deleted an ACTIVE session's dir — directory mtime does not reflect subdirectory writes. Do not re-run that predicate. →persistent-memory.d/2026-09-21-disk-triage-and-the-rescued-adapter.md -
[2026-09-21]Draupnir geometry engine provisioned on irv-ml1 and acceptance-tested — (34c4179,e574b91,2e08edc). →persistent-memory.d/2026-09-21-draupnir-geometry-engine-provisioned-on-irv-ml1-and.md -
[2026-09-21]Backup alarm verdict split:STALE(exit 1) vsERRORED-JOBS(exit 3) — (7fe4102). →persistent-memory.d/2026-09-21-backup-alarm-verdict-split-stale-exit-1-vs-errored-jobs.md -
[2026-09-21]vh/forgefirmmirrored — fromgithub.com/openglow-org/forgefirm, following the house convention read off the existing 17:vh/… →persistent-memory.d/2026-09-21-vh-forgefirm-mirrored.md -
[2026-09-20]ravenpen.comREGISTERED by the operator — the 09-18 hold is discharged and hamr-dev is answered. — infra-ops deliberately did NOT execute this twice, because it was a non-refundable purchase reaching us as a… →persistent-memory.d/2026-09-20-ravenpen-com-registered-by-the-operator-the-09-18-hold-is.md -
[2026-09-19]FV is a dedicated 20 A circuit carrying fv-ml1 AND the R420 running OPNsense, and nothing else (operator-confirmed). →persistent-memory.d/2026-09-19-fv-is-a-dedicated-20-a-circuit-carrying-fv-ml1-and-the-r420.md -
[2026-09-19]althing 3.7.0 rolled out to both infra-ops surfaces — and the rollout broke the claim tooling built that morning. →persistent-memory.d/2026-09-19-althing-3-7-0-rolled-out-to-both-infra-ops-surfaces-and-the.md -
[2026-09-19]Three agents commit as one git author, and closing that gap took three instruments to get right. — An unattributable commit (e43e262) appeared in the push set between two of mine — unidentifiable from git… →persistent-memory.d/2026-09-19-three-agents-commit-as-one-git-author-and-closing-that-gap.md -
[2026-09-19]The backup alarm had no wire for three weeks, and fixing it exposed two more instruments that pass without looking. →persistent-memory.d/2026-09-19-the-backup-alarm-had-no-wire-for-three-weeks-and-fixing-it.md -
[2026-09-19]elway evaluatedwhen:/creates:/removes:/changed_when:WITHOUT the step's sudo, and it fails silently in the dangerous direction. →persistent-memory.d/2026-09-19-elway-evaluated-when-creates-removes-changedwhen-without.md -
[2026-09-19]ESH VM 102 (esh-vm-workstation) excluded from the nightly backup job — operator ruling. — It is a Windows 11 Parsec/RDP sandbox (no password, no state to recover), and its vzdump had failed… →persistent-memory.d/2026-09-19-esh-vm-102-esh-vm-workstation-excluded-from-the-nightly.md -
[2026-09-19]The ops log is BUILT —scripts/ops-log, automatic writers, and a detector for the path they cannot cover. →persistent-memory.d/2026-09-19-the-ops-log-is-built-scripts-ops-log-automatic-writers-and.md -
[2026-09-19]infra-hermesis this session's ASSISTANT, and the division of labour is now standing policy. — infra-ops keeps improving infrastructure tooling plus the hard calls; infra-hermes does **day-to-day… →persistent-memory.d/2026-09-19-infra-hermes-is-this-session-s-assistant-and-the-division.md -
[2026-09-18]⭐⭐⭐ NH3↔Anaheim had been running over a throttled DERP relay, not a direct path — 78 GB of fleet traffic on someone else's free infrastructure. Four additive objects on ana-gw gave ana-scale a stable inbound UDP 41641 endpoint;tailscale ping373–522 ms → 6 ms direct, cross-site HTTP 1.2 s → 0.015 s, STT via the ANA gateway 1.4 s → 0.25 s. ⚠ That box runscentral-nat, so a policydstaddris the REAL internal address, not the VIP. No OOB access — back up withshowto a local file and make additive changes ONLY. irv-ml1 still relayed. →persistent-memory.d/2026-09-18-nh3-ana-derp-relay.md -
[2026-09-18]⭐⭐⭐.internalDNS was failing ~10% of lookups fleet-wide, from two independent causes. A PUBLIC resolver was the fallback for a PRIVATE zone (Cloudflare answers NXDOMAIN authoritatively, so a transient miss became a hard failure) — now a cross-site ring, each site local-first with a different site as backup. Then the root cause: all three AdGuards shippedratelimit: 20shared across an entire /24, silently dropping queries at 5 s each. Set to 0. Hard failures 3/40 → 0/40; burst timeouts 40/60 → 0/60. ⚠resolv.confis DHCP-managed — change it at the UDM/FortiGate, not the file. →persistent-memory.d/2026-09-18-fleet-dns-ring-and-ratelimit.md -
[2026-09-18]⭐⭐ SearXNG had ONE working general web engine and every health check said fine. 7 of 55 were enabled-by-default and six of those are dictionary/translation engines —inactive: falseonly makes an engine SELECTABLE,disabled: falseputs it in the DEFAULT set. Now seven. ⭐ This stack tracks:latestON PURPOSE (upstream ships engine-handler fixes continuously; a pin freezes breakage). ⭐ The Brave key is committed in plaintext by explicit operator decision — scoped to one low-value credential, NOT a change to the no-secrets rule. →persistent-memory.d/2026-09-18-searxng-one-engine-to-seven.md -
[2026-09-18]⭐⭐ althing search returned ZERO for every hyphenated query — silently, for every handle, and on this fleet that is most of our hostnames. FTS5 read the hyphen as a column filter;search()caught the error, its probe passed, and it returned[]. Routed to forseti (they own the code, I own rollout) → v3.6.3 deployed, 17 s bus outage. The reported symptom, a message that never arrived, was NOT a defect: the reporter's evidence file was truncated at 8 KiB. ⚠ No CI builds the post-office image. →persistent-memory.d/2026-09-18-althing-363-hyphen-search.md -
[2026-09-18]⭐ FleetTools: one capability index, autoloaded by Claude, Codex and Grok from a single symlinked file.docs/fleettools/+~/FLEETTOOLS.md; absolute detail paths because a non-Claude agent cats them. Rule zero is query-live-inventories-never-a-written-list. The global CLAUDE.md tools section went 231 lines → 39, keeping only the three rules that govern behaviour rather than lookup. →persistent-memory.d/2026-09-18-fleettools-agent-index.md -
[2026-09-18]Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy — (operator-approved). →persistent-memory.d/2026-09-18-miranda-s-hermes-plugin-install-is-now-a-symlink-to-the.md -
[2026-09-18]Worldtree'senv.shsecrets are vaulted — 10 entries underworldtree/(gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal… →persistent-memory.d/2026-09-18-worldtree-s-env-sh-secrets-are-vaulted.md -
[2026-09-17]dragonfireacoustics.comexpires 2026-10-30 — six weeks — at eNom with NO transfer lock, and its sibling domain was already lost exactly this way. →persistent-memory.d/2026-09-17-dragonfireacoustics-com-expires-2026-10-30-six-weeks-at.md⚠ If it is ever transferred, DNS does NOT come with the registration — the nameservers are eNom'sname-services.comand the zone must be recreated first or mail dies. The whole zone is two facts plus a landmine:*(WILDCARD) → 199.250.192.76 which is dead (no HTTP at all, and it is what the apex/mail/webmail/admin/ftp all answer with),www→ 38.120.12.45 (us), and 7 Google Workspace MX records that must not be lost. No DNSSEC (delegationSigned: false), so no transfer complication. ⚠ Also found: no SPF and no DMARC at the apex on a Google Workspace domain — a live deliverability problem independent of everything else. Transfer gate is the TAC/EPP code from the eNom account, not the lock; the missing lock is not authorization. 60-day rule is satisfied (last changed 2025-10-24). -
[2026-09-17]dragonfireacoustics.comIS configured onpfi-ana-webhost, and the whole thing is dead — a forgotten public-facing VM. →persistent-memory.d/2026-09-17-dragonfireacoustics-com-is-configured-on-pfi-ana-webhost.md -
[2026-09-17]PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the… →persistent-memory.d/2026-09-17-parked-lv-krakauer-and-the-reason-is-a-selection-criterion.md -
[2026-09-17]The althing route-declaring SessionStart hook is documented but NOT installed on nh3-dev —dev_launch.pyhas zero occurrences of "route", no hook declares one, and every live route was hand-declared… →persistent-memory.d/2026-09-17-the-althing-route-declaring-sessionstart-hook-is-documented.md -
[2026-09-17]PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the… →persistent-memory.d/2026-09-17-parked-lv-krakauer-and-the-reason-is-a-selection-criterion-2.md -
[2026-09-15]⚠⚠--gpu-memory-utilizationDOES NOT PREDICT RESIDENT VRAM — measure it, never compute it. Wrong in both directions on fv-ml1:vllm-cyberprevutil 0.40 (expect ~39,155 MiB) holds 47,124 (+8 GB over);vllm-gen-smallutil 0.48 (expect ~46,986) holds 36,942 (−10 GB under). Planning a placement off the fractions would have been 8 GB wrong. Readnvidia-smi --query-compute-apps. Full per-seat residency table + the breeze shuffle arithmetic →persistent-memory.d/2026-09-15-breeze-placement-sizing.md -
[2026-09-15]breeze-tts stays on irv-ml1; the TTS-stack move to fv-ml1 is PARKED (park id 75,move-the-tts-stack-breeze-tts-bragi-tts-gateway), triggered on evacuating embed/rerank/reward. ⚠ Trigger as stated says "gpu0" but those three are on GPU 1 (~0.16 util, ~15.7 GB; GPU 1 is the tight card at 0.975 / 4,336 MiB free) — confirm which he meant before executing. All three services move together because onlybreeze-ttsis GPU-resident (~10.3 GiB, growing) whilebragiandtts-gatewayare CPU proxies, and co-location is what avoids a cross-site hop per TTS call. breeze-tts sizing — original recommendation NOT to move it. ~10.3 GiB measured under load at 53 min uptime, up from 9.2 GiB shortly after warm-up (it grows; n=2, plateau unmeasured) — so GPU 0's 11,982 MiB free is a 1.7 GB margin and shrinking, on the live chat serving path. ⚠ Two measurement traps: it reports nothing at idle on the wrong card (BREEZE_GPU_DEVICES=0= the 3090, not the A6000), and an early reading understates it. ⭐ The real objection is topology:tts-gatewayis on irv-ml1 and reaches it same-box, so moving breeze alone adds a cross-site hop to every TTS call against a 478 ms first-sample budget. GPU 3 would fit it but spends the reserve. →persistent-memory.d/2026-09-15-breeze-placement-sizing.md -
[2026-09-15]Hermes bearer rotation hold released — svos-dev split their HS256 signing key off the shared value (svos7165272) →persistent-memory.d/2026-09-15-hermes-bearer-rotation-hold-released-svos-dev-split-their.md -
[2026-09-13]STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint. →persistent-memory.d/2026-09-13-standing-policy-operator-cap-gpu-power.md -
[2026-08-25]Fused MoE kernel path — DEFERRED, tracked at parkfused-moe-kernel-path-for-gemma-4-moe-training(id 47). →persistent-memory.d/2026-08-25-fused-moe-kernel-path-deferred-tracked-at-park-fused-moe.md -
[2026-08-24]nconnect=8on/mnt/smithy— approved but DEFERRED at operator instruction. brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread01M0R46SFYF83099N16WD67KGD.
167 older entries archived to archival-memory.md.
Tried and abandoned
-
[2026-10-01]Size + mtime scans to pick reclaimable "staging" on a model store (infra-hermes, irv-ml1)._inbound/retro-diffusionlooked like staging leftovers but is the TARGET of 47 model-store symlinks, holding a purchased bundle used by a live commission. Resolve symlinks INTO any staging dir before proposing to delete it. -
[2026-09-30]Shorter Parakeet slices as Scriberr's memory fix. I predicted ~6× less memory from the attention-matrix arithmetic. Measured, 300 → 120 s only went 9,384 → 6,510 MiB: a ~5.6 GB fixed floor dominates.expandable_segments:Truewas the real lever (5,496). Measure the process peak; never extrapolate it from one tensor. -
[2026-09-30]Whole-file local-attention Parakeet in Scriberr (context 255/255). OOM past 16 GB on a 35-min file. Local attention inside chunks is also non-deterministic run to run. -
[2026-09-30]Start-time midpoint stitching of overlapped Parakeet chunks. It duplicated a word at 26 of 108 stitches, because Parakeet timestamps a post-pause word anywhere inside the pause. Hand over at a word both chunks agree on instead. -
[2026-09-30]int8 ONNX (sherpa-onnx) as the low-latency Parakeet runtime. The int8 graph runs on ONE CPU thread with the GPU at 2–9%. Unified-en int8 was slower than the seat; fp32 ONNX was 4–12× faster, and NeMo was fastest. -
[2026-09-30]GPU budgets computed as total − used. nvidia-smiFreeis ~640 MiB lower per card (driver reserve). Budget fromFree. -
[2026-09-27]The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel. Every tool worked exceptget_viewport_screenshot, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. →stacks/blender/README.md -
[2026-09-27]A TCP connect as the "is Blender ready" probe. docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on toping. The same trap applies to any service behind a published port. -
[2026-09-27]log.exception()in a GPU failure path — the record keepsexc_info, so any retaining handler (pytest's capture does) pins the traceback's frames and tensors. Logtraceback.format_exc()text instead (semif-serveengine._guard). -
[2026-09-27]fla/triton in a slim image without gcc — triton compiles its CUDA driver shim at runtime ("Failed to find C compiler"). The warm-up died and startup failed closed. The semif image now installs gcc + libc6-dev with the fast extra. -
[2026-09-27]uv export→ requirements.txt →uv pip installfor a torch-from-cu128 project — the export drops uv's per-package index routing; with--extra-index-url … --index-strategy unsafe-best-match, triton came from the pytorch index and failed the lock's hash. Keepuv syncand blank the project version in a deps-only stage instead (services/semif-serve/Dockerfile). -
[2026-09-27]Translating a CUDA OOM by raising insideexcept(orfrom exc) — the chained exception's traceback pins the failed call's frames, and their GPU tensors (11.9 GiB after the 503). Raise after the block, unchained,gc.collect()beforeempty_cache(). -
[2026-09-27]deploy-stack.sh --yes/dns-sync.pyfed a blindy— the harness refuses a blind apply. Review withecho n |first, then apply withecho y |. -
[2026-09-27]Relying on the build cache surviving on esh-ml1 — the augaman v0.1.3 build missed its dependency layer there (fv-ml1 hit it; reason unexplained, suspect thedocker builder pruneruns) and the rootfs peaked at 90%. Check for ≥15 GB free, or remove the old image, before building there. -
[2026-09-26]Greedy exact-match as parity for a small generative seat (coder) — same-host repeats matched only 30–38% (run B hit the prefix cache, so a different numeric path; near-tie… →persistent-memory.d/2026-09-26-greedy-exact-match-as-parity-for-a-small-generative-seat.md -
[2026-09-26]zsh traps in ad-hoc test loops:set -- $adoes NOT word-split, so alert POSTs went out with empty fields; andlocal path=…inside a function CLOBBERS$PATH(zsh's tied array), so the failover test ran nothing while TEI sat stopped. Use other names and${=var}. -
[2026-09-25]Concluding "not on the UDM" from a port table read 15 s after link-up — UniFi polls (~60 s); AMT was on port 6 as Prime said. Wait a poll, and prefer a which-port-lists-the-MAC check. auto-memoryfeedback_polled_stats_lag_the_event. -
[2026-09-25]The esh-pve NVIDIA DKMS recipe on nh3-pve — it built, then the kernel refused to load it: nh3-pve has SECURE BOOT ON (esh-pve, the same MS-01, has it off). The installer rolled back. The playbook now pre-flightsmokutil --sb-state(7ddd116). -
[2026-09-25]elway against an unpinned host-key name — everywhen:hit ssh rc 255, so all 9 steps showed SKIPPED, not failed. Only the verify phase went red. Pin the name (the IP key matched) or use the IP. Untracked elway defect: rc 255 inwhen:should fail. -
[2026-09-24]Testing "Restore AC Power Loss" with an OS shutdown — a shutdown stays off BY DESIGN; only pulling and restoring AC tests it. A community README claimed the two are indistinguishable; trusting it cost Prime two trips to the pfi-gx10 power button. -
[2026-09-23]booth link --help— there is no help flag; it posts--helpto the operator's link board as a link. Readboothwith no args for usage. -
[2026-09-21]Using directory mtime as a liveness test when pruning session scratchpads —find /tmp/claude-1000 -mindepth 2 -maxdepth 2 -type d -mtime +7deleted an ACTIVE session's working… →persistent-memory.d/2026-09-21-using-directory-mtime-as-a-liveness-test-when-pruning.md -
[2026-09-18]Routing SearXNG's egress through a SOCKS5 proxy on esh-scale — one day live, reverted. It fixed nothing, and the reason I gave for reverting it was itself wrong: the rollback was the counterfactual and it falsified my own published claim. Kept reverted on its own merits (no measurable gain, added a hard ESH dependency for all fleet search). →persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md -
[2026-09-18]api_key: !ENV SEARXNG_BRAVE_API_KEYin searxng settings — this build has NO!ENVYAML constructor, so the file was unparseable and the container crash-looped ten… →persistent-memory.d/2026-09-18-apikey-env-searxngbraveapikey-in-searxng-settings.md
118 older entries archived to archival-memory.md.