memory: snapshot — Zigbee2MQTT + HA-MQTT route fix, Blender on GPU 3 (MCP per task + blender-run), semif 0.1.4, Worldtree reward config live and pushed; 3 entries archived; next: 2026-09-28 restic freshness check

This commit is contained in:
vh
2026-09-28 09:57:41 -07:00
parent 6f0c480693
commit 1756b891c6
5 changed files with 38 additions and 31 deletions
+27
View File
@@ -10629,3 +10629,30 @@ Canonical configs/restic/esh-vm-db and playbooks/esh-vm-db-restic-repair.yaml.
No DB restarts or auth-policy changes; changes saved locally, no commit. No DB restarts or auth-policy changes; changes saved locally, no commit.
_Archived 2026-09-26._ _Archived 2026-09-26._
# `[2026-09-13]` FV SITE DARK — every Fountain Valley address including the BMC went unreachable ~2.5 min into a two-card load
⚠⚠⚠ **FV SITE DARK — every Fountain Valley address including the BMC went unreachable ~2.5 min into a two-card load test; all other sites healthy. Operator's leading hypothesis: the 1500 VA Eaton UPS overloaded and DIED.** It fits better than a breaker trip because a UPS's output rating sits far below the circuit's, making it the first protective device to give — which explains why the site let go at **two** cards loaded rather than four, and why the ~25 W OPNsense box died with it. ⚠ **Will not self-recover** (tripped needs a human, dead needs replacing) — do NOT poll FV. ⚠ **Do NOT use surge-only outlets to exceed a UPS rating**: both banks share one NEMA 5-15P inlet rated 12 A total; the surge bank bypasses the inverter, not the current limit. ⭐ **Recover `/tank/aimodels/flash-next-mtp-bench/power.log` FIRST** — all four cards every 10 s to the cut, on `/tank` not in a container, and the ONLY load measurement that exists. ⚠ **19 of 30 gateway aliases dark and NO local fallback** — every free local model was on fv-ml1, irv-ml1 runs no chat seat at all; the only non-fv chat backends are paid, and any coverage must be a NEW opt-in alias, never a silent repoint. ⭐ **OOB design gap**: OPNsense-as-subnet-router covers box-down/gateway-up and nothing for a site-wide loss, since the BMC's only route out is that gateway. ⚠ Recovery hazard: ten `restart: unless-stopped` vLLM containers will all load at once on power-up — mask Docker first, then `compose up -d` seat by seat (which also finishes the stale-homepage-label fix, since labels attach only at creation). → `docs/runbooks/fv-site-dark-20260913.md`, `persistent-memory.d/2026-09-13-flash-next-seat-and-fv-outage.md`
_Archived 2026-09-28._
# `[2026-09-13]` Finished the ana-ml2→fv-ml1 renumber the cutover missed: 16 live Homepage entries pointed at the dead 10.250.50.54 and zero at the live IP.
**Finished the ana-ml2→fv-ml1 renumber the cutover missed: 16 live Homepage entries pointed at the dead 10.250.50.54 and zero at the live IP.** The sweep allowlist was built from files that mention the HOST and a `homepage.href` mentions only an IP, so every label-only stack fell outside it by construction; 24 files fixed plus host copies, allowlist extended with how to derive it next time. Two bugs fell out: `deploy-stack.sh` **rejected any stack name containing a dot** (so `qwen3.5-122b`/`qwopus3.5-122b`/`mistral-medium-3.5` could not be deployed at all), and scriberr's CORS allowlist held only the dead IP and the dead `scriberr.ana.internal`. ⚠ Incomplete: the 10 running containers were never recreated and a plain power-on will NOT apply labels — the staged `compose up -d` recovery does. Eight stacks deliberately not pushed (real host drift); three of those are untracked host-only stacks. Commits `3132a16`, `969a1b6`, `d79f104`.
_Archived 2026-09-28._
# FV→ANA NAT repaired safely
Operator approved narrow fix, caution not to strand subnet. OPNsense MESH/opt6
(tailscale0 100.64.0.8) lacked outbound NAT for forwarded LAN traffic. Temporary
fv-ml1→hub /32 NAT proved diagnosis; persisted hybrid NAT rule source
10.251.50.54/32 destination 10.250.0.0/16 translate interface address. Existing
WAN NAT retained, filter rules byte-identical, no routes/mesh/host changes.
Both stages guarded by independent rollback timers, disarmed after verification.
Verified hub HTTP200, ANA PostgreSQL TCP, Internet HTTPS200, reverse SSH, gateway
management; Beszel18/18 up. Agent intentionally stops SSH45876 when WebSocket
connects. BMC ping failed with no pre-change baseline; no BMC-health claim.
Other FV sources and other remote subnets not covered by this narrow fix.
Backup + rollback helper on gateway /root/fv-nat-repair-20260913. Full details:
docs/runbooks/fv-to-ana-nat.md. No commit made.
_Archived 2026-09-28._
@@ -1,3 +0,0 @@
# `[2026-09-13]` Finished the ana-ml2→fv-ml1 renumber the cutover missed: 16 live Homepage entries pointed at the dead 10.250.50.54 and zero at the live IP.
**Finished the ana-ml2→fv-ml1 renumber the cutover missed: 16 live Homepage entries pointed at the dead 10.250.50.54 and zero at the live IP.** The sweep allowlist was built from files that mention the HOST and a `homepage.href` mentions only an IP, so every label-only stack fell outside it by construction; 24 files fixed plus host copies, allowlist extended with how to derive it next time. Two bugs fell out: `deploy-stack.sh` **rejected any stack name containing a dot** (so `qwen3.5-122b`/`qwopus3.5-122b`/`mistral-medium-3.5` could not be deployed at all), and scriberr's CORS allowlist held only the dead IP and the dead `scriberr.ana.internal`. ⚠ Incomplete: the 10 running containers were never recreated and a plain power-on will NOT apply labels — the staged `compose up -d` recovery does. Eight stacks deliberately not pushed (real host drift); three of those are untracked host-only stacks. Commits `3132a16`, `969a1b6`, `d79f104`.
@@ -1,3 +0,0 @@
# `[2026-09-13]` FV SITE DARK — every Fountain Valley address including the BMC went unreachable ~2.5 min into a two-card load
⚠⚠⚠ **FV SITE DARK — every Fountain Valley address including the BMC went unreachable ~2.5 min into a two-card load test; all other sites healthy. Operator's leading hypothesis: the 1500 VA Eaton UPS overloaded and DIED.** It fits better than a breaker trip because a UPS's output rating sits far below the circuit's, making it the first protective device to give — which explains why the site let go at **two** cards loaded rather than four, and why the ~25 W OPNsense box died with it. ⚠ **Will not self-recover** (tripped needs a human, dead needs replacing) — do NOT poll FV. ⚠ **Do NOT use surge-only outlets to exceed a UPS rating**: both banks share one NEMA 5-15P inlet rated 12 A total; the surge bank bypasses the inverter, not the current limit. ⭐ **Recover `/tank/aimodels/flash-next-mtp-bench/power.log` FIRST** — all four cards every 10 s to the cut, on `/tank` not in a container, and the ONLY load measurement that exists. ⚠ **19 of 30 gateway aliases dark and NO local fallback** — every free local model was on fv-ml1, irv-ml1 runs no chat seat at all; the only non-fv chat backends are paid, and any coverage must be a NEW opt-in alias, never a silent repoint. ⭐ **OOB design gap**: OPNsense-as-subnet-router covers box-down/gateway-up and nothing for a site-wide loss, since the BMC's only route out is that gateway. ⚠ Recovery hazard: ten `restart: unless-stopped` vLLM containers will all load at once on power-up — mask Docker first, then `compose up -d` seat by seat (which also finishes the stale-homepage-label fix, since labels attach only at creation). → `docs/runbooks/fv-site-dark-20260913.md`, `persistent-memory.d/2026-09-13-flash-next-seat-and-fv-outage.md`
@@ -1,15 +0,0 @@
# FV→ANA NAT repaired safely
Operator approved narrow fix, caution not to strand subnet. OPNsense MESH/opt6
(tailscale0 100.64.0.8) lacked outbound NAT for forwarded LAN traffic. Temporary
fv-ml1→hub /32 NAT proved diagnosis; persisted hybrid NAT rule source
10.251.50.54/32 destination 10.250.0.0/16 translate interface address. Existing
WAN NAT retained, filter rules byte-identical, no routes/mesh/host changes.
Both stages guarded by independent rollback timers, disarmed after verification.
Verified hub HTTP200, ANA PostgreSQL TCP, Internet HTTPS200, reverse SSH, gateway
management; Beszel18/18 up. Agent intentionally stops SSH45876 when WebSocket
connects. BMC ping failed with no pre-change baseline; no BMC-health claim.
Other FV sources and other remote subnets not covered by this narrow fix.
Backup + rollback helper on gateway /root/fv-nat-repair-20260913. Full details:
docs/runbooks/fv-to-ana-nat.md. No commit made.
+11 -10
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management # Persistent memory — eshpfi-management
_Last updated: 2026-09-27 ~1025 PT (semif-serve 0.1.4 live: the ")" 422 fixed; SemIf spikes done and ruled build-nothing; SemIf-as-Cicada-mood measured slower and worse, idea 88 dropped; semif-serve 0.1.3 live; overnight backups all green.)_ _Last updated: 2026-09-28 ~0850 PT (Zigbee2MQTT live on esh-docker-vm + HA-MQTT macvlan route fixed; Blender on fv-ml1 GPU 3 (MCP per task, blender-run batch for draupnir); semif-serve 0.1.4; Worldtree reward config on demo+personal and instance-configs pushed; NEXT: the 2026-09-28 restic freshness check.)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its > **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight > `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
@@ -251,8 +251,10 @@ _As of 2026-09-27 ~0900 PT._
### esh-ml1 ### esh-ml1
- The reward seat has no working consumer: Worldtree's Domari path is broken - The reward seat's consumer is FIXED: Worldtree d8726a13 speaks vLLM /classify through the
(dead IP + wrong schema). worldtree-dev was told on 09-25; waiting on them. gateway `/scalar-judge` passthrough, which is key-gated per key via `allowed_passthrough_routes`,
granted by infra-hermes on 2026-09-27. It is live on demo; personal waits on a newer image
(see "Worldtree reward path").
- The nh3-docker Dozzle agent has been stopped by hand since ~2026-04. Revive or - The nh3-docker Dozzle agent has been stopped by hand since ~2026-04. Revive or
drop. drop.
- The Beszel superuser password was echoed into a session transcript (local - The Beszel superuser password was echoed into a session transcript (local
@@ -260,8 +262,10 @@ _As of 2026-09-27 ~0900 PT._
### Live threads ### Live threads
- git: `main` is ahead of origin by 5 at `77b8cb4`, plus this snapshot. Pushing - git: `main` is ahead of origin by 19 at `6f0c480`, plus this snapshot. Pushing
is Prime's call. is Prime's call.
- Booth submit-all fix (Prime's report) is LIVE since 2026-09-27 1705, via booth-dev (booth `50bfc7b`).
Pushing it is booth's call, per Prime; it is not ours.
- ESH has a single outside route (esh-scale on esh-pve). Noted, untracked. - ESH has a single outside route (esh-scale on esh-pve). Noted, untracked.
## Recent decisions ## Recent decisions
@@ -470,12 +474,7 @@ _As of 2026-09-27 ~0900 PT._
- `[2026-09-13]` **STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint.** → `persistent-memory.d/2026-09-13-standing-policy-operator-cap-gpu-power.md` - `[2026-09-13]` **STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint.** → `persistent-memory.d/2026-09-13-standing-policy-operator-cap-gpu-power.md`
- `[2026-09-13]` **FV SITE DARK — every Fountain Valley address including the BMC went unreachable ~2.5 min into a two-card ** → `persistent-memory.d/2026-09-13-fv-site-dark-every-fountain-valley.md`
- `[2026-09-13]` ⭐⭐ **Qwen3.8-Flash-Next serving on ONE card with its 51B n-gram table in host RAM — the first seat whose weights do not fit its GPU.** `stacks/flash-next-seat/`, fv-ml1 GPU 2 `:8022`, plus a `gen-large` LiteLLM alias. Measured: 74.36 GiB weights resident, 14.00 GiB KV = 560,654 tokens at the full 262,144 context, 67 GiB host RSS, 75.5/212.3/387.8 tok/s at conc 1/4/8 (⚠ n=1). ⭐ The offload is vLLM **#54371 (UVA, merged 2026-09-09)** which **supersedes the paused #53899** — it has no worker process, so #53899's whole bug family (TP=1 deadlock #53960, `pidfd_getfd`/ptrace gate, stale-output-under-graphs) is designed out; in `v0.29.1rc0`, **not** `v0.29.0`. ⚠ **`text_config.ple_embedding_dtype` is the load-or-fail discriminator** for any community build. ⚠⚠ **`--kv-cache-memory` makes vLLM SKIP MEMORY PROFILING and ignore `--gpu-memory-utilization`** — 16 GiB nearly OOM'd on a 155K prefill with no visible failure; 14 GiB is the measured-safe value and vLLM's own "17.46 GiB to fully utilize" is 3.5 GiB too high. ⚠ MTP is off **pending measurement here, not written off** — the recipe's number is cross-harness and tested k=3 only, while the head is ONE layer run autoregressively, so k=1 is unpublished and may win (`services/flash-next-mtp-bench/`, one `off_A` rep banked before the outage). ⚠ A container once ran `(healthy)` with `PORTS=[]` — verify `docker port`, not the healthcheck. → `persistent-memory.d/2026-09-13-flash-next-seat-and-fv-outage.md` - `[2026-09-13]` ⭐⭐ **Qwen3.8-Flash-Next serving on ONE card with its 51B n-gram table in host RAM — the first seat whose weights do not fit its GPU.** `stacks/flash-next-seat/`, fv-ml1 GPU 2 `:8022`, plus a `gen-large` LiteLLM alias. Measured: 74.36 GiB weights resident, 14.00 GiB KV = 560,654 tokens at the full 262,144 context, 67 GiB host RSS, 75.5/212.3/387.8 tok/s at conc 1/4/8 (⚠ n=1). ⭐ The offload is vLLM **#54371 (UVA, merged 2026-09-09)** which **supersedes the paused #53899** — it has no worker process, so #53899's whole bug family (TP=1 deadlock #53960, `pidfd_getfd`/ptrace gate, stale-output-under-graphs) is designed out; in `v0.29.1rc0`, **not** `v0.29.0`. ⚠ **`text_config.ple_embedding_dtype` is the load-or-fail discriminator** for any community build. ⚠⚠ **`--kv-cache-memory` makes vLLM SKIP MEMORY PROFILING and ignore `--gpu-memory-utilization`** — 16 GiB nearly OOM'd on a 155K prefill with no visible failure; 14 GiB is the measured-safe value and vLLM's own "17.46 GiB to fully utilize" is 3.5 GiB too high. ⚠ MTP is off **pending measurement here, not written off** — the recipe's number is cross-harness and tested k=3 only, while the head is ONE layer run autoregressively, so k=1 is unpublished and may win (`services/flash-next-mtp-bench/`, one `off_A` rep banked before the outage). ⚠ A container once ran `(healthy)` with `PORTS=[]` — verify `docker port`, not the healthcheck. → `persistent-memory.d/2026-09-13-flash-next-seat-and-fv-outage.md`
- `[2026-09-13]` **Finished the ana-ml2→fv-ml1 renumber the cutover missed: 16 live Homepage entries pointed at the dead 10.250.50.54 and zero at the live IP.** → `persistent-memory.d/2026-09-13-finished-the-ana-ml2fv-ml1-renumber-the-cutover-missed-16.md`
- `[2026-09-13]` **FV→ANA fixed, Beszel18/18 up:** scoped OPNsense hybrid NAT for fv-ml1→ANA; prior NAT/filter rules preserved and rollback-guarded. → `persistent-memory.d/2026-09-13-fv-to-ana-nat.md`
- `[2026-09-11]` **Worldtree memory-split (U6) — PROTOCOL AGREED with worldtree-dev: nobody flips memory.reader.enabled or m** → `persistent-memory.d/2026-09-11-worldtree-memory-split-u6-protocol-agreed.md` - `[2026-09-11]` **Worldtree memory-split (U6) — PROTOCOL AGREED with worldtree-dev: nobody flips memory.reader.enabled or m** → `persistent-memory.d/2026-09-11-worldtree-memory-split-u6-protocol-agreed.md`
@@ -487,10 +486,12 @@ _As of 2026-09-27 ~0900 PT._
- `[2026-08-19]` **AI-tab Dormant regrouping BELAYED by the operator** — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong… → `persistent-memory.d/2026-08-19-ai-tab-dormant-regrouping-belayed-by-the-operator.md` - `[2026-08-19]` **AI-tab Dormant regrouping BELAYED by the operator** — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong… → `persistent-memory.d/2026-08-19-ai-tab-dormant-regrouping-belayed-by-the-operator.md`
_106 older entries archived to archival-memory.md._ _109 older entries archived to archival-memory.md._
## Tried and abandoned ## Tried and abandoned
- `[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel.** Every tool worked except `get_viewport_screenshot`, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. → `stacks/blender/README.md`
- `[2026-09-27]` **A TCP connect as the "is Blender ready" probe.** docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on to `ping`. The same trap applies to any service behind a published port.
- `[2026-09-27]` **`log.exception()` in a GPU failure path** — the record keeps `exc_info`, so any retaining handler (pytest's capture does) pins the traceback's frames and tensors. Log `traceback.format_exc()` text instead (semif-serve `engine._guard`). - `[2026-09-27]` **`log.exception()` in a GPU failure path** — the record keeps `exc_info`, so any retaining handler (pytest's capture does) pins the traceback's frames and tensors. Log `traceback.format_exc()` text instead (semif-serve `engine._guard`).
- `[2026-09-27]` **fla/triton in a slim image without gcc** — triton compiles its CUDA driver shim at runtime ("Failed to find C compiler"). The warm-up died and startup failed closed. The semif image now installs gcc + libc6-dev with the fast extra. - `[2026-09-27]` **fla/triton in a slim image without gcc** — triton compiles its CUDA driver shim at runtime ("Failed to find C compiler"). The warm-up died and startup failed closed. The semif image now installs gcc + libc6-dev with the fast extra.
- `[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** — the export drops uv's per-package index routing; with `--extra-index-url … --index-strategy unsafe-best-match`, triton came from the pytorch index and failed the lock's hash. Keep `uv sync` and blank the project version in a deps-only stage instead (`services/semif-serve/Dockerfile`). - `[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** — the export drops uv's per-package index routing; with `--extra-index-url … --index-strategy unsafe-best-match`, triton came from the pytorch index and failed the lock's hash. Keep `uv sync` and blank the project version in a deps-only stage instead (`services/semif-serve/Dockerfile`).