memory: snapshot — dev-backup retention broken since 07-18 (log rotated, errors counted; Prime to rule); gen-small KV + graphs-off left as is

This commit is contained in:
vh
2026-10-01 05:29:27 -07:00
parent 0d9bb109b8
commit 9af5a16d07
+14 -2
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-10-01 ~0420 PT (Parakeet seat → unified-en under NeMo LIVE + audited; gen-small util 0.36 → .env 0.33; leftover bench weights + spike dirs deleted; Scriberr no-upstream; eshpfi + worldtree-instance-configs pushed. Prior: U11a off, SemIf → intern-decision, Scriberr GPU 3 + patches.)_
_Last updated: 2026-10-01 ~0446 PT (gen-small OOM fixed by parakeet-nemo 0.1.1, with a hard cap and cache return; Prime: leave the gen-small KV and graphs-off as is; arbo webhook fixed (Gitea allow-list); leftovers deleted, repos pushed. Prior: Parakeet seat → unified-en LIVE, U11a off, SemIf → intern-decision, Scriberr GPU 3 + patches.)_
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
@@ -115,7 +115,7 @@ no longer deployed sidecars here. See Recent decisions.)
## Current state / in-flight
_As of 2026-10-01 ~0420 PT._
_As of 2026-10-01 ~0446 PT._
### Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01)
@@ -125,6 +125,7 @@ _As of 2026-10-01 ~0420 PT._
- Result: the seat rests at 2,108 MiB and peaks at 3,028 during a 12-min file. gen-small is steady at 36,116 (its runtime growth landed at init, with zero movement on later requests). GPU 0 Free is ~1.05 GB at rest.
- Cost: graphs-off is ~2–8 ms slower at short clips and ~27 ms at 20–60 s (35 / 40 / 50 / 98 ms against 33 / 36 / 42 / 71), still 4–15× faster than the old seat.
- **Lesson: a shared-card tenant must carry a HARD cap, whether or not it returns memory.**
- **Prime ~0445 2026-10-01: leave both as is.** gen-small keeps its 8 GiB KV pin (no insurance trim), and CUDA graphs stay OFF (the few-ms cost is accepted; no graph-pool rework).
- **Rollback:** `docker stop parakeet-nemo && docker start parakeet`. The old container and image are kept.
- **GPU 0 is FULL:**
- The seat's steady state is **3,582 MiB** (its cached window peak); Free is **385 MiB**.
@@ -133,6 +134,16 @@ _As of 2026-10-01 ~0420 PT._
- Seat invariants (in its README): cast to bf16 AFTER change_attention_model; uvicorn pinned with `--http h11` (httptools 0.8.0 emits `HTTP/1.1 200\x00OK`, which LiteLLM/httpx rejects).
- NVIDIA Open Model License accepted for internal use. → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md`
### dev-backup (nh3-dev `~/development` → nh3-nas hourly): retention BROKEN since 2026-07-18 (found 2026-10-01)
- **Its retention prune has failed every run since 2026-07-18** with "Permission denied", because read-only dirs were copied from the source. **1,788 snapshots** sit in `nh3-nas:/volume1/Backup/nh3-dev-development` against a design of 48. The script still logged "snapshot OK": an instrument that passes in both states.
- The log had reached 4.6 GB (25.7M error lines) and tripped the nh3-dev 85% disk alert. I rotated it to `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` (99 MB) and made the script COUNT prune errors instead of dumping them (`dev-backup.sh.bak-20261001`). nh3-dev is now at 84%; nh3-nas /volume1 is at 78% of 42 TB.
- **Awaiting Prime (one-way for any deleted history):**
- fix the prune (`chmod -R u+w` before `rm`) and prune to the designed 48;
- OR switch to a tiered retention (48 hourly + 30 daily + 12 weekly) and prune the rest;
- OR keep everything for now.
- Until he rules, the prune keeps failing and its error count shows in the log. Nothing is deleted.
### Worldtree U11 memory cutover (demo + personal)
- **Legacy plane OFF since 0115/0120 PT 2026-09-30** (config repo 63cf268; personal /metrics b6fdd81; the repo is pushed). → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
@@ -289,6 +300,7 @@ _As of 2026-10-01 ~0420 PT._
## Recent decisions
- `[2026-10-01]` **Prime: gen-small keeps its 8 GiB KV pin, and parakeet-nemo keeps CUDA graphs OFF** (~2–8 ms at short clips, accepted). GPU 0's ~1 GB of spare memory is enough with the seat's hard 3,840 MiB cap.
- `[2026-10-01]` **⚠ Gitea's `[webhook] ALLOWED_HOST_LIST = external, 10.0.0.0/8` silently REJECTS headscale mesh IPs (100.64.0.0/10).** The test API still returns 204 and nothing arrives. The vh/arbo hook therefore targets irv-ml1's LAN `10.6.110.50:9009`; the secret was re-set and a delivery is verified (deploy ran 04:41). A Gitea webhook PATCH without `secret` and `branch_filter` drops both, so always resend them.
- `[2026-10-01]` **vh/arbo push webhook repointed from the retired wg0 lifeline `10.100.79.3:9009` to irv-ml1's mesh address `100.64.0.6:9009`.** It had been dead since 09-06 (comfy-dev report). No other of the 109 repos' hooks pointed at 10.100.79.x. The Actions runner `irv-ml1-arbo` was fine; run 677 failed, likely colliding with a hand deploy. The secret was not resent: if the next push shows an HMAC rejection, re-set it.
- `[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push.** DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; **unified-en KEPT, the live seat mounts it**), plus `/tank/spikes/scriberr-slicer` (including the private copies of Prime's recordings) and `/tank/spikes/parakeet-ab`. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed.