memory: ravenpen.com registered — registrar, expiry, and the scope boundary I can't cross
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
# Persistent memory — eshpfi-management
|
||||
|
||||
_Last updated: 2026-09-19 ~08:55 PT (⭐ **FV power budget recorded** — dedicated 20 A circuit shared by fv-ml1 + the OPNsense R420, ~85% of the 1920 W continuous ceiling with the 275 W GPU caps in place; KEEP THEM, and note the router shares the breaker with the thing most likely to trip it. ⭐ The ops log is BUILT and caught its own tooling's failure during a live bus upgrade. ⭐ althing **3.7.0** deployed, forseti-verified. ⭐ The backup alarm had been unable to notify ANYONE since 2026-08-28 — fixed, plus coverage-awareness and vzdump task-status checking. ⚠ lv-mccarthy's run outcome STILL UNVERIFIED.)_
|
||||
_Last updated: 2026-09-21 ~05:55 UTC (⭐ **ravenpen.com registered** by the operator — Cloudflare registrar, expires 2028-09-20, zone active but BARE; hamr-dev answered, the 09-18 hold discharged. ⭐ Booth gained keep-both-ways + per-item cosmetic blur. ⭐ Backup alarm verdict split: STALE (exit 1) vs ERRORED-JOBS (exit 3) — it was crying STALE over a fleet whose every body was fresh. ⭐ The ops log is BUILT; commit attribution now works end-to-end, both controls measured. ⚠ lv-mccarthy's run outcome STILL UNVERIFIED.)_
|
||||
|
||||
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
|
||||
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
|
||||
@@ -221,6 +221,8 @@ nothing touched. Full context in the 09-17 Recent decisions entries.
|
||||
|
||||
## Recent decisions
|
||||
|
||||
- `[2026-09-20]` **`ravenpen.com` REGISTERED by the operator — the 09-18 hold is discharged and hamr-dev is answered.** infra-ops deliberately did NOT execute this twice, because it was a non-refundable purchase reaching us as a peer relay with no registrar credential on our side; the standing instruction was to post registrar + expiry back on thread `01M2SERDB3DR3RV7JS1J0GMF0H` only once the operator bought it himself. Done. Facts from **RDAP (Verisign, authoritative)** rather than a dashboard: registrar **Cloudflare, Inc.** (IANA 1910), registered 2026-09-20T20:12:30Z, **expires 2028-09-20** (two-year), `clientTransferProhibited`. Zone `2df4c5eb4ea4b9410423bdebcb6c5192` active on `vh@phasefinal`, activated 0.4 s after creation — registered THROUGH Cloudflare Registrar, which is why the zone's `original_registrar` is null. ⚠ **The zone is BARE — zero DNS records**, so the name resolves to nothing and mail to it bounces; correct for bought-not-built, but say so before anyone points at it. ⚠ **Scope boundary measured, not assumed:** the fleet `infra-ops` Cloudflare token is Zone·DNS·Edit and **403s on the Registrar API** — I can build records in the zone, and I can NOT read auto-renew state, renew, or transfer. **Auto-renew is therefore UNCONFIRMED**; do not let anyone assume it. ⚠ The token is **vaulted, not on disk** — `secret get 'nh3-dev/.config/cloudflare/infra-ops-dns-token'`; the memory's "vaulted at nh3-dev/..." names a VAULT KEY, and reading it as a filesystem path wastes a step.
|
||||
|
||||
- `[2026-09-19]` ⭐⭐ **FV is a dedicated 20 A circuit carrying fv-ml1 AND the R420 running OPNsense, and nothing else (operator-confirmed).** The ceiling is **1920 W**, not 2400 — a GPU host running for hours is a continuous load, so NEC's 80% rule governs. Measured: BMC `dcmi power reading` 390 W instantaneous / 461 W max over 2423 s **at idle GPUs**, so the non-GPU baseline is ~313 W (EPYC 9254 24C/96T, 5+ drives, 4 PSUs). GPUs are capped **275 W** each against a **300 W** stock TGP (`power.default_limit`) — a 100 W saving across four cards, NOT the 200 W you get by measuring against the 325 W firmware ceiling; the operator corrected me on exactly that. Derived worst case **~1625 W capped (85%)** vs **~1725 W at stock (90%)**. ⚠ **Keep the caps** — 90% leaves nothing for a heavier-than-estimated R420, PSU efficiency, or a warm day. ⚠⚠ **The coupling is worse than the trip:** OPNsense IS the FV edge and shares the breaker with the thing most likely to trip it, so an overload takes the router with it and removes the remote path needed to recover. Four PSUs do not help — they are all downstream of one breaker. ⚠ **NOT measured:** fv-ml1 under real 4-GPU load (the 461 W max is idle-ish), whether the BMC reports AC or DC (±10% ≈ 150 W at load), and the R420's actual draw. Recorded in `servers/fv-ml1/README.md` § Power. ⚠ Also fixed there: the README claimed **2x** GPUs; `nvidia-smi` reports **four**.
|
||||
|
||||
- `[2026-09-19]` ⭐⭐ **althing 3.7.0 rolled out to both infra-ops surfaces — and the rollout broke the claim tooling built that morning.** Post office on nh3-docker (built from `althing@6db955f`, manifest `sha256:df0709b3`, **10.9 s** recreate, volume preserved) and the nh3-extdev system wheel + herald. forseti independently verified both. ⚠ Everything CONTENT-verified, never tag-verified: `postbox --version` inside the image before the push and inside the running container after; data continuity proven by reading forseti's own message back out of the running 3.7.0 store. ⚠ **`deploy-stack.sh` rsyncs with `--delete`, so the host-side `.bak-<version>` compose convention is GONE** — it converges rather than accretes; rollback is git history + the retained 3.6.3 registry digest. ⚠ **The claim bug:** a 45-minute operation claim was refreshed *and then released* by deploy-stack.sh's exit trap, silently dropping the protection mid-rollout. Fixed 3e7d3a3 — `ops-log claim` exits **10** when the claim is already the caller's and leaves the holder file UNTOUCHED (a refresh would overwrite the reason and TTL the original claimant chose).
|
||||
|
||||
Reference in New Issue
Block a user