Commit Graph
1536 Commits
Author SHA1 Message Date
vh 4b2a81bf39 feat(pve-nag-patch): remove the subscription popup on the PVE hosts, re-applied by an apt hook 2026-10-03 13:27:11 -07:00
vh 22224b5335 ops(nh3-pve-2): register host, network/AMT/disk layout and the re-IP lesson 2026-10-03 13:17:46 -07:00
vh 9b1cfa7344 feat(nh3-pve-2): drop stale 1 TB boot entry, wipe it into LVM-thin storage vmstore 2026-10-03 13:17:28 -07:00
vh 00a6fbd849 feat(cira-tunnel-watchdog): auto-drop stale AMT CIRA tunnels on MeshCentral's MPS 2026-10-03 12:34:35 -07:00
vh fe1d472aae dns: nh3-pve-2 -> 10.100.250.62 (10G port, nh3-mgmt reservation) 2026-10-03 12:32:44 -07:00
vh e4fcd1398d docs(meshcentral): stale CIRA tunnel, third occurrence (nh3-pve-2) 2026-10-03 12:19:42 -07:00
vh 73da3175e2 docs(albok-service): pushed tag verified to match the deployed image 2026-10-03 10:00:03 -07:00
vh f8290f64bf memory: worldtree gate duty reframed as post-deletion regression check (infra-ops 01M413W4) 2026-10-03 07:50:47 -07:00
vh 9b71f6c2f3 memory: first post-deletion Worldtree gate batch PASS 2026-10-03 07:49:44 -07:00
vh 3a6d997c99 ops(esh-nas-pve): tank verified healthy after two scrubs; disk kept; PVE 9 blocker cleared 2026-10-03 06:28:48 -07:00
vh 6a274a412a docs(albok-nemi): nemi now prunes reports and takes a run lock 2026-10-03 06:00:23 -07:00
vh 70fa0fbd53 memory: Nemi timer installed 2026-10-03 05:58:12 -07:00
vh e887c6a4bc feat(albok-nemi): hourly Nemi walk timer on nh3-dev (albok-dev ask), with OnFailure hook 2026-10-03 05:57:54 -07:00
vh 8a4dbf2341 fix(albok-service): 0.1.2 on nh3-docker (fd leak); add restricted wing personal/agent-feedback; Nemi credentials 2026-10-03 02:54:18 -07:00
vh 1fb36b32d2 memory: ESH tank first scrub done, second scrub running 2026-10-03 01:18:24 -07:00
vh 3a91eff2e6 memory: ESH tank scrub repairing the Aug 20 replacement's unrepaired blocks 2026-10-02 23:33:07 -07:00
vh 36c52d8860 docs(meshcentral): stale CIRA tunnel reproduced on a single shutdown 2026-10-02 22:53:30 -07:00
vh ad0322bf83 ops(meshcentral): diagnose and fix HW Connect stuck at Setup (stale CIRA tunnel); add relay probe 2026-10-02 22:49:27 -07:00
vh 55004090d1 ops(esh-pve-2): register host, AMT phoning home to MeshCentral; note MeshCentral first-CIRA crash race 2026-10-02 22:25:37 -07:00
vh e707d87713 feat(esh-pve-2): drop stale 1 TB boot entry, wipe it into LVM-thin storage vmstore 2026-10-02 22:08:19 -07:00
vh 6315ce38d5 memory: ESH tank verified residual, cleared, full scrub running 2026-10-02 19:26:30 -07:00
vh de143364fd docs: fleet Proxmox host inventory snapshot 2026-10-02 (all sites) 2026-10-02 19:23:58 -07:00
vh 62254023f3 docs: esh-pve-cluster PVE 8->9 upgrade plan (not executed); ESH tank DEGRADED finding 2026-10-02 19:20:28 -07:00
vh 5439d5fdc1 albok-service: 0.1.1 (health fix) on nh3-docker 2026-10-02 18:04:57 -07:00
vh d3958655c1 feat(albok-service): deploy the fleet knowledgebase service on nh3-docker
albok-service 0.1.0 (vh/albok 0b37431), image pfi/albok-service pinned by
digest, published on host port 8392 because 8390 is the post office.
Host prep playbook creates the fixed ids (albok 1500, albok-read 1510,
albok-personal 1511) and the local store/private roots; the container
gets a mounted /etc/group and group_add so the service can resolve and
chgrp its wing dirs. The config carries a LiteLLM key scoped to
qwen3-embedding and lives outside the deploy-synced conf dir. DNS name
albok.nh3.internal.
2026-10-02 17:56:34 -07:00
vh 9c233ca255 memory: nh3-pve-2 end-of-day state and on-site reinstall checklist 2026-10-02 17:40:50 -07:00
vh 56c372dd6f docs(pfi-tacticalrmm): rmm-mesh nginx upload limit raised to 4G (was the 1 MB default) 2026-10-02 17:08:21 -07:00
vh 5c0d5f0a73 docs(pfi-tacticalrmm): MeshCentral site-admin account lkraven 2026-10-02 17:06:19 -07:00
vh 34659ae9e7 ops(nh3-pve-2): AMT 21 configured (KVM, no-consent, listener) and phoning home to MeshCentral 2026-10-02 14:54:56 -07:00
vh e73adfe7ae dns+memory: nh3-pve-2-amt 10.100.250.63 (UDM port 5 on nh3-mgmt, reservation) 2026-10-02 14:43:16 -07:00
vh 59c9c0715a memory: MS-03 roles — esh-dev inherits nh3-dev's sessions; nh3-pve-2 purpose TBD by design 2026-10-02 09:17:39 -07:00
vh 536230170e memory: one MS-03 deploys as nh3-pve-2 (AMT on DHCP for phone-home; addressing plan) 2026-10-02 09:10:31 -07:00
vh 9dba6baccf homepage: NH3-PVE-AMT card opens MeshCentral (AMT now phones home; LAN management is off) 2026-10-02 09:06:49 -07:00
vh 1a2d763de4 ops(nh3-pve): AMT phones home to MeshCentral (CIRA); AMT on DHCP; LAN management now dark by design 2026-10-02 09:06:24 -07:00
vh 9bcb9417fb feat(amt): amt-cira-setup.py — configure Intel AMT phone-home (CIRA) to MeshCentral over WS-Man
MeshCentral only pushes CIRA through an on-host agent, so this mirrors its
amtmanager.js sequence directly over the LAN: trust MeshCentral's root,
add the MPS (username = 16-char meshid prefix), a periodic policy,
BIOS+OS user-initiated connections and a random environment-detection
domain. Idempotent, with a read-only state report. Enumerate uses
Enumerate+Pull (AMT 16 ignores OptimizeEnumeration; caught by a
positive control that first read 0 instances of a class that has one).

Applied to nh3-pve's AMT 2026-10-02 alongside MeshCentral mpsPass and an
ana-gw VIP/policy for 4433; the AMT does not dial out yet (static IP).
2026-10-02 08:50:47 -07:00
vh 61ba384d28 memory: demo fix-forward 7ab6ae40 verified; deploy-gate gap filed as Worldtree #423 2026-10-02 08:24:33 -07:00
vh 1219cfaca9 memory: demo rolled back from broken c2d87263 to fa8bc51c; b193 retired-key list for demo 2026-10-02 08:17:24 -07:00
vh 8b81579bca ops(pfi-tacticalrmm): MeshCentral to hybrid mode; nh3-pve AMT added and connected 2026-10-02 08:14:19 -07:00
vh a967bb95a6 docs(pfi-tacticalrmm): MeshCentral facts — WAN-only mode silently drops AMT adds; CLI access via the vaulted login token 2026-10-02 08:09:02 -07:00
vh cd30592e55 fix(wt-memory-gate-batch): pass --legacy-mode only when the deployed harness has it
b193 (U11b) retired the legacy plane and removed --legacy-mode from the
harness; passing it there fails the daily batch. The wrapper now greps
core/memory_acceptance at the deployed sha and omits the flag when it is
gone, so it is correct on both sides of the b192 -> b193 deploy (checked:
7a83f2f -> passes it, c2d87263 -> omits it).
2026-10-02 07:59:03 -07:00
vh 1cc8ad3292 memory: U11b step 5 done — live legacy memory deleted on demo and personal (H2 0), stamp sent 2026-10-02 07:54:50 -07:00
vh b0789f4b29 memory: U11b — gate-batch wrapper drops --legacy-mode at b193 2026-10-02 07:50:38 -07:00
vh 6e15f610bc memory: U11b gate streak 3 of 3; step 5 held for Prime's direct go 2026-10-02 07:50:12 -07:00
vh 59fb6b77a8 ops(nh3-dev): seat healthz watchdog — the 5-day pull-only outage must page, not age
althing-seat-daemon took an outside SIGTERM on 2026-09-27 and exited 0, so
Restart=on-failure never revived it; hermes-gateway sat pull-only for five
days with a yellow verdict nobody was watching. Fixed the death mode with a
Restart=always drop-in (also on the jekyll twin, same latent bug) and added
this watchdog for every other death mode: 5-min timer, alert on the second
consecutive non-200 (~10 min sustained), re-alert at most hourly, recovery
mail on green, exit 2 distinguishes a broken alarm wire from a down seat
(backup-freshness precedent). Lifecycle exercised end-to-end against a dead
port before install.
2026-10-02 07:47:17 -07:00
vh 902e16630f feat(dev-backup): add daily and weekly retention (48 hourly + 30 daily + 12 weekly)
Prime's ruling 2026-10-02. retention.py picks the snapshots to delete:
the newest 48, plus the newest of each of the last 30 days and of each of
the last 12 ISO weeks, counting only days and weeks that have snapshots.
Names that are not exactly YYYY-MM-DD_HHMM are never selected, and the
NAS side refuses any path outside that pattern. The unit fails unless
the number kept equals the number expected. Live run: deleted 1, 0
errors, 48 kept as expected.
2026-10-02 07:42:45 -07:00
vh 49c8fdf383 memory: nh3-dev snapshot waits for a natural reboot (Prime); post-boot checks listed 2026-10-02 07:36:15 -07:00
vh 6ec76bf35c ops(nh3-dev): grow root into the 378 GB disk online; swap moved to /swapfile
Prime grew VM 102's scsi0 from 250 to 378 GB after the root alert hit 85%
twice in 14 h. The swap partition sat right after sda1 and blocked growth,
so the playbook moves swap to a 4 GB /swapfile, deletes sda5/sda2, grows
sda1 in place (start sector unchanged) and ext4 online, and sets initramfs
RESUME=none so boots don't wait for the vanished swap. Root is 372 GB, 58%.
2026-10-02 07:33:36 -07:00
vh 01c356e7fb docs(ana-docker): record the RustDesk server key and where the client setup steps live 2026-10-01 15:31:32 -07:00
vh abcf4c3166 fix(scriberr): serve over HTTPS via the fleet TLS caddy so the browser recorder works
The in-browser recorder calls getUserMedia, which browsers refuse on http://
origins, so it sat at "Initializing recorder...". scriberr.nh3.phasefinal.com
is now fronted by the fleet TLS caddy on nh3-dev (wildcard cert) and added to
Scriberr's ALLOWED_ORIGINS; the Homepage link points at it. The plain
http://10.251.50.54:8080 URL keeps working except for recording.
2026-10-01 13:07:18 -07:00
vh 474f78811c memory: U11b — writer/reader default to enabled at b193 (no instance action) 2026-10-01 12:35:57 -07:00