From 6ec76bf35c4c65c3e4a82bb6c13acfd681069390 Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Fri, 2 Oct 2026 07:33:36 -0700 Subject: [PATCH] ops(nh3-dev): grow root into the 378 GB disk online; swap moved to /swapfile Prime grew VM 102's scsi0 from 250 to 378 GB after the root alert hit 85% twice in 14 h. The swap partition sat right after sda1 and blocked growth, so the playbook moves swap to a 4 GB /swapfile, deletes sda5/sda2, grows sda1 in place (start sector unchanged) and ext4 online, and sets initramfs RESUME=none so boots don't wait for the vanished swap. Root is 372 GB, 58%. --- persistent-memory.md | 1 + playbooks/nh3-dev-grow-root.yaml | 104 +++++++++++++++++++++++++++++++ servers/nh3-dev/README.md | 7 +++ 3 files changed, 112 insertions(+) create mode 100644 playbooks/nh3-dev-grow-root.yaml diff --git a/persistent-memory.md b/persistent-memory.md index fbf990e..75303db 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -291,6 +291,7 @@ _As of 2026-10-01 ~0446 PT._ ## Recent decisions +- `[2026-10-02]` **nh3-dev root grown 250 → 378 GB (Prime resized scsi0; I grew the guest online, no reboot).** Root 372 GB, 58%, 150 GB free, after the 85% alert fired twice in 14 h (uv cache + agent venvs). Swap moved to a 4 GB `/swapfile` (the old `sda5` blocked growth), `RESUME=none`, all initrds rebuilt (`playbooks/nh3-dev-grow-root.yaml`). ⚠ **TODO: delete VM snapshot `pre-rootgrow-20261002` on nh3-pve after the next clean reboot** (it pins ZFS space as the disk churns). - `[2026-10-01]` **MIA (Make-It-Animatable v2) auto-rigger LIVE on fv-ml1 GPU 3, one-shot (Prime via dread-dev).** `scripts/mia-run --job DIR -- name=in.glb ...`, image `local/mia:0.1.0` (10.5 GB), weights in the shared HF cache at pinned revisions (sha256-verified). Acceptance: 3.9–4.7 s a mesh (median of 3), 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical; GPU-vs-CPU distances the same size as sampling noise (positive control: unseeded GPU). Skeleton template is a SUBSTITUTE (gated HF dataset `jasongzy/Mixamo`, terms not accepted; Prime's call). ⚠ fv-ml1 zroot at 85% after the build. → `stacks/mia/README.md` - `[2026-10-01]` **Prime: "clean up the 1700+ backups" — dev-backup retention FIXED, 1,740 husks removed.** They were never 1,788 real snapshots: the prune had deleted everything except one 0555 dir (`vastblue/praxis/references/PraxisPM_Rev0_07.11.26`, 17 files hardlinked into every snapshot), so each old dir was a 22-entry husk and only the newest 48 were complete. Removing them lost no history and freed ~0 bytes (same inodes, link count 1,788). Deletion ran from a generated list of 1,740 literal paths, each verified as a husk first, with 0 errors. The script now runs `chmod -R u+w` before `rm`, logs the error count and the kept count, and fails the unit (exit 3) on a bad prune or exit 1 on a failed rsync. Positive control: a manual run at 0556 pruned the full snapshot `2026-09-29_0602` with 0 errors, leaving 48. Retention stays the designed 48 hourly; the 4.6 GB log rotated at 0530 is kept, compressed to 99 MB, at `~/.config/dev-backup/dev-backup.log.20261001-0530.gz` until someone deletes it. - `[2026-10-01]` **Prime: gen-small keeps its 8 GiB KV pin, and parakeet-nemo keeps CUDA graphs OFF** (~2–8 ms at short clips, accepted). GPU 0's ~1 GB of spare memory is enough with the seat's hard 3,840 MiB cap. diff --git a/playbooks/nh3-dev-grow-root.yaml b/playbooks/nh3-dev-grow-root.yaml new file mode 100644 index 0000000..f819f37 --- /dev/null +++ b/playbooks/nh3-dev-grow-root.yaml @@ -0,0 +1,104 @@ +# Grow nh3-dev's root filesystem into space added to its disk in Proxmox (VM 102 on nh3-pve). +# First used 2026-10-02: Prime grew scsi0 250G -> 378G after the 85% root alert fired twice in 14 h. +# +# The obstacle: the Debian installer put swap (sda5, inside extended sda2) right AFTER sda1, so +# sda1 cannot grow until that swap partition is gone. This playbook moves swap to /swapfile, deletes +# sda5 and sda2, grows sda1 in place (start sector unchanged) and grows ext4 online. No reboot. +# +# Take a VM snapshot on nh3-pve FIRST (qm snapshot 102 ). The partition table is also dumped to +# /var/backups/nh3-dev-rootgrow-/sda-ptable.sfdisk, which `sfdisk /dev/sda < dump` restores (sda1's data is untouched +# either way: only its end moves). +# +# RESUME: initramfs-tools points resume-from-hibernation at the old swap UUID. With that partition +# gone, every boot would wait ~30 s for it, so RESUME is set to none and all initrds are rebuilt. +# +# scripts/elway infra-ops@10.100.10.50 --playbook playbooks/nh3-dev-grow-root.yaml +vars: + stamp: "20261002" + old_swap_uuid: 97bca850-bf72-4881-ae33-23d9b68315b6 + swapfile_size: 4G + +steps: + - name: Back up the partition table, fstab and the resume config + sudo: true + shell: | + set -e + install -d -m 755 /var/backups/nh3-dev-rootgrow-{{ stamp }} + cp -p /etc/fstab /etc/initramfs-tools/conf.d/resume /var/backups/nh3-dev-rootgrow-{{ stamp }}/ + /usr/sbin/sfdisk -d /dev/sda > /var/backups/nh3-dev-rootgrow-{{ stamp }}/sda-ptable.sfdisk + # guards run as infra-ops, so the marker lives where infra-ops can see it (not under /root) + creates: /var/backups/nh3-dev-rootgrow-{{ stamp }}/sda-ptable.sfdisk + + - name: Install growpart (cloud-guest-utils) + sudo: true + shell: DEBIAN_FRONTEND=noninteractive apt-get install -y -q cloud-guest-utils + creates: /usr/bin/growpart + + - name: Create and enable /swapfile ({{ swapfile_size }}) BEFORE dropping the swap partition + sudo: true + shell: | + set -e + fallocate -l {{ swapfile_size }} /swapfile + chmod 600 /swapfile + /usr/sbin/mkswap /swapfile + /usr/sbin/swapon /swapfile + creates: /swapfile + + - name: Add /swapfile to fstab + sudo: true + shell: echo '/swapfile none swap sw 0 0' >> /etc/fstab + when: "! grep -q '^/swapfile ' /etc/fstab" + + - name: Turn off the old swap partition (its pages move to RAM or /swapfile) + sudo: true + shell: /usr/sbin/swapoff /dev/sda5 + when: "/usr/sbin/swapon --show=NAME --noheadings | grep -qx /dev/sda5" + + - name: Drop the old swap partition from fstab + sudo: true + shell: sed -i '/^UUID={{ old_swap_uuid }} /d' /etc/fstab + when: "grep -q '^UUID={{ old_swap_uuid }} ' /etc/fstab" + + - name: Point initramfs resume at nothing and rebuild every initrd + sudo: true + shell: | + set -e + echo 'RESUME=none' > /etc/initramfs-tools/conf.d/resume + update-initramfs -u -k all + when: "! grep -qx 'RESUME=none' /etc/initramfs-tools/conf.d/resume" + + - name: Delete partitions 5 (old swap) and 2 (the extended partition holding it) + sudo: true + shell: | + set -e + /usr/sbin/sfdisk --no-reread --force --delete /dev/sda 5 + /usr/sbin/sfdisk --no-reread --force --delete /dev/sda 2 + partx -d --nr 5 /dev/sda || true + partx -d --nr 2 /dev/sda || true + when: "sudo -n /usr/sbin/sfdisk -d /dev/sda | grep -q '^/dev/sda2 '" + + - name: Grow partition 1 to the end of the disk (start sector unchanged) + sudo: true + shell: growpart /dev/sda 1 + when: "sudo -n growpart --dry-run /dev/sda 1 >/dev/null 2>&1" + + - name: Grow the ext4 filesystem online + sudo: true + shell: /usr/sbin/resize2fs /dev/sda1 + +verify: + - name: Only sda1 remains, still starting at sector 2048 + sudo: true + shell: | + test "$(/usr/sbin/sfdisk -d /dev/sda | grep -c '^/dev/sda')" = 1 && /usr/sbin/sfdisk -d /dev/sda | grep -q '^/dev/sda1 : start= *2048,' + changed_when: "false" + - name: Root filesystem is now over 350 GB + shell: test "$(df -BG --output=size / | tail -1 | tr -dc 0-9)" -gt 350 + changed_when: "false" + - name: /swapfile is the active swap and the old partition is gone from fstab + sudo: true + shell: /usr/sbin/swapon --show=NAME --noheadings | grep -qx /swapfile && ! grep -q '{{ old_swap_uuid }}' /etc/fstab + changed_when: "false" + - name: Root is still mounted by its UUID in fstab + shell: grep -q '^UUID=689f8c0c-1230-43a5-93c2-6baf76867d70 / ' /etc/fstab + changed_when: "false" diff --git a/servers/nh3-dev/README.md b/servers/nh3-dev/README.md index 3c559b3..1f59eb9 100644 --- a/servers/nh3-dev/README.md +++ b/servers/nh3-dev/README.md @@ -10,6 +10,13 @@ Code sessions; **not** a Docker-stack host in the `stacks/` sense. in auto-memory). Note: Claude Code sessions often run **natively on this box**, so local Bash already executes here — no SSH-to-self needed for non-privileged work. +**Disk (2026-10-02):** VM 102 on nh3-pve, `scsi0` on `local-zfs`, **378 GB**, one partition (`sda1`, ext4, +372 GB). Swap is a **4 GB `/swapfile`**; there is no swap partition, and initramfs `RESUME=none`. +It was 250 GB with swap in `sda5` right after root, which blocked growth; the root alert hit 85% +twice in 14 h. To grow it again: resize `scsi0` in Proxmox, snapshot, then +`scripts/elway infra-ops@10.100.10.50 --playbook playbooks/nh3-dev-grow-root.yaml` (the swap and +partition steps skip themselves now; growpart + resize2fs do the work, online). + ## What runs here - **NH3 egress proxy — RETIRED 2026-09-06** (replaced by headscale exit nodes; `danted` disabled, config `.retired`). Was: durable internal-only SOCKS5 `socks5h://10.100.10.50:1080`