fix(restic): vm-esh-nas off env-file too; infra-ops now provisioned there

Prime bootstrapped infra-ops on vm-esh-nas with playbooks/bootstrap-infra-ops-user.yaml.
It got the fleet-pinned uid/gid 850, NOPASSWD sudo with log_output, the docker
group and a 0700 home. That let playbooks/restic-repository-file.yaml migrate
the last restic host: the live profile matched the repo's pre-change sha,
its units no longer carry the URL, its secrets are vaulted, and the live and repo
profiles now match (a5ea75ea). All eight restic hosts are clean.

The staged helper script is gone, both from Prime's home on the host and from the
repo. The docs that described vm-esh-nas as lkraven-only are updated.
This commit is contained in:
vh
2026-09-27 01:56:04 -07:00
parent d775a01856
commit 30f2c977b7
6 changed files with 12 additions and 42 deletions
+1 -1
View File
@@ -33,7 +33,7 @@ On each server, deployed to `/etc/restic/`:
- **No secrets in committed config.** `profiles.yaml` references `RESTIC_PASSWORD_FILE=/etc/restic/password` and reads the repository URL through `repository-file: /etc/restic/repository`. Both files live only on the host, root 0400. Both are also vaulted as `<host>/etc/restic/{password,repository}` (2026-09-27).
- **Creds go in the URL, not netrc.** restic's rest backend doesn't consult `~/.netrc` — HTTP basic-auth has to be embedded in the repository URL.
- ⚠ **Never `env-file` for the URL.** Until 2026-09-27 the profiles loaded it with `env-file: /etc/restic/restic.env`. But `resticprofile schedule` copies env-file values into the generated systemd units, which are 0644, so every local user could read the rest-server password. All hosts except vm-esh-nas were moved with `playbooks/restic-repository-file.yaml`. vm-esh-nas has no infra-ops account; its script is `configs/restic/vm-esh-nas/migrate-repository-file.sh`, and Prime runs it with sudo. `restic.env` is kept (root 0600, not a leak) because the per-host READMEs and the freshness probe source it. **On a password rotation, update the vault, `restic.env` AND `repository`.**
- ⚠ **Never `env-file` for the URL.** Until 2026-09-27 the profiles loaded it with `env-file: /etc/restic/restic.env`. But `resticprofile schedule` copies env-file values into the generated systemd units, which are 0644, so every local user could read the rest-server password. Every host was moved with `playbooks/restic-repository-file.yaml` (vm-esh-nas once Prime had bootstrapped infra-ops there, the same day). `restic.env` is kept (root 0600, not a leak) because the per-host READMEs and the freshness probe source it. **On a password rotation, update the vault, `restic.env` AND `repository`.**
- **Pre-hook runs DB dumps into a staging dir**, then `restic backup` includes that dir alongside the regular paths. One snapshot = one point-in-time.
- **`restic forget` is scheduled; `restic prune` is not.** Rest-server's `--append-only` blocks prune from the client side by design. Prune is a manual ceremony (flip the flag, run prune, flip back).
@@ -1,30 +0,0 @@
#!/bin/bash
# vm-esh-nas: move restic from env-file to repository-file (2026-09-27).
# The same change playbooks/restic-repository-file.yaml made on the seven hosts
# where infra-ops has sudo. vm-esh-nas has no infra-ops account, so run it as:
# ssh -t vm-esh-nas 'sudo bash ~/restic-repofile-migrate.sh'
# It edits the env-file line IN PLACE (the live profile cannot be diffed without
# root, so any drift is preserved), and restores the old profile if a check fails.
set -eu
cd /etc/restic
grep -q '^ *env-file: /etc/restic/restic.env' profiles.yaml \
|| { echo "no env-file line: already migrated, or the profile differs; nothing done"; exit 1; }
val=$(sh -c 'set -a; . /etc/restic/restic.env; printf %s "$RESTIC_REPOSITORY"')
case "$val" in rest:http*) ;; *) echo "RESTIC_REPOSITORY is not a rest: URL; nothing done"; exit 1;; esac
umask 077
printf '%s\n' "$val" > repository.new
chown root:root repository.new; chmod 0400 repository.new; mv repository.new repository
unset val
cp -p profiles.yaml profiles.yaml.bak-20260927-envfile
trap 'cp -p profiles.yaml.bak-20260927-envfile profiles.yaml; resticprofile --no-ansi --config /etc/restic/profiles.yaml --name default schedule >/dev/null 2>&1 || true; echo "FAILED: old profile restored"' ERR
sed -i 's|^\( *\)env-file: /etc/restic/restic.env.*$|\1repository-file: /etc/restic/repository # not env-file: schedule copies env-file values into world-readable units|' profiles.yaml
resticprofile --no-ansi --config /etc/restic/profiles.yaml --name default cat config >/dev/null
resticprofile --no-ansi --config /etc/restic/profiles.yaml --name default schedule >/dev/null
for u in backup check; do
f=/etc/systemd/system/resticprofile-$u@profile-default.service
test -f "$f"
! grep -q 'rest:http' "$f"
done
systemctl is-active --quiet resticprofile-backup@profile-default.timer
trap - ERR
echo "vm-esh-nas migrated: units carry no repository URL, backup timer active"
+4 -6
View File
@@ -178,7 +178,7 @@ the stop→umount→rm-ghost→remount→start variant.
## Known gaps / TODO
### ✅ `resticprofile schedule` published the repo credential (found and fixed 2026-09-27, except vm-esh-nas)
### ✅ `resticprofile schedule` published the repo credential (found and fixed on all eight hosts, 2026-09-27)
On a host whose profile loads `RESTIC_REPOSITORY` from `env-file:
/etc/restic/restic.env`, `resticprofile schedule` copies that value, **including
@@ -200,11 +200,9 @@ containing `rest:http`. nh3-docker's scheduled unit then ran a real backup (snap
`a29b889d`). All seven hosts' URLs and passphrases are now vaulted as
`<host>/etc/restic/{repository,password}`.
**Still open: vm-esh-nas.** infra-ops has no account there, so its unit still
leaks. The migration script is staged at `~lkraven/restic-repofile-migrate.sh`
(repo copy `configs/restic/vm-esh-nas/migrate-repository-file.sh`), and Prime runs
it with `ssh -t vm-esh-nas 'sudo bash ~/restic-repofile-migrate.sh'`. Its secrets are
not vaulted yet, because that needs root there.
**vm-esh-nas followed the same day.** Prime bootstrapped infra-ops there
(`playbooks/bootstrap-infra-ops-user.yaml`), and the same playbook then migrated it.
Its secrets are vaulted too. **Zero restic hosts now leak.**
The passwords themselves were readable until the move, so the **rotation below is
still the real fix**. On rotation, update the vault, `restic.env` and `repository`
+2 -2
View File
@@ -148,8 +148,8 @@ _As of 2026-09-26 ~1620 PT._
so the systemd units no longer carry the rest-server password. Secrets are vaulted as
`<host>/etc/restic/{repository,password}`. `restic.env` is KEPT for manual snippets, so
**a rotation must update the vault, `restic.env` and `repository`.**
- **vm-esh-nas is still leaking**, because infra-ops has no account there. Prime runs
`ssh -t vm-esh-nas 'sudo bash ~/restic-repofile-migrate.sh'`. Its secrets are not vaulted.
- **vm-esh-nas done too** (Prime bootstrapped infra-ops there at uid 850, NOPASSWD). **All 8 hosts are clean
and vaulted.**
- The passwords were readable until today, so the rotation (Prime's) remains the real fix.
### augaman: face recognition for Cicada
+3 -1
View File
@@ -9,7 +9,9 @@ Dozzle agent, and Dockge.
- **LAN IP:** 10.0.50.154
- **FQDN:** `vm-esh-nas.esteban.net`
- **SSH:** `lkraven@vm-esh-nas` (key auth; `/etc/hosts` entry on the workstation)
- **SSH:** `infra-ops@10.0.50.154`: the fleet ops identity, uid/gid 850, NOPASSWD sudo, docker group.
Bootstrapped by Prime on 2026-09-27. `lkraven@vm-esh-nas` (the `vm-esh-nas` alias) is the human
account, key auth, with password sudo only.
## Hardware (from latest snapshot)
+2 -2
View File
@@ -34,8 +34,8 @@ Started by hand, then **fixed the same day**: `hosts/nh3-dev.yaml` no longer bin
the NAS shares and `BESZEL_EXTRA_FS` is empty. Their capacity is nh3-nas's own
volume, which nh3-nas's agent reports.
Use `infra-ops@<ip>` with passwordless sudo, except vm-esh-nas:
`lkraven@10.0.50.154` has Docker access. Irvine's hub address is
Use `infra-ops@<ip>` with passwordless sudo. vm-esh-nas has it too since 2026-09-27.
Before that date only `lkraven@10.0.50.154`, with Docker access, worked there. Irvine's hub address is
`100.64.0.6`; its retired `10.100.79.3` address caused silent loss of monitoring.
## Filesystems and deployment