fix(elway): probe NOPASSWD with sudo -n true, never sudo -n -v
`sudo -v` refreshes the auth timestamp, and a NOPASSWD-only rule creates no timestamp to refresh, so on sudo >= 1.9.15 `sudo -n -v` returns non-zero while every real command runs passwordless. Measured: pfi-gx10 sudo 1.9.15p5 sudo -n -v rc=1 sudo -n true rc=0 nh3-docker sudo 1.9.13p3 sudo -n -v rc=0 sudo -n true rc=0 ana-docker sudo 1.9.13p3 sudo -n -v rc=0 sudo -n true rc=0 irv-ml1 sudo 1.9.13p3 sudo -n -v rc=0 sudo -n true rc=0 Only pfi-gx10 is new enough to hit it today, but every host does as it moves past 1.9.13, and the failure mode is bad: elway prompts for a password on a host with working NOPASSWD sudo, which in a non-interactive run is an EOFError partway through a playbook. The same probe in my own notes cost this session directly. gx10 looked like a fleet exception with no NOPASSWD sudo when it had it from account creation, and the operator was asked for a password that was never needed. Corrected in auto-memory too. Also lands the gx10 privileged outfit playbook, now green at 5/5: NOPASSWD sudo, nvcc, docker group, a CUDA container seeing the GB10, and the userspace torch stack still working afterward.
This commit is contained in:
@@ -181,7 +181,62 @@ Consequences:
|
|||||||
- ⚠ **Any library shipping cubin-only binaries with no PTX will simply fail here.** That is the
|
- ⚠ **Any library shipping cubin-only binaries with no PTX will simply fail here.** That is the
|
||||||
real shape of the "per-arch unknown" risk — not source builds, but prebuilt kernels.
|
real shape of the "per-arch unknown" risk — not source builds, but prebuilt kernels.
|
||||||
|
|
||||||
### 🔴 BLOCKED ON SUDO — the privileged half is not done
|
### ✅ PRIVILEGED HALF DONE — and I WAS NEVER BLOCKED
|
||||||
|
|
||||||
|
⚠⚠ **The block was my own bad test.** `infra-ops` has had **NOPASSWD sudo since account
|
||||||
|
creation** (`/etc/sudoers.d/infra-ops`, installed 11:31 with the account). I probed with
|
||||||
|
`sudo -n -v`, which returns non-zero on **sudo >= 1.9.15** because `-v` refreshes an auth
|
||||||
|
timestamp that a NOPASSWD-only rule never creates — and asked the operator for a password that
|
||||||
|
was never needed.
|
||||||
|
|
||||||
|
pfi-gx10 sudo 1.9.15p5 sudo -n -v rc=1 sudo -n true rc=0
|
||||||
|
nh3-docker sudo 1.9.13p3 sudo -n -v rc=0 sudo -n true rc=0
|
||||||
|
ana-docker sudo 1.9.13p3 sudo -n -v rc=0 sudo -n true rc=0
|
||||||
|
irv-ml1 sudo 1.9.13p3 sudo -n -v rc=0 sudo -n true rc=0
|
||||||
|
|
||||||
|
**Always `sudo -n true`.** Fixed in `scripts/elway` (which carried the same probe and would
|
||||||
|
have started falsely prompting fleet-wide as hosts pass sudo 1.9.13) and in auto-memory
|
||||||
|
`reference_infra_ops_sudo_identity`, which *recommended the broken test*.
|
||||||
|
[[feedback_assert_effective_value_not_substring]]
|
||||||
|
|
||||||
|
⚠ **Same shape twice in one session.** I also reported "nvcc: ABSENT" from `command -v nvcc`
|
||||||
|
when the CUDA toolkit was installed all along at `/usr/local/cuda/bin/nvcc`, just not on PATH.
|
||||||
|
Both times I checked a proxy and reported it as the thing itself.
|
||||||
|
|
||||||
|
**Credential:** the temp password was `infra-ops`'s, not `lkraven`'s. Rotated to 32 random
|
||||||
|
chars and **vaulted at `pfi-gx10/infra-ops-password`** (round-trip verified). Kept rather than
|
||||||
|
locked — deviating from the canonical bootstrap's no-password posture **deliberately, because
|
||||||
|
this box has no iDRAC/BMC**, so a console break-glass credential is the only way back if key
|
||||||
|
auth breaks. ⚠ `lkraven` (uid 1000, GUI desktop user, `sudo` group) has a password nobody
|
||||||
|
knows; out of scope, but it is a gap.
|
||||||
|
|
||||||
|
### ⚠ DOCKER'S BUILDKIT DB WAS ALREADY CORRUPT — the restart surfaced it, did not cause it
|
||||||
|
|
||||||
|
`nvidia-ctk runtime configure` + `systemctl restart docker` failed with
|
||||||
|
`error initializing buildkit: error creating buildkit instance: invalid database`. **Not the
|
||||||
|
nvidia config** — `daemon.json` was valid and `nvidia-container-runtime` was present. dockerd
|
||||||
|
had been running on a process whose on-disk buildkit state was already bad; it simply could not
|
||||||
|
come back. Remediated by removing `/var/lib/docker/buildkit` (108 KB of **cache**, and the box
|
||||||
|
had **0 images, 0 containers**), then `daemon-reload` + `reset-failed` + start.
|
||||||
|
|
||||||
|
⚠ **A long-running daemon can be unable to restart for hours before anyone finds out.** Nothing
|
||||||
|
reported unhealthy until something asked it to restart.
|
||||||
|
|
||||||
|
### ✅ FINAL STATE — playbook green, 5/5 verifies
|
||||||
|
|
||||||
|
playbooks/gx10-outfit-privileged.yaml -> steps 1 ok / 1 changed / 0 failed / 6 skipped
|
||||||
|
verify 5 ok / 0 failed
|
||||||
|
|
||||||
|
NOPASSWD sudo OK
|
||||||
|
nvcc /usr/local/cuda/bin/nvcc (NOT on default PATH)
|
||||||
|
docker socket as infra-ops OK (docker group)
|
||||||
|
CUDA container sees the GB10 OK (nvidia/cuda:13.0.0-base -> nvidia-smi -L reports GB10)
|
||||||
|
userspace torch still sees GPU OK
|
||||||
|
docker Runtimes nvidia runc io.containerd.runc.v2
|
||||||
|
|
||||||
|
**The box is fully outfitted for inference + training, bare metal, containers available.**
|
||||||
|
|
||||||
|
### Historical note — what the block looked like before it was disproved
|
||||||
|
|
||||||
⚠ **`infra-ops` does NOT have NOPASSWD sudo on pfi-gx10** (`sudo -n -v` -> "a password is
|
⚠ **`infra-ops` does NOT have NOPASSWD sudo on pfi-gx10** (`sudo -n -v` -> "a password is
|
||||||
required"), unlike the rest of the fleet — see [[reference_infra_ops_sudo_identity]], which
|
required"), unlike the rest of the fleet — see [[reference_infra_ops_sudo_identity]], which
|
||||||
|
|||||||
+10
-2
@@ -133,8 +133,16 @@ _As of 2026-09-01 — **the GX10 is on the operator's desk, NOT racked. Standing
|
|||||||
**first use of every kernel pays a JIT compile** and cubin-only libraries will fail.
|
**first use of every kernel pays a JIT compile** and cubin-only libraries will fail.
|
||||||
⚠⚠ **ALWAYS WARM UP BEFORE TIMING** — an unwarmed benchmark read 27 TFLOP/s (it was timing the
|
⚠⚠ **ALWAYS WARM UP BEFORE TIMING** — an unwarmed benchmark read 27 TFLOP/s (it was timing the
|
||||||
JIT) and nearly became a phantom "tensor cores are broken" report.
|
JIT) and nearly became a phantom "tensor cores are broken" report.
|
||||||
🔴 **BLOCKED ON SUDO** for nvcc / container toolkit / docker group: **infra-ops has NO NOPASSWD
|
✅ **PRIVILEGED HALF DONE 2026-09-01** — nvcc, NVIDIA Container Toolkit, docker group; playbook
|
||||||
sudo on gx10** (fleet exception) and **no credential is vaulted**.
|
`gx10-outfit-privileged.yaml` green 5/5 (a CUDA container sees the GB10). **Bare metal, fully
|
||||||
|
outfitted.** ⚠⚠ **I was NEVER blocked** — infra-ops had NOPASSWD sudo all along; I probed with
|
||||||
|
**`sudo -n -v`, which is WRONG on sudo >= 1.9.15** (it refreshes a timestamp NOPASSWD never
|
||||||
|
creates) and asked the operator for a password that did not exist to need. **Always
|
||||||
|
`sudo -n true`.** Fixed in `scripts/elway` (same probe, would have broken fleet-wide as hosts
|
||||||
|
pass sudo 1.9.13) and in auto-memory. Password rotated + vaulted at
|
||||||
|
`pfi-gx10/infra-ops-password` (kept not locked: **no iDRAC on this box**, so it is the only
|
||||||
|
console break-glass). ⚠ Docker's buildkit DB was **already corrupt** — the restart surfaced it,
|
||||||
|
did not cause it; cleared `/var/lib/docker/buildkit` (cache, 0 images/containers).
|
||||||
→ `persistent-memory.d/2026-09-01-pfi-gx10-onboarding.md`
|
→ `persistent-memory.d/2026-09-01-pfi-gx10-onboarding.md`
|
||||||
- **▶ ADA INFERENCE SERVER (R750xa `JPJ1ZP3`) — awaiting an iDRAC inventory.** 2× RTX 6000
|
- **▶ ADA INFERENCE SERVER (R750xa `JPJ1ZP3`) — awaiting an iDRAC inventory.** 2× RTX 6000
|
||||||
Ada to be fitted; ComfyUI's new home at NH3. **Order regardless: 8× `M04W6` RDIMM (restores
|
Ada to be fitted; ComfyUI's new home at NH3. **Order regardless: 8× `M04W6` RDIMM (restores
|
||||||
|
|||||||
+18
-2
@@ -426,9 +426,25 @@ def ensure_sudo(ctx: SSHContext) -> None:
|
|||||||
if ctx.sudo_probed:
|
if ctx.sudo_probed:
|
||||||
return
|
return
|
||||||
ctx.sudo_probed = True
|
ctx.sudo_probed = True
|
||||||
# Is sudo configured NOPASSWD for this user? `sudo -n -v` exits 0 if so.
|
# Is sudo configured NOPASSWD for this user?
|
||||||
|
#
|
||||||
|
# ⚠ This MUST be `sudo -n true`, never `sudo -n -v`. `-v` refreshes the
|
||||||
|
# auth timestamp, and a NOPASSWD-only rule creates no timestamp to refresh,
|
||||||
|
# so on **sudo >= 1.9.15** `-n -v` returns non-zero even though every real
|
||||||
|
# command runs passwordless. Measured 2026-09-01:
|
||||||
|
#
|
||||||
|
# pfi-gx10 sudo 1.9.15p5 sudo -n -v rc=1 sudo -n true rc=0
|
||||||
|
# nh3-docker sudo 1.9.13p3 sudo -n -v rc=0 sudo -n true rc=0
|
||||||
|
# ana-docker sudo 1.9.13p3 sudo -n -v rc=0 sudo -n true rc=0
|
||||||
|
# irv-ml1 sudo 1.9.13p3 sudo -n -v rc=0 sudo -n true rc=0
|
||||||
|
#
|
||||||
|
# The old probe cost a whole session on gx10: elway prompted for a password
|
||||||
|
# on a host that had working NOPASSWD sudo the entire time, and in a
|
||||||
|
# non-interactive run that prompt is an EOFError mid-playbook. Every host
|
||||||
|
# will hit this as it moves past sudo 1.9.13. Assert the effective
|
||||||
|
# behaviour (can I run a command?), not a proxy for it.
|
||||||
p = subprocess.run(
|
p = subprocess.run(
|
||||||
ctx.ssh_cmd("sudo -n -v"),
|
ctx.ssh_cmd("sudo -n true"),
|
||||||
stdout=subprocess.DEVNULL,
|
stdout=subprocess.DEVNULL,
|
||||||
stderr=subprocess.DEVNULL,
|
stderr=subprocess.DEVNULL,
|
||||||
)
|
)
|
||||||
|
|||||||
Reference in New Issue
Block a user