fix(beszel): GPU-LXC temperature alerts watch the GPU, not the hypervisor CPU
nh3-ml1's first Temperature alert (2026-09-25 2105, 88.8 C) was nh3-pve's CPU
package during the nightly vzdump; the GPU sat at 47 C. hwmon is not namespaced,
so an LXC's agent reports its host's coretemp sensors under the LXC's name.
- hosts/{nh3,esh}-ml1.yaml: SENSORS=-coretemp_*,acpitz (blacklist); GPU and
NVMe temperatures verified still reported.
- Hub: Temperature >95 C / 5 m on nh3-pve and esh-pve for the CPUs (i9-13900H
TjMax 100 C; nh3-pve 2-hour averages reach 88 C on busy days).
- nh3-pve README: runs warmer than esh-pve; workload confounds it, check
airflow on the next visit.
This commit is contained in:
@@ -117,6 +117,16 @@ not power back on by itself.
|
||||
oops and has run fine, so it is benign so far: only the BT worker dies. The fix
|
||||
is to blacklist `btusb` on both hypervisors. That is not done, because it only
|
||||
takes effect at the next boot.
|
||||
- **Runs warmer than esh-pve, its twin (Beszel, 2026-09-15 → 09-25).**
|
||||
- nh3-pve's CPU package: median of 2-hour averages 76 °C, peak 2-hour average
|
||||
88 °C (2026-09-23, CPU ~15%).
|
||||
- esh-pve: medians 52–55 °C, peak 20-min average 78 °C.
|
||||
- During the 2026-09-25 2100 vzdump, the 1-minute samples reached 90 °C at
|
||||
13–21% CPU. They were back to 61 °C once the job finished.
|
||||
- Workload is a confound: nh3-pve usually carries more load, so this does not
|
||||
prove a cooling fault. Worth checking airflow and dust on the next visit.
|
||||
- TjMax is 100 °C. The Beszel CPU alert on this host is set at >95 °C for
|
||||
5 minutes.
|
||||
- **lxc-pve 6.0.0-1 → 6.0.0-2** (2026-09-25 1533, that one package only). This is
|
||||
Proxmox fix #7006. Without it, runc 1.5 inside a nesting CT fails every
|
||||
`docker run`. `playbooks/gpu-lxc.yaml` now upgrades it first.
|
||||
|
||||
Reference in New Issue
Block a user