Files
esh-pfi-infrastructure/servers/pfi-tacticalrmm/README.md
T

109 lines
7.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# pfi-tacticalrmm
Tactical RMM (remote monitoring + management) server at the Anaheim colo.
## Network
- **LAN IP:** 10.250.50.57
- **SSH:** `lkraven@pfi-tacticalrmm`
## Infrastructure
- **Hypervisor:** `pfi-pve` (VMID **111**)
- **Type:** Linux VM
- **Site:** Anaheim (PFI colo)
## Role
[TacticalRMM](https://tacticalrmm.com/) — open-source RMM platform.
Monitors and manages endpoints, pushes patches, runs scripts, etc.
## MeshCentral (bundled with TacticalRMM) — facts checked 2026-10-02
- Runs natively (`meshcentral.service`, user `tactical`, `/meshcentral`), Node 18.20.8, MeshCentral 1.2.0,
postgres-backed. Public at `https://rmm-mesh.phasefinal.com` (nginx terminates TLS, `tlsOffload`), MPS
(Intel AMT CIRA) at `rmm-mesh.phasefinal.com:4433`. 2FA is NOT forced (`force2factor` unset).
- **HYBRID since 2026-10-02 0812 (Prime: "hybrid, make it so").** `settings.WANonly` set false (backup
`config.json.bak-20261002-hybrid`); the log says "Hybrid (LAN + WAN) mode". All 14 connected agents came
back. **nh3-pve's AMT is in `PFI-AMT` as `nh3-pve-amt`** (10.100.250.61, TLS, admin): MeshCentral reached it
at once (AMT 16.1.25, activated, power on), so the old-TLS worry did not apply. ⚠ The path still runs
through nh3-scale (CT 107 ON nh3-pve): with nh3-pve down, this AMT is unreachable from here. KVM needs
an active iGPU output: the NanoKVM serves today; fit the 1080p dummy plug before it moves.
- **Phone-home (CIRA) plumbing, 2026-10-02 (Prime go):** `settings.mpsPass` set (vaulted
`pfi-tacticalrmm/meshcentral-mpspass`; without it the MPS checked only the 16-char username). FortiGate
ana-gw: service `MeshCentral-MPS-4433`, VIP `mps-to-tacticalrmm` (38.120.12.46:4433 → 10.250.50.57) and
policy 76 (wan1→servers, accept) — 4433 verified open from the internet, 4434 closed as a control. The AMT
side is `scripts/amt-cira-setup.py`. **nh3-pve's AMT phones home since 0859** after moving it from static IP
to DHCP (Intel: CIRA does not work on a static IP; the FortiGate sniffer had shown zero attempts before).
The CIRA entries are `nh3-pve-amt` (old LAN entry removed) and `nh3-pve-2-amt` (MS-03, AMT 21.0.6, from 2026-10-02 1449;
configured on DHCP from the start, so it phoned home the moment its environment detection was set). Two things it needed: the AMT
credentials set on the device (`changedevice` intelamt user/pass) and `intelamt.tls` = 1. Then a
MeshCentral restart re-ran its AMT manager, which logged in at once (16.1.25, power on). ⚠ MeshCentral
1.2.0 `amtmanager.js` picks TLS-vs-not over CIRA with `boundPorts.indexOf('16992')` used as a boolean
(−1 is truthy), so a TLS-only AMT can be tried without TLS. Set `tls` explicitly as above.
- **Third CIRA device: `esh-pve-2-amt`** (MS-03 at ESH, AMT 21.0.6, tunnel from ESH's public IP 128.177.138.182),
2026-10-02 2218. Same recipe. Instead of restarting MeshCentral after setting credentials, only that device's
tunnel was dropped (`sudo ss -K -tn "dst [::ffff:<public-ip>]:664"`). AMT reconnected within seconds, and the AMT
manager logged in with the new credentials. The other tunnels and the agents were untouched.
- ⚠ **A device's FIRST CIRA connection can crash MeshCentral 1.2.0** (it did once, 2026-10-02 22:18:07; the launcher
restarted it at 22:18:14 and all 14 agents came back). Read at source: in `mpsserver.js`, password-auth path for
a device not yet in the DB, `socket.tag.meshid` is set only inside the reverse-DNS callback, while
`addCiraConnection()` runs at once and its 300 ms timer reads `meshid`. `SetConnectivityState` →
`NotifyUserOfDeviceStateChange` then calls `meshid.split` on undefined, which is uncaught, so the whole server
restarts. The likely trigger is a slow PTR lookup on the source IP: ESH's resolves at zayo.com, while NH3's
has none and its two first connects did not crash. That is inferred, not measured. Reconnects of
an existing device take the other branch and are safe. **Onboard new AMTs at a quiet hour**; expect one ~7 s
MeshCentral restart.
- ⚠ **"HW Connect" stuck at "Setup…" = a stale CIRA tunnel** (2026-10-02 2247, esh-pve-2-amt; Prime power-cycled
the host several times). The AMT opened a new tunnel without closing the old one. MeshCentral then held two:
the live one (power actions and the AMT manager worked through it) and a dead one, `backoff:14`, nothing
received for 782 s. `webrelay.ashx` (KVM, SOL, IDE-R, MeshCommander) picked the dead one. A relay TLS failure
is only logged at debug level and never closes the websocket, so the browser waits forever. MPS's 90 s idle
timeout did not fire, because MeshCentral's own writes to the dead socket kept resetting it. TCP gives up on
its own after ~15–30 min of retransmits.
**Fix:** `sudo ss -tnio state established "( sport = :4433 )"`. Find the device's public IP with a large
`lastrcv` (ms) or a `backoff`. Healthy tunnels show `lastrcv` of seconds. Then
`sudo ss -K -tn "dst [::ffff:<ip>]:<its port>"`. Verify with `scripts/meshcentral-amt-relay-probe.js`, using
another AMT as a positive control. Likely whenever a host whose AMT shares its NIC (esh-pve-2) power-cycles.
Seen once so far.
- Before 2026-10-02: `"WANonly": true` (TacticalRMM's install default). In that mode MeshCentral SILENTLY DROPS
"Add Intel AMT computer": `meshuser.js` line 2682, `if (args.wanonly == true) return;`. No error, no
event. LAN-mode AMT needs `WANonly` false (hybrid) + a service restart; CIRA works in WAN mode. TacticalRMM's
`update.sh` only touches the compression keys of `config.json`, so a WANonly change survives updates.
- Device groups: `TC2-MacMini` (agent group), `PFI-AMT` (mtype 1, Intel AMT only; created by Prime
2026-10-02, empty). 24 devices visible to Prime's account, 7 of them report Intel AMT.
- **Site-admin account `lkraven` (2026-10-02, Prime):** full site admin (users, server files), full rights on
`PFI-AMT` and `TC2-MacMini`; password vaulted `pfi-tacticalrmm/meshcentral-lkraven-password`. Created with
`node node_modules/meshcentral --createaccount lkraven --hashpass <salt,hash>` + `--adminaccount lkraven` while
the service was stopped (the CLI writes the DB directly). TacticalRMM's mesh sync only manages `name___N`
accounts, so it leaves this one alone. TRMM's own MeshCentral superuser is `fosxggiw`. ⚠ Not in 2FA yet.
- **Uploads ("My Files") need nginx's limit raised:** `/etc/nginx/sites-enabled/meshcentral.conf` had no
`client_max_body_size`, so nginx's 1 MB default 413'd every real upload. Since 2026-10-02:
`client_max_body_size 4G` + `proxy_request_buffering off` (backup `.bak-20261002-upload`). Site admins have no
MeshCentral quota. ⚠ Re-check after a TacticalRMM update or reinstall rewrites the nginx sites.
- CLI access: `meshctrl.js` on this host with Prime's login token, vaulted
`pfi-tacticalrmm/meshcentral-login-token` (one line: `username: ~t:… password: …`). Example:
`sudo -n -u tactical node /meshcentral/node_modules/meshcentral/meshctrl.js listdevicegroups --url wss://rmm-mesh.phasefinal.com --loginuser <u> --loginpass <p>`.
Feed the credentials over stdin; never put them in a file or a log.
## Backup coverage
- **VM-image:** ✅ vzdump on pfi-pve (daily)
- **File-level restic:** ❌ not yet configured
TacticalRMM state (Postgres DB with inventory + automation history,
MeshCentral config, agent registrations) is critical if used in
production. Worth setting up app-consistent DB dumps + file-level
restic if it's the primary management plane.
## Refresh state
```bash
scripts/refresh-server-info.sh pfi-tacticalrmm
```
## Discovered via
`scripts/discover-fortigate.sh 10.250.250.1` on 2026-04-21 —
FortiGate DHCP lease (MAC `ba:fa:65:f6:46:25`).