diff --git a/docs/runbooks/esh-media-dac-10g.md b/docs/runbooks/esh-media-dac-10g.md new file mode 100644 index 0000000..0611f20 --- /dev/null +++ b/docs/runbooks/esh-media-dac-10g.md @@ -0,0 +1,95 @@ +# ESH-Media 10G uplink — the DAC that negotiated 1G, and the fix + +Fixed 2026-09-04. The link now runs at 10G by an explicit rate override on the +UDM side. **Read the pinned-port warning below before touching that port.** + + link ESH-Media SFP+1 (port 11) <--DAC--> ESH-UDMPM SFP+2 (port 11, eth10) + cable OEM "SFP-H10GB-CU3M", serial CSC200803000004, 3 m passive twinax + before both ends 1000 Mbps, autoneg on + after both ends 10000 Mbps + +This link backhauls the USW Flex XG "ESH-Office-Vuong", the E7-ESH-Media access +point, and the Zigbee coordinator — the office switch and the AP both connect at +10G and were squeezing through a 1G uplink, so this was a real bottleneck rather +than a cosmetic one. + +## The diagnosis, and what made it certain + +The cable was linking *cleanly* at the wrong speed — **zero rx/tx errors, no +LOS, no tx/rx fault, satisfaction 100**. That rules out a damaged or marginal +cable, which links at 10G and throws CRCs rather than dropping a rate. Both ends +were *choosing* 1G. + +The reason showed on the switch side: + + sfp_compliance: Unknown + +The switch reads the cable's EEPROM but cannot parse its compliance codes. +`SFP-H10GB-CU3M` is a Cisco part number and the vendor string is literally +`OEM` — a third-party cable wearing Cisco coding. With no recognised 10G +compliance byte, autonegotiation falls back to the safe rate. + +**The control case is what settled it.** On the *same UDM*, one port over: + + SFP+1 (eth9) TP-Link TL-SM5220-1M -> 10000 + SFP+2 (eth10) OEM SFP-H10GB-CU3M -> 1000 + +Same `speed_caps`, autoneg on both. So the port hardware, the firmware and +autoneg were all exonerated before anything was changed, and the cable was the +only remaining variable. + +## The fix + +An explicit rate override on the **UDM** end only: + + autoneg False speed 10000 full_duplex True setting_preference manual + +ESH-Media's end was left on autoneg and **followed to 10000 by itself** — which +is the proof the cable was always electrically capable of 10G and only the +EEPROM-driven negotiation was wrong. + +## ⚠ Why the UDM end and not the switch end + +The switch end is the unrecoverable one. ESH-Media reaches the controller +*through* this link, so a forced rate that fails to link leaves the switch with +no uplink and therefore no way to receive the revert — a physical visit. + +The UDM end is safe because the path to the controller does not cross this link: + + nh3-dev -> 10.100.10.1 -> 10.0.0.1 two hops, site VPN, terminates on the UDM + +Verified with `ip route get` / `traceroute` **before** the change, and the exact +pre-change `port_overrides` array was written to a revert payload before the +forward one was sent. If a future attempt needs *both* ends forced, that is a +different risk class and needs the operator's sign-off, not an agent's. + +## ⚠⚠ THAT PORT IS NOW PINNED — a 1G module will not link + +UDM SFP+2 is `autoneg False, speed 10000`. Plug anything that is not a 10G +module into it and the link stays down with no obvious cause. If the DAC is ever +replaced with a properly-coded one, **remove the override** and let autoneg do +its job rather than leaving a pin nobody remembers. + +## What the link bounce cost + +~60 seconds. ESH-Media, the AP and the office switch all returned `state=1` at +`uplink_speed=10000` with their **uptimes unbroken** (the switch did not reboot, +only the link renegotiated). The Zigbee coordinator and Home Assistant's path to +it were re-verified afterwards. + +## Errors: 200 at link-up, then flat + +The switch end logged ~200 rx / ~196 tx errors **during the renegotiation +window** and none afterwards. Counters were flat across the following two +minutes while real traffic flowed, and a longer soak was run to confirm. + +⚠ **This is the number to watch.** A forced rate on a cable the platform did not +want to run at 10G is exactly the case that can be marginal. Errors climbing +steadily — rather than a one-off burst at link-up — mean the cable really cannot +carry 10G cleanly and the override should be reverted and the cable replaced. + +## Rollback + +Set UDM SFP+2 back to `autoneg True, setting_preference auto`, clearing `speed` +and `full_duplex`. Remember `port_overrides` is a **whole-array PUT**: the UDM +carries 2 overrides and anything omitted is deleted.