Files
esh-pfi-infrastructure/docs/runbooks/esh-media-dac-10g.md
T
vh 054c098777 fix(esh): revert the forced 10G — the link was flaking hard, operator called it
Both ends are back on autoneg at 1000 Mbps and stable; UDM SFP+2 restored to
autoneg True / setting_preference auto from the payload staged before the
original change. ~50 s of renegotiation, nothing rebooted.

The force failed progressively and I called it too early. Errors on the
ESH-Media side went 200 at link-up, 221 at 42 minutes, 416 by four hours, with
flapping the operator noticed before my soak had accumulated enough to raise
it. I read a two-minute flat window as a plateau and said so; a marginal link
declares itself over hours, not minutes.

Do not re-force this port. The cable is being replaced at the weekend and a
correctly-coded DAC negotiates 10G on its own.
2026-09-04 14:41:02 -07:00

5.4 KiB

ESH-Media 10G uplink — the DAC that negotiated 1G, and the fix

2026-09-04. The link runs at 10G by an explicit rate override on the UDM side. Read the pinned-port warning below before touching that port.

CURRENT STATE: REVERTED to autoneg / 1G. The force is gone.

2026-09-04 14:38 — the operator reported the 10G link "flaking hard" and I reverted it. Both ends are back at 1000 Mbps, UDM SFP+2 is autoneg True / setting_preference auto, ESH-Media is reachable and stable. The revert took about 50 s of renegotiation and nothing rebooted.

The forced-10G experiment failed, and the failure was progressive. Error counts on the ESH-Media side over the life of the force:

10:03   200      link-up burst
10:48   221      +21 in 42 min  -- the first sign it was not plateauing
14:38   416      +195 over the next 4 h, plus user-visible flapping

The lesson, and it is mine to own: I read a two-minute flat window as a plateau and reported it as one. The counter simply had not moved yet. A marginal link declares itself over hours, not minutes — and the operator felt the flapping before the soak I left running had accumulated enough to call it.

Do not re-force this port. The cable is being replaced at the weekend; a correctly-coded DAC negotiates 10G on its own, which is the actual fix.

The diagnosis, and what made it certain

The cable was linking cleanly at the wrong speed — zero rx/tx errors, no LOS, no tx/rx fault, satisfaction 100. That rules out a damaged or marginal cable, which links at 10G and throws CRCs rather than dropping a rate. Both ends were choosing 1G.

The reason showed on the switch side:

sfp_compliance:  Unknown

The switch reads the cable's EEPROM but cannot parse its compliance codes. SFP-H10GB-CU3M is a Cisco part number and the vendor string is literally OEM — a third-party cable wearing Cisco coding. With no recognised 10G compliance byte, autonegotiation falls back to the safe rate.

The control case is what settled it. On the same UDM, one port over:

SFP+1 (eth9)   TP-Link TL-SM5220-1M    ->  10000
SFP+2 (eth10)  OEM SFP-H10GB-CU3M      ->   1000

Same speed_caps, autoneg on both. So the port hardware, the firmware and autoneg were all exonerated before anything was changed, and the cable was the only remaining variable.

The fix

An explicit rate override on the UDM end only:

autoneg False   speed 10000   full_duplex True   setting_preference manual

ESH-Media's end was left on autoneg and followed to 10000 by itself — which is the proof the cable was always electrically capable of 10G and only the EEPROM-driven negotiation was wrong.

⚠ Why the UDM end and not the switch end

The switch end is the unrecoverable one. ESH-Media reaches the controller through this link, so a forced rate that fails to link leaves the switch with no uplink and therefore no way to receive the revert — a physical visit.

The UDM end is safe because the path to the controller does not cross this link:

nh3-dev -> 10.100.10.1 -> 10.0.0.1     two hops, site VPN, terminates on the UDM

Verified with ip route get / traceroute before the change, and the exact pre-change port_overrides array was written to a revert payload before the forward one was sent. If a future attempt needs both ends forced, that is a different risk class and needs the operator's sign-off, not an agent's.

The port is NOT pinned any more — the override was removed 2026-09-04

While the force was in place, UDM SFP+2 was autoneg False, speed 10000 and a non-10G module would not have linked there. That is no longer true: the override is gone and the port is back on autoneg. Recorded because the warning stood in this runbook for four hours and someone may remember it.

~60 seconds. ESH-Media, the AP and the office switch all returned state=1 at uplink_speed=10000 with their uptimes unbroken (the switch did not reboot, only the link renegotiated). The Zigbee coordinator and Home Assistant's path to it were re-verified afterwards.

The switch end logged ~200 rx / ~196 tx errors during the renegotiation window and none afterwards. Counters were flat across the following two minutes while real traffic flowed, and a longer soak was run to confirm.

This is the number to watch, and it went the wrong way. A two-hour soak showed the errors do not plateau: 200 → 221 rx and 196 → 199 tx over 42 minutes, monotonic, about 1 rx error per 23.2 MB — plus one link drop at 10:12 that recovered on its own with no device reboot.

⚠⚠ So an earlier claim in this runbook's own history was wrong and is corrected here: a clean, zero-error link at 1G does NOT rule out a marginal cable. It only proves the cable is clean at 1G. Clean at 1G and marginal at 10G is exactly what a 3 m OEM twinax with unparseable compliance codes turns out to be — which means the platform's autoneg fallback was protecting something real, not being fussy about vendor coding. For comparison, the TP-Link TL-SM5220 on the adjacent UDM port carries 780 GB with 4 errors.

The fix is the cable, not the configuration.

Rollback

Set UDM SFP+2 back to autoneg True, setting_preference auto, clearing speed and full_duplex. Remember port_overrides is a whole-array PUT: the UDM carries 2 overrides and anything omitted is deleted.