An infrastructure day with no training work, and the through-line is that every fault was invisible to monitoring. NH3↔Anaheim had been crossing a throttled DERP relay rather than a direct path for long enough to carry 78 GB; `.internal` DNS was failing roughly one lookup in ten from two independent causes; SearXNG had exactly one working general web engine. Nothing alarmed on any of it. All three surfaced because tts-dev had a 1545 ms voice-loop budget and chose to measure rather than adapt around the problem. Also landed: althing v3.6.3, which makes hyphenated search work for the first time on a fleet whose hostnames are nearly all hyphenated; FleetTools, a capability index autoloaded by Claude, Codex and Grok from one symlinked file; Miranda's Hermes plugin moved from a copy to a repo symlink; Worldtree's env.sh secrets vaulted. Six detail files. The in-flight section is rewritten and shrinks 142 lines to 64 — it opens on the one thing this session did NOT verify, lv-mccarthy's run outcome, which was left untouched and must not be assumed good. Two foot-guns recorded, both mine: the ESH egress experiment reverted on a diagnosis the rollback itself falsified, and `!ENV` in searxng settings, which has no constructor in that build and crash-looped the container ten times. No archival this run. 165 of 169 dated entries are under the 14-day guard and the remaining four all carry open deferred pointers, so the index stays over the soft cap at 480 lines — an over-cap file that keeps live decisions beats a scannable one that lost one.
2.5 KiB
[2026-09-18] SearXNG's ESH SOCKS5 egress — one day live, reverted, and it fixed nothing
On 2026-09-17, at operator request, SearXNG's search egress was moved from
nh3-docker's direct residential path to socks5h://10.0.50.65:1080 on esh-scale
(CT 108 on esh-pve) — an application-level proxy, no host route or exit-node change.
Reverted 2026-09-18.
Why it was reverted, and why the STATED reason was wrong
It was reverted on a diagnosis that turned out to be false: that the move had cost
three of four search engines to CAPTCHAs. Reverting to direct NH3 egress produced a
byte-identical result — same three engines down — which falsified it. A live
!ddg bang probe on a freshly restarted container also CAPTCHA'd, ruling out stale
suspension timers too.
The real cause was elsewhere (see 2026-09-18-searxng-one-engine-to-seven).
The revert still stands on its own merits: ESH egress bought no measurable improvement while making all fleet search depend on ESH WAN and mesh availability. The simpler configuration is the better one. It is simply not the fix for the engines.
What was built, because it was built correctly
The proxy side was not at fault and is worth keeping as a pattern. microsocks ran
as nobody under searxng-egress.service, bound 10.0.50.65:1080 only, and allowed
source 10.100.50.40 alone — everything else had to supply a password regenerated at
each start and never distributed. Allowed-host egress and denied-host rejection were
both tested. The unit is kept at configs/esh-scale/searxng-egress.service; the
service is stopped and disabled on esh-scale, and tailscaled there was not
touched (CT 108 is ESH's whole-site mesh SPOF).
Measured egress was 128.177.138.182 in use and 154.50.58.126 at cutover — ESH's
WAN address moves and nothing pins it, so the number in the stack README is historical.
The transferable lesson
⚠ I asserted causation from correlation with no baseline. The only evidence that residential egress avoided CAPTCHAs was a config comment dated 2026-09-03, which was no longer true of that address. Five samples of the post-change state and zero of the working state is not a comparison. The rollback WAS the counterfactual, and it falsified the claim I had already published in a commit message.
This was the first of three wrong causal attributions in a single afternoon. The common shape: measure after a change, attribute the delta to my change, never check what else moved.
Commits 156e126 (applied), 1a35181 (reverted).