Files
esh-pfi-infrastructure/stacks/searxng/conf/searxng-settings.yml
T
vh 1a35181b67 revert(searxng): return search egress to direct NH3
Reverts the outgoing.proxies block added in 156e126. Canonical restored from
that commit's parent and verified byte-identical to the host's
searxng-settings.yml.pre-esh-20260917 backup, then deployed via
scripts/deploy-stack.sh so canonical and host converge rather than drift. The
esh-scale searxng-egress.service is stopped and disabled; tailscaled on that
container was not touched.

⚠ THE ROLLBACK DID NOT RESTORE THE ENGINES, WHICH FALSIFIES THE REASON GIVEN
FOR IT. 156e126 recorded that moving egress to ESH had cost three of four
engines. Measured after this revert, with egress confirmed back on
70.230.226.88 and the same instrument used for the before-measurement, the
result is identical: brave and startpage suspended, duckduckgo CAPTCHA, google
cse the only engine answering. Per-engine bang probes confirm duckduckgo is
CAPTCHA-ing the residential address live, so this is not a stale suspension
timer.

The engine failures therefore have some other cause and predate or are
independent of the ESH move. The claim in 156e126 asserted causation from a
correlation without measuring the pre-change state; the only evidence for
"residential egress avoids CAPTCHAs" was a comment dated 2026-09-03, which is
no longer true of this address.

The revert still stands on its own merits: ESH egress bought no measurable
improvement while adding a hard dependency on ESH WAN and mesh availability
for all fleet search, so the simpler configuration is the better one. It is
simply not the fix for the engines.

README rewritten to match: direct NH3 is documented as current, the ESH
attempt is kept as history with its measured outcome, and the health script's
blind spot is called out — scripts/searxng-health.sh prints a passing result
while three engines are blocked, because it gates on "any results returned"
and treats failed engines as informational. That script needs to fail on
blocked engines before any future egress change, or the next regression is
equally invisible.
2026-09-18 12:34:52 -07:00

80 lines
2.9 KiB
YAML

# SearXNG — PFI fleet meta-search. Deployed on nh3-docker (10.100.50.40:9996).
#
# ⚠ WHY NH3 AND NOT THE COLO. Measured 2026-09-03:
# ana-docker egress 38.120.12.42 (datacenter) -> DuckDuckGo + Startpage CAPTCHA
# nh3-docker egress 70.230.226.88 (residential) -> no CAPTCHA
# Search engines gate datacenter ranges. Same reason the fleet keeps a
# residential SOCKS5 egress proxy on nh3-dev for yt-dlp. Running the search
# aggregator from a residential-egress site removes the problem at the source
# rather than proxying around it.
use_default_settings:
engines:
remove:
# Onion engines: no Tor proxy is configured here, so they only ever
# contribute timeouts.
- ahmia
- torch
# ⚠ Removal keys must match the engine's REAL name, spaces and all.
# `karmasearch.videos` (dotted) did NOT match on the old instance and the
# engine kept appearing in unresponsive_engines despite being "removed".
# The name is "karmasearch videos".
- karmasearch
- karmasearch videos
general:
instance_name: "SearXNG"
instance_about_url: false
contact_url: false
debug: false
# Public metrics page off — smaller attack surface on an unauthenticated
# internal service.
enable_metrics: false
search:
safe_search: 0
autocomplete: ""
default_lang: "auto"
# `json` is what makes this usable as a tool rather than only a web page.
# Removing it breaks every non-browser consumer, including Claude sessions.
formats:
- html
- json
# 3s is too tight for slower engines; 8s covers them without hanging the UI.
request_timeout: 8.0
# Ban an engine only briefly when it raises suspended-time. The default 86400
# means one bad afternoon silences an engine for a day.
ban_time_on_fail: 60
max_ban_time_on_fail: 600
server:
# secret_key comes from SEARXNG_SECRET in the environment — never hardcode it
# here. Generated + vaulted at nh3-docker/searxng-secret.
bind_address: "0.0.0.0"
port: 8080
# Enable ONLY with a limiter.toml AND a proxy that forwards X-Real-IP;
# otherwise it logs "X-Forwarded-For nor X-Real-IP header is set!" forever.
limiter: false
public_instance: false
# ⚠ Kept in sync with BASE_URL in compose.yaml. The old instance still said
# `https://searxng.pfi.local/` here — a name retired on 2026-08-19 — while the
# environment said something else. The env wins, so nothing broke, and the
# file quietly lied to everyone who read it.
base_url: "http://10.100.50.40:9996/"
method: "GET"
compression: true
image_proxy: false
# Outgoing pool — low-traffic private instance.
outgoing:
request_timeout: 6.0
max_request_timeout: 12.0
pool_connections: 100
pool_maxsize: 20
enable_http2: true
# No proxy needed: this host already egresses residentially (see header).
# If that ever changes, the fleet's NH3 SOCKS5 proxy is the fallback:
# proxies:
# all://:
# - socks5h://10.100.10.50:1080