0f748ea54e
The ana-docker instance was returning zero results for every query while reporting healthy — 4.5 months stale (2026.4.17 against a current 2026.9.3), its engine scrapers rotted against sites that had changed. /healthz proves the web app answers and says nothing about whether search works, so seven days of green sat on top of a search box that found nothing. Moved to nh3-docker rather than updated in place, because the colo egress is the other half of the problem: 38.120.12.42 is a datacenter address that DuckDuckGo and Startpage CAPTCHA, while nh3-docker egresses residentially at 70.230.226.88. Same reasoning as the fleet's residential proxy for yt-dlp, applied at the source instead of around it. Config corrected along the way: base_url said searxng.pfi.local, a name retired on 2026-08-19, while the environment said something else — the env won so nothing broke and the file quietly lied. The karmasearch.videos removal key never matched, because the engine's real name has a space. scripts/searxng-health.sh asserts results > 0 across three unrelated queries. That is the check that would have caught this, and the only kind that can: the mechanism was healthy throughout. services/searxng-mcp exposes it as `web_search` at user scope, so every Claude Code session has it. Zero results raise rather than returning an empty list — an empty list is indistinguishable from a broken aggregator, which is precisely how this hid. Old instance stopped and removed; DNS alias repointed to searxng.nh3.internal.
77 lines
3.4 KiB
Markdown
77 lines
3.4 KiB
Markdown
# searxng
|
|
|
|
Privacy-respecting metasearch. **Runs on nh3-docker** (`10.100.50.40:9996`,
|
|
`searxng.nh3.internal`), and is exposed to every Claude Code session on nh3-dev
|
|
as the `web_search` MCP tool.
|
|
|
|
## ⚠ Why NH3 and not the colo
|
|
|
|
Measured 2026-09-03:
|
|
|
|
ana-docker egress 38.120.12.42 datacenter -> DuckDuckGo/Startpage CAPTCHA
|
|
nh3-docker egress 70.230.226.88 residential -> no gate on that account
|
|
|
|
Search engines gate datacenter ranges. Running the aggregator from the
|
|
residential-egress site removes the problem at the source rather than proxying
|
|
around it — the same reason the fleet keeps a residential SOCKS5 proxy on
|
|
nh3-dev for yt-dlp. If NH3's egress ever changes, `outgoing.proxies` in
|
|
`conf/searxng-settings.yml` has the fallback commented in place.
|
|
|
|
⚠ It is **not** a complete fix: `brave`, `duckduckgo` and `startpage` still
|
|
CAPTCHA from here. `google cse` carries general search at ~20 results/query, and
|
|
`yandex`, `wiby`, `github`, `stackoverflow` and `marginalia` all work. General
|
|
search is therefore effectively single-engine — if `google cse` breaks, the
|
|
instance goes quiet, which is exactly the failure below.
|
|
|
|
## ⚠ /healthz cannot tell you whether search works
|
|
|
|
On 2026-09-03 the old ana-docker instance was found returning **zero results for
|
|
every query**, for an unknown number of weeks, while:
|
|
|
|
- the container reported `healthy` for 7 straight days, 0 restarts;
|
|
- the Homepage card was green;
|
|
- `/healthz` returned 200 every 30 seconds.
|
|
|
|
It was running **2026.4.17 while current was 2026.9.3** — 4.5 months of engine
|
|
scrapers rotting against sites that had changed their markup. `:latest` means
|
|
"latest at pull time", and nothing re-pulls on its own.
|
|
|
|
Proven by experiment before touching anything: the new image, same settings
|
|
file, same host, same query returned **20 results where the running one returned
|
|
0**.
|
|
|
|
**`scripts/searxng-health.sh` asserts the property, not the mechanism** — three
|
|
unrelated queries must each return results > 0. Run it after any change, and
|
|
periodically; it is the only thing that catches rot.
|
|
|
|
## Update
|
|
|
|
ssh infra-ops@nh3-docker 'cd /opt/docker/compose/searxng && \
|
|
sudo docker compose pull && sudo docker compose up -d'
|
|
scripts/searxng-health.sh
|
|
|
|
## The MCP tool
|
|
|
|
`services/searxng-mcp/` — installed with `uv tool install` and registered at
|
|
**user scope** (`claude mcp add --scope user searxng searxng-mcp`), so every
|
|
Claude Code session gets `web_search` with no per-project setup.
|
|
|
|
⚠ **Zero results raise an error rather than returning an empty list.** An empty
|
|
list is indistinguishable from a broken aggregator, and that ambiguity is what
|
|
let the old instance fail silently for weeks. Same reasoning as althing's
|
|
"an unreachable post office is an OUTAGE, never an empty inbox".
|
|
|
|
⚠ Written against **mcp 2.x**, where `FastMCP` became `MCPServer` and the v1
|
|
`@app.list_tools()` decorator API is gone. v1 examples fail at import with
|
|
`'Server' object has no attribute 'list_tools'`.
|
|
|
|
⚠ `uv tool install --force` alone served a **cached build** and silently
|
|
reinstalled the old code — the installed file still had the v1 API after the
|
|
source no longer did. `--reinstall --no-cache` was required, and the check that
|
|
caught it was `md5sum` of source vs installed.
|
|
|
|
## Secret
|
|
|
|
`SEARXNG_SECRET` lives in `/opt/docker/compose/searxng/.env` (0600, root) on the
|
|
host and is vaulted at `nh3-docker/searxng-secret`. Never in git.
|