feat(searxng): move to nh3-docker, update, and expose as an MCP tool
The ana-docker instance was returning zero results for every query while reporting healthy — 4.5 months stale (2026.4.17 against a current 2026.9.3), its engine scrapers rotted against sites that had changed. /healthz proves the web app answers and says nothing about whether search works, so seven days of green sat on top of a search box that found nothing. Moved to nh3-docker rather than updated in place, because the colo egress is the other half of the problem: 38.120.12.42 is a datacenter address that DuckDuckGo and Startpage CAPTCHA, while nh3-docker egresses residentially at 70.230.226.88. Same reasoning as the fleet's residential proxy for yt-dlp, applied at the source instead of around it. Config corrected along the way: base_url said searxng.pfi.local, a name retired on 2026-08-19, while the environment said something else — the env won so nothing broke and the file quietly lied. The karmasearch.videos removal key never matched, because the engine's real name has a space. scripts/searxng-health.sh asserts results > 0 across three unrelated queries. That is the check that would have caught this, and the only kind that can: the mechanism was healthy throughout. services/searxng-mcp exposes it as `web_search` at user scope, so every Claude Code session has it. Zero results raise rather than returning an empty list — an empty list is indistinguishable from a broken aggregator, which is precisely how this hid. Old instance stopped and removed; DNS alias repointed to searxng.nh3.internal.
This commit is contained in:
+21
-55
@@ -1,38 +1,23 @@
|
||||
services:
|
||||
searxng:
|
||||
# ⚠ `:latest` means "latest AT PULL TIME", and nothing re-pulls on its own.
|
||||
# On 2026-09-03 this instance was found running 2026.4.17 — 4.5 months old —
|
||||
# while reporting `healthy` and returning ZERO results for every query,
|
||||
# because SearXNG engine scrapers rot as upstream sites change their markup
|
||||
# and the project ships near-daily releases to keep up. `docker compose pull
|
||||
# && up -d` is the update; scripts/searxng-health.sh is what tells you it is
|
||||
# needed, because /healthz cannot.
|
||||
image: searxng/searxng:latest
|
||||
container_name: searxng
|
||||
restart: unless-stopped
|
||||
# ------------------------------------------------------------------
|
||||
# Port binding — 9996 on all interfaces.
|
||||
# Change to "127.0.0.1:9996:8080" to restrict to localhost only.
|
||||
# Traefik handles public routing and TLS via the labels below.
|
||||
# ------------------------------------------------------------------
|
||||
ports:
|
||||
- 9996:8080
|
||||
# ------------------------------------------------------------------
|
||||
# Volumes
|
||||
# Config: settings.yml bind-mounted read-only into the container.
|
||||
volumes:
|
||||
- /opt/docker/conf/searxng/searxng-settings.yml:/etc/searxng/settings.yml:ro
|
||||
# ------------------------------------------------------------------
|
||||
# Environment — see https://docs.searxng.org/admin/settings/index.html
|
||||
# SEARXNG_SECRET — required for cryptographic signing (cookies, etc.)
|
||||
# BASE_URL — public URL SearXNG reports in pages/RSS/OPDS
|
||||
# INSTANCE_NAME — shown in the page title / footer
|
||||
# ------------------------------------------------------------------
|
||||
environment:
|
||||
- SEARXNG_SECRET=${SEARXNG_SECRET}
|
||||
# searxng.ana.internal, not the old searxng.pfi.local (migrated
|
||||
# 2026-08-19). `.local` is reserved for mDNS, so the old name was a
|
||||
# standards collision that happened to work; `.internal` is ICANN-
|
||||
# reserved for exactly this. The name is served by the fleet's AdGuard
|
||||
# resolvers from dns/internal.yaml — see scripts/dns-sync.py.
|
||||
- BASE_URL=https://searxng.ana.internal/
|
||||
- BASE_URL=http://10.100.50.40:9996/
|
||||
- INSTANCE_NAME=SearXNG
|
||||
# ------------------------------------------------------------------
|
||||
# Resource limits — tune for VM 102's available RAM/CPU
|
||||
# ------------------------------------------------------------------
|
||||
deploy:
|
||||
resources:
|
||||
limits:
|
||||
@@ -40,26 +25,18 @@ services:
|
||||
cpus: "1.0"
|
||||
reservations:
|
||||
memory: 128M
|
||||
# ------------------------------------------------------------------
|
||||
# Health check — SearXNG /healthz is the canonical liveness probe.
|
||||
# ⚠ `--tries=1` MUST keep its `=1`. Written as two argv entries
|
||||
# (`- --tries` / `- --spider`) wget consumes `--spider` as the VALUE of
|
||||
# `--tries`, spider mode never engages, and every probe DOWNLOADS the
|
||||
# response to a file. On the ana-docker instance that left 295,287
|
||||
# `healthz.N` files in the container's writable layer, one per probe since
|
||||
# April; wget scanning them to pick the next free name is what blew the 10s
|
||||
# timeout and made the dashboard card flap. Self-worsening — each probe made
|
||||
# the next slower. Watch for `docker exec searxng ls | wc -l` climbing.
|
||||
#
|
||||
# ⚠️ `--tries=1` MUST keep its `=1`. This read `- --tries` / `- --spider`
|
||||
# as two separate argv entries until 2026-08-18, and in that form wget
|
||||
# consumed `--spider` as the VALUE of `--tries` — so spider mode never
|
||||
# engaged and every probe DOWNLOADED the response to a file instead of
|
||||
# just checking it. By the time it was caught the container's working
|
||||
# directory held 295,287 `healthz.N` files, one per probe since April,
|
||||
# and wget had to scan all of them to pick the next free filename. That
|
||||
# scan is what intermittently blew the 10s timeout and made the card on
|
||||
# the dashboard flap UNHEALTHY while the service itself was fine. It was
|
||||
# self-worsening: every probe made the next one slower.
|
||||
#
|
||||
# The junk lived in the container's writable layer (the only volume here
|
||||
# is the read-only settings mount), so recreating the container cleared
|
||||
# it. Symptom to watch for if this regresses: `docker exec searxng ls |
|
||||
# wc -l` climbing, and health log entries reading
|
||||
# "Health check exceeded timeout (10s)".
|
||||
# ------------------------------------------------------------------
|
||||
# ⚠ AND KNOW WHAT THIS PROBE DOES NOT TELL YOU: /healthz proves the web app
|
||||
# answers. It says nothing about whether any engine returns a result. Seven
|
||||
# days of `healthy` sat on top of a search box that found nothing.
|
||||
healthcheck:
|
||||
test:
|
||||
- CMD
|
||||
@@ -75,23 +52,12 @@ services:
|
||||
networks:
|
||||
- tnet
|
||||
labels:
|
||||
# Traefik configuration — auto-discovery via Docker provider
|
||||
- traefik.enable=true
|
||||
# Both names during the migration: `.internal` is the real one now, and
|
||||
# the old `.pfi.local` is kept as a fallback so anything still pointing
|
||||
# at it (a bookmark, a hardcoded config elsewhere) does not break the
|
||||
# day the name changes. Drop the second Host() once nothing uses it —
|
||||
# the Traefik access log will tell you when that is.
|
||||
- traefik.http.routers.searxng.rule=Host(`searxng.ana.internal`) || Host(`searxng.pfi.local`)
|
||||
- traefik.http.routers.searxng.entrypoints=websecure
|
||||
- traefik.http.routers.searxng.tls=true
|
||||
- traefik.http.routers.searxng.service=searxng
|
||||
- traefik.http.services.searxng.loadbalancer.server.port=8080
|
||||
- homepage.group=Daily
|
||||
- homepage.name=SearXNG
|
||||
- homepage.icon=si-searxng
|
||||
- homepage.description=Privacy-respecting meta-search
|
||||
- homepage.href=http://10.250.50.70:9996
|
||||
- homepage.href=http://10.100.50.40:9996
|
||||
|
||||
networks:
|
||||
tnet:
|
||||
name: traefik-net
|
||||
|
||||
Reference in New Issue
Block a user