Files
esh-pfi-infrastructure/persistent-memory.d/2026-08-19-unmanaged-stacks-searxng-seafile.md
T
vh bc3aada73a memory: snapshot — .internal DNS live, waterland containerised, homepage themed
Captures a long infra session: fleet *.internal DNS (git-sourced, 42 names,
three resolvers including a new colo one), waterland studio containerised on
irv-ml1, Homepage cleaned up and themed with Australis Skyfall over an
Arbo-generated background, and four unmanaged stacks adopted into stacks/.

Four detail files added. Auto-archived 4 entries to archival-memory.md
(Recent decisions 2, Tried and abandoned 2); 5 held back by the open-deferred
guard rather than moved.

Also records three operator-owned open items: the colo DNS repoint, the
static-v6 convention, and the deliberately belayed AI-tab Dormant regrouping.
2026-08-19 01:20:53 -07:00

78 lines
3.8 KiB
Markdown

# `[2026-08-19]` Four unmanaged stacks found on live hosts — and two of them were quietly broken
Commits `42c594c`, `dc3e47b`, plus `uptimekuma` in `9d92c4b`.
## The pattern worth remembering
Chasing two bad-looking cards on the dashboard turned up **four stacks running
on fleet hosts with no canonical copy anywhere**: `uptimekuma` and (already
known) the two AdGuards on esh-docker-vm, `searxng` and `seafile` on
ana-docker, and `heretic2-charrp-reasoning` on ana-ml2 (untracked in git).
**A dashboard card is a cheap census of what is actually running.** When
something on it looks wrong, check whether the stack behind it is even in
`stacks/` before debugging the symptom — twice here the answer was "no", and
the fix belonged in version control as much as on the host.
Adopted: `stacks/uptimekuma/`, `stacks/searxng/`, `stacks/seafile/`,
`stacks/heretic2-charrp-reasoning/`. ESH/NH3 AdGuard compose files were
**deliberately left unmanaged** — adopting three live resolvers while also
introducing a new DNS naming system is two risky changes at once.
## SearXNG — the healthcheck was eating itself
Card flapped UNHEALTHY; the container was fine the whole time. The compose
passed `--tries` and `--spider` as **two separate argv entries**, so wget
consumed `--spider` as the *value* of `--tries`. Spider mode never engaged,
which means every probe since April **downloaded** the healthz response to a
file:
```
295,287 healthz.N files in the container's working directory
```
With that many files, wget's scan for the next free filename is what
intermittently blew the 10s timeout. **Self-worsening — every probe made the
next one slower.** Restored `--tries=1`; the junk lived in the writable layer
so the recreate cleared it. Now `healthy`, `fails=0`, 200 in 0.16s.
Lesson: an argv list in YAML has no shell to catch a missing `=`. A flag that
silently swallows the next argument turns a liveness probe into a workload.
## SeaFile — not broken, never restarted
Card showed EXITED for three months. **None of the three services declared a
restart policy**, so Docker defaulted them to `no`. On
**2026-05-06T21:27:45Z** the daemon stopped all three within 200ms of each
other — a daemon restart or host reboot — and nothing brought them back.
⚠️ **Exit code `255` is a red herring**: it is what a container that ignores
SIGTERM reports when the daemon stops it, **not** evidence of a crash. Reading
it as one sends you hunting a bug that does not exist. The tell was all three
services stopping within 200ms.
Added `restart: unless-stopped` to all three; brought up; mariadb gated on its
healthcheck exactly as the existing `depends_on` comments intended, seahub
started without the race, `302` → login page. Data was in local named volumes,
not on the ana-nas NFS, so nothing was at risk.
Three months of silent downtime whose only signal was a card nobody read as an
outage — the argument for semantic status colour on the dashboard (see
[[2026-08-19-homepage-skyfall-theme]], where amber EXITED pills made six
mis-grouped AI seats obvious at a glance).
## heretic2-charrp-reasoning — tracked, with its shim
The `char-rp-reasoning` seat (NEO-CODE Heretic2 27B, modelopt NVFP4 + grafted
BF16 MTP head, ~77 tok/s via `qwen3_5_mtp` spec-decode) had been running
untracked. Now in `stacks/`, including
`conf/mtp-workaround/sitecustomize.py`, which is **not optional**: vLLM 0.24.0
does not propagate modelopt `exclude_modules` to the spec-decode **draft**
model, so the BF16 MTP head gets quantized and the engine dies at load. Both
the mount and `PYTHONPATH` are load-bearing.
Added the two files house convention expects and the directory lacked — a
`.env.example` naming every knob (all values are compose defaults; the host
overrides only the three VRAM ones) and a README pointing at
`docs/runbooks/heretic2-nvfp4-mtp-seat.md` rather than duplicating it.