feat(alerts): generalize the althing bridge, wire Kuma to it, retire chamber

THE BRIDGE. `beszel-althing` hardcoded a "[Beszel] " subject prefix and a
Beszel hub footer from when Beszel was its only caller. Routing Uptime Kuma
through it unchanged would have delivered Kuma outages labelled [Beszel],
pointing the reader at the wrong dashboard -- an alert that lies about its own
source is worse than no alert.

Now a route registry: /beszel and /kuma, each with its own prefix, footer and
payload parser, because the tools do not agree on a shape (Beszel sends
{title, message}; Kuma sends {heartbeat, monitor, msg}). Generalising cost a
dict; a sibling service would have cost a second unit, a second port and a
second thing to notice had died.

Renamed beszel-althing -> althing-alert-bridge with it. A service named after
one consumer that carries two is the invisible coupling that sends a future
session looking in the wrong place.

⚠ /beszel IS FROZEN and this refactor proves it rather than claiming it. The
three original tests were kept BYTE-UNCHANGED -- including the one asserting
the exact postbox argv -- and deliver() still defaults to the Beszel route so
they exercise it. A new test asserts the Kuma footer never leaks into a Beszel
body or vice versa. Verified live after the rename: a real POST to /beszel
landed as "[Beszel] BRIDGE RENAME CHECK" with the correct hub footer, read back
from the thread rather than trusted from the receipt.

Payload shapes are parsed HERE, not via Kuma's custom-webhook-body feature,
because Kuma's notification config lives in its own database -- and that
database was destroyed and rebuilt from scratch hours ago. Anything living only
in a tool's DB is lost on the next rebuild; format knowledge belongs in git,
next to a test.

parse_kuma also handles the monitorless case. testNotification and cert-expiry
alerts carry no monitor and no heartbeat, and the first cut fabricated "unknown
monitor is ?" from them -- caught by sending a real one and reading the subject,
not by the suite. Fixed, pinned, and the earlier test asserting the bad
behaviour was corrected rather than worked around.

KUMA IS NOW WIRED. scripts/kuma gained notification support and the channel is
in monitors.yaml, seeded BEFORE the monitors and with applyExisting, so a
rebuild restores alerting and not just detection. Ground truth from the DB:
13 of 13 monitors carry the channel.

⚠ A THIRD instance of the same class of bug, worth naming: notifications() is
pushed as `notificationList` at LOGIN ONLY -- there is no event to ask with. The
first cut cleared the captured value before waiting, discarding the only copy it
would ever be sent, then blocked for the full timeout and reported an empty
list. That reads exactly like "no channels configured" and is a lie. Same family
as the getMonitorList ack-vs-push trap, different shape.

End-to-end, both shapes, read back from the inbox:
  [Uptime Kuma] Homepage is DOWN          + target + board link
  [Uptime Kuma] althing (infra-ops) Testing   (no fabricated subject)
  [Beszel] BRIDGE RENAME CHECK            + hub footer, unchanged

ALTHING CHAMBER RETIRED (operator). Three of its four containers had never
started -- created 2026-09-19, StartedAt epoch-zero, 0 restarts -- so :7881
refused, and nothing was watching it. Only its valkey was running, on the
project's own network with no external consumer. Stack, compose/build/conf dirs
and the local image removed; the Homepage card went with the label.
This commit is contained in:
vh
2026-09-21 23:11:54 -07:00
parent 3a85a6bce1
commit 6f0a9b9fae
19 changed files with 818 additions and 927 deletions
+33
View File
@@ -0,0 +1,33 @@
steps:
# The service was `beszel-althing` until 2026-09-21, when it gained a /kuma
# route and the Beszel-specific name became a lie. Stop and remove the old
# unit FIRST: both bind 10.100.10.50:8096, so they cannot run side by side.
- name: Stop and remove the superseded beszel-althing unit
sudo: true
shell: >-
systemctl disable --now beszel-althing.service 2>/dev/null || true;
rm -f /etc/systemd/system/beszel-althing.service;
rm -rf /opt/beszel-althing;
systemctl daemon-reload
- name: Install fleet alert bridge
sudo: true
upload:
src: services/althing-alert-bridge/bridge.py
dest: /opt/althing-alert-bridge/bridge.py
mode: '0644'
- name: Install fleet alert bridge unit
sudo: true
upload:
src: services/althing-alert-bridge/althing-alert-bridge.service
dest: /etc/systemd/system/althing-alert-bridge.service
mode: '0644'
- name: Start fleet alert bridge
sudo: true
shell: systemctl daemon-reload && systemctl enable althing-alert-bridge.service && systemctl restart althing-alert-bridge.service
verify:
- name: Verify bridge process and BOTH routes are advertised
shell: >-
systemctl is-active althing-alert-bridge.service &&
curl --retry 5 --retry-connrefused --retry-delay 1 -fsS http://10.100.10.50:8096/healthz | grep -q '"/kuma"' &&
curl -fsS http://10.100.10.50:8096/healthz | grep -q '"/beszel"'
changed_when: 'false'
-20
View File
@@ -1,20 +0,0 @@
steps:
- name: Install Beszel alert bridge
sudo: true
upload:
src: services/beszel-althing/bridge.py
dest: /opt/beszel-althing/bridge.py
mode: '0644'
- name: Install Beszel alert bridge unit
sudo: true
upload:
src: services/beszel-althing/beszel-althing.service
dest: /etc/systemd/system/beszel-althing.service
mode: '0644'
- name: Start Beszel alert bridge
sudo: true
shell: systemctl daemon-reload && systemctl enable beszel-althing.service && systemctl restart beszel-althing.service
verify:
- name: Verify bridge process
shell: systemctl is-active beszel-althing.service && curl --retry 5 --retry-connrefused --retry-delay 1 -fsS http://10.100.10.50:8096/healthz
changed_when: 'false'
-147
View File
@@ -1,147 +0,0 @@
# Deploy althing-chamber (https://gitea.phasefinal.com/vh/althing) to a
# Docker host following the PFI /opt/docker/ convention (ana-docker by
# default, but the playbook works against any host with Docker in place).
#
# Brings up FOUR compose services:
# althing-chamber — FastAPI/HTMX web UI, port 7881 host → 8000 container
# althing-forseti — moderator daemon, no port
# althing-agent-runner — Phase 2 worldtree-driver dispatcher, no port
# althing-valkey — Phase 3.1 valkey 8 alpine, pub/sub bridge (internal only)
#
# The three althing services use the same image (built from vh/althing);
# valkey is a stock upstream image. SQLite-backed state shares between the
# three althing services via the bind-mount under /opt/docker/conf/althing-chamber/data;
# streaming events ride pub/sub on the docker default network via valkey.
#
# Idempotent: rerunning is safe. Creates-gates skip work that's already
# done; `docker compose up -d` is itself idempotent (no restart unless
# compose content or env changed).
#
# Usage:
# scripts/elway ana-docker --playbook playbooks/deploy-althing-chamber.yaml
# scripts/elway ana-docker --playbook playbooks/deploy-althing-chamber.yaml --var ref=v0.1.0
#
# Prereqs on the target host:
# - Docker + docker compose plugin
# - Target user (lkraven) has git SSH access to gitea.phasefinal.com
# — either SSH key authorized in gitea, or the repo is HTTPS-reachable
# if you swap `repo_url` below.
# - Target user is in the `docker` group.
vars:
repo_url: git@gitea.phasefinal.com:vh/althing.git
ref: main
build_dir: /opt/docker/build/althing-chamber
image_tag: althing-chamber:local
compose_dir: /opt/docker/compose/althing-chamber
data_dir: /opt/docker/conf/althing-chamber/data
host_port: "7881"
steps:
# ── host-side directory prep ─────────────────────────────────────────
- name: Ensure /opt/docker/build parent exists
shell: mkdir -p /opt/docker/build
sudo: true
creates: /opt/docker/build
- name: Chown /opt/docker/build to lkraven (only if mkdir'd by root above)
shell: chown lkraven:lkraven /opt/docker/build
sudo: true
when: '[ "$(stat -c %U /opt/docker/build)" != lkraven ]'
# ── fetch / sync source ─────────────────────────────────────────────
- name: Clone althing repo if absent
# Auto-accept the first-run host key so the playbook doesn't hang
# prompting for yes/no.
shell: GIT_SSH_COMMAND="ssh -o StrictHostKeyChecking=accept-new" git clone {{ repo_url }} {{ build_dir }}
creates: "{{ build_dir }}/.git"
- name: Fetch from origin
shell: cd {{ build_dir }} && git fetch --quiet origin
- name: Reset working tree to {{ ref }}
# Accept either a branch name (resolves via origin/<ref>) or a
# full/short SHA (resolves directly). CI passes the triggering
# commit SHA via --var ref=${{ github.sha }}; manual runs pass
# branch names like main / v0.1.0.
shell: |
cd {{ build_dir }}
if sha=$(git rev-parse --verify --quiet "origin/{{ ref }}^{commit}"); then :;
elif sha=$(git rev-parse --verify --quiet "{{ ref }}^{commit}"); then :;
else echo "elway: ref not found: {{ ref }}" >&2; exit 1; fi
git reset --hard "$sha"
# Report ok (no-change) when the tree was already at the requested
# ref — saves a noisy CHANGED status line on no-op reruns.
changed_when: '[ "$(cd {{ build_dir }} && git rev-parse HEAD)" != "$(cd {{ build_dir }} && (git rev-parse --verify --quiet "origin/{{ ref }}^{commit}" || git rev-parse --verify --quiet "{{ ref }}^{commit}"))" ]'
# ── image build ─────────────────────────────────────────────────────
- name: Build image {{ image_tag }}
shell: cd {{ build_dir }} && docker build -t {{ image_tag }} .
# Docker build reuses layer cache and is fast on reruns, but it
# always runs — we can't cheaply know up-front whether anything
# downstream has changed. Leave it in the always-run lane; Docker
# itself handles the no-op efficiently.
# ── compose + data dirs ─────────────────────────────────────────────
- name: Ensure compose dir exists
shell: mkdir -p {{ compose_dir }}
creates: "{{ compose_dir }}"
- name: Ensure data dir exists
# Single bind-mount shared between chamber + forseti. Created as
# lkraven (uid 1000 on these hosts), matching the container's `app`
# user — no chown dance needed.
shell: mkdir -p {{ data_dir }}
creates: "{{ data_dir }}"
# ── deploy compose files ────────────────────────────────────────────
- name: Upload compose.yaml
upload:
src: stacks/althing-chamber/compose.yaml
dest: "{{ compose_dir }}/compose.yaml"
mode: "0644"
- name: Seed .env from template (only if absent)
upload:
src: stacks/althing-chamber/.env.example
dest: "{{ compose_dir }}/.env"
mode: "0644"
when: "[ ! -f {{ compose_dir }}/.env ]"
# ── bring up + wait for ready ───────────────────────────────────────
- name: docker compose up -d
shell: cd {{ compose_dir }} && docker compose up -d
- name: Wait for chamber /health to respond
# Chamber's healthcheck is internal (inside the container's network);
# this host-side poll confirms the published port is reachable too.
# Short retry loop — docker compose up returns before the FastAPI
# app finishes booting.
shell: |
for i in $(seq 1 30); do
curl -sf -o /dev/null http://localhost:{{ host_port }}/health && exit 0
sleep 2
done
exit 1
changed_when: "false"
verify:
- name: chamber /health returns 200
shell: curl -sf -o /dev/null http://localhost:{{ host_port }}/health
changed_when: "false"
- name: chamber container running
shell: docker ps --filter name=^/althing-chamber$ --format '{{.Status}}' | grep -q '^Up'
changed_when: "false"
- name: forseti container running
shell: docker ps --filter name=^/althing-forseti$ --format '{{.Status}}' | grep -q '^Up'
changed_when: "false"
- name: agent-runner container running
shell: docker ps --filter name=^/althing-agent-runner$ --format '{{.Status}}' | grep -q '^Up'
changed_when: "false"
- name: valkey container running + healthy
shell: docker ps --filter name=^/althing-valkey$ --format '{{.Status}}' | grep -q 'healthy'
changed_when: "false"