Files
vh c488eadc31 memory: snapshot — althing v3 fleet-wide at 3.1.1, sec on GPU0
The in-flight section was a day stale: it still described run 3c as the
live subject on a box where nothing had moved. Rewritten around what is
actually true now -- v3 deployed fleet-wide, the post office relocated
to nh3-docker, sec serving on GPU0, run 3c still held on power.

Six new decision entries, three of which carry findings that outlive
their incident: the inbound half of the handle-resolution bug (a stale
ALTHING_HANDLE reads another agent's mailbox and reports it empty,
which is a second route into the failure v3 exists to prevent), the OOM
attribution to Claude Code sessions, and the operator's two explicit
belays recorded so a later session does not re-raise them as new.

Auto-archival fired at 301 lines and moved exactly one entry. Three
others were old enough and every one carries a still-open deferred
pointer -- the parked CI flip, muninn-gate's submit path, and the triton
backend deferred to the Ada refresh. Held back per the guards; an
over-cap file that keeps live decisions beats a scannable one that lost
them. The entry that did move had its deferred item closed today:
nh3-extdev's staged v2.1.2 wheel is moot now that the box runs 3.1.1.
2026-08-28 22:22:32 -07:00

13 KiB

[2026-08-28] althing v3.0.0 flag day (U9b) — the post office replaced the P2P bus, one-way

Operator-authorised, executed by infra-ops. v2 is gone from both boxes: every v2 command was deleted, not deprecated. No rollback was designed or tested; failures are fixed forward.

post office   ONE container on nh3-dev, http://10.100.50.40:8390
              the only stateful component. SQLite + FTS + the typed API + the operator page.
herald        althing-po-herald, ONE PER BOX, supervised. Dials out, opens no port,
              holds no state. Refuses to start if another herald holds the node.
waiter        althing-listen, one per session. Holds a FIFO, blocks, exits when poked.
client        postbox  (+ althing-mcp for the stdio tool surface)

althing-cli -> postbox                althing-wake-listener -> althing-listen
althing-light-monitor -> GONE         althing-receiver      -> GONE
althing-herald -> althing-po-herald

Every session needs BOTH env vars: ALTHING_POST_OFFICE=http://10.100.50.40:8390 and ALTHING_HANDLE (dev-launch sets the latter per pane). postbox has no default post-office address — a bare postbox status errors out rather than guessing.

⚠ An unreachable post office is an OUTAGE, never an empty inbox

v2 could not distinguish those. v3 can, and the distinction only pays if it is honoured: if postbox says it could not reach the post office, that is the fault. Do not read it as "no mail".

Deployment facts worth not rediscovering

  • network_mode: host is deliberate — do NOT "fix" it into a bridged container with -p. api.py's resolve_bind_host refuses any address that resolves to a wildcard and tests the RESOLVED PROPERTY rather than matching strings, so ALTHING_BIND_HOST=0.0.0.0 is structurally impossible. Reachability on the private network IS the authorisation story; there is no login. Bridging would move access control to a -p flag the application cannot see.
  • Compose schema is 2.4 on purpose. nh3-dev has docker-compose 1.29.2 (v1 only, no docker compose subcommand), where mem_limit is honoured only under 2.x; under 3.x it moves to deploy: which is swarm-only and SILENTLY IGNORED. Verified honoured: docker inspect -> 536870912.
  • nh3-extdev is the box a git pull cannot move. althing is a system WHEEL at /opt/uv-tools/althing-core with entry points in /usr/local/bin; its v2 daemons were SYSTEM units, not user units; and uv is NOT on lkraven's PATH there — it lives at /home/infra-ops/.local/bin/uv. Install path: build the wheel on nh3-dev, stage to /tmp, sudo env UV_TOOL_DIR=/opt/uv-tools UV_TOOL_BIN_DIR=/usr/local/bin <uv> tool install --force. Playbook: playbooks/nh3-extdev-althing-v3.yaml.
  • nh3-dev's herald is a user unit at ~/.config/systemd/user/althing-po-herald.service (Environment=ALTHING_POST_OFFICE, Restart=always); nh3-extdev's is a system unit at /etc/systemd/system/althing-po-herald.service with User=lkraven.

⚠ THE FIFO LEAK WAS STRUCTURAL, AND v3 FIXED IT

5,043 orphaned wake FIFOs were deleted from ~/.althing/wake/. The reason there were five thousand: v2 named them per SESSION with the PID (advisor-dev-listener-2448686.fifo) and never reaped them, so every pane that ever armed left one behind permanently. v3 names them per HANDLE (infra-ops.fifo). Bounded at 73 by construction. That is a fix, not a cleanup.

The roster: 73, and the authoritative source is the CLI, not the DB

Seeded via POST /tool/declare with an X-Althing-Handle header (a request without one is rejected bad_request). declare is an OPERATOR verb — not in postbox, never will be.

⚠ The v2 database held 91 agent rows; the v2 CLI listed 73. The extra 18 were legacy — superseded handles (bifrost -> bifrost-dev), one typo (galdrbok), an underscore variant, and two machine-qualified handles (ldp-dev@nh3-extdev, mailman@nh3-extdev). v3 abolishes @-qualification: identity no longer has a home, so an @-qualified handle is a category error. Verified after seeding by SET DIFFERENCE in both directions, not by counting — forseti's "72 rendered vs 73 declared" was a line-count artifact and the sets are identical.

v2 history is inert, not migrated

~/.althing/althing.db — 95 MB, 12,437 messages — is untouched on disk. No import path exists and none should be improvised. Plain SQLite if something must be recovered by hand.

Two-agent flag day: the collision worth remembering

forseti and I both ran scripts/sync_skill.sh 65 seconds apart (13:29:18Z / 13:30:23Z). My --check said "in sync" BEFORE I ran it — that was forseti a minute earlier, and I read it as "already done at some point" rather than "someone is working in here right now." Cost was one redundant backup. Two agents worked the same checklist with no ownership marked per line. Next flag day: name an owner per item.

Peers notified individually (operator-directed), NOT broadcast

eitri-smithy-dev, dvalin-smithy-dev, bil-smithy-dev carried their own 2.1.x SKILL.md copies; regin-smithy-dev was named by the operator though no copy is visible on nh3-dev. Operator's ruling on a general fleet announcement: pointless in both directions — anyone already on v3 knows, anyone not on v3 cannot receive it. Consistent with the standing never-broadcast-unsolicited directive.

sync_skill.sh deliberately does NOT write into peer repos — althing owns canonical (vh/althing @ v3.0.0 : skills/althing/SKILL.md) and peers pull. Nobody does it for them.


[2026-08-28, same day] MOVED to nh3-docker — and the WAL nearly ate the mail

Operator: "I want it on the docker machine — that was always the goal." The flag-day deployment put the post office on nh3-dev, which was wrong on three counts:

  • our own server table calls nh3-dev "not a Docker-stack host"; NH-site non-GPU → nh3-docker

  • nh3-dev had three OOM events in fourteen days, interval halving, and the confirmed hog is CC sessions at 5-18 GB — the box's actual job. See 2026-08-28-nh3-dev-oom-attribution.

  • mem_limit: 512m protects the fleet FROM the post office. It does nothing to protect the post office from the box: oom_score_adj was 0, an ordinary kill candidate, and the 08-28 sweep took althing-herald and uvicorn. A sweep taking the post office takes all 73 handles.

    NOW http://10.100.50.40:8390 nh3-docker, oom_score_adj=-500, mem 512m verified honoured WAS http://10.100.10.50:8390 nh3-dev (address now refuses) canon stacks/althing-post-office/compose.yaml -> /opt/docker/compose/ on nh3-docker

⚠ THE DURABLE FINDING — docker stop does NOT checkpoint the SQLite WAL

post_office.db       155 KB  mtime 15:09
post_office.db-wal   4.1 MB  mtime 16:56    <- every recent message lived HERE

I expected a clean container stop to checkpoint. It did not — after docker stop the WAL was still 4,124,152 bytes, unchanged. An explicit PRAGMA wal_checkpoint(TRUNCATE) was required, after which the .db grew 155 KB → 163,840 B and the WAL/-shm vanished.

A docker cp of post_office.db alone would have produced a database that opens cleanly, passes PRAGMA integrity_check, serves the full 73-handle roster — and is missing the day's mail. Nothing would have errored. Only a row count distinguishes those two outcomes.

Procedure for moving any WAL-mode SQLite service: stop → EXPLICIT checkpoint → verify counts → copy → verify counts again on the far side, before deleting anything. Not stop → copy. Verified 8 messages / 73 handles / 2 nodes / 8 recipients at source, in the staged copy, and after seeding.

Repoint list (everything that names the address)

~/.config/systemd/user/althing-po-herald.service        nh3-dev herald (user unit)
/etc/systemd/system/althing-po-herald.service           nh3-extdev herald (system unit)
~/.claude/statusline-command.sh                         the hardcoded statusline fallback
every session's ALTHING_POST_OFFICE                     + re-arm althing-listen

Do NOT blind-sed 10.100.10.50:8390 across the memory tree. reference_corviduo_dev_emergency_ops carries that exact string as a Bifrost "affect" plane entry in the personal Worldtree's BIFROST_CLIENT_ALLOWED_HOSTS — an unrelated service that happens to share the port. Incidentally, moving the post office off nh3-dev:8390 also cleared a latent collision with it.

Evidence the outage semantics work under a real outage

During the gap the nh3-dev herald logged, verbatim: "push outage on nh3-dev: the post office did not answer, so push is DOWN on this box. This is an outage, not an empty poke list." Then after the repoint: 10:00:27 INFO poked infra-ops on nh3-dev via fifo (rung 0) — poke path re-verified end to end.

Open follow-up

The image has no registry push. It moves by docker save | ssh | docker load, so a rebuild means repeating that by hand. The fleet pattern (skaldsong, soong-lab) is a gitea registry pull; this should join it. The old nh3-dev volume is left in place untouched — not a rollback path (the operator ruled that out), just not deleting the only other copy on the day of the move.


[2026-08-28, later] The release train: 3.0.1 → 3.1.1 in one afternoon, and the registry

Six releases landed the same day as the cutover. The container was touched exactly once (the nh3-docker move); every other release was herald- or client-side, established each time by forseti's import-graph argument — asking what althing/post_office/* IMPORTS rather than reading the diff. A diff tells you what moved; an import graph tells you what can be affected.

3.0.1  pane routes (a pane agent can register its own route)
3.0.2  the docs are now checked against the CLI, not against each other  — NO redeploy
3.0.3  PANE_SETTLE_S=0.3, the write/submit race
3.1.0  herald writes $ALTHING_ROOT/post-office; dev-launch reads it  — BOTH binaries
3.1.1  warn when ALTHING_HANDLE disagrees with launch-history

Deploy recipe per release: uv build --wheel on nh3-dev → uv tool install --force . locally → stage the wheel to nh3-extdev and install under UV_TOOL_DIR=/opt/uv-tools with /home/infra-ops/.local/bin/uv → restart both althing-po-herald units. playbooks/nh3-extdev-althing-v3.yaml does the extdev half.

⚠ THE REPEATED DEFECT — a verify that half-passes, four costumes in one day

Every one of these was written correctly for its first run and silently wrong on the next:

1. install step gated on `postbox` not existing   -> right for the cutover, would SKIP
                                                      every release after and report success
2. content check pinned to the PREVIOUS release's markers -> passes forever, asserts nothing
3. ONE file:marker pair against a TWO-file release  -> asserts half a release
4. a Go template `{{ }}` inside elway, whose own substitution ate it -> FAILED on a green deploy

The check now matches on PRESENCE (grep -q), never a count, and takes a LIST of file:marker pairs bumped per release. 3.1.1's own release note said grep -c handles_launched_at dev_launch.py # 2+; the real count there is 1 (the definition, with 2 in postbox.py), so a count assertion would have reported FAILED on a byte-perfect install.

The registry, and why the namespace is claude-bot

gitea.phasefinal.com/claude-bot/althing-post-office:3.0.0@sha256:410fed41...

Digest-pinned, not tag-floating — a tag is a mutable pointer on a registry anyone can re-push and this container is the whole bus. ⚠ claude-bot's token carries write:package and docker login SUCCEEDS, but package namespaces are owned: pushing to vh/ returns unauthorized: authentication required AFTER a successful login — an ownership refusal wearing a credential error's clothes. Publishing under claude-bot's own namespace also satisfies the standing directive to stop reusing the operator's personal credentials.

The deployed CC plugin copies are a release step NOBODY owns

sync_skill.sh covers the SKILL, not the plugin. Nothing in the repo reaches ~/.local/share/althing-plugin/ or ~/.claude/plugins/cache/althing/althing/0.0.1/. Both must be rsync -a --delete'd from the repo's plugin/ on every release, by hand, or they carry the previous release's bugs into the live surface — which happened: forseti's new docs-vs-CLI check found postbox reply --to twice in plugin/commands/inbox.md, the file a CC session reads every time it drains its inbox, and both my deployed copies had it.

A running CC session keeps the plugin text it loaded at startup. Files being right is necessary and not sufficient — the session has to restart. I synced the copies at 09:30 and was handed the v2 /althing:monitor text an hour later, calling three deleted binaries.

Peer skill copies: I stopped hand-syncing, deliberately

I wrote into eitri/dvalin/bil's repos three times (operator-authorised, and right while they were dark and could not pull). Once they were awake and pulling, it became a race I was losing — canonical moved three times in an hour and I was chasing a one-line version-banner lag. It also cost provenance: dvalin had to correct their own account of their file because my write looked like a pre-existing partial. Two writers, no lock. sync_skill.sh deliberately does not write into peer trees; althing owns canonical and peers pull. Respect that boundary once they can.