The in-flight section was a day stale: it still described run 3c as the live subject on a box where nothing had moved. Rewritten around what is actually true now -- v3 deployed fleet-wide, the post office relocated to nh3-docker, sec serving on GPU0, run 3c still held on power. Six new decision entries, three of which carry findings that outlive their incident: the inbound half of the handle-resolution bug (a stale ALTHING_HANDLE reads another agent's mailbox and reports it empty, which is a second route into the failure v3 exists to prevent), the OOM attribution to Claude Code sessions, and the operator's two explicit belays recorded so a later session does not re-raise them as new. Auto-archival fired at 301 lines and moved exactly one entry. Three others were old enough and every one carries a still-open deferred pointer -- the parked CI flip, muninn-gate's submit path, and the triton backend deferred to the Ada refresh. Held back per the guards; an over-cap file that keeps live decisions beats a scannable one that lost them. The entry that did move had its deferred item closed today: nh3-extdev's staged v2.1.2 wheel is moot now that the box runs 3.1.1.
13 KiB
[2026-08-28] althing v3.0.0 flag day (U9b) — the post office replaced the P2P bus, one-way
Operator-authorised, executed by infra-ops. v2 is gone from both boxes: every v2 command was deleted, not deprecated. No rollback was designed or tested; failures are fixed forward.
post office ONE container on nh3-dev, http://10.100.50.40:8390
the only stateful component. SQLite + FTS + the typed API + the operator page.
herald althing-po-herald, ONE PER BOX, supervised. Dials out, opens no port,
holds no state. Refuses to start if another herald holds the node.
waiter althing-listen, one per session. Holds a FIFO, blocks, exits when poked.
client postbox (+ althing-mcp for the stdio tool surface)
althing-cli -> postbox althing-wake-listener -> althing-listen
althing-light-monitor -> GONE althing-receiver -> GONE
althing-herald -> althing-po-herald
Every session needs BOTH env vars: ALTHING_POST_OFFICE=http://10.100.50.40:8390 and
ALTHING_HANDLE (dev-launch sets the latter per pane). postbox has no default post-office
address — a bare postbox status errors out rather than guessing.
⚠ An unreachable post office is an OUTAGE, never an empty inbox
v2 could not distinguish those. v3 can, and the distinction only pays if it is honoured: if postbox says it could not reach the post office, that is the fault. Do not read it as "no mail".
Deployment facts worth not rediscovering
network_mode: hostis deliberate — do NOT "fix" it into a bridged container with-p.api.py'sresolve_bind_hostrefuses any address that resolves to a wildcard and tests the RESOLVED PROPERTY rather than matching strings, soALTHING_BIND_HOST=0.0.0.0is structurally impossible. Reachability on the private network IS the authorisation story; there is no login. Bridging would move access control to a-pflag the application cannot see.- Compose schema is
2.4on purpose. nh3-dev has docker-compose 1.29.2 (v1 only, nodocker composesubcommand), wheremem_limitis honoured only under 2.x; under 3.x it moves todeploy:which is swarm-only and SILENTLY IGNORED. Verified honoured:docker inspect -> 536870912. - nh3-extdev is the box a
git pullcannot move. althing is a system WHEEL at/opt/uv-tools/althing-corewith entry points in/usr/local/bin; its v2 daemons were SYSTEM units, not user units; anduvis NOT on lkraven's PATH there — it lives at/home/infra-ops/.local/bin/uv. Install path: build the wheel on nh3-dev, stage to /tmp,sudo env UV_TOOL_DIR=/opt/uv-tools UV_TOOL_BIN_DIR=/usr/local/bin <uv> tool install --force. Playbook:playbooks/nh3-extdev-althing-v3.yaml. - nh3-dev's herald is a user unit at
~/.config/systemd/user/althing-po-herald.service(Environment=ALTHING_POST_OFFICE, Restart=always); nh3-extdev's is a system unit at/etc/systemd/system/althing-po-herald.servicewithUser=lkraven.
⚠ THE FIFO LEAK WAS STRUCTURAL, AND v3 FIXED IT
5,043 orphaned wake FIFOs were deleted from ~/.althing/wake/. The reason there were five
thousand: v2 named them per SESSION with the PID (advisor-dev-listener-2448686.fifo) and
never reaped them, so every pane that ever armed left one behind permanently. v3 names them
per HANDLE (infra-ops.fifo). Bounded at 73 by construction. That is a fix, not a cleanup.
The roster: 73, and the authoritative source is the CLI, not the DB
Seeded via POST /tool/declare with an X-Althing-Handle header (a request without one is
rejected bad_request). declare is an OPERATOR verb — not in postbox, never will be.
⚠ The v2 database held 91 agent rows; the v2 CLI listed 73. The extra 18 were legacy —
superseded handles (bifrost -> bifrost-dev), one typo (galdrbok), an underscore variant,
and two machine-qualified handles (ldp-dev@nh3-extdev, mailman@nh3-extdev). v3 abolishes
@-qualification: identity no longer has a home, so an @-qualified handle is a category error.
Verified after seeding by SET DIFFERENCE in both directions, not by counting — forseti's
"72 rendered vs 73 declared" was a line-count artifact and the sets are identical.
v2 history is inert, not migrated
~/.althing/althing.db — 95 MB, 12,437 messages — is untouched on disk. No import path
exists and none should be improvised. Plain SQLite if something must be recovered by hand.
Two-agent flag day: the collision worth remembering
forseti and I both ran scripts/sync_skill.sh 65 seconds apart (13:29:18Z / 13:30:23Z). My
--check said "in sync" BEFORE I ran it — that was forseti a minute earlier, and I read it as
"already done at some point" rather than "someone is working in here right now." Cost was one
redundant backup. Two agents worked the same checklist with no ownership marked per line.
Next flag day: name an owner per item.
Peers notified individually (operator-directed), NOT broadcast
eitri-smithy-dev, dvalin-smithy-dev, bil-smithy-dev carried their own 2.1.x SKILL.md
copies; regin-smithy-dev was named by the operator though no copy is visible on nh3-dev.
Operator's ruling on a general fleet announcement: pointless in both directions — anyone
already on v3 knows, anyone not on v3 cannot receive it. Consistent with the standing
never-broadcast-unsolicited directive.
⚠ sync_skill.sh deliberately does NOT write into peer repos — althing owns canonical
(vh/althing @ v3.0.0 : skills/althing/SKILL.md) and peers pull. Nobody does it for them.
[2026-08-28, same day] MOVED to nh3-docker — and the WAL nearly ate the mail
Operator: "I want it on the docker machine — that was always the goal." The flag-day deployment put the post office on nh3-dev, which was wrong on three counts:
-
our own server table calls nh3-dev "not a Docker-stack host"; NH-site non-GPU → nh3-docker
-
nh3-dev had three OOM events in fourteen days, interval halving, and the confirmed hog is CC sessions at 5-18 GB — the box's actual job. See 2026-08-28-nh3-dev-oom-attribution.
-
mem_limit: 512mprotects the fleet FROM the post office. It does nothing to protect the post office from the box:oom_score_adjwas 0, an ordinary kill candidate, and the 08-28 sweep took althing-herald and uvicorn. A sweep taking the post office takes all 73 handles.NOW http://10.100.50.40:8390 nh3-docker, oom_score_adj=-500, mem 512m verified honoured WAS http://10.100.10.50:8390 nh3-dev (address now refuses) canon stacks/althing-post-office/compose.yaml -> /opt/docker/compose/ on nh3-docker
⚠ THE DURABLE FINDING — docker stop does NOT checkpoint the SQLite WAL
post_office.db 155 KB mtime 15:09
post_office.db-wal 4.1 MB mtime 16:56 <- every recent message lived HERE
I expected a clean container stop to checkpoint. It did not — after docker stop the WAL was
still 4,124,152 bytes, unchanged. An explicit PRAGMA wal_checkpoint(TRUNCATE) was required,
after which the .db grew 155 KB → 163,840 B and the WAL/-shm vanished.
A docker cp of post_office.db alone would have produced a database that opens cleanly,
passes PRAGMA integrity_check, serves the full 73-handle roster — and is missing the day's
mail. Nothing would have errored. Only a row count distinguishes those two outcomes.
Procedure for moving any WAL-mode SQLite service: stop → EXPLICIT checkpoint → verify counts → copy → verify counts again on the far side, before deleting anything. Not stop → copy. Verified 8 messages / 73 handles / 2 nodes / 8 recipients at source, in the staged copy, and after seeding.
Repoint list (everything that names the address)
~/.config/systemd/user/althing-po-herald.service nh3-dev herald (user unit)
/etc/systemd/system/althing-po-herald.service nh3-extdev herald (system unit)
~/.claude/statusline-command.sh the hardcoded statusline fallback
every session's ALTHING_POST_OFFICE + re-arm althing-listen
⚠ Do NOT blind-sed 10.100.10.50:8390 across the memory tree.
reference_corviduo_dev_emergency_ops carries that exact string as a Bifrost "affect" plane
entry in the personal Worldtree's BIFROST_CLIENT_ALLOWED_HOSTS — an unrelated service that
happens to share the port. Incidentally, moving the post office off nh3-dev:8390 also cleared a
latent collision with it.
Evidence the outage semantics work under a real outage
During the gap the nh3-dev herald logged, verbatim: "push outage on nh3-dev: the post office did
not answer, so push is DOWN on this box. This is an outage, not an empty poke list." Then
after the repoint: 10:00:27 INFO poked infra-ops on nh3-dev via fifo (rung 0) — poke path
re-verified end to end.
Open follow-up
The image has no registry push. It moves by docker save | ssh | docker load, so a rebuild
means repeating that by hand. The fleet pattern (skaldsong, soong-lab) is a gitea registry pull;
this should join it. The old nh3-dev volume is left in place untouched — not a rollback path
(the operator ruled that out), just not deleting the only other copy on the day of the move.
[2026-08-28, later] The release train: 3.0.1 → 3.1.1 in one afternoon, and the registry
Six releases landed the same day as the cutover. The container was touched exactly once (the
nh3-docker move); every other release was herald- or client-side, established each time by
forseti's import-graph argument — asking what althing/post_office/* IMPORTS rather than
reading the diff. A diff tells you what moved; an import graph tells you what can be affected.
3.0.1 pane routes (a pane agent can register its own route)
3.0.2 the docs are now checked against the CLI, not against each other — NO redeploy
3.0.3 PANE_SETTLE_S=0.3, the write/submit race
3.1.0 herald writes $ALTHING_ROOT/post-office; dev-launch reads it — BOTH binaries
3.1.1 warn when ALTHING_HANDLE disagrees with launch-history
Deploy recipe per release: uv build --wheel on nh3-dev → uv tool install --force .
locally → stage the wheel to nh3-extdev and install under UV_TOOL_DIR=/opt/uv-tools with
/home/infra-ops/.local/bin/uv → restart both althing-po-herald units.
playbooks/nh3-extdev-althing-v3.yaml does the extdev half.
⚠ THE REPEATED DEFECT — a verify that half-passes, four costumes in one day
Every one of these was written correctly for its first run and silently wrong on the next:
1. install step gated on `postbox` not existing -> right for the cutover, would SKIP
every release after and report success
2. content check pinned to the PREVIOUS release's markers -> passes forever, asserts nothing
3. ONE file:marker pair against a TWO-file release -> asserts half a release
4. a Go template `{{ }}` inside elway, whose own substitution ate it -> FAILED on a green deploy
The check now matches on PRESENCE (grep -q), never a count, and takes a LIST of
file:marker pairs bumped per release. 3.1.1's own release note said
grep -c handles_launched_at dev_launch.py # 2+; the real count there is 1 (the definition,
with 2 in postbox.py), so a count assertion would have reported FAILED on a byte-perfect install.
The registry, and why the namespace is claude-bot
gitea.phasefinal.com/claude-bot/althing-post-office:3.0.0@sha256:410fed41...
Digest-pinned, not tag-floating — a tag is a mutable pointer on a registry anyone can re-push and
this container is the whole bus. ⚠ claude-bot's token carries write:package and docker login SUCCEEDS, but package namespaces are owned: pushing to vh/ returns
unauthorized: authentication required AFTER a successful login — an ownership refusal wearing a
credential error's clothes. Publishing under claude-bot's own namespace also satisfies the
standing directive to stop reusing the operator's personal credentials.
The deployed CC plugin copies are a release step NOBODY owns
sync_skill.sh covers the SKILL, not the plugin. Nothing in the repo reaches
~/.local/share/althing-plugin/ or ~/.claude/plugins/cache/althing/althing/0.0.1/. Both must be
rsync -a --delete'd from the repo's plugin/ on every release, by hand, or they carry the
previous release's bugs into the live surface — which happened: forseti's new docs-vs-CLI check
found postbox reply --to twice in plugin/commands/inbox.md, the file a CC session reads
every time it drains its inbox, and both my deployed copies had it.
⚠ A running CC session keeps the plugin text it loaded at startup. Files being right is
necessary and not sufficient — the session has to restart. I synced the copies at 09:30 and was
handed the v2 /althing:monitor text an hour later, calling three deleted binaries.
Peer skill copies: I stopped hand-syncing, deliberately
I wrote into eitri/dvalin/bil's repos three times (operator-authorised, and right while they were
dark and could not pull). Once they were awake and pulling, it became a race I was losing —
canonical moved three times in an hour and I was chasing a one-line version-banner lag. It also
cost provenance: dvalin had to correct their own account of their file because my write looked
like a pre-existing partial. Two writers, no lock. sync_skill.sh deliberately does not write
into peer trees; althing owns canonical and peers pull. Respect that boundary once they can.