Files
esh-pfi-infrastructure/docs/runbooks/althing-deploy.md
T

13 KiB

althing deploy — six surfaces

scripts/deploy-althing.sh does all four. --check reports drift without touching anything. This page is the why.

Deploying althing means updating four independent surfaces on nh3-dev. Three were known; the fourth had no step in any runbook and drifted for five days before anyone noticed, on 2026-09-01.

# surface what it is how it drifts
1 binaries uv tool install --force . — the 7 console scripts silently, see below
2 herald systemctl --user restart althing-po-herald new guards live here; nothing takes effect until it restarts
3 skill scripts/sync_skill.sh~/.agents/skills/althing/SKILL.md covered by its own --check
4 plugin repo plugin/ → marketplace dir → Claude Code cache two hops, neither was automated
5 per-seat route althing-route declare on each seat a channel change needs a re-declare; the plugin's SessionStart hook does it for CC seats
6 ~/.claude/settings.json crossSessionInbound: "accept" outside the althing repo entirely — no version can carry it

uv tool install . without --force is a silent no-op

$ uv tool install .
`althing-core @ file:///home/lkraven/development/althing` is already installed
$ echo $?
0

uv matches on the source spec, not its contents. On a box that already installed from that path it declines and reports success. An operator following a runbook literally would restart the herald, see everything green, and wonder why the new binary was missing — with every command exiting 0.

⚠ Surface 4 is the one that bites, and it ate a hook

The chain is:

repo plugin/  →  ~/.local/share/althing-plugin/  →  ~/.claude/plugins/cache/althing/althing/<ver>/
              ^^^ nothing synced this hop

On 2026-09-01 the marketplace directory was a frozen copy from Aug 28:

deployed 0.0.1   hooks: ['UserPromptSubmit']
repo     0.1.1   hooks: ['UserPromptSubmit', 'SessionStart', 'SessionEnd']
                 + scripts/pane-route.sh   (absent from the deployment entirely)

So "CC seats re-declare their pane route automatically at next SessionStart" was never true on this box. The hook existed and worked upstream; it was never deployed. That is why every seat — including infra-ops' own Claude Code seat — had to be hand-declared with a pid someone measured by hand, and why the idle_cursor pin from 3.2.1 would not have self-applied either.

⚠ Compare the HOOK LIST, not the version string

A version number cannot tell you what a stale plugin actually cost you. 0.0.1 and 0.1.1 differ by two hooks and a script:

for d in ~/.claude/plugins/cache/althing/althing/*/; do
  printf "  %-8s hooks: %s  pane-route.sh: %s\n" "$(basename "$d")" \
    "$(python3 -c "import json;print(list(json.load(open('$d/.claude-plugin/plugin.json'))['hooks'].keys()))")" \
    "$(test -f "$d/scripts/pane-route.sh" && echo yes || echo NO)"
done

Expect ['UserPromptSubmit', 'SessionStart', 'SessionEnd'] and pane-route.sh: yes. Anything less and pane routes are not being declared or released automatically.

⚠ WHY THIS HID FOR FIVE DAYS — the shape to recognise

A missing deploy surface does not present as an error. It presents as "the migration needs manual work" — and we had a ready explanation for that, because four of five seats were non-Claude and genuinely did need hand-holding.

The seat that falsified the story was infra-ops' own: a Claude Code seat that should have self-declared and didn't, and it looked exactly like the other four. Neither infra-ops nor forseti asked why the automatic path had not fired for the one seat it was supposed to work on.

When a migration needs manual intervention, check whether the automatic path was ever deployed before concluding it does not apply to your case.

Use the supported CLI for the second hop. claude plugin update althing (also install / uninstall / list / details / validate / marketplace). ⚠ Do not hand-edit ~/.claude/plugins/installed_plugins.json or fabricate a cache directory — that is Claude Code's own bookkeeping, and a subtle mistake there breaks the plugin in a way that looks like an upstream bug.

A Claude Code restart is required for new plugin hooks. They are read at session start; a running session keeps the old ones.

Verifying a seat is actually reachable

postbox status reports mode: push, but until 3.2.4 a failed push declaration was silent — _declare_push devnulled both streams with check=False, so a missing binary, a non-zero exit and a server-side rejection rendered identically as nothing. A seat could have been silently pull-only since 3.1.2.

Audit for "looks armed but is not", cross-referencing local waiter locks against what the post office believes:

for f in ~/.althing/wake-listener-*.lock; do
  h=$(basename "$f" .lock); h=${h#wake-listener-}
  pid=$(cat "$f" 2>/dev/null)
  [ -n "$pid" ] && [ -r "/proc/$pid/cmdline" ] || continue
  # IDENTITY, not liveness: a lock left by a reaped listener names a pid the
  # kernel is free to hand to anything, so `kill -0` alone reports a stranger
  # as a live waiter. Match the same cmdline segment `althing-listen --stop`
  # requires before it will signal anything.
  ours=0
  while IFS= read -r -d '' seg; do [ "$seg" = "--_route=$h" ] && ours=1; done \
    < "/proc/$pid/cmdline"
  [ "$ours" = 1 ] || { printf "  %-24s pid %-8s STRANGER (recycled pid)\n" "$h" "$pid"; continue; }
  printf "  %-24s waiter %-8s mode: %s\n" "$h" "$pid" \
    "$(postbox status --handle "$h" 2>/dev/null | grep -oP 'mode: \K\w+')"
done

A live waiter reporting mode: pull is a seat that will never be poked. From 3.2.4 onward $ALTHING_ROOT/listen.log records failed declarations directly.

Liveness is not identity, and this loop is the place that gets it wrong. The earlier kill -0 form would print a phantom waiter for any handle whose stale lock happens to name a recycled pid — and a phantom waiter is exactly what sends a false "you are unreachable" notice to a seat that is fine. session_listener.sh refuses to SIGTERM on liveness alone for precisely this reason; an audit that only reads has no excuse for a weaker standard than the one that kills. (Adopted 2026-09-02 after the regin-smithy-dev cross-reference, where the pid happened to be genuine and the weaker check happened to be right.)

A stale wake-listener-*.lock is NOT a fault. The gate is flock -n on an open fd, which the kernel releases when the holder dies, so a lock file left by a reaped listener is inert and exit 3 only fires against a genuinely live holder. The pid in the file is read by --stop alone. (Recorded because infra-ops claimed the opposite, untested, on 2026-09-01; forseti measured it.)

⚠ The plugin cache does not take a content-only change

claude plugin update althing matches on the version in plugin.json and declines when it is unchanged. A release that edits hook or script CONTENT without bumping the plugin version therefore leaves the Claude Code cache stale while every version check reports success — the marketplace dir gets the new bytes, the cache does not, and update says "already at the latest version".

Measured on the 3.3.0 deploy (2026-09-02): marketplace and live cache both read 0.1.1, and pane-route.sh + README.md differed. That delta was documentation-only and harmless. The next one need not be.

scripts/deploy-althing.sh now diffs the marketplace tree against the live cache dir on every run and on --check, so the drift is reported rather than inferred. The fix is a plugin version bump upstream, not a hand-edit~/.claude/plugins/installed_plugins.json and the cache directory are Claude Code's own bookkeeping.

The 2026-09-01 lesson was compare the hook list, not the version string. This is the same lesson one turn deeper: here the hook list was identical too, and only the file contents moved.

⚠ The cc channel is an undocumented interface, taken deliberately

From 3.3.0 a Claude Code seat is poked over its own message socket ($XDG_RUNTIME_DIR/cc-socks/<pid>.sock) rather than by typing into its pane. No process to reap, nothing near the input line, and delivery lands at the receiver's next turn boundary.

It is not a published interface and is expected to break on some future Claude Code release. Accepted on the operator's reasoning: the FIFO poker it replaces was also an unsanctioned hack of the background-watcher system, so this is a better instance of a dependency class we already had, not a new one.

When it breaks the failure mode is a seat going pull-only with a logged reason — not lost mail; the post office holds it either way. Recovery is one command:

althing-route declare --pid <pid>     # prefers cc, falls back to pane

Why cc is worth an undocumented dependency, in one measurement. A pane poke is typed into the session and submitted, so Claude Code records it in the operator's prompt history exactly as if he had written it — 8 herald-written entries accumulated across three projects, permanently in up-arrow recall. A cc poke enters none: 2 delivered, 0 entries. Verified from the herald log against the history on 2026-09-02. Channel sequence on one seat that day, all three in order:

09-01 21:44  poked via fifo            the old way, later reaped
09-02 08:34  could NOT reach via pane  guard 4 refusing: he was typing
09-02 09:07  poked via pane            typed into his input line
09-02 09:14  poked via cc              arrived as a peer message
09-02 09:31  poked via cc              after crossSessionInbound accept

Count herald pokes by the whole display string, not by grepping ALTHING-HERALD. The substring matches the operator's own messages when he pastes a hold notice into one, so the count runs high and the excess reads as new activity. See the identity-not-resemblance note in the audit section.

Ordering constraint at any upgrade that changes the channel set: a pre-3.3.0 herald refuses channel=cc at parse, so a seat that declares the new channel before the herald restarts goes silently pull-only. uv tool install --force then the herald restart, and only then let anything re-declare. deploy-althing.sh already runs them in that order and updates the plugin (whose SessionStart hook declares the route) last, which is what keeps the window to a couple of seconds.

⚠ Surface 6 — crossSessionInbound, and it is not in the althing repo

A cc poke to a default-configured seat is HELD, not delivered. Claude Code auto-delivers an inbound cross-session message only when the sender's permission-mode class matches the receiver's, and a sender that asserts no class is held. The herald is a daemon and asserts none, deliberately. What the operator sees instead of a delivery:

Held peer message — from an unidentified session [verified pid <herald>];
preview: «ALTHING-HERALD althing: you have mail ...» — not delivered to
Claude (1 held). The sender did not attest its permission mode and this
session bypasses prompts.

Claude Code verified the herald through SO_PEERCRED and then correctly declined to let it speak, because it would not say what it was.

The fix is one key in ~/.claude/settings.json:

"crossSessionInbound": "accept"

Set on nh3-dev 2026-09-02. The operator's reasoning is the part to keep: the herald can only reach local seats, so bypass is the correct authorization type; when the guard was not there, that was our default posture, including the pane poke. A pane poke types into a session and presses Enter — bypass-level access by any measure, and what we had been doing all along. The socket channel is strictly narrower than what it replaces, so accepting here states the existing trust boundary rather than widening it.

There is no attestation the herald could send instead. Four probes established that Claude Code identifies a sender by verified pid against the session registry and reads that session's live runtime permission mode; a daemon has none and the registry has no field for one. A from_mode field on a type:"user" frame is not consulted at all — it belongs to the control actions. Adding it would be shipping a field nothing reads.

Cost without the setting is smaller than it looks: the correspondent record is in-memory session state, so the hold is first-contact, not per-message — one approval per seat per session lifetime, and a long-lived pane pays it once.

Why this surface is the dangerous one. A seat without it is declared, reachable and green, and its pokes go to a human instead of to the session. That is the same shape as the SessionStart hook that was never deployed: everything reports success, nothing arrives. deploy-althing.sh reads the key and reports it on every run and on --check — and deliberately never sets it. That file is the operator's permission configuration; a deploy script that edits it is a deploy script granting itself trust.

Rollback

uv tool install althing-core==3.1.2

Routes written by later versions stay parseable — the old reader ignores unknown keys — so nothing is stranded.