Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-10-althing-362-rollout.md
T
vh b8dbe71a1c memory: snapshot — R49 1-epoch pilot complete and awaiting adjudication; MeroMero A4B quantized, v2 blocked; althing 3.6.2 rolled
Ships the two pending code changes alongside the memory that describes them:
train_voice_lora.py gains --eval-steps/--save-steps (the 3-epoch pilot overfit
with per-epoch eval and save_strategy="no", so the minimum was neither visible
nor recoverable), and the althing post-office compose is pinned to 3.6.2.

Index rewritten: in-flight compressed from 176 lines to the live state, three
detail files added for R49 D1-D3, the MeroMero four-failure chain, and the
althing rollout. Seven closed pre-08-27 entries archived; the guards held back
the rest because they carry open deferred pointers.
2026-09-10 10:25:52 -07:00

3.6 KiB

[2026-09-10] althing 3.6.2 rolled — post office + both heralds, and it was TWO nodes not seven

forseti shipped 3.6.2 to fix a Claude Code dedupe: the harness drops an inbound peer frame byte-identical to the previous one from the same sender within 30 s, and the herald's poke text was a single constant — so two pokes to a seat inside 30 s lost the second, and that seat waited for rung 1 five minutes later. 3.6.2 stamps From <handle>, poked at HH:MM:SS. onto every poke so two can never share a body.

What I rolled

post office  nh3-docker   3.6.0 -> 3.6.2   built + pushed to gitea, digest-pinned
nh3-dev      herald       3.6.1 -> 3.6.2
nh3-extdev   herald       3.1.1 -> 3.6.2

Every version read back off the RUNNING install, not the tag. Both nodes heartbeating after.

⚠ It was two boxes, not the seven the instruction assumed

The rollout said "the seven heralds." The post office's own nodes table has exactly two rows — nh3-dev and nh3-extdev. The other five have no install, no unit, no heartbeat. Ask the post office, not the fleet inventory.

Three traps, all mine to avoid next time

My survey said "not-installed" everywhere because I ran it as infra-ops and the heralds run as lkraven. Same shape as the ana-docker sudo trap but on the USER axis. Had I trusted it I would have concluded the fleet had no heralds.

nh3-extdev is worse than "needs the right prefix." It is a SYSTEM unit (not --user), running as lkraven, off /opt/uv-tools/althing-core — a venv with no pip and no uv, so uv tool install there is not merely ineffective, it is impossible. Path: build a wheel on nh3-dev, ensurepip into the venv, force-reinstall.

I took the bus down ~12 minutes, and it was self-inflicted. The documented backup is stop → checkpoint → verify counts → copy → verify counts. My PRAGMA wal_checkpoint(TRUNCATE) DELETED the -wal/-shm files, and the volume directory is root-owned 755 while the app runs as uid 1000 — so WAL mode could not recreate them and the container crash-looped on "attempt to write a readonly database". Fix: chown 1000:1000 the volume directory. ⚠ The runbook says checkpoint; it does not warn that on this volume that is a one-way trip until ownership is fixed.

Then I misread the recovery. docker logs --tail 25 without --timestamps showed the accumulated crash history from BEFORE the fix and I called it a current failure — nearly rolling back a working deploy. Rollback would not have helped anyway: the fault was in the volume, not the image.

The backup lesson that paid for itself

My first copy took only post_office.db and its sha256 MATCHED the source — while 7.2 MB of WAL sat uncopied. Restoring it would have silently lost recent mail. The runbook insists on ROW COUNTS rather than hashes for exactly this, and it was right: the hash agreed while the backup was incomplete. Post-deploy counts identical on all six tables (handles 76, messages 1722, recipients 1749, nodes 2, sessions 24, notifications 23).

The drop-count instrument, corrected

I reported the defect "reproduces on nh3-dev" with 3 drops in 21 s. It does not. Grepping transcripts for the notice WRITES the notice into the transcripts — my grep output and my own quoting of it. Genuine notices carry "type":"system". True count 26 in 15 sessions; my naive count was 109; the session I claimed 3 drops in had zero. forseti's own first count had the same contamination. BEFORE baselines captured: nh3-dev 27 in 18, nh3-extdev 0. Auto-memory: feedback_grep_over_a_log_that_records_your_greps