feat(u4): a booth's lifetime is derived from its state, not from a boolean

`.forever` was the only way to say three different things — "this is durable",
"I have not answered yet", "I am still looking" — and the census said it was
carrying all three: 17 of 24 live booths (70%, up from 54% the day before).
Three of the four booths in the fleet awaiting an answer had been pinned by
hand as well, and 10 of the 17 were younger than the TTL, so the sentinel had
bought them nothing and was pressed pre-emptively.

Only the first meaning is what `keep` means. The other two are facts the
service already held and did not consult.

    KEPT       `.forever` present                      never swept  (unchanged)
    HELD       an open pick, or marks we cannot read   never swept  (new)
    EPHEMERAL  everything else                         24h          (unchanged)

Viewing is activity: a deliberately-served response from a booth's own page
route writes `.viewed`, which is a dotfile and not a `.lock` dotfile, so
`_newest_mtime` already counts it. There is no new arithmetic — `booth_age_seconds`,
`is_expired` and `expires_in` are unchanged. Machine reads are excluded on
purpose: an agent must not be able to hold its own booth open by polling for
the answer it is waiting on.

The hold is unbounded, and what makes that safe is visibility plus two exits
that already existed. Every surface whose chrome the Booth owns says
`held until answered` where the countdown was, and `booth rm` / the UI x /
`DELETE /b/<n>` take a held booth exactly as they take a kept one. A hold is
protection from the timer, never from the operator.

Three cross-frontier panels ran and each found a class the others could not:

  * the paraphrase panel found that two reads of one file are not one read of
    one state — the contract's `is_held(marks_for(c), read_error(c))` could
    resolve to `([], None)`, the pair that deletes. `hold_read` is one read.
  * the code-review panel found, 4-of-4, that the booth header's board branch
    rendered no lifetime at all; and that five of seven invariant tests passed
    under the change that defeats them.
  * the bug-hunt panel found four more paths where a failed read still
    authorized a delete, and a `record_view` that followed a planted symlink.

`is_held` became `hold_reason`, which returns the reason rather than a bool
beside a string that can disagree with it.

Prediction, to re-count on or after 2026-10-06: the `.forever` rate falls to
the booths that are genuinely durable references. Only 4 booths carry marks at
all, so this rests on both halves of the unit; a null result cannot distinguish
a wrong diagnosis from a habit that outlived its need.

406 tests (341 before). Contract: docs/contracts/u4_derived_lifetime.contract.md
This commit is contained in:
vh
2026-09-22 09:44:25 -07:00
parent d37b81ab9f
commit c3a97c1b64
22 changed files with 2339 additions and 71 deletions
@@ -0,0 +1,44 @@
# The `.forever` diagnosis got a live positive control
_2026-09-22 · booth_
The U4 diagnosis was that `.forever` is the only way to say three different
things — "this is durable", "I have not answered yet", "I am still looking" —
and that only the first is what keep means. That was an argument. **On
2026-09-22 it stopped being one.**
Census of `~/booth-data`, whole population, every value a deterministic file
fact:
| | |
|---|---|
| live booths | 24 |
| carrying `.forever` | 17 (70%, up from 54% on 2026-09-21) |
| carrying `.marks.json` at all | 4 |
| of those, with an open pick | **4 of 4** |
| **open pick AND `.forever`** | **3** |
**Three of the four booths in the entire fleet that were waiting on an answer
had also been pinned by hand.** That is the "not yet" case caught in the act,
not inferred from a rate.
The staleness distribution says it from the other side: **10 of the 17 kept
booths were under one day old** — younger than the TTL, so the sentinel had
bought them nothing and was pressed pre-emptively. Only 4 were old enough
(2.4-4.6 days) that keep is the reason they still existed.
⚠ **A number I got wrong, caught by a cross-frontier arm, kept here because the
class repeats.** The contract first said "12 are under 1.5 days old — younger
than the TTL". The TTL is 24 hours. 1.5 days is not younger than 24 hours. The
measurement was sound and the sentence was not; the claim only holds at the
one-day line, where it is 10 rather than 12. Nobody on the Claude side caught
it, including the author twice.
⚠ **The hold's live blast radius is SMALL** — only 4 booths have marks at all —
so the `.forever` re-count prediction rests on BOTH halves of U4 and on the
sentinel becoming unnecessary rather than forbidden. **RE-COUNT A FORTNIGHT
AFTER U4 LANDS**, i.e. on or after **2026-10-06**. If the rate does not move,
the honest readings are "the diagnosis was wrong" OR "the habit outlived the
need", and a bare re-count cannot tell those apart. **The three
open-pick-plus-`.forever` booths are the ones to watch**, because for them the
mechanism is now unambiguous.
@@ -0,0 +1,57 @@
# Four independent paths to one fail-open delete
_2026-09-22 · booth_
The U4 bug-hunt panel declared invariant was **"a deletion decision must never
be made from a read that failed"**. The panel found **four independent paths
through it, and no single arm found all four.** That is the strongest argument
yet for running the panel rather than one arm.
1. **An entry-level hydration error lost its hold** (the round's best finding).
`.marks.json` parses; one mark fails normalization; `_hydrate_safe` returns a
`Mark` carrying `error`; `_is_open` returns False for an errored pick — **on
purpose**, because a broken pick can never be answered. So the booth read as
not-held and **swept**, while the panel beside it rendered the broken mark in
full. The fail-safe had been built for FILE-level damage and missed
ENTRY-level. A mark we cannot read is judgment we cannot see; deleting the
booth it belongs to is the one thing we must not do with it.
2. **A present-but-blank `.marks.json` swept.** `_read_raw_strict` early-returns
for whitespace-only content — correct for the WRITE path it was written for
(a blank file is safe to overwrite), wrong for the DELETE path. Fixed with a
`blank_is_corrupt=True` flag used only by `hold_read`. ⚠ The near-regression
worth remembering: a **valid document with an empty `marks` list** is what
deleting the last mark leaves behind, and holding on THAT would make every
finished booth immortal. Blank bytes are damage; an empty list is an answer.
3. **`_newest_mtime` returned 0.0 when the booth's own stat failed**, which made
it maximally ancient and therefore the FIRST thing the sweeper takes — a
permissions problem resolving to a deletion. Now returns `now`: not knowing a
booth's age is a reason to leave it alone. ⚠ Per-entry `FileNotFoundError`
stays a skip, because a dangling symlink raises it and has no mtime worth
counting; only OTHER stat errors mean "something is here we cannot read".
4. **`is_kept` collapsed a stat failure into not-kept.** `Path.exists()` maps
ELOOP and EACCES to False. Now `lstat`, with any non-ENOENT error reading as
KEPT, and a `.forever` symlink counting dangling or not.
**`is_held` was replaced by `hold_reason`, which returns the REASON** —
`"open"`, `"unreadable"`, or None — rather than a bool beside a separate error
string. Two representations of one state drift; Regin independently flagged that
the display could not tell the two holds apart. One value, read by the sweeper
and by all four rendering surfaces.
**Convergent 3-of-4, and the one with teeth beyond lifetime:** `record_view`
used `Path.touch()`, which FOLLOWS an existing symlink. A booth carrying a
planted `.viewed -> /anywhere` turned every page view into an mtime write at an
arbitrary path under the service uid — and **any fleet session can write into a
booth, because making a folder is the whole API.** Now `os.open(..., O_NOFOLLOW)`
plus `os.utime(fd)`; a planted link raises ELOOP into the existing swallow.
⚠ **THE CAPTURE TOOLING FAILED SILENTLY AND THE PEER CAUGHT IT, NOT US.** The
snapshot `files/` tree shipped to the arms was EMPTY. The loop was
`for f in $IN` over a multi-line variable — and **zsh does not word-split
unquoted parameter expansions the way bash does**, so it iterated once against a
path that was the entire list. jekyll recovered by re-applying the bundled diff
to HEAD and verified every file byte-identical, so the round was sound. **The
failure mode is the dangerous one: an empty bundle reads exactly like a clean
result.** Quote-and-split explicitly (`print -r -- $IN | while read f`) or build
the list as a real array. Same family as `[[2026-09-22-vacuous-falsifiers]]` —
an instrument that cannot fail loudly will fail quietly.
@@ -0,0 +1,42 @@
# The third one-branch template miss — this repo's recurring blind spot
_2026-09-22 · booth_
**All four arms of the U4 code-review panel found the same drift, independently.**
That is the strongest convergence either panel has produced here.
The booth header's sub-line forks on `{% if board %}`, and the U4 lifetime macro
had been added only to the `{% else %}`. So **a booth carrying `links.md`
rendered a link count and nothing at all about its lifetime** — no countdown, no
hold — while INV-4 said the templates have no path that renders neither. The
standing board being kept by construction (`booth link` drops `.forever` on
first use) is what hid it; a **released** board or a hand-made `links.md` booth
is a live non-kept booth on that path, and both are reachable from the UI.
**This is the third of the same shape in this repo's short history:**
1. `blurtoggle` — the blur only patched the image/video `<figure>`; inline docs
render through their OWN branch and shipped unblurred. Suite green; a live
look caught it.
2. verbatim chrome — a verbatim booth's own `index.html` is served untouched, so
the inline marks panel never renders there. Found by looking at the live
service during U4, not by the suite.
3. the board branch — this one.
**The pattern: the suite renders the surface the author was thinking about.**
Every one of these was a second branch of a conditional the author had already
satisfied once and stopped reading. A cold reader with no idea which branch was
"the real one" finds them; the author does not, and neither does a test the
author wrote.
**Practical consequence for this repo.** When a template gains a fact, grep the
template for `{% if %}` in the block you edited and render EVERY branch in a
test — one test per branch, each rendering only its own surface, or the passing
test on branch A will mask the omission on branch B. U4 now has one per surface
(index card, booth header, board header, marks page) for exactly this reason.
Declined, and worth recording: Regin and Kimi both recommended amending INV-4 to
carve the board header out, on the grounds that board layout belongs to U7.
**Cutting an invariant down to fit an implementation gap is the wrong direction
when the fix is one template edit**, and U7 owns navigation and section layout —
not whether a header states a lifetime.
@@ -0,0 +1,35 @@
# Two reads of one file are not one read of one state
_2026-09-22 · booth_
**The one finding across both U4 panels that changed code rather than prose,
and it came from Hulda (Codex) on the CONTRACT-paraphrase round — before any
code existed.**
The contract specified the hold check as:
is_held(marks_for(child), read_error(child))
Two reads of `.marks.json`, presented as one answer. They are not. A write or a
repair landing between them yields a pair that described the booth at **no
instant**, and the losing pair is `([], None)` — no marks, no error — which is
**exactly the pair that deletes**. A lenient reader plus a strict reader, each
correct on its own, compose into a fail-open delete.
The fix is `booth.marks.hold_read(booth) -> (marks, error)`: ONE strict read
answering both questions. `sweep_once` now does one read per booth per tick
instead of two. And because `_read_raw_strict` **raises rather than dropping an
entry**, a non-raising strict read returns exactly what the lenient read would —
so the index uses that same one read for its badge too and falls back to
`marks_for` only on the error path, where leniency is the point. Better than the
original in both correctness and cost.
**The generalisable class, in heid's words: a two-read seam presented as one
answer is a TOCTOU race even when nothing on the page looks concurrent.** Worth
looking for anywhere two reader functions with different strictness feed one
decision — especially when that decision ends in `rmtree`.
Related: `[[2026-09-21-marks-write-wiped-judgment]]` is the same
reads-lenient/writes-strict asymmetry; U4 extends it to the reaper with
"deletes strict", whose scope is **the sweeper only** — a hand delete is never
strict, which is what gives an unreadable-marks hold an exit at all.
@@ -0,0 +1,50 @@
# U4 landed — lifetime is derived, not declared
_2026-09-22 · booth_
**A booth's lifetime stopped being a boolean somebody remembered to press.**
Three states now, and `sweep_once` is the only thing that honours the first two:
KEPT `.forever` present never swept (unchanged)
HELD an open pick, or marks we cannot read never swept (new)
EPHEMERAL everything else 24h (unchanged)
Plus **viewing is activity**: a deliberately-served response from a booth's own
page route writes `.viewed`. That dotfile is not a `.lock` dotfile, so
`_newest_mtime` already counts it — **there is no new arithmetic anywhere**.
`booth_age_seconds`, `is_expired` and `expires_in` are byte-for-byte what they
were. A view is one more thing in the tree, which is the same trick `.booth.json`
used in U5.
**What counts as a view, and why the exclusions matter more than the inclusions.**
`/b/<n>/` (gallery, verbatim report, `?download=1` zip), `/b/<n>/view` and
`/b/<n>/marks` count. `/b/<n>/marks.json`, asset GETs, `/`, `/healthz` and a
zoom URL that 404s do NOT. The marks.json exclusion is load-bearing: **an agent
must not be able to hold its own booth open by polling for the answer it is
waiting on.** `/b/<n>/asks` is a 308 into `/marks` and records through it — one
call, not two.
Checked because it would have been silent: **nothing in the fleet polls a booth
page.** Homepage's `siteMonitor` for the Booth is `/healthz`, which is on the
not-a-view list. Had it been pointed at a booth URL, every booth would have
become immortal on deploy and nothing would have reported it.
**The hold is unbounded and that is the point** — unanswered is unfinished. What
makes it safe is visibility plus two exits that already existed: the card and
every Booth-owned header say `held until answered` where the countdown was, and
`booth rm` / the UI x / `DELETE /b/<n>` take a held booth exactly as they take a
kept one. **A hold is protection from the timer, never from the operator.**
**Release is activity, stated rather than accidental.** Releasing a kept board
still buys a full TTL — unchanged — but now because `booth_unkeep` calls
`record_view`, which is a rule, and no longer because unlinking a file happened
to bump a directory's mtime, which is not. The CLI warning against
"unkeep and let it expire" stays and stays true.
⚠ **Running `scripts/layout-probe.py` over booth pages resets every booth's
clock**, because a GET of a booth page is a view and the probe is not exempt
from its own rule. Harmless, recoverable, and noted in the probe so nobody
debugs it later as a sweeper that stopped working.
Contract: `docs/contracts/u4_derived_lifetime.contract.md`. Both heid panels ran
and the bug hunt after them; see the sibling entries.
@@ -0,0 +1,40 @@
# Five of seven INV falsifiers did not falsify anything
_2026-09-22 · booth_
The U4 contract carried seven invariants, each with a *Falsifiable:* line, and
each had a test. **The code-review panel showed that five of the seven tests
would still pass under a change that defeats the invariant they name.** Gróa's
"per INV entry, what would still pass" section is the single most useful thing
either panel produced on this unit.
| INV | what the test asserted | what still passed |
|---|---|---|
| 1 (no new arithmetic) | the clock moved after a view | special-casing `.viewed` inside `_newest_mtime` — the exact new arithmetic INV-1 forbids |
| 3 (`is_held` is pure) | the right answer, once | `is_held` doing I/O, or `return True` unconditionally |
| 4 (every surface says why) | a substring on `GET /` | dropping the line from the booth header, the marks page, or the board branch |
| 5 (a view cannot fail a request) | `record_view` did not raise | a second `touch` outside the guard, 500ing all three routes |
| 6 (unreadable marks hold) | the corrupt booth survived | a sweeper that deletes nothing at all (no doomed sibling in the fixture) |
| 7 (machine reads do not hold) | `.viewed` was absent | a handler writing any other non-dot file, holding the booth open just as well |
**The shape of the error is the same every time: the test asserted the OUTCOME
the author was thinking about, not the DISCRIMINATOR the invariant names.** A
green test proved the happy path and nothing about the invariant. Writing the
falsifiable line in the contract did not produce a falsifying test — it produced
a test that *cited* one.
Fixed by rewriting each to fail under the change that defeats it: same-mtime
equivalence with an arbitrary non-lock dotfile (plus a `.lock` that must NOT
count); `is_held` called with marks belonging to a booth that does not exist on
disk; one test per rendered surface, each rendering only its own; the three
routes GET against a chmod'd booth; a doomed sibling; the AGE asserted rather
than the marker. **The board-header pair was verified RED against the pre-fix
template rather than assumed** — which is the step that makes "fixed, not
amended" trustworthy.
**The method to keep: for each invariant, name a change that defeats it and ask
whether the test goes red.** If you cannot name one, the invariant is not
falsifiable yet. Regin and Kimi independently proposed this as a contract-time
"vacuity pass"; heid rates this round the strongest evidence for it so far, and
it is a `/heid*` skill proposal sitting with the operator, not a change to this
repo.