feat(u4): a booth's lifetime is derived from its state, not from a boolean

`.forever` was the only way to say three different things — "this is durable",
"I have not answered yet", "I am still looking" — and the census said it was
carrying all three: 17 of 24 live booths (70%, up from 54% the day before).
Three of the four booths in the fleet awaiting an answer had been pinned by
hand as well, and 10 of the 17 were younger than the TTL, so the sentinel had
bought them nothing and was pressed pre-emptively.

Only the first meaning is what `keep` means. The other two are facts the
service already held and did not consult.

    KEPT       `.forever` present                      never swept  (unchanged)
    HELD       an open pick, or marks we cannot read   never swept  (new)
    EPHEMERAL  everything else                         24h          (unchanged)

Viewing is activity: a deliberately-served response from a booth's own page
route writes `.viewed`, which is a dotfile and not a `.lock` dotfile, so
`_newest_mtime` already counts it. There is no new arithmetic — `booth_age_seconds`,
`is_expired` and `expires_in` are unchanged. Machine reads are excluded on
purpose: an agent must not be able to hold its own booth open by polling for
the answer it is waiting on.

The hold is unbounded, and what makes that safe is visibility plus two exits
that already existed. Every surface whose chrome the Booth owns says
`held until answered` where the countdown was, and `booth rm` / the UI x /
`DELETE /b/<n>` take a held booth exactly as they take a kept one. A hold is
protection from the timer, never from the operator.

Three cross-frontier panels ran and each found a class the others could not:

  * the paraphrase panel found that two reads of one file are not one read of
    one state — the contract's `is_held(marks_for(c), read_error(c))` could
    resolve to `([], None)`, the pair that deletes. `hold_read` is one read.
  * the code-review panel found, 4-of-4, that the booth header's board branch
    rendered no lifetime at all; and that five of seven invariant tests passed
    under the change that defeats them.
  * the bug-hunt panel found four more paths where a failed read still
    authorized a delete, and a `record_view` that followed a planted symlink.

`is_held` became `hold_reason`, which returns the reason rather than a bool
beside a string that can disagree with it.

Prediction, to re-count on or after 2026-10-06: the `.forever` rate falls to
the booths that are genuinely durable references. Only 4 booths carry marks at
all, so this rests on both halves of the unit; a null result cannot distinguish
a wrong diagnosis from a habit that outlived its need.

406 tests (341 before). Contract: docs/contracts/u4_derived_lifetime.contract.md
This commit is contained in:
vh
2026-09-22 09:44:25 -07:00
parent d37b81ab9f
commit c3a97c1b64
22 changed files with 2339 additions and 71 deletions
+40 -1
View File
@@ -217,12 +217,21 @@ def _fingerprint(entries: list[dict]) -> str:
return json.dumps(entries, sort_keys=True, ensure_ascii=False)
def _read_raw_strict(booth: Path) -> list[dict]:
def _read_raw_strict(booth: Path, *, blank_is_corrupt: bool = False) -> list[dict]:
"""Like `_read_raw`, but RAISES `MarksCorrupt` on a file it cannot parse.
Absent, empty and valid-but-empty are all "no marks yet" and are fine — the
distinction that matters is bytes-present-but-unreadable, because that is the
case where writing would destroy something.
`blank_is_corrupt` is the DELETE path's reading of a present-but-whitespace
file, and only the delete path's: this writer never produces a blank marks
document, so a blank one that exists is something that went wrong, and
`rmtree` is not the response to that. The write path keeps the lenient
reading — a blank file is safe to overwrite, which is the question
`_Locked` is asking. A VALID document with an empty `marks` list is not
blank and never holds: that is what deleting the last mark leaves behind,
and it must stay sweepable.
"""
path = Path(booth) / MARKS_FILE
try:
@@ -245,6 +254,8 @@ def _read_raw_strict(booth: Path) -> list[dict]:
except (OSError, UnicodeDecodeError, MemoryError) as exc:
raise MarksCorrupt(f"{path} cannot be read: {exc}") from exc
if not text.strip():
if blank_is_corrupt:
raise MarksCorrupt(f"{path} is present but holds no marks document")
return []
try:
raw = json.loads(text)
@@ -545,6 +556,34 @@ def open_marks(marks: Sequence[Mark]) -> list[Mark]:
return [m for m in marks if _is_open(m)]
def hold_read(booth: Path) -> tuple[list[Mark], str | None]:
"""ONE read of `.marks.json`, answering both questions the LIFETIME rule asks:
what is still open, and whether the file could be read at all.
U4 decides whether a booth may be SWEPT from those two facts. Asking them
with two calls — `marks_for` then `read_error` — reads the file twice, and
two reads of one file are not one read of one state: a write or a repair
landing between them yields a pair that never described the booth at any
instant. The losing pair is `([], None)` — no marks, no error — which is
exactly the one that deletes. Cross-frontier review (2026-09-22) found it;
that is why this exists rather than the obvious two calls.
When the file reads clean the marks are byte-identical to `marks_for`'s:
`_read_raw_strict` raises rather than dropping an entry, so a non-raising
strict read returns the same entries the lenient read would, hydrated and
sorted the same way. The caller can therefore use this ONE read for the
display too, and fall back to `marks_for` only on the error path, where
leniency is the point.
"""
try:
entries = _read_raw_strict(booth, blank_is_corrupt=True)
except MarksCorrupt as exc:
return [], str(exc)
marks = [_hydrate_safe(e) for e in entries]
marks.sort(key=lambda m: (m.created, m.id))
return marks, None
def marks_for_target(marks: Sequence[Mark], rel: str | None) -> list[Mark]:
"""The marks attached to one item, or to the booth itself for None."""
return [m for m in marks if m.target == rel]