fix(manifest)!: the size cap opened a service-wide hang; close it

The diff-scoped bug-hunt panel, four arms, artifact-only. Its strongest
finding is one I created two hours earlier while hardening the reader.

`stat` reports size 0 for a FIFO and 0 for a symlink to /dev/zero, so both
sail under the byte cap added for the RecursionError round — and then
`read_text` either blocks in read() with no EOF, so the except never runs,
or allocates until the kernel intervenes. `list_booths` reads every booth
on every GET / and /healthz, so ONE such file stalls the front page for the
whole service, with no error and no recovery short of a restart.
Reproduced before believing it (timeout returned 124). S_ISREG is checked
BEFORE the size in both modules now; verified against the live service with
two FIFOs planted, which answered 200 in 36ms.

The shape worth carrying: st_size answers a different question than "can
this be read", and a bound that trusts it inherits everything it does not
mean. A hardening fix opened a worse hole than the one it closed.

THE UPLOAD PATH WROTE ABOVE ITS OWN CLEANUP GUARD (4/4)

A failed manifest write orphaned a .uploaded half-booth with no files in
it — and because the temp name now carries a random suffix, nothing ever
overwrote the leak, and .booth.json.<hex>.tmp is not a .lock, so
_newest_mtime counted it and kept that empty booth past every sweep. The
uniqueness fix from the previous round is what made the leak permanent.
Both writes moved inside the guard; the temp is removed on every exit path.

DAMAGED BYTES ARE KEPT, NOT REPLACED (4/4, INV-6)

Marks made this explicit in v0.2.1 and this write path contradicted it: a
manifest that failed on ONE field lost the others with it, including a why
the re-announcer may never have kept anywhere. It diverges from marks in
HOW it honours the rule — marks refuse and answer 409 because the
operator's judgment is not restatable; a manifest quarantines and proceeds,
because refusing would fail `booth add` and lose the files it was copying.

ONE OPENNESS PREDICATE, AS U2 SAID (2/4)

`booth answer` spelled out `if m.answer is None` while `booth marks` asked
`open_marks`, so a partially-answered pick read as done to one verb and
open to the other — at the same instant, on the same booth. U2's INV-2 put
openness in one function precisely so they could not drift. The mirror case
is fixed too: a pick that hydrates broken is refused by the web route, so
`answer --wait` polled an hour on a form nothing could ever land.

ALSO

- now_stamp was whole-second while the importer had moved to microseconds,
  and '-' sorts before '.', so a later mark came out ahead of an earlier
  import inside the same second. One format; the previous round's ordering
  fix had opened this one.
- `_broken` was the third of three directory-name fallbacks and the one
  still handing a raw name into a card's sub-line.
- An identical re-announce rewrote the file and reset the TTL. `booth link`
  does this on every post to the standing board.
- The importer's return went through the bare _hydrate, not _hydrate_safe.
- A marks document could be written larger than it can be read back, and
  then read as no marks at all. Refused at the write instead.
- `choice` reached the answer builder raw while `notes` beside it did not.

AND ONE FINDING DELIBERATELY NOT FULLY CLOSED

The mtime-restore race is real. The clean fix — ignore a booth directory's
own mtime whenever the booth holds anything — also silently retires the
documented rule that releasing a kept board resets its clock, which the CLI
header, the README and a deliberately-written test all pin. That is a TTL
doctrine change, not a bug fix, and an existing test caught the attempt.
The concrete half is fixed (a failing os.utime escaped and 500'd the
route); the race is stated in the code where the next reader will meet it.

341 tests. Live service restarted, 24/24 booth pages verified.
This commit is contained in:
Vuong Hoang
2026-09-22 02:27:18 -07:00
parent f3193fb054
commit 95beede3c3
11 changed files with 616 additions and 61 deletions
+17 -8
View File
@@ -25,15 +25,24 @@ surface) and therefore the safest thing to land first or in parallel.
**U1, U2 and U5 are landed.** U3 and U4 are unblocked and unstarted; U6 remains
independent and unstarted; U7 waits on the rest.
**U5's adoption is a measured prediction, not a finished result.** The operator
declined a fleetwide announcement so that adoption could be told apart from
design: the convention propagates through the README alone, and the count of
booths carrying a `.booth.json` gets re-measured on **2026-09-29** against a
baseline of **0 of 26** at landing. A near-zero count means nobody heard about
it — an adoption failure, fixed by announcing — which is a different thing from
nobody wanting it. Same instrument as U4's `.forever` prediction below.
**U5's adoption is a measured prediction, not a finished result**, and it is
TWO predictions rather than one. The operator declined a fleetwide announcement
so that adoption could be told apart from design; within fifty minutes of the
deploy a peer that had been told nothing (`comfy-dev`) created a booth and it
announced itself with a handle and an empty `why`. That is the split:
find ~/booth-data -maxdepth 2 -name .booth.json | wc -l
- **The handle rides for free.** It is written by `booth new` and `booth add`,
so every existing caller starts announcing without learning anything.
- **The `why` has to be learned.** It needs someone to know the flag exists.
Both get re-measured on **2026-09-29**:
find ~/booth-data -maxdepth 2 -name .booth.json | wc -l # free
grep -l '"why": "[^"]' ~/booth-data/*/.booth.json | wc -l # learned
A high first count with a near-zero second is the predicted shape of "nobody was
told" — an adoption failure fixed by announcing, which is a different thing from
nobody wanting it. Same instrument as U4's `.forever` prediction below.
### Cross-cutting invariant — deterministic order, everywhere