fix(manifest)!: the size cap opened a service-wide hang; close it
The diff-scoped bug-hunt panel, four arms, artifact-only. Its strongest finding is one I created two hours earlier while hardening the reader. `stat` reports size 0 for a FIFO and 0 for a symlink to /dev/zero, so both sail under the byte cap added for the RecursionError round — and then `read_text` either blocks in read() with no EOF, so the except never runs, or allocates until the kernel intervenes. `list_booths` reads every booth on every GET / and /healthz, so ONE such file stalls the front page for the whole service, with no error and no recovery short of a restart. Reproduced before believing it (timeout returned 124). S_ISREG is checked BEFORE the size in both modules now; verified against the live service with two FIFOs planted, which answered 200 in 36ms. The shape worth carrying: st_size answers a different question than "can this be read", and a bound that trusts it inherits everything it does not mean. A hardening fix opened a worse hole than the one it closed. THE UPLOAD PATH WROTE ABOVE ITS OWN CLEANUP GUARD (4/4) A failed manifest write orphaned a .uploaded half-booth with no files in it — and because the temp name now carries a random suffix, nothing ever overwrote the leak, and .booth.json.<hex>.tmp is not a .lock, so _newest_mtime counted it and kept that empty booth past every sweep. The uniqueness fix from the previous round is what made the leak permanent. Both writes moved inside the guard; the temp is removed on every exit path. DAMAGED BYTES ARE KEPT, NOT REPLACED (4/4, INV-6) Marks made this explicit in v0.2.1 and this write path contradicted it: a manifest that failed on ONE field lost the others with it, including a why the re-announcer may never have kept anywhere. It diverges from marks in HOW it honours the rule — marks refuse and answer 409 because the operator's judgment is not restatable; a manifest quarantines and proceeds, because refusing would fail `booth add` and lose the files it was copying. ONE OPENNESS PREDICATE, AS U2 SAID (2/4) `booth answer` spelled out `if m.answer is None` while `booth marks` asked `open_marks`, so a partially-answered pick read as done to one verb and open to the other — at the same instant, on the same booth. U2's INV-2 put openness in one function precisely so they could not drift. The mirror case is fixed too: a pick that hydrates broken is refused by the web route, so `answer --wait` polled an hour on a form nothing could ever land. ALSO - now_stamp was whole-second while the importer had moved to microseconds, and '-' sorts before '.', so a later mark came out ahead of an earlier import inside the same second. One format; the previous round's ordering fix had opened this one. - `_broken` was the third of three directory-name fallbacks and the one still handing a raw name into a card's sub-line. - An identical re-announce rewrote the file and reset the TTL. `booth link` does this on every post to the standing board. - The importer's return went through the bare _hydrate, not _hydrate_safe. - A marks document could be written larger than it can be read back, and then read as no marks at all. Refused at the write instead. - `choice` reached the answer builder raw while `notes` beside it did not. AND ONE FINDING DELIBERATELY NOT FULLY CLOSED The mtime-restore race is real. The clean fix — ignore a booth directory's own mtime whenever the booth holds anything — also silently retires the documented rule that releasing a kept board resets its clock, which the CLI header, the README and a deliberately-written test all pin. That is a TTL doctrine change, not a bug fix, and an existing test caught the attempt. The concrete half is fixed (a failing os.utime escaped and 500'd the route); the race is stated in the code where the next reader will meet it. 341 tests. Live service restarted, 24/24 booth pages verified.
This commit is contained in:
+17
-8
@@ -25,15 +25,24 @@ surface) and therefore the safest thing to land first or in parallel.
|
||||
**U1, U2 and U5 are landed.** U3 and U4 are unblocked and unstarted; U6 remains
|
||||
independent and unstarted; U7 waits on the rest.
|
||||
|
||||
**U5's adoption is a measured prediction, not a finished result.** The operator
|
||||
declined a fleetwide announcement so that adoption could be told apart from
|
||||
design: the convention propagates through the README alone, and the count of
|
||||
booths carrying a `.booth.json` gets re-measured on **2026-09-29** against a
|
||||
baseline of **0 of 26** at landing. A near-zero count means nobody heard about
|
||||
it — an adoption failure, fixed by announcing — which is a different thing from
|
||||
nobody wanting it. Same instrument as U4's `.forever` prediction below.
|
||||
**U5's adoption is a measured prediction, not a finished result**, and it is
|
||||
TWO predictions rather than one. The operator declined a fleetwide announcement
|
||||
so that adoption could be told apart from design; within fifty minutes of the
|
||||
deploy a peer that had been told nothing (`comfy-dev`) created a booth and it
|
||||
announced itself with a handle and an empty `why`. That is the split:
|
||||
|
||||
find ~/booth-data -maxdepth 2 -name .booth.json | wc -l
|
||||
- **The handle rides for free.** It is written by `booth new` and `booth add`,
|
||||
so every existing caller starts announcing without learning anything.
|
||||
- **The `why` has to be learned.** It needs someone to know the flag exists.
|
||||
|
||||
Both get re-measured on **2026-09-29**:
|
||||
|
||||
find ~/booth-data -maxdepth 2 -name .booth.json | wc -l # free
|
||||
grep -l '"why": "[^"]' ~/booth-data/*/.booth.json | wc -l # learned
|
||||
|
||||
A high first count with a near-zero second is the predicted shape of "nobody was
|
||||
told" — an adoption failure fixed by announcing, which is a different thing from
|
||||
nobody wanting it. Same instrument as U4's `.forever` prediction below.
|
||||
|
||||
### Cross-cutting invariant — deterministic order, everywhere
|
||||
|
||||
|
||||
Reference in New Issue
Block a user