18 Commits
Author SHA1 Message Date
vh c75d7a2797 fix: four defects the U4 bug-hunt panel found in code it did not add
All four pre-date U4 and sit in files it touched, which is why a diff-scoped
robustness lens saw them. They are separated from the unit's own commit so the
feature history stays readable; the release tags both.

* A booth name reached a JS string context. The confirm dialogs interpolated
  the name into a string literal inside `onsubmit`. Jinja's autoescape is
  HTML-attribute escaping, not JS-string escaping: the browser decodes the
  entity back to a quote before the JS parser sees it, so a name crafted to
  close the string executed on submit. Booth names are agent-authored — making
  a folder under the data dir is the whole API — so this was a live path, not a
  theoretical one. The name now travels as a data attribute to a delegated
  handler, where escaping is escaping.

* An unreadable `links.md` returned 500 for the whole booth page. `is_file()`
  then an unguarded `read_text()`. The board is one tile on that page, and a
  page that will not load is worse than one missing a tile — the posture
  `read_blurred`, `marks_for` and `read_manifest` already take.

* The index order had no tie-breaker, which violates the deterministic-order
  invariant. Equal-mtime booths fell back to whatever `iterdir()` yielded, and
  two booths landed by one `rsync` batch share an mtime exactly. Now
  `(mtime, name)` reverse: newest first, then name. The operator refers to
  cards positionally, so a sequence that moves between renders misfiles his
  judgment rather than crashing.

* `/b/<n>/marks.json` reported damage as empty success. `booth marks` exits 3
  on an unreadable file precisely so a caller can tell "not yet" from "broken";
  the HTTP mirror — the only reader a remote session has — returned the same
  empty list for both. It now carries `error` and `detail`. The status stays
  200 deliberately: reads are lenient here, and a pinned status code is a
  promise to remote clients this fix has no business breaking.

Each has a regression test. 410 tests.
2026-09-22 09:51:14 -07:00
vh c3a97c1b64 feat(u4): a booth's lifetime is derived from its state, not from a boolean
`.forever` was the only way to say three different things — "this is durable",
"I have not answered yet", "I am still looking" — and the census said it was
carrying all three: 17 of 24 live booths (70%, up from 54% the day before).
Three of the four booths in the fleet awaiting an answer had been pinned by
hand as well, and 10 of the 17 were younger than the TTL, so the sentinel had
bought them nothing and was pressed pre-emptively.

Only the first meaning is what `keep` means. The other two are facts the
service already held and did not consult.

    KEPT       `.forever` present                      never swept  (unchanged)
    HELD       an open pick, or marks we cannot read   never swept  (new)
    EPHEMERAL  everything else                         24h          (unchanged)

Viewing is activity: a deliberately-served response from a booth's own page
route writes `.viewed`, which is a dotfile and not a `.lock` dotfile, so
`_newest_mtime` already counts it. There is no new arithmetic — `booth_age_seconds`,
`is_expired` and `expires_in` are unchanged. Machine reads are excluded on
purpose: an agent must not be able to hold its own booth open by polling for
the answer it is waiting on.

The hold is unbounded, and what makes that safe is visibility plus two exits
that already existed. Every surface whose chrome the Booth owns says
`held until answered` where the countdown was, and `booth rm` / the UI x /
`DELETE /b/<n>` take a held booth exactly as they take a kept one. A hold is
protection from the timer, never from the operator.

Three cross-frontier panels ran and each found a class the others could not:

  * the paraphrase panel found that two reads of one file are not one read of
    one state — the contract's `is_held(marks_for(c), read_error(c))` could
    resolve to `([], None)`, the pair that deletes. `hold_read` is one read.
  * the code-review panel found, 4-of-4, that the booth header's board branch
    rendered no lifetime at all; and that five of seven invariant tests passed
    under the change that defeats them.
  * the bug-hunt panel found four more paths where a failed read still
    authorized a delete, and a `record_view` that followed a planted symlink.

`is_held` became `hold_reason`, which returns the reason rather than a bool
beside a string that can disagree with it.

Prediction, to re-count on or after 2026-10-06: the `.forever` rate falls to
the booths that are genuinely durable references. Only 4 booths carry marks at
all, so this rests on both halves of the unit; a null result cannot distinguish
a wrong diagnosis from a habit that outlived its need.

406 tests (341 before). Contract: docs/contracts/u4_derived_lifetime.contract.md
2026-09-22 09:44:25 -07:00
Vuong Hoang d37b81ab9f memory: snapshot — U5 released at v0.3.0, and the index goes two-tier
The two dated log sections had never been split, so every one of their 29
entries sat inline and the startup index had grown to 372 lines — which is
the cost the two-tier scheme exists to remove, paid on every session that
reads the file. 27 entries were over threshold. All 29 now have a detail
file under persistent-memory.d/ and a one-line index entry that routes
rather than restates. Index: 372 -> 93 lines.

No archival. The soft cap fired, but every entry in this repo is dated
2026-09-21 or later, so the under-14-days guard held all of them back — and
the split alone took the index well under the target without moving
anything out of the active file.

The in-flight section is rewritten for the post-release state: nothing is
in flight, no gate is outstanding, and the next unit is explicitly recorded
as the operator's undecided call rather than as a plan. The session's
recommendation (U4, on three grounds) is written down so it does not have
to be re-derived, alongside the two alternatives and why they are
alternatives.

Two dated predictions are carried forward with their dates and their
instruments: the U5 adoption re-measure on 2026-09-29, which already reads
3 of 24 announced and 2 with a why from peers told nothing, and the
.forever re-count a fortnight AFTER U4 lands, which is U4's own success
criterion and is destroyed by running it early.
2026-09-22 08:20:51 -07:00
Vuong Hoang 95beede3c3 fix(manifest)!: the size cap opened a service-wide hang; close it
The diff-scoped bug-hunt panel, four arms, artifact-only. Its strongest
finding is one I created two hours earlier while hardening the reader.

`stat` reports size 0 for a FIFO and 0 for a symlink to /dev/zero, so both
sail under the byte cap added for the RecursionError round — and then
`read_text` either blocks in read() with no EOF, so the except never runs,
or allocates until the kernel intervenes. `list_booths` reads every booth
on every GET / and /healthz, so ONE such file stalls the front page for the
whole service, with no error and no recovery short of a restart.
Reproduced before believing it (timeout returned 124). S_ISREG is checked
BEFORE the size in both modules now; verified against the live service with
two FIFOs planted, which answered 200 in 36ms.

The shape worth carrying: st_size answers a different question than "can
this be read", and a bound that trusts it inherits everything it does not
mean. A hardening fix opened a worse hole than the one it closed.

THE UPLOAD PATH WROTE ABOVE ITS OWN CLEANUP GUARD (4/4)

A failed manifest write orphaned a .uploaded half-booth with no files in
it — and because the temp name now carries a random suffix, nothing ever
overwrote the leak, and .booth.json.<hex>.tmp is not a .lock, so
_newest_mtime counted it and kept that empty booth past every sweep. The
uniqueness fix from the previous round is what made the leak permanent.
Both writes moved inside the guard; the temp is removed on every exit path.

DAMAGED BYTES ARE KEPT, NOT REPLACED (4/4, INV-6)

Marks made this explicit in v0.2.1 and this write path contradicted it: a
manifest that failed on ONE field lost the others with it, including a why
the re-announcer may never have kept anywhere. It diverges from marks in
HOW it honours the rule — marks refuse and answer 409 because the
operator's judgment is not restatable; a manifest quarantines and proceeds,
because refusing would fail `booth add` and lose the files it was copying.

ONE OPENNESS PREDICATE, AS U2 SAID (2/4)

`booth answer` spelled out `if m.answer is None` while `booth marks` asked
`open_marks`, so a partially-answered pick read as done to one verb and
open to the other — at the same instant, on the same booth. U2's INV-2 put
openness in one function precisely so they could not drift. The mirror case
is fixed too: a pick that hydrates broken is refused by the web route, so
`answer --wait` polled an hour on a form nothing could ever land.

ALSO

- now_stamp was whole-second while the importer had moved to microseconds,
  and '-' sorts before '.', so a later mark came out ahead of an earlier
  import inside the same second. One format; the previous round's ordering
  fix had opened this one.
- `_broken` was the third of three directory-name fallbacks and the one
  still handing a raw name into a card's sub-line.
- An identical re-announce rewrote the file and reset the TTL. `booth link`
  does this on every post to the standing board.
- The importer's return went through the bare _hydrate, not _hydrate_safe.
- A marks document could be written larger than it can be read back, and
  then read as no marks at all. Refused at the write instead.
- `choice` reached the answer builder raw while `notes` beside it did not.

AND ONE FINDING DELIBERATELY NOT FULLY CLOSED

The mtime-restore race is real. The clean fix — ignore a booth directory's
own mtime whenever the booth holds anything — also silently retires the
documented rule that releasing a kept board resets its clock, which the CLI
header, the README and a deliberately-written test all pin. That is a TTL
doctrine change, not a bug fix, and an existing test caught the attempt.
The concrete half is fixed (a failing os.utime escaped and 500'd the
route); the race is stated in the code where the next reader will meet it.

341 tests. Live service restarted, 24/24 booth pages verified.
2026-09-22 02:27:18 -07:00
Vuong Hoang f3193fb054 fix(probe): the disclosure-opening loop was manufacturing its own findings
`page.locator("details:not([open])").all()` hands back POSITIONAL locators
that re-resolve against the current DOM, and `:not([open])` stops matching
an element the moment it is opened — so opening them one at a time shrinks
the set underneath the indices and leaves some closed. Those then report
OCCLUDED, which is exactly the false-positive class the block was added to
remove. One on booth-redesign, three on cr123a-to-d-sleeve, one on
denoise-first-run, and invisible as a bug because a false positive is
shaped like a finding.

Measured both hypotheses rather than guessing between them: per-element
loop against a single document-wide evaluate, at 150 ms and 1000 ms settle.
The loop reports them at either wait; the single pass reports none at
either. The variable was the method, not the timing.

One evaluate over the whole document now. All three pages clean.

Also carries the ROADMAP U5 row, the two-panel record in
persistent-memory.d/, and the memory index line for it.
2026-09-22 01:39:43 -07:00
Vuong Hoang c015a917ee fix(manifest): fold in both cross-frontier panels — and a live hole in v0.2.2
Two four-arm artifact-only rounds landed together: the contract paraphrase
(against the pre-seam-review capture) and the code-vs-contract conformance
review (against the amended one), correctly firewalled from each other.
The conformance round found ZERO drift in the strict sense — the code is a
clause-for-clause implementation of the contract — and the weight of both
rounds landed one layer down, in what green tests structurally cannot
report. Full triage in persistent-memory.d/.

A LIVE HOLE IN RELEASED CODE, FOUND ON THE SIBLING MODULE

v0.2.2 adopted the RecursionError finding from the bug-hunt round and
closed half of it: `_hydrate_safe` guards hydration, but `json.loads` runs
above it in `_read_raw`, whose catch list covers neither RecursionError nor
MemoryError. A 400 KB file of nothing but brackets in any ONE booth
therefore still returned 500 for `/` and `/healthz` across every booth on
the service. Confirmed by running it before believing it.

Both modules now bound the read by `stat` before touching the bytes and
catch both classes anyway, so raising a bound later cannot quietly re-open
the hole. The strict half of the marks asymmetry refuses everything the
lenient half tolerates, or a file that reads as "no marks" gets replaced by
a write that believed it.

THE WHY-WIPE

`booth new x --why "..."` then `booth add x out/*.png` erased the sentence
the first command existed to record. Omitted flags meant empty strings and
empty strings overwrote. Two arms predicted it from the contract's wording
alone; every test here passed --why on both calls and so could not see it.
Omitted now means unchanged and an explicit --why "" still clears — the
shell carries the distinction by leaving the variable UNSET, not empty.

--title WAS WRITE-ONLY

Stored, flag-surfaced, rendered nowhere. 4/4, and independently top-ranked
by every arm of the paraphrase round. It lands on the booth page heading
with the directory name beside it, because the directory name is the
identity the operator navigates by and refers to positionally.

THREE TESTS THAT COULD NOT FAIL

- test_the_write_is_atomic asserted no *.tmp survived, which a plain
  write_text passes. It asserts the inode changes now. (The first
  replacement was ALSO vacuous — it spied on os.open, which Path.write_text
  reaches through io.open in C and never touches. Recorded in the test,
  because writing a second vacuous test while fixing the first is exactly
  the failure this round is about.)
- The INV-3 preservation test passed against an implementation that
  regenerated `created` every time, because _now() is whole-second
  resolution and back-to-back writes share a stamp. Seeded from 2019 now.
- test_announcing_is_activity passed whether or not _newest_mtime counted
  the manifest, because writing it bumps the directory mtime either way.
  The directory's clock is put back, leaving the file as the only thing
  that can keep the booth alive.

ALSO

- The title fallback skipped the normalizer the explicit value gets; a
  directory name may legally carry a newline and run to 255 bytes.
- Every writer derived the same .booth.json.tmp. Marks are protected from
  that by their flock; the manifest has none, so uniqueness stands in.
- test_stdlib_only was blind to relative imports in all four modules.
- INV-1 had no guard at all; INV-5 named two different promises; the
  negative render states were asserted on the index only.

Contract amended throughout: the 4 GB case is a stat-checked bound rather
than a return constraint, every field of an error-carrying record has a
stated value, INV-1 no longer contradicts INV-3, repo-wide rules are named
in words instead of by a colliding number, and touches admits the macro
partial the implementation added.

329 tests.
2026-09-22 01:29:27 -07:00
Vuong Hoang fac83de8f4 docs(u5): name the verbatim-booth boundary as U3's, not a gap
Five live booths serve the author's HTML raw and the Booth owns no
header there to put a provenance line into — it reaches those pages
through six regexes injected into arbitrary markup, which is the defect
U3 exists to fix. Their index cards carry provenance like everything
else. Written down so the boundary reads as a boundary rather than as
something this unit forgot. Verified on pewpew-ui-brief.
2026-09-22 01:05:05 -07:00
Vuong Hoang aa61fcf5fd test(probe): teach the layout probe about <details>, and guard the flag parser
THE PROBE. A control inside a CLOSED <details> is laid out but sits
outside its collapsed parent's box, so elementFromPoint at its centre
returns an ancestor and it reports OCCLUDED — 23 of them on sindra-set,
every one a false positive. Verified both ways before believing it:
closed, elementFromPoint returns div.gallery; opened, the button itself,
and a real trial click lands on it.

Opening every <details> rather than skipping them is the deliberate
choice. Skipping would make the probe quiet by declaring put-away
controls out of scope, and the add-note button inside
details.item-addnote is exactly the class of control this instrument
exists to check. Fourth false-positive class this probe has grown a
guard for; the other three are already in its header.

THE FLAG PARSER. Three cases that silently break and are cheap to
pin: the flags on either side of the glob (a session should not have to
remember which), a why carrying quotes, an em-dash, a newline and
non-ASCII, and --why with no value after it, which must produce usage
rather than eating the booth name and creating a booth called nothing.

313 tests.
2026-09-22 01:01:20 -07:00
Vuong Hoang c9a175ba4a memory: U5 adoption is a prediction with a re-measure date
Operator declined the fleetwide announcement (2026-09-22) and chose to
let the convention propagate through the README alone, specifically so
adoption can be distinguished from design. Baseline 0 of 26 booths at
landing; re-count 2026-09-29. Near-zero means nobody heard about it,
which is a different failure from nobody wanting it.
2026-09-22 00:56:25 -07:00
Vuong Hoang 75dca53483 docs(probe): the probe covers the index only, and says so now
The docstring claimed a no-argument run probes 'the booth index and
every booth linked from it'. main() probes argv[1:] or the default URL
and follows nothing — so a coverage claim that reads as 26 pages has
always been one. A probe that overstates its reach is worse than one
that states a small reach honestly, because this is the instrument
standing in for a class of bug the test suite structurally cannot see.

Also records the zsh trap that hid it: an unquoted $URLS holding twelve
space-separated URLs arrives as ONE argument, and the probe cheerfully
reports '2 page(s)' while covering two.
2026-09-22 00:54:54 -07:00
Vuong Hoang 67ab7d1cd5 docs(readme): --why, on the page the 17 consuming handles actually read
The quickstart is where a session learns the CLI, so the announcement
verb has to be in the first code block rather than in a section further
down that nobody scrolls to. States the trade plainly: optional, nothing
breaks without it, and a booth that cannot say what it is has no way to
ask for attention except by posting its URL somewhere else.
2026-09-22 00:50:02 -07:00
Vuong Hoang ac35f2441f docs(u5): state what the contract deliberately leaves out
The out-of-scope block is load-bearing for the cross-frontier review
gates — without negative constraints their signal-to-noise drops sharply,
and both /heid-code-review and /heid-bug-hunt refuse to fire without one.
Written for the reviewer, but it is the same list the roadmap gate
produced: the what-landed feed is parked for v1.1, nothing enforces that
a booth must announce itself (rsync is a documented path and never runs
the CLI), and the manifest describes rather than decides — U4 owns
lifetime.
2026-09-22 00:49:22 -07:00
Vuong Hoang a48ef83ef5 feat(manifest): U5 — booths that say who posted them and why
The index card showed a name, an item count and a countdown, and nothing
the poster chose. An agent with something to show therefore had no way to
make the booth say "look at this" and posted a URL to the link board
instead — which is why 145 of that board's 210 rows (69%) ended up
pointing at booths that had already been swept. The board was absorbing a
job it was never shaped for. This is the shape.

Each booth carries `.booth.json` — {handle, title, why, created} — written
by the CLI from $ALTHING_HANDLE, and the provenance line renders on both
index lanes and on the booth page header.

WHAT IS WHERE

- booth/manifest.py, stdlib-only and importing nothing from booth.* either:
  scripts/booth imports it under the system python3 with no venv, and a
  cross-import between two stdlib-only modules is a second way for that
  invariant to break. It joins the shared test_stdlib_only list and keeps
  a stricter copy of its own.
- The read is lenient and cannot raise. list_booths touches every booth on
  every index load, so a manifest that cannot be parsed costs that booth's
  provenance and nothing else. That is the v0.2.2 lesson applied before the
  same mistake rather than after it.
- Absent and damaged render differently — `unannounced` and `unreadable`.
  Folding "cannot be read" into "never said" would hide the one case
  somebody has to go and fix.
- Re-announcing preserves `created`. A second `booth add` sharpening the
  why is not a second appearance of the booth.
- The write is atomic (invariant 5); the temp file is itself a dotfile, so
  no listing can see it mid-write.

THREE OPERATOR CALLS, 2026-09-22

Flags on the existing new/add verbs rather than a separate `announce` verb
(a second step is the step that gets forgotten, which is the rot's own
mechanism). Unannounced booths get a quiet marker rather than nothing — the
convention is only adoptable if the gap is visible. U5 adds provenance only
and does NOT add a second index ordering keyed on announcement time; that
is a different surface needing its own stated rule, parked for v1.1.

NO EXEMPTION LIST

A pickup booth and the standing link board are created by the service, so
they announce themselves with handle `booth`, which is true rather than
manufactured. One rule — a booth with no manifest is unannounced — instead
of a growing set of special cases.

ALSO

tests/test_booth.py's keep/release assertion was slicing the page on the
bare word `boothhead`, which has lived in the stylesheet far longer than
the assertion has; it was reading CSS and passing on luck, and went red the
first time a new rule landed above the old one. Same assertion, aimed at
the markup. A U5 test had the mirror-image bug: pytest derives tmp_path
from the test name and the index renders data_dir, so a test named
`test_an_unannounced_booth_says_so` put the needle in the haystack itself
and passed against a template that did not yet exist.

310 tests (304 before this unit's CLI half). Live service restarted, 26/26
booth pages verified 200, end-to-end smoke through the real CLI.

NOT TAGGED. The cold contract-review panel is still in flight and the
code-review and bug-hunt gates have not run. Tagging with a gate
outstanding is what made v0.2.0 premature.
2026-09-22 00:48:41 -07:00
Vuong Hoang 109190b0d6 docs(u5): contract for self-announcing booths, plus its seam review
Blast-radius pass first (graphify explain list_booths + grep over every
mkdir and every dotfile skip), then the contract, then a caller-side seam
review against the real sibling module surfaces.

The seam review earned its place again: the contract asserted that the
upload path's filename dedupe set must gain MANIFEST_FILE or an uploaded
file could collide with the manifest. safe_upload_name strips leading
dots, so that collision is unreachable — and the UPLOAD_MARKER entry
already sitting in that set has never been able to matter either. A
scope item the contract reasoned its way into and the sibling refutes.

Operator calls settled 2026-09-22: flags on the existing new/add verbs
rather than a second announce verb; unannounced booths get a quiet
marker rather than nothing; U5 adds provenance only and does not add a
second index ordering keyed on announcement time (parked for v1.1).

Cold contract panel dispatched to heid before this landed; its findings
fold in before any code ships.
2026-09-22 00:37:35 -07:00
Vuong Hoang 026a1fc392 fix(marks): v0.2.2 — nine findings from the cross-frontier bug-hunt panel
`/heid-bug-hunt` on U2's diff, four arms, artifact-only. Eight findings were
real against live code; a ninth was already closed by v0.2.1 and is recorded as
declined. Full triage in persistent-memory.d/2026-09-22-bug-hunt-panel.md.

THE LOCK LIFECYCLE (4/4 convergent, and two defects in one place)

`_Locked.__exit__` unlinked `.marks.lock` on the no-op path so a booth that had
never been marked was left exactly as it was found. `flock` binds to an INODE:
unlinking it under a blocked waiter leaves that waiter holding an exclusive
lock on a deleted file while the next writer creates a fresh lock and takes it
immediately. Two processes then run the read-modify-write concurrently, the
later os.replace drops the earlier one's mark, and both obeyed the protocol.

The cleanup existed to protect the booth's TTL, and was failing at that too:
creating or removing a directory entry bumps the DIRECTORY's mtime, which is
what `_newest_mtime` seeds from. The guard's comment reasons about the lock
file's own mtime and misses that the directory moved underneath it.

One fix: never unlink the lock, exempt `.<name>.lock` dotfiles from
`_newest_mtime`, and restore the directory's mtime after creating one.

THE READ PATH'S BLAST RADIUS

`_clean_text` did `(text or "").replace(...)` and `marks_for` sorts on
`(created, id)`, so a stored `text` that was a dict or a `created` that was a
number raised out of the read path. `list_booths` reads every booth's marks on
every index load, so one hand-edited file returned 500 for `/` and `/healthz`
across all 25 booths. Guarded in two layers — a named type check and a
`_hydrate_safe` backstop that cannot raise — and an unreadable mark now renders
as ⚠ broken rather than as an empty note.

ALSO

- import_legacy_asks stamped `created` at whole-second resolution, so two
  sidecars from the same second lost the ordering the importer had just
  established and re-sorted alphabetically. Microseconds, per the stated
  `(mtime, name)` rule.
- The five mark-write routes ran a blocking flock on the event loop; they now
  dispatch through run_in_threadpool, asserted structurally like INV-1.
- `/answer` 500'd on a non-string `notes` form value where `/note` handled it.
- The inline-doc tile had a flag control and no note field.
- The marks panel was suppressed on any booth carrying a links.md.
- The viewer's arrow keys and Escape threw away a note being typed.

CLI

`booth marks` printed a traceback and exited 0 on a failed read, and `--wait`
emitted a whole JSON document per poll. `booth answer --wait` read a damaged
file as "not yet" and spun the full hour. Both now use real exit codes —
0 ok, 1 unanswered/timed-out, 2 no such pick, 3 unreadable — and `--wait`
prints once. `marks.read_error()` lets the CLI ask what the page must not: the
browser stays lenient, the machine consumer gets the truth.

`scripts/booth` had no tests; it has five now, run against the real script
under the system python3, which also makes them a live check on INV-1.

275 tests (253 before). Live service restarted, 25/25 booth pages verified 200.
2026-09-22 00:20:58 -07:00
vh 70fb15886b memory: snapshot — U1 and U2 released at v0.2.1, U5 next
Records what this session learned that the code does not say on its own: the
read-lenient/write-strict asymmetry and why pointing both at one reader silently
collapses them; that the seam review and the cold contract panel had zero overlap
in BOTH directions on one unit, so neither substitutes for the other; that every
code-changing panel finding came from the ambiguity pass rather than the
paraphrase; and the timing lesson that a tag waits for an outstanding gate.

In-flight is set up for U5 with the two things already settled about it, so the
next session does not re-derive them: .booth.json is a dotfile and so is already
excluded by booth_items, and the deterministic-order invariant applies to whatever
it adds to the index card.

No version bump — memory snapshot, on the SemVer skip list.
2026-09-21 23:58:25 -07:00
vh a0448bdc24 fix(booth): list the deprecated asks alias in the usage string
Reported by draupnir. The v0.2.0 note told consumers the alias survives, and the
usage line is exactly where a session checks that claim — a deprecated-but-live
verb that is invisible at its own discovery surface reads as removed.
2026-09-21 23:56:10 -07:00
vh 5e41108cd3 fix(marks): a write over a damaged mark file was wiping the booth's judgment
Three defects and a missing test, all surfaced by the cross-frontier contract
panel dispatched before implementation and triaged after it (heid, four arms,
artifact-only, thread 01M33VSNFER4N1554G0Y0VC9C8). v0.2.0 was already tagged and
announced to fifteen handles when they landed, which is the argument for running
the gate at all.

DATA LOSS. `marks_for` is deliberately lenient — an unparseable `.marks.json`
reads as "no marks" so a review page still loads. The write path inherited that
leniency through the same reader, so one flag click appended a single entry to an
empty list and atomically replaced the file: every mark in the booth gone,
silently, from a click. Reproduced first, then fixed.

The fix is an asymmetry, not a retreat from leniency. Reads stay lenient; writes
go strict through `_read_raw_strict`, which distinguishes bytes-present-but-
unreadable from absent and valid-but-empty, and raises `MarksCorrupt`. The
damaged bytes are left on disk. Routes answer 409 rather than 500 — the service
is fine and the request was well-formed, the state on disk is not — and the body
says what to do, because the alternative the operator reaches for otherwise is
deleting the file, which is the thing being protected. The CLI says it in one
line instead of a traceback.

A PICK COULD NOT TARGET AN ITEM. `Mark.target` carried one, `marks_for_target`
retrieved by it, and the panel already rendered "on <item>" — but `declare_pick`
had no parameter for it, so no session could produce one. A question about one
artifact is the whole point of the 2026-09-09 inline-placement ruling; the door
was simply missing.

THE IMPORTER STRANDED AN ANSWER. A stem already present as a mark was skipped
wholesale. If a session had re-declared that stem through marks while the
operator's choice sat in the legacy sidecar, that choice was lost permanently —
reads are forbidden from looking at sidecars. The declaration is still skipped
(idempotence holds) but a legacy answer is now adopted when the existing mark is
an unanswered pick, and an answer made through marks is never overwritten.

INV-3 NAMED A SURFACE NOTHING TESTED. All four arms converged on it: the rule
protects gallery tile, zoom view and doc view; the falsifiable check covered one.
The doc view was implemented and untested, so shipping it unmarked would have
passed. Three tests now, one per surface.

The contract carries the full triage, including two findings accepted and NOT
closed: INV-2's and INV-5's checks comply in letter — openness can be re-derived
without spelling the grepped pattern, and importlib inside a function defeats the
AST walk. Both describe a future careless change, and the honest statement is
that these checks raise the cost of drifting rather than making it impossible.
Recorded rather than papered over.

Also pins the three prose ambiguities the panel found, normatively and once each:
what counts as open, the three distinct broken-declaration cases, and INV-6,
which had named a helper that does not exist and forbidden the calls that helper
must make.

253 tests.
2026-09-21 23:54:42 -07:00
64 changed files with 6418 additions and 321 deletions
+3 -2
View File
@@ -62,8 +62,9 @@ test is the only thing standing here.
No database. `ls ~/booth-data` tells you everything the service knows.
Per-booth operator state is a **dotfile inside the booth**: `.forever` (keep),
`.blurred` (one rel per line), `.marks.json` + `.marks.lock` (judgment), `.pins`
(link-board pin ids), `.uploaded` (upload-booth marker). `booth_items()` skips `name.startswith(".")`, so a new
`.viewed` (last deliberate look — U4's "viewing is activity"), `.blurred` (one
rel per line), `.marks.json` + `.marks.lock` (judgment), `.pins` (link-board pin
ids), `.uploaded` (upload-booth marker). `booth_items()` skips `name.startswith(".")`, so a new
dotfile costs nothing in item counts, galleries or zips. That skip is why the
dotfile is the right shape for new operator state — use it rather than
inventing a sidecar-per-item.
+86 -10
View File
@@ -15,16 +15,18 @@ filesystem *is* the state.
- **Live:** http://10.100.10.50:8090/ (nh3-dev) · linked from Homepage → *Apps → The Booth*
- **Data dir:** `~/booth-data/` on nh3-dev (one subfolder per booth)
- **TTL:** 24h, measured from the newest mtime in a booth's tree (it lives while
you're touching it, self-destructs 24h after you stop)
you're touching it, self-destructs 24h after you stop). Two things hold a booth
open past that: the `.forever` sentinel, and **an unanswered question** — see
*Lifetime* below. **Opening a booth page is activity**; polling it is not.
## How a session posts
A booth is **just a folder** under the data dir. Three ways, cheapest first:
```bash
# 1. On nh3-dev — the helper (services/booth/scripts/booth):
booth add my-run out/a.png out/b.png # creates booth + copies, prints URL
booth new my-run # empty booth, then cp/mv into ~/booth-data/my-run/
# 1. On nh3-dev — the helper (scripts/booth):
booth add my-run out/a.png out/b.png --why "pick the denoiser, v3 on the left"
booth new my-run --why "..." # empty booth, then cp/mv into ~/booth-data/my-run/
booth url my-run # just print the URL
booth ls # list booths
booth rm my-run # wipe now (TTL would anyway)
@@ -39,6 +41,25 @@ rsync -a ./out/ nh3-dev:booth-data/my-run/
Then hand the operator `http://10.100.10.50:8090/b/my-run/`.
### Say what it is — `--why`
**`--why` is one line telling the operator what he is looking at and why.** It
lands on the index card and on the booth page next to your handle (taken from
`$ALTHING_HANDLE`), stored as `.booth.json` in the booth.
It is optional and nothing breaks without it — a booth with no announcement
renders as `unannounced`, which is also what every booth created by `rsync` or
a bare `mkdir` looks like. But a booth that cannot say what it is has no way to
ask for attention except by posting its URL somewhere else, and that is exactly
how the link board ended up 69% dead rows. **The booth is the place to say it.**
```bash
booth add r18-ab out/*.png --why "which denoiser — v3 left, v4 right" --title "R18 A/B"
```
A second `new` or `add` on the same booth updates the why and keeps the
original creation stamp: the booth appeared once.
## Checking that controls can actually be clicked
```bash
@@ -140,7 +161,46 @@ is exactly why the direct `×` was worth adding.
an ephemeral booth could only be kept from a shell. The `/keep` route and the
CLI verb both already existed; only the button was missing.
## Kept boards — the one exception to the 24h rule
## Lifetime — derived, not declared
A booth is in exactly one of three states, and only the first is a button you
press:
| state | what puts it there | swept? |
|---|---|---|
| **kept** | you pressed `keep` / dropped `.forever` | never |
| **held** | an **unanswered pick**, or a `.marks.json` the service cannot read | not while that holds |
| **ephemeral** | everything else | 24h after the last activity |
**An open question holds its own booth.** A session that runs `booth ask` does
not also need to `keep` the booth — the booth cannot be swept while the operator
still owes it an answer, and it is released automatically when he answers. A
*partially* answered multi-question pick still counts as open, so a review in
flight is never swept out from under him. The index card and the booth header
say `held until answered` where the countdown would be, so a booth that has
stopped counting down always tells you why.
**Viewing is activity.** A deliberate GET of a booth's own page — the gallery, a
verbatim report, the zoom view, the marks page, a zip download — resets the
clock. If the operator is still looking at it, it is still alive. Browsing the
index does **not** count, and neither does a session polling `marks.json` or
`booth marks --wait`: machine reads are deliberately excluded, so an agent
cannot hold its own booth open by waiting on it.
**A held booth is still yours to delete.** The hold is protection from the
timer, never from you: `booth rm`, the UI ×, and `DELETE /b/<name>` all work
exactly as before. `sweep_once` is the only thing that honours a hold, exactly
as it is the only thing that honours `.forever`.
**Why this exists:** `.forever` used to be the only way to say three different
things — "this is durable", "I haven't answered yet", and "I'm still looking at
it" — and the measurement showed it carrying all three. On 2026-09-22, 17 of 24
live booths (70%) held the sentinel, up from 54% the day before; three of the
four booths in the fleet awaiting an answer had been pinned by hand as well.
Only the first meaning is what `keep` means. The other two the service already
knew and did not consult.
### Kept boards — the explicit pin
A booth containing a **`.forever`** dotfile is **never swept**, and renders in
its own **Kept** lane at the top of the index (blue top edge, `★ kept` badge, no
@@ -149,6 +209,7 @@ still ephemeral, so nobody inherits a cleanup chore they didn't ask for.
```bash
booth keep my-board # drop the sentinel — exempt from the sweep, forever
# (NOT for "waiting on an answer" — the pick holds it)
booth unkeep my-board # release the pin — the board rejoins the sweep
booth rm my-board # delete it NOW (works on kept boards; says so when it was kept)
@@ -258,6 +319,14 @@ declare_pick(booth, "batch", {
# "unanswered", "complete", "notes", "answered_at", "answered_by"}
```
**A pick can be about ONE item, not just the booth.** Pass `target` — an item's
booth-relative path — and the question renders beside that artifact:
```python
declare_pick(booth, "which-crop", {"prompt": "Which crop?", "options": ["tight", "wide"]},
target="v3/DSC03389.jpg")
```
**A partial answer is recorded, not refused.** A question left blank is a
deliberate outcome — "none of these", "not yet", "ask me later" — so it lands in
`unanswered`, stays absent from `answers` unless it carried a note, and
@@ -272,6 +341,14 @@ files (`<stem>.ask.json` / `<stem>.answer.json`) are imported, never deleted:
booth marks-import r18-ab # idempotent; the sidecars stay on disk
```
If the stem is already a mark the declaration is skipped, but a legacy answer
still gets adopted, so the operator's recorded choice is never stranded on disk.
**If a booth's `.marks.json` is damaged**, reads degrade to "no marks" so the page
still loads, and every WRITE refuses with a 409 rather than replacing the file —
which would otherwise wipe every judgment in that booth. Repair or move the file
by hand; nothing deletes it for you.
## Upload for pickup
The reverse direction — put files in through the web, pick them up by id:
@@ -403,11 +480,10 @@ wipe it from there. Release is reversible — press keep again and nothing was
lost. From the CLI, `booth rm <name>` deletes a kept board immediately and
tells you it was kept.
**Do not "unkeep and let it expire."** Removing the sentinel *bumps the booth
directory's mtime*, and a booth's age is the newest mtime in its tree — so a
released board's clock **resets** and it survives another full TTL.
Unkeep-and-wait is a 24-hour delay, not a delete. Use the × or `booth rm` when
you mean now.
**Do not "unkeep and let it expire."** **Releasing a board is activity** — you
just touched it — so a released board's clock **resets** and it survives another
full TTL. Unkeep-and-wait is a 24-hour delay, not a delete. Use the × or
`booth rm` when you mean now.
## Ops
+31 -5
View File
@@ -1,7 +1,7 @@
# The Booth — roadmap
Design: [`docs/design/information-architecture.md`](docs/design/information-architecture.md).
Current version: `0.2.0` (U1 + U2 landed; extracted from eshpfi 2026-09-21).
Current version: `0.4.0` (U1, U2, U4 and U5 landed; extracted from eshpfi 2026-09-21).
## v1 target
@@ -13,8 +13,8 @@ defect — not a wish. The measurements are in the IA doc.
| 1 | ~~**One item record**~~ — **landed `ce598b3`** | captions never reach the zoom view (never sent, not lost) | U1 |
| 2 | ~~**Marks**~~ — **landed `c7f9437`, released `v0.2.0`** | 5 mechanisms for 1 job; operator→session loop runs through chat | U2 |
| 3 | **Declared embed seam** — `/_booth/embed.js`, chrome mounts via DOM | 6 regexes injected into arbitrary author HTML, load-bearing for asks | U3 |
| 4 | **Derived lifetime** — open marks pin; viewing is activity | 54% of booths on the `.forever` escape hatch | U4 |
| 5 | **Self-announcing booths** — `.booth.json`, provenance on the index | job 5 had no home, so it lived on the link board as 145 dead rows | U5 |
| 4 | ~~**Derived lifetime**~~ — **landed `c3a97c1`, released `v0.4.0`** | 70% of booths on the `.forever` escape hatch (54% when first counted) | U4 |
| 5 | ~~**Self-announcing booths**~~ — **landed `c015a91`, released `v0.3.0`** | job 5 had no home, so it lived on the link board as 145 dead rows | U5 |
| 6 | **Benches** — registry, identity, enforced rule, migration | 69% link-board rot; the same bench posted 5× | U6 |
| 7 | **Navigation at 270 items** — sections, rail, filters, grid keyboard | one flat wall; subfolder structure discarded at render | U7 |
@@ -22,8 +22,27 @@ Ordering is dependency-driven, not priority-driven: **U1 → U2 → {U3, U4, U5}
U7**, with **U6 independent** of all of them (different storage, different
surface) and therefore the safest thing to land first or in parallel.
**U1 and U2 are landed**, which unblocks U3, U4 and U5 — all three read marks.
**U5 is next** (operator, 2026-09-21). U6 remains independent and unstarted.
**U1, U2, U4 and U5 are landed.** U3 is unblocked and unstarted; U6 remains
independent and unstarted; U7 waits on the rest.
**U5's adoption is a measured prediction, not a finished result**, and it is
TWO predictions rather than one. The operator declined a fleetwide announcement
so that adoption could be told apart from design; within fifty minutes of the
deploy a peer that had been told nothing (`comfy-dev`) created a booth and it
announced itself with a handle and an empty `why`. That is the split:
- **The handle rides for free.** It is written by `booth new` and `booth add`,
so every existing caller starts announcing without learning anything.
- **The `why` has to be learned.** It needs someone to know the flag exists.
Both get re-measured on **2026-09-29**:
find ~/booth-data -maxdepth 2 -name .booth.json | wc -l # free
grep -l '"why": "[^"]' ~/booth-data/*/.booth.json | wc -l # learned
A high first count with a near-zero second is the predicted shape of "nobody was
told" — an adoption failure fixed by announcing, which is a different thing from
nobody wanting it. Same instrument as U4's `.forever` prediction below.
### Cross-cutting invariant — deterministic order, everywhere
@@ -53,6 +72,13 @@ Where it already binds, and what the rule is in each case:
| marks in a booth | `(created, id)` — time, with the id as tie-break so two marks written in the same second cannot swap |
| legacy ask import | `(mtime, name)`, which is the order `list_asks` gave them |
| link board rows | pinned first, then newest-first |
| a booth's announcement | not a collection — one flat record per booth, nothing to order (U5) |
U4 added no ordered collection — a booth's lifetime is one state per booth,
not a sequence — so the rule above did not need a new row. The three lifetime
surfaces (index card, booth header, marks page) render through ONE macro
precisely so they cannot disagree, which is the same property stated for
ordering: one rule, one place, every surface reading it.
Where it is still to be decided, and must be before the unit ships: **U7's
section ordering and its compare pairing** (sections need a stated order among
+404 -44
View File
@@ -7,15 +7,26 @@ Model (deliberately dead-simple, no database):
* GET /b/<name>/ -> if <name>/index.html exists, serve it verbatim; otherwise
auto-render a gallery of the images / webm-videos / audio in it.
* GET /b/<name>/<file> -> serve a file out of the booth (also feeds a custom index.html's assets).
* 24h TTL: a background sweeper wipes any booth untouched for TTL hours. A booth's
age is measured from the *newest* mtime in its tree, so it lives while it's being
worked on and self-destructs TTL hours after the last activity.
* KEPT BOOTHS: a booth containing the KEEP_MARKER dotfile (`.forever`) is exempt
from the sweep and renders in its own lane above the ephemeral grid. That is the
home for durable operator-facing boards — chiefly the standing link board agent
sessions post to, whose whole purpose is to survive longer than the scrollback
it replaces. Opt-in per booth, so the ephemeral default is unchanged and nobody
inherits a cleanup chore; `rm` the sentinel and the booth rejoins the sweep.
* LIFETIME IS DERIVED, not set by a boolean (U4). Three states, and `sweep_once`
is the only thing that honours the first two:
KEPT `.forever` present. Never swept, own lane at the top of the index.
Durable operator-facing boards — chiefly the standing link board,
whose whole purpose is to outlive the scrollback it replaces.
HELD an open pick in `.marks.json`, or marks that cannot be read at all.
A booth the operator still owes an answer to is not the sweeper's
to take, and one whose judgment we failed to READ is certainly not.
EPHEMERAL everything else: wiped TTL hours after the last activity. Age is the
*newest* mtime in the tree, so a booth lives while it is being
worked on and self-destructs once it stops.
* VIEWING IS ACTIVITY. A deliberate GET of a booth's own page writes VIEW_MARKER,
which the age rule already counts — if the operator is still looking at it, it
is still alive. Browsing the index is not a view, and neither is a session
polling `marks.json`: an agent must not be able to hold its own booth open.
* WHY DERIVED. `.forever` was the ONLY way to say three different things, and the
measurement showed it carrying all of them — 17 of 24 live booths on 2026-09-22
(70%, up from 54%), with three of the four booths awaiting an answer ALSO pinned
by hand. Only "this is durable" is what keep means. The other two are facts the
service already held and did not consult.
State is the filesystem — `ls ~/booth-data` tells you everything. That is the whole point.
"""
@@ -34,6 +45,7 @@ import time
import zipfile
from contextlib import asynccontextmanager
from pathlib import Path
from typing import Sequence
from urllib.parse import quote, unquote
from fastapi import FastAPI, File, Form, HTTPException, Request, UploadFile
@@ -45,6 +57,7 @@ from fastapi.responses import (
Response,
)
from fastapi.templating import Jinja2Templates
from starlette.concurrency import run_in_threadpool
from jinja2 import Environment, FileSystemLoader, select_autoescape
try:
@@ -84,6 +97,14 @@ from booth.items import ( # noqa: E402,F401
# flag to remember, no state anywhere but the filesystem.
KEEP_MARKER = ".forever"
# Records the last deliberate look at a booth (U4). A dotfile for the same two
# reasons KEEP_MARKER is one — `booth_items` and `zip_booth` skip it, so it
# costs nothing in counts, galleries or zips — and NOT a `.lock` dotfile, so
# `_newest_mtime` COUNTS it and the existing age rule picks the view up with no
# new arithmetic. That is the whole integration: a view is one more thing in
# the tree, not a second term in the formula.
VIEW_MARKER = ".viewed"
# ⚠⚠ BLUR IS COSMETIC, NOT ACCESS CONTROL. The file is still served at its own
# URL, still in the zip, still on disk. This hides an item from a glance — a
# shoulder, a screen-share, a scroll past something you did not want to see
@@ -123,11 +144,14 @@ from booth.asks import ( # noqa: E402
)
from booth.marks import ( # noqa: E402
MARKS_FILE,
Mark,
MarksCorrupt,
answer_pick,
as_dict,
declare_pick,
delete_mark,
import_legacy_asks,
hold_read,
marks_for,
marks_for_target,
open_marks,
@@ -139,6 +163,12 @@ from booth.inline import ( # noqa: E402
has_placeholders,
place as place_asks,
)
from booth.manifest import ( # noqa: E402
MANIFEST_FILE,
SERVICE_HANDLE,
read_manifest,
write_manifest,
)
from booth.links import ( # noqa: E402
LINK_LOCK,
LINKS_FILE,
@@ -168,16 +198,47 @@ def human_dur(seconds: float) -> str:
def _newest_mtime(path: Path) -> float:
"""Newest mtime among a folder and everything under it."""
"""Newest mtime among a folder and everything under it — OUR LOCKS EXCEPT.
A booth's age is how long since somebody touched it, and a lock sidecar is
machinery: `marks.py` and `links.py` each create one on the way into a
read-modify-write, including one that turns out to change nothing. Counting
it made reading-through-a-write-path look like activity, and a no-op mark
POST on a dead booth reset its clock.
The exclusion is `.<something>.lock` — a DOTfile, which is the Booth's own
namespace. An agent that posts a real artifact called `build.lock` still
gets its clock counted. Everything else counts too, dotfiles included,
because `.marks.json`, `.blurred` and `.pins` are the operator doing
something.
⚠ A STAT WE CANNOT DO READS AS *FRESH*, NEVER AS EPOCH-OLD. This function
feeds `is_expired`, which feeds `rmtree`. Returning 0.0 for a booth whose
own stat fails made it maximally ancient and therefore the FIRST thing the
sweeper takes — a permissions or ELOOP problem resolving to a deletion. The
bug-hunt panel found this as one of four paths into the same shape. Not
knowing a booth's age is a reason to leave it alone.
`FileNotFoundError` on an entry is the exception, and it stays a skip: a
dangling symlink and a file removed mid-scan both raise it, and neither is
a thing with an mtime worth counting. Any OTHER per-entry OSError means we
could not read something that IS there, so the age is unknowable and the
booth reads as fresh.
"""
now = time.time()
try:
newest = path.stat().st_mtime
except OSError:
return 0.0
return now
for p in path.rglob("*"):
if p.name.startswith(".") and p.name.endswith(".lock"):
continue
try:
m = p.stat().st_mtime
except FileNotFoundError:
continue # dangling symlink, or gone mid-scan
except OSError:
continue
return now # cannot read it — cannot judge the age
if m > newest:
newest = m
return newest
@@ -199,8 +260,102 @@ def is_expired(path: Path, ttl_seconds: float, now: float | None = None) -> bool
def is_kept(path: Path) -> bool:
"""True if this booth carries the keep sentinel and must never be swept."""
return (path / KEEP_MARKER).exists()
"""True if this booth carries the keep sentinel and must never be swept.
`lstat`, not `Path.exists()`, and an unreadable answer counts as KEPT. The
old form collapsed ELOOP and EACCES into False, so a kept booth whose
sentinel could not be stat'd became eligible for the sweep — a failed read
authorizing a delete, which is the shape the bug-hunt panel found four ways
into. `lstat` also means a `.forever` SYMLINK counts, dangling or not:
somebody put it there to mean keep.
"""
try:
(path / KEEP_MARKER).lstat()
return True
except FileNotFoundError:
return False
except OSError:
return True
def record_view(booth: Path) -> None:
"""Note that somebody deliberately looked at this booth (U4).
Touches VIEW_MARKER and lets `_newest_mtime` do the rest — a view enters
the age rule as a file in the tree, not as a new term in the arithmetic.
NEVER RAISES. A read-only mount, a booth owned by another uid, a full disk,
a booth deleted between the route's resolve and this call: every one of
those costs the timestamp, not the page. The same trade `_Locked.__enter__`
makes on its `os.utime`, and for the same stated reason — not recording the
look is a cost this service can absorb, not answering the request is not.
A booth whose view cannot be recorded simply ages on its content mtime,
which is what every booth did before this existed.
"""
# O_NOFOLLOW, not `Path.touch()`. `touch` on an existing symlink follows it,
# so a booth carrying a planted `.viewed -> /anywhere` turned EVERY page
# view into an mtime write at an arbitrary path under the service uid — and
# any fleet session can write into a booth, because making a folder is the
# whole API. Three of four bug-hunt arms found it independently. A symlink
# here now raises ELOOP into the swallow below: view-recording quietly stops
# for that booth, which is the right way to lose this argument.
#
# O_CREAT alone does not move the mtime of a file that already exists, so
# the utime is not decoration: the marker must read as NOW or the whole
# mechanism is a file nobody's clock looks at.
try:
fd = os.open(booth / VIEW_MARKER,
os.O_WRONLY | os.O_CREAT | os.O_NOFOLLOW, 0o644)
try:
os.utime(fd)
finally:
os.close(fd)
except OSError:
pass
HOLD_UNREADABLE = "unreadable"
HOLD_OPEN = "open"
def hold_reason(marks: Sequence[Mark], error: str | None) -> str | None:
"""WHY this booth must not be swept, or None if it may be. THE hold predicate.
Returns a reason rather than a bool so the surface that has to say why can
read it off the same value the sweeper acts on. A boolean plus a separate
error string is two representations of one state, and they drift.
PURE — it takes the result of a read and does none of its own, so the index
card and the sweeper cannot answer differently about the same booth. That is
U1's rule (one resolver, every surface reads the record) applied to lifetime.
FAIL-SAFE ON BOTH LEVELS OF DAMAGE, which is the correction the bug-hunt
panel forced (2026-09-22). `marks_for` is lenient because a review page that
will not load is worse than one missing an annotation — the right trade for
a RENDER and the wrong one for a DELETE, where the same leniency wipes the
booth whose judgment we had just failed to read, artifacts and all. The
first cut of this caught FILE-level damage only:
* file-level — `.marks.json` will not parse at all. `hold_read` reports it.
* ENTRY-level — the document parses, but one mark fails normalization and
`_hydrate_safe` hands back a `Mark` carrying `error`. `_is_open` returns
False for an errored pick, ON PURPOSE (a broken pick can never be
answered; the CLI spells that exit code 4) — so such a booth read as
`not held` and SWEPT, while the panel beside it rendered the broken mark
in full. Four arms found four ways into that shape; this was the worst.
A mark that cannot be read is judgment we cannot see. Deleting the booth it
belongs to is the one thing we must not do with it.
Openness itself is `open_marks` and nothing else (U2 INV-2): a partially
answered pick is STILL open and still holds, which is the reading that
makes this rule correct rather than one that sweeps a review in flight.
"""
if error is not None or any(m.error is not None for m in marks):
return HOLD_UNREADABLE
if open_marks(marks):
return HOLD_OPEN
return None
def sweep_once(data_dir: Path, ttl_seconds: float, now: float | None = None) -> list[str]:
@@ -209,10 +364,25 @@ def sweep_once(data_dir: Path, ttl_seconds: float, now: float | None = None) ->
Only ever removes direct children of data_dir (never data_dir itself), and
skips dotfolders so a stray control dir can opt out.
TWO exemptions, and this is the only function that honours either.
A booth carrying KEEP_MARKER is exempt no matter how stale it is. That is
the one escape hatch from the 24h contract, and it is opt-in per booth: the
default stays ephemeral, so nobody inherits a cleanup chore they did not ask
for. Removing the sentinel hands the booth straight back to the sweeper.
the explicit escape hatch, opt-in per booth: the default stays ephemeral, so
nobody inherits a cleanup chore they did not ask for. Removing the sentinel
hands the booth straight back to the sweeper.
A booth that is HELD — an open pick, or marks we cannot read — is exempt for
as long as that holds (U4). This is the derived half: the operator was
pressing `.forever` to mean "not yet" because nothing else could say it, and
the service already knew. 17 of 24 live booths carried the sentinel on
2026-09-22, and three of the four booths in the fleet awaiting an answer
carried it too — the "not yet" case, caught in the act.
Reading the marks costs ONE strict read per booth per tick — `hold_read`,
which answers both halves of the hold question at once. It is deliberately
not two calls: two reads of one file are not one read of one state, and the
pair that loses that race is the pair that deletes. Do not "optimize" this
back into `marks_for` plus `read_error`.
"""
wiped: list[str] = []
if not data_dir.is_dir():
@@ -223,6 +393,8 @@ def sweep_once(data_dir: Path, ttl_seconds: float, now: float | None = None) ->
try:
if is_kept(child):
continue
if hold_reason(*hold_read(child)): # ONE read — see hold_read
continue
if is_expired(child, ttl_seconds, now):
shutil.rmtree(child)
wiped.append(child.name)
@@ -255,7 +427,25 @@ def list_booths(data_dir: Path, ttl_seconds: float, now: float | None = None) ->
# flag a booth that is waiting on the operator. ONE file read per booth
# — which is why marks live in one file per booth rather than a sidecar
# per mark. This loop runs on every index page load.
marks = marks_for(child)
# ONE read for BOTH the badge and the lifetime decision. It has to be
# one: `marks_for` is lenient, so an unreadable `.marks.json` reads as
# no marks — fine for a card, wrong for the reaper, which would then
# delete the booth whose judgment it had just failed to read. And
# asking the two questions with two reads is not one read of one state:
# a write landing between them yields `([], None)`, the pair that
# deletes. `hold_read` answers both from one read; the lenient reader
# comes back only on the error path, where leniency is the point.
held_marks, read_err = hold_read(child)
# The DECISION comes from that one read and nothing else. The lenient
# re-read below is for DISPLAY only — feeding it back into the predicate
# would rebuild the two-read seam this call exists to close.
hold = hold_reason(held_marks, read_err)
marks = held_marks if read_err is None else marks_for(child)
# The booth's own announcement — who posted it and why. One more small
# read per booth, beside the marks read already here, and `read_manifest`
# cannot raise for the same reason `marks_for` must not: this loop runs
# over EVERY booth on every index page load.
manifest = read_manifest(child)
kinds = {"image": 0, "video": 0, "audio": 0, "other": 0}
thumb_url = None
thumb_blurred = False
@@ -273,6 +463,7 @@ def list_booths(data_dir: Path, ttl_seconds: float, now: float | None = None) ->
{
"name": child.name,
"name_url": quote(child.name, safe=""),
"manifest": manifest,
"count": len(items),
"kinds": kinds,
"thumb_url": thumb_url,
@@ -285,11 +476,23 @@ def list_booths(data_dir: Path, ttl_seconds: float, now: float | None = None) ->
# tested `answer is None`, so a half-answered pick read as closed
# here while the panel beside it rendered `◐ partial`.
"marks_open": len(open_marks(marks)),
# U4: WHY this booth is or is not counting down. The card must
# never just stop the clock silently — `.forever` was at least
# visible as a lane, and an invisible rule would be worse than
# the boolean it replaces.
# WHY it is or is not counting down — the reason, not a bool
# beside a string that can disagree with it.
"hold": hold,
"expires_in": max(0.0, ttl_seconds - (now - mtime)),
"mtime": mtime,
}
)
booths.sort(key=lambda b: b["mtime"], reverse=True)
# Newest first, NAME as the tie-break. Sorting on mtime alone left equal-mtime
# booths ordered by whatever `iterdir()` yielded, which is not a rule — and
# invariant 6 is not "usually stable", it is a sentence you can write down.
# Two booths created by one `rsync` batch share an mtime exactly, and the
# operator refers to cards positionally.
booths.sort(key=lambda b: (b["mtime"], b["name"]), reverse=True)
return booths
@@ -477,6 +680,18 @@ PICKUP_WORDS = (
).split()
def _form_text(form, key: str) -> str:
"""One form field as text, or "" for anything that is not text.
A multipart FILE part named `notes` parses to an UploadFile, not a str, and
every downstream cleaner calls `.replace` on what it is handed. Coercing
here keeps that decision in one place instead of one `isinstance` per call
site — which is how `/note` came to have the guard and `/answer` not to.
"""
value = form.get(key)
return value if isinstance(value, str) else ""
def safe_upload_name(name: str, fallback: str) -> str:
"""Reduce a client-supplied filename to a safe basename (no path, no hidden)."""
base = (name or "").replace("\\", "/").split("/")[-1].strip()
@@ -578,6 +793,25 @@ def create_app(
# test needs a handle on the env that the app actually renders with.
app.state.templates = templates
@app.exception_handler(MarksCorrupt)
async def _marks_corrupt(request: Request, exc: MarksCorrupt):
"""A write was refused because the booth's mark file is damaged.
409, not 500: the service is fine and the request was well-formed — the
state on disk is not, and the refusal is deliberate. Says what to do,
because the alternative the operator will otherwise reach for is
deleting the file, which is the thing being protected.
"""
return JSONResponse(
status_code=409,
content={
"error": "this booth's .marks.json cannot be read, so nothing was written",
"detail": str(exc),
"why": "writing would replace every mark in the booth with just this one",
"fix": "repair or move the file by hand; the marks panel still renders as empty",
},
)
ttl_display = int(ttl_hours) if float(ttl_hours).is_integer() else ttl_hours
base_ctx = {
"ttl_hours": ttl_display,
@@ -629,6 +863,10 @@ def create_app(
@app.get("/b/{name}/", response_class=HTMLResponse)
def booth_view(request: Request, name: str, download: int = 0):
booth = resolve_booth(name)
# U4: viewing is activity. ABOVE both early returns — the zip download
# and the verbatim-index.html branch are looks at this booth too, and a
# verbatim report is the shape the operator stares at longest.
record_view(booth)
if download:
# whole-booth zip — the download path for a verbatim index.html booth
# (which has no gallery/per-file chrome), and a "download all" for any.
@@ -661,7 +899,9 @@ def create_app(
it for it in build_gallery(booth)
if not ((booth / LINKS_FILE).is_file() and it["name"] == LINKS_FILE)
]
marks = marks_for(booth)
held_marks, read_err = hold_read(booth) # ONE read; see list_booths
hold = hold_reason(held_marks, read_err)
marks = held_marks if read_err is None else marks_for(booth)
return templates.TemplateResponse(
request,
"booth.html",
@@ -679,13 +919,13 @@ def create_app(
# Ordered pinned-first then newest-first, each row stamped with a
# `pinned` flag. Empty list for every other booth, so the template
# branch simply does not fire.
"board": (
order_for_display(
parse_link_entries((booth / LINKS_FILE).read_text()),
read_pins(booth),
)
if (booth / LINKS_FILE).is_file() else []
),
# `is_file()` then an UNGUARDED read was a 500 waiting on a
# mode change or an EIO: the board is one tile on this page, and
# a page that will not load is worse than one missing a tile —
# the same posture `read_blurred`, `marks_for` and
# `read_manifest` already take. A booth whose `links.md` cannot
# be read renders as a booth with no board.
"board": _board_rows(booth),
# Marks: operator judgment attached to this booth or to one of
# its items — a session's question (`pick`), the operator's own
# remark (`note`), the operator's selection (`flag`). Rendered
@@ -700,10 +940,33 @@ def create_app(
},
"booth_marks": marks_for_target(marks, None),
"uploaded": (booth / UPLOAD_MARKER).exists(),
# The same provenance line the index card carries. Deliberate:
# a booth URL handed to the operator lands HERE, never on the
# index, and job 5 is "operator, look at this".
"manifest": read_manifest(booth),
# The lifetime line, same three states as the index card: a
# booth URL handed to the operator lands HERE, not on the index,
# so "why is this not counting down" has to be answerable here.
"hold": hold,
"expires_in": max(0.0, ttl_seconds - booth_age_seconds(booth)),
},
)
def _board_rows(booth: Path) -> list[dict]:
"""The link board's rows, or [] for a board that cannot be read.
NEVER RAISES, for the reason every other read on this page does not:
one damaged file must cost its own tile, not the booth page."""
try:
if not (booth / LINKS_FILE).is_file():
return []
return order_for_display(
parse_link_entries((booth / LINKS_FILE).read_text()),
read_pins(booth),
)
except (OSError, ValueError, UnicodeDecodeError):
return []
def _mark_redirect(name: str, form, anchor: str) -> RedirectResponse:
"""Land where the form was: the standalone marks page for a verbatim
booth (its own index.html cannot show the recorded judgment), else the
@@ -740,13 +1003,27 @@ def create_app(
if spec.error is not None:
raise HTTPException(status_code=400, detail=spec.error)
who = request.client.host if request.client else ""
# `notes` is whatever the form parser yielded. A multipart FILE part
# named `notes` is an UploadFile, and `_clean_notes` calls `.replace` on
# it — a 500 on hostile-but-legal input, where the sibling `/note` route
# returns 400 for exactly the same class of value. Same parser, same
# question, one answer.
notes = _form_text(form, "notes")
try:
if spec.multi:
choice = {q["key"]: form.get(f"choice.{q['key']}") for q in spec.questions}
qnotes = {q["key"]: form.get(f"notes.{q['key']}") for q in spec.questions}
answer_pick(booth, mark_id, choice, form.get("notes", ""), who=who, qnotes=qnotes)
choice = {q["key"]: _form_text(form, f"choice.{q['key']}")
for q in spec.questions}
qnotes = {q["key"]: _form_text(form, f"notes.{q['key']}")
for q in spec.questions}
await run_in_threadpool(answer_pick, booth, mark_id, choice, notes,
who=who, qnotes=qnotes)
else:
answer_pick(booth, mark_id, form.get("choice"), form.get("notes", ""), who=who)
# `choice` through the same reader as `notes`. It was raw, so a
# multipart FILE part named `choice` reached the answer builder
# as an UploadFile — the asymmetry that had already been fixed
# once on the field beside it.
await run_in_threadpool(answer_pick, booth, mark_id,
_form_text(form, "choice"), notes, who=who)
except AskError as exc:
raise HTTPException(status_code=400, detail=str(exc))
return _mark_redirect(name, form, f"mark-{quote(mark_id, safe='')}")
@@ -765,8 +1042,9 @@ def create_app(
target = raw_target if isinstance(raw_target, str) and raw_target else None
text = form.get("text")
try:
mark = write_note(booth, target, text if isinstance(text, str) else "",
who=request.client.host if request.client else "")
mark = await run_in_threadpool(
write_note, booth, target, text if isinstance(text, str) else "",
who=request.client.host if request.client else "")
except AskError as exc:
raise HTTPException(status_code=400, detail=str(exc))
return _mark_redirect(name, form, f"mark-{quote(mark.id, safe='')}")
@@ -786,7 +1064,9 @@ def create_app(
raise HTTPException(status_code=400, detail="a flag needs a target")
on = str(form.get("on", "1")) not in ("0", "", "false", "off")
try:
set_flag(booth, target, on, who=request.client.host if request.client else "")
await run_in_threadpool(
set_flag, booth, target, on,
who=request.client.host if request.client else "")
except AskError as exc:
raise HTTPException(status_code=400, detail=str(exc))
return _mark_redirect(name, form, f"item-{quote(target, safe='')}")
@@ -800,7 +1080,7 @@ def create_app(
mark_id = form.get("mark")
if not isinstance(mark_id, str) or not mark_id:
raise HTTPException(status_code=400, detail="which mark?")
delete_mark(booth, mark_id)
await run_in_threadpool(delete_mark, booth, mark_id)
return _mark_redirect(name, form, "marks")
@app.post("/b/{name}/import-asks")
@@ -812,7 +1092,7 @@ def create_app(
migrated from the page you are already looking at.
"""
booth = resolve_booth(name)
import_legacy_asks(booth)
await run_in_threadpool(import_legacy_asks, booth)
form = await request.form()
return _mark_redirect(name, form, "marks")
@@ -892,13 +1172,29 @@ def create_app(
place a verbatim-index.html booth can show its marks — that page is served
untouched by design, so the inline panel never renders there."""
booth = resolve_booth(name)
marks = marks_for(booth)
# U4: for a verbatim booth this IS the booth page. `/b/<n>/asks` is a
# 308 into here, so the legacy URL records through this call and must
# not get one of its own.
record_view(booth)
held_marks, read_err = hold_read(booth) # ONE read; see list_booths
hold = hold_reason(held_marks, read_err)
marks = held_marks if read_err is None else marks_for(booth)
return templates.TemplateResponse(
request,
"marks.html",
{**base_ctx, "name": name, "name_url": quote(name, safe=""),
"marks": marks, "marks_open": len(open_marks(marks)),
"booth_marks": marks_for_target(marks, None), "marks_page": True},
"booth_marks": marks_for_target(marks, None), "marks_page": True,
# U4 INV-4, and this page is WHY the invariant needs a third home.
# A verbatim booth's own index.html is served untouched, so it has
# no Booth-rendered header to carry the lifetime line — this page
# is the only surface besides the index card where the Booth owns
# the chrome. Without it, the booths most likely to be held (a
# report that ASKS something is the archetype) would be the ones
# that never say they are.
"kept": is_kept(booth),
"hold": hold,
"expires_in": max(0.0, ttl_seconds - booth_age_seconds(booth))},
)
@app.get("/b/{name}/asks", include_in_schema=False)
@@ -921,12 +1217,29 @@ def create_app(
the whole booth instead of one question at a time.
"""
booth = resolve_booth(name)
marks = marks_for(booth)
return JSONResponse({
marks, read_err = hold_read(booth)
body = {
"booth": name,
"marks": [as_dict(m) for m in marks],
"open": [m.id for m in open_marks(marks)],
})
}
# A DAMAGED file used to come back as an empty list and nothing else,
# which is indistinguishable from "you were never asked anything" — and
# this endpoint is the ONLY reader a remote session has. Its filesystem
# sibling has told the truth since U2: `booth marks` exits 3 on an
# unreadable file precisely so a caller can tell "not yet" from
# "broken". One question, two surfaces, two answers.
#
# The STATUS stays 200 and that is deliberate. Reads are lenient here —
# the same rule that keeps a poisoned booth from 500ing the index — and
# a pinned status code is a promise to remote clients this fix has no
# business breaking. The information goes in the body instead: a client
# that wants the CLI's exit-3 parity reads `error`, and one that does
# not behaves exactly as it does today.
if read_err is not None:
body["error"] = "this booth's .marks.json cannot be read"
body["detail"] = read_err
return JSONResponse(body)
@app.get("/b/{name}/view", response_class=HTMLResponse)
def booth_view_file(request: Request, name: str, f: str):
@@ -948,6 +1261,14 @@ def create_app(
items = booth_items(booth)
item = find_item(items, f)
# U4: a bookmarked zoom URL is somebody looking — but only once we know
# there is an ITEM to look at. Below the 404s, and gated on the record,
# because `f` is any path that stats inside the booth: the bug-hunt
# panel pointed `?f=.marks.lock` at this and held a booth open with a
# file the service created itself. A dotfile is not an item, and a view
# of a thing that is not an item is not a view of the booth.
if item is not None:
record_view(booth)
marks = marks_for(booth)
item_marks = marks_for_target(marks, f)
common = {
@@ -1025,11 +1346,22 @@ def create_app(
booth_id = generate_pickup_id(lambda n: (data_dir / n).exists())
dest = data_dir / booth_id
dest.mkdir(parents=True)
(dest / UPLOAD_MARKER).write_text("") # stamp as an upload (dotfile, not listed)
total = 0
used: set = {UPLOAD_MARKER}
# Both markers are belt-and-braces: `safe_upload_name` strips leading
# dots, so an uploaded file can never be named either of them. Listed
# anyway so the set says what the directory already contains.
used: set = {UPLOAD_MARKER, MANIFEST_FILE}
try:
(dest / UPLOAD_MARKER).write_text("") # dotfile, not listed
# A booth the SERVICE made says so, rather than being exempted from
# the unannounced marker. INSIDE the guard, with the marker: both
# sat above it, so a failure here left a half-booth on disk with no
# files in it — and the manifest's unique temp name meant a leaked
# `.booth.json.<hex>.tmp` was never overwritten, was not a `.lock`,
# and so kept that empty booth alive past every sweep. Found 4/4.
write_manifest(dest, SERVICE_HANDLE, title=booth_id,
why="browser upload, for pickup")
for i, f in enumerate(files):
name = _dedupe_name(safe_upload_name(f.filename, f"file-{i + 1}"), used)
used.add(name)
@@ -1118,7 +1450,35 @@ def create_app(
@app.post("/b/{name}/unkeep")
def booth_unkeep(name: str, next: str = Form("/")):
# missing_ok: releasing an already-released board is a no-op, not a 500.
(resolve_booth(name) / KEEP_MARKER).unlink(missing_ok=True)
booth = resolve_booth(name)
marker = booth / KEEP_MARKER
try:
marker.unlink()
released = True
except FileNotFoundError:
released = False # already released: a no-op, not a 500
except OSError:
# A `.forever` that is a DIRECTORY raised IsADirectoryError straight
# through this route and 500'd it, which made the card's release
# button permanently dead for that booth. Pre-existing; the panel
# re-exposed it. Removing it is still best-effort, and failing to is
# not worth refusing the request over.
released = False
if not released:
return RedirectResponse(url=_safe_next(next), status_code=303)
# U4: RELEASE IS ACTIVITY, and now it is a rule rather than an accident.
# A released board already survived another full TTL, because unlinking
# a file bumps the directory's mtime — behaviour the note above calls
# "not intuitive" precisely because nothing declared it. The behaviour
# is unchanged; its reason is now stated. Releasing a board is somebody
# touching it, so it gets one full TTL, the same as any other look.
#
# ONLY when something was actually released, which is the correction the
# bug-hunt panel forced: an unconditional call made POSTing release at
# an already-released booth an endless TTL refresh, contradicting this
# route's own no-op promise and diverging from the CLI, which `rm`s the
# sentinel without recording anything.
record_view(booth)
return RedirectResponse(url=_safe_next(next), status_code=303)
@app.post("/b/{name}/blur")
+274
View File
@@ -0,0 +1,274 @@
"""A booth's own announcement — who posted it, and why.
U5. The index card used to show a name, an item count and a countdown, and
nothing the poster chose. An agent with something to show therefore had no way
to make the booth say "look at this" and posted a URL to the link board
instead — which is why 145 of that board's 210 rows (69%) ended up pointing at
booths that had already been swept. The board was absorbing a job it was never
shaped for. This is the shape.
.booth.json -> {"handle": ..., "title": ..., "why": ..., "created": ...}
⚠ STDLIB ONLY, and it imports nothing from `booth.*` either.
`scripts/booth` — the CLI every fleet session uses — imports this module
directly under the system `python3` with no venv, through a `python3 -c`
heredoc no AST extractor can see. A single third-party import here breaks
`booth new` and `booth add` on every host, and the failure surfaces in an
agent's session rather than in ours. The ban extends to sibling `booth` modules:
importing `marks` to reuse its atomic write would drag marks' own import list
into this one's, so the four-line pattern is copied instead. `test_stdlib_only`
in tests/test_manifest.py is the only thing standing here.
Contract: docs/contracts/u5_booth_manifest.contract.md.
"""
from __future__ import annotations
import json
import os
import secrets
import stat as statmod
from dataclasses import dataclass
from datetime import datetime
from pathlib import Path
MANIFEST_FILE = ".booth.json"
# A `why` renders inside a card's sub-line, so it is one line by construction
# rather than by convention — enforced at the WRITE so nothing downstream has to
# remember. The caps are display budgets, not storage limits.
HANDLE_MAX = 64
TITLE_MAX = 120
WHY_MAX = 200
CREATED_MAX = 64
# A manifest is four short fields. Anything near this is not one, and reading it
# into memory to find that out is the wrong order of operations: `list_booths`
# calls the reader once per booth on every index load, so an unbounded read is
# the service-wide outage the lenient reader exists to prevent, arriving in a
# different costume. Checked by `stat`, before the bytes are touched.
MANIFEST_MAX_BYTES = 64 * 1024
# Where bytes that could not be read go when a re-announcement replaces them.
# ONE fixed name, deliberately: a timestamped quarantine accumulates forever in
# a folder nothing prunes, and the most recent damage is the only copy anybody
# would look at. A dotfile, so it is invisible to every listing and zip.
QUARANTINE_FILE = ".booth.json.broken"
# The handle a booth created by the service itself carries. A pickup booth and
# the standing link board are made by the Booth, not by an agent, and saying so
# is true rather than manufactured — which is the whole reason there is no
# exemption list. One rule: a booth with no manifest is unannounced.
SERVICE_HANDLE = "booth"
@dataclass(frozen=True)
class Manifest:
"""One booth's announcement.
`handle` is an althing agent handle, or `SERVICE_HANDLE` for a booth the
Booth made. `error` is a read-time verdict and is never stored.
"""
handle: str
title: str
why: str
created: str
error: str | None = None
def _one_line(value, limit: int) -> str:
"""One line, bounded. Collapses ALL runs of whitespace, not only newlines —
a tab or a forty-space indent in a `why` renders as badly inside a card's
sub-line as a newline does, and the field is one line by construction."""
if not isinstance(value, str):
return ""
return " ".join(value.split())[:limit]
def _temp_path(booth: Path) -> Path:
"""A scratch name no other writer will pick.
Every writer used to derive the same `.booth.json.tmp`, so two `booth add`
calls on one booth could interleave through a stale descriptor into the
published path. Marks are protected from that by their flock; the manifest
deliberately has none — it is written once at creation, not read-modify-
written per click — so uniqueness is what stands in for the lock. Still a
dotfile, so no listing, gallery or zip can see it mid-write.
"""
return booth / f"{MANIFEST_FILE}.{secrets.token_hex(4)}.tmp"
def _as_doc(m: "Manifest") -> dict:
"""The stored shape of a record, for the no-op comparison."""
return {"handle": m.handle, "title": m.title, "why": m.why, "created": m.created}
def _now() -> str:
return datetime.now().astimezone().isoformat(timespec="seconds")
def read_manifest(booth: Path) -> Manifest | None:
"""This booth's announcement, or None if it never made one.
LENIENT, AND IT NEVER RAISES (INV-2). `list_booths` calls this once per
booth on every index page load, so a read that can raise is a service-wide
outage wearing a single-booth bug's clothes. That is not hypothetical: a
poisoned `.marks.json` did exactly that to `/` and `/healthz` across all 25
live booths, and the fix shipped in v0.2.2. Same posture, applied before the
same mistake rather than after it.
Absent -> None. Present but unreadable -> a Manifest carrying `error`, so a
card can say `unreadable` instead of quietly showing the same thing as a
booth that never announced (INV-5). Folding the two together would hide the
one case somebody has to go and fix.
Only `handle` is required. A hand-written manifest is a supported input —
the file is plain JSON in a folder the operator owns, and half the point of
the Booth is that a booth is just a directory.
"""
booth = Path(booth)
path = booth / MANIFEST_FILE
# BOUNDED BEFORE THE READ. "Never raises" was not true of an unbounded one:
# a 4 GB file raises MemoryError and a deeply nested document raises
# RecursionError out of `json.loads`, and neither is an OSError or a
# ValueError. Both escape into `list_booths`, which calls this per booth on
# every index load — so one file returns 500 for the whole front page. Size
# first, by `stat`; then catch the two classes anyway, because a bound that
# is one day raised should not quietly re-open the hole.
try:
st = path.stat()
except FileNotFoundError:
return None
except OSError as exc:
return _broken(booth, f"cannot be read: {exc}")
# ⚠ REGULAR-FILE FIRST, then size. `st_size` answers a different question
# than "can this be read": it is 0 for a FIFO and 0 for /dev/zero, so both
# sail under the cap, and then `read_text` either blocks forever with no EOF
# or allocates until the kernel intervenes. The bound ABOVE is what made
# this reachable — a cap that trusts st_size inherits everything st_size
# does not mean. One such file stalls every `GET /` and `/healthz`.
if not statmod.S_ISREG(st.st_mode):
return _broken(booth, "is not a regular file")
if st.st_size > MANIFEST_MAX_BYTES:
return _broken(booth, f"is too large to be a manifest ({st.st_size} bytes)")
try:
text = path.read_text(encoding="utf-8")
except FileNotFoundError:
return None
except (OSError, UnicodeDecodeError, MemoryError) as exc:
return _broken(booth, f"cannot be read: {exc}")
if not text.strip():
return _broken(booth, "is empty")
try:
raw = json.loads(text)
except (ValueError, RecursionError, MemoryError) as exc:
return _broken(booth, f"is not valid JSON: {type(exc).__name__}")
if not isinstance(raw, dict):
return _broken(booth, "is not a JSON object")
handle = _one_line(raw.get("handle"), HANDLE_MAX)
if not handle:
return _broken(booth, "names no handle")
return Manifest(
handle=handle,
# `or booth.name` goes THROUGH the normalizer too. A directory name may
# legally carry a newline on POSIX and may run to 255 bytes, and the
# fallback used to hand either straight into a card's sub-line.
title=_one_line(raw.get("title"), TITLE_MAX) or _one_line(booth.name, TITLE_MAX),
why=_one_line(raw.get("why"), WHY_MAX),
created=_one_line(raw.get("created"), CREATED_MAX),
)
def _broken(booth: Path, reason: str) -> Manifest:
# The directory name goes through the normalizer here too. This was the
# THIRD fallback of three; the write path's and the read path's were fixed a
# round earlier and this one was missed, with the same consequence — a
# newline or 255 bytes of directory name straight into a card's sub-line.
return Manifest(handle="", title=_one_line(booth.name, TITLE_MAX), why="",
created="", error=f"{MANIFEST_FILE} {reason}")
def write_manifest(booth: Path, handle: str, *, title: str | None = None,
why: str | None = None) -> Manifest:
"""Announce a booth, atomically (CLAUDE.md invariant 5).
Temp file + `os.replace`, because the CLI writes this in one process while
the browser reads it in another — a reader must never see a half-written
document. The temp file is itself a dotfile, so no listing, gallery or zip
can see it mid-write either.
OMITTED MEANS UNCHANGED; `""` MEANS CLEAR. `title` and `why` default to
None, not to the empty string, because the ordinary sequence is
`booth new x --why "..."` and then `booth add x out/*.png` — and while
omission meant empty, that second command silently erased the sentence the
first one existed to record. Two arms of the contract panel predicted it
from the wording alone; every test written for this module passed `--why`
on both calls and so could not see it.
RE-ANNOUNCING PRESERVES `created` (INV-3). It is when the booth APPEARED,
and saying something more about it later is not a second appearance. A
`created` that cannot be read back is replaced rather than guessed at: a
stamp that is silently wrong is worse than one that is silently new.
An empty `handle` becomes `SERVICE_HANDLE` rather than being refused — a
manifest with no handle does not read back at all, and an unreadable file is
the worse outcome. Unreachable from the CLI, whose fallback chain always
yields something; callers of this function directly should pass a real one.
"""
booth = Path(booth)
booth.mkdir(parents=True, exist_ok=True)
prior = read_manifest(booth)
usable = prior if prior and not prior.error else None
created = usable.created if usable and usable.created else _now()
record = Manifest(
handle=_one_line(handle, HANDLE_MAX) or SERVICE_HANDLE,
title=(_one_line(title, TITLE_MAX) if title is not None
else (usable.title if usable else "")) or _one_line(booth.name, TITLE_MAX),
why=(_one_line(why, WHY_MAX) if why is not None
else (usable.why if usable else "")),
created=created,
)
path = booth / MANIFEST_FILE
doc = {"handle": record.handle, "title": record.title,
"why": record.why, "created": record.created}
# A write that changes nothing is not activity and must not reset the
# booth's TTL — the rule marks learned in v0.2.0, applied here because
# `booth link` re-announces the standing board on EVERY post to it.
if prior is not None and not prior.error and _as_doc(prior) == doc:
return record
# NOTHING THAT COULD NOT BE READ IS DESTROYED. Reads stay lenient, writes
# go strict, damaged bytes stay on disk — the doctrine marks made explicit
# in v0.2.1, which this write path contradicted by replacing them outright.
# A file that fails on ONE field still holds the others, and a `why` the
# re-announcer never kept anywhere is exactly what went missing.
#
# QUARANTINED rather than REFUSED, which is where this diverges from marks:
# refusing would fail `booth add` and lose the files it was mid-way through
# copying, and a booth's own description is restatable in a way the
# operator's judgment is not.
if prior is not None and prior.error:
try:
os.replace(path, booth / QUARANTINE_FILE)
except OSError:
pass # nothing to preserve beats failing the write
tmp = _temp_path(booth)
try:
tmp.write_text(
json.dumps(doc, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
os.replace(tmp, path)
except BaseException:
# A leaked temp is worse here than it would be with a fixed name: the
# unique suffix means nothing ever overwrites it, and it is not a
# `.lock`, so `_newest_mtime` counts it and it keeps a dead booth alive
# forever. Cleaning up is the price of the uniqueness.
tmp.unlink(missing_ok=True)
raise
return record
+304 -24
View File
@@ -42,6 +42,7 @@ from __future__ import annotations
import fcntl
import json
import os
import stat as statmod
from dataclasses import asdict, dataclass, field
from datetime import datetime
from pathlib import Path
@@ -59,6 +60,28 @@ from booth.asks import (
valid_stem,
)
class MarksCorrupt(RuntimeError):
"""The mark file exists but cannot be parsed, and a WRITE was attempted.
The read path is deliberately lenient — `marks_for` returns [] so a review
page still loads. The write path must not inherit that leniency: reading a
damaged file as "no marks" and then atomically replacing it destroys every
judgment in the booth from one click, silently. Shipped in v0.2.0 and found
by a cross-frontier contract panel, not by the suite.
A page that renders without an annotation is recoverable. A file that
overwrote the operator's judgment is not.
"""
# A booth's whole judgment lives in one document, so this is generous — a
# 270-item booth flagged throughout, with notes, is far under it. What it rules
# out is the case that is not marks at all: an unbounded read raises MemoryError
# and a deeply nested one raises RecursionError out of `json.loads`, neither of
# which is an OSError or a ValueError, and `list_booths` calls the reader once
# per booth on every index load. Bounded by `stat`, before the bytes are read.
MARKS_MAX_BYTES = 4 * 1024 * 1024
MARKS_FILE = ".marks.json"
MARKS_LOCK = ".marks.lock"
SCHEMA_VERSION = 1
@@ -112,7 +135,17 @@ class Mark:
def now_stamp() -> str:
return datetime.now().astimezone().isoformat(timespec="seconds")
"""ONE stamp format across every writer in this module.
MICROSECONDS, matching `import_legacy_asks`. They diverged when the
importer was moved to sub-second precision to stop same-second sidecars
re-sorting — and the divergence opened a fresh ordering bug in the other
direction, because `-` (0x2D) sorts before `.` (0x2E): a whole-second stamp
lands ahead of ANY fractional stamp in the same second, so a later mark came
out before an earlier import. Marks sort on `(created, id)`; one format is
what makes that rule statable.
"""
return datetime.now().astimezone().isoformat(timespec="microseconds")
def _clean_text(text) -> str:
@@ -159,9 +192,17 @@ def _read_raw(booth: Path) -> list[dict]:
for the same reason: a review surface that will not load is worse than one
that has lost an annotation.
"""
path = Path(booth) / MARKS_FILE
try:
raw = json.loads((Path(booth) / MARKS_FILE).read_text(encoding="utf-8"))
except (OSError, ValueError, UnicodeDecodeError):
st = path.stat()
# Regular-file first, then size. `st_size` is 0 for a FIFO and 0 for a
# symlink to /dev/zero, so both pass a byte cap and then `read_text`
# either blocks with no EOF or allocates until the kernel intervenes.
# This loop runs over EVERY booth on every index load.
if not statmod.S_ISREG(st.st_mode) or st.st_size > MARKS_MAX_BYTES:
return []
raw = json.loads(path.read_text(encoding="utf-8"))
except (OSError, ValueError, UnicodeDecodeError, RecursionError, MemoryError):
return []
if not isinstance(raw, dict):
return []
@@ -176,14 +217,94 @@ def _fingerprint(entries: list[dict]) -> str:
return json.dumps(entries, sort_keys=True, ensure_ascii=False)
def _read_raw_strict(booth: Path, *, blank_is_corrupt: bool = False) -> list[dict]:
"""Like `_read_raw`, but RAISES `MarksCorrupt` on a file it cannot parse.
Absent, empty and valid-but-empty are all "no marks yet" and are fine — the
distinction that matters is bytes-present-but-unreadable, because that is the
case where writing would destroy something.
`blank_is_corrupt` is the DELETE path's reading of a present-but-whitespace
file, and only the delete path's: this writer never produces a blank marks
document, so a blank one that exists is something that went wrong, and
`rmtree` is not the response to that. The write path keeps the lenient
reading — a blank file is safe to overwrite, which is the question
`_Locked` is asking. A VALID document with an empty `marks` list is not
blank and never holds: that is what deleting the last mark leaves behind,
and it must stay sweepable.
"""
path = Path(booth) / MARKS_FILE
try:
st = path.stat()
except FileNotFoundError:
return []
except OSError as exc:
raise MarksCorrupt(f"{path} cannot be read: {exc}") from exc
# The strict half has to refuse everything the lenient half tolerates, or a
# file that reads as "no marks" gets replaced by a write that believed it.
if not statmod.S_ISREG(st.st_mode):
raise MarksCorrupt(f"{path} is not a regular file")
if st.st_size > MARKS_MAX_BYTES:
raise MarksCorrupt(
f"{path} is too large to be a marks document ({st.st_size} bytes)")
try:
text = path.read_text(encoding="utf-8")
except FileNotFoundError:
return []
except (OSError, UnicodeDecodeError, MemoryError) as exc:
raise MarksCorrupt(f"{path} cannot be read: {exc}") from exc
if not text.strip():
if blank_is_corrupt:
raise MarksCorrupt(f"{path} is present but holds no marks document")
return []
try:
raw = json.loads(text)
except (ValueError, RecursionError, MemoryError) as exc:
raise MarksCorrupt(
f"{path} is not valid JSON: {type(exc).__name__}") from exc
if not isinstance(raw, dict) or not isinstance(raw.get("marks"), list):
raise MarksCorrupt(f"{path} is not a marks document")
entries = [e for e in raw["marks"] if isinstance(e, dict) and isinstance(e.get("id"), str)]
if len(entries) != len(raw["marks"]):
raise MarksCorrupt(f"{path} holds entries this version cannot read")
return entries
def read_error(booth: Path) -> str | None:
"""Why this booth's marks cannot be read, or None if they can.
`marks_for` is lenient on purpose — a review page that will not load is
worse than one missing an annotation — and that leniency turns an
unreadable file into "no marks". For a BROWSER that is the right trade. For
the CLI it is not: a session that asked a question and is told "no such
pick" will conclude the question was never posted, when in fact the file
holding it is damaged. A machine consumer can act on the difference, so it
gets to ask.
"""
try:
_read_raw_strict(booth)
except MarksCorrupt as exc:
return str(exc)
return None
def _write_raw(booth: Path, entries: list[dict]) -> None:
"""Atomic replace, so a reader never sees a half-written document and a
crash mid-write cannot truncate the file into a shorter — and therefore
quieter — set of marks."""
path = Path(booth) / MARKS_FILE
doc = {"version": SCHEMA_VERSION, "marks": entries}
body = json.dumps(doc, ensure_ascii=False, indent=2) + "\n"
# The read bound is on the STORED bytes and `indent=2` grows them, so a
# document that fits in memory can land over the limit on disk and then read
# back as no marks at all. Refuse loudly instead: a write that fails is
# recoverable, a file that silently empties is not.
if len(body.encode("utf-8")) > MARKS_MAX_BYTES:
raise MarksCorrupt(
f"{path} would be larger than this version can read back "
f"({len(body.encode('utf-8'))} bytes)")
tmp = path.with_suffix(path.suffix + ".tmp")
tmp.write_text(json.dumps(doc, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
tmp.write_text(body, encoding="utf-8")
os.replace(tmp, path)
@@ -207,15 +328,58 @@ class _Locked:
self.booth.mkdir(parents=True, exist_ok=True)
lock = self.booth / MARKS_LOCK
# `touch(exist_ok=True)` on an EXISTING file bumps its mtime, and a
# booth's TTL is measured from its newest mtime including dotfiles — so
# an unconditional touch would keep a booth alive just for being read
# through a write path. Create it only when it is not there.
# booth's TTL is measured from its newest mtime — so an unconditional
# touch would keep a booth alive just for being read through a write
# path. Create it only when it is not there.
#
# ONCE CREATED, THE LOCK FILE IS NEVER REMOVED (see __exit__).
if not lock.exists():
# Creating a directory entry bumps the DIRECTORY's mtime, which is
# what `_newest_mtime` reads. An earlier version put the clock back
# with `os.utime` — which closed the bug and opened a race: the
# restore ran before the flock, so anything landing in the window
# between the stat and the utime had its bump rolled backward. An
# `rsync -a` batch is the case that bites, because it PRESERVES
# source mtimes and so has only the directory's freshness to look
# alive by. It could also raise OSError on a read-only directory
# and take the route down with it.
#
# THE RESTORE STAYS, and the honest reason is that the alternative
# was worse. Ignoring a booth directory's own mtime whenever the
# booth holds anything would close the race outright — and would
# also silently retire the documented behaviour that RELEASING a
# kept board resets its clock, which the CLI header, the README and
# a deliberate test all pin. That is a TTL doctrine change, not a
# bug fix, and it does not belong in one.
#
# ⚠ RESIDUAL RACE, stated rather than papered over: between the stat
# and the utime, another writer's directory-entry change can be
# rolled backward. The case that bites is an `rsync -a` batch, which
# preserves source mtimes and so has only the directory's freshness
# to look alive by. The window is the two syscalls below and the
# booth must also be one being written to at that instant.
#
# The concrete half IS fixed: a failing utime (read-only directory,
# a booth whose owner we are not) used to escape and take the whole
# route down with a 500. Not putting the clock back is a cost this
# module can absorb; not answering the request is not.
before = self.booth.stat()
lock.touch()
self._made_lock = True
try:
os.utime(self.booth, (before.st_atime, before.st_mtime))
except OSError:
pass
self._lf = lock.open("r+")
fcntl.flock(self._lf, fcntl.LOCK_EX)
self.entries = _read_raw(self.booth)
try:
# STRICT here, lenient in marks_for — see MarksCorrupt.
self.entries = _read_raw_strict(self.booth)
except MarksCorrupt:
fcntl.flock(self._lf, fcntl.LOCK_UN)
self._lf.close()
self._lf = None
raise
self._before = _fingerprint(self.entries)
return self
@@ -233,10 +397,17 @@ class _Locked:
# would otherwise keep a dead booth alive forever.
if exc_type is None and _fingerprint(self.entries) != self._before:
_write_raw(self.booth, self.entries)
elif self._made_lock and not (self.booth / MARKS_FILE).exists():
# Nothing was written and this booth had no marks before: do not
# leave a lock file behind as the only trace of a no-op.
(self.booth / MARKS_LOCK).unlink(missing_ok=True)
# THE LOCK FILE IS NEVER UNLINKED. It used to be, on the no-op path,
# so a booth that had never been marked was left exactly as it was
# found. That tidiness cost mutual exclusion outright: `flock` binds
# to an INODE, so unlinking the lock while a second writer is blocked
# on it leaves that writer holding an exclusive lock on a deleted
# file, and the NEXT writer creates a fresh lock and takes it at
# once. Two processes then run the read-modify-write concurrently,
# the later `os.replace` drops the earlier one's mark, and both of
# them obeyed the protocol. A zero-byte dotfile is the cheaper
# thing to leave behind — `booth_items` skips it, the zip skips it,
# and `_newest_mtime` exempts it so it cannot hold a booth open.
finally:
fcntl.flock(lf, fcntl.LOCK_UN)
lf.close()
@@ -250,6 +421,25 @@ class _Locked:
# ---- read -------------------------------------------------------------------
def _entry_type_error(entry: dict) -> str | None:
"""The stored scalars this module refuses to guess at.
`_clean_text` did `(text or "").replace(...)` and `marks_for` sorts on
`(created, id)` — so a stored `text` that is a dict, or a `created` that is a
number, raised AttributeError or TypeError out of the READ path. That is not
a marks bug, it is an INDEX bug: `list_booths` reads every booth's marks on
every page load and `/healthz` does the same, so one hand-edited or
foreign-written file took down the front page for every booth on the
service. A wrong type is a broken mark, and this module already knows how to
render one of those.
"""
for name in ("created", "by", "text", "error"):
value = entry.get(name)
if value is not None and not isinstance(value, str):
return f"{name} is {type(value).__name__}, not a string"
return None
def _hydrate(entry: dict) -> Mark:
"""One stored entry -> one Mark, declarations normalized.
@@ -261,6 +451,15 @@ def _hydrate(entry: dict) -> Mark:
"""
mid = entry["id"]
shape = entry.get("shape") if entry.get("shape") in SHAPES else NOTE
bad = _entry_type_error(entry)
if bad is not None:
# `created` is dropped rather than coerced, which sorts the entry to the
# TOP of the booth's marks: a mark nobody can read is the one that wants
# looking at, and burying it under 270 items' worth of notes is how it
# stays unnoticed. Deterministic, and stated — `("", id)` against
# `(created, id)`.
return Mark(id=mid, shape=shape, target=None, created="",
error=f"unreadable mark: {bad}")
target = entry.get("target")
if not _valid_target(target):
target = None
@@ -308,11 +507,26 @@ def _hydrate(entry: dict) -> Mark:
return Mark(**base, text=_clean_text(entry.get("text")))
def _hydrate_safe(entry: dict) -> Mark:
"""`_hydrate`, with the promise that it cannot raise.
`_entry_type_error` covers the shapes we know how to name; this is the
backstop for the ones we do not, and it exists because of WHERE this runs.
One unreadable mark must cost that mark, never the page — and on the index
it is not even that booth's page, it is all of them.
"""
try:
return _hydrate(entry)
except Exception as exc: # noqa: BLE001 - deliberate
return Mark(id=str(entry.get("id", "")), shape=NOTE, target=None,
created="", error=f"unreadable mark: {exc}")
def marks_for(booth: Path) -> list[Mark]:
"""Every mark in a booth, oldest first, declarations normalized and answers
folded in. ONE file read — which is the whole point of the storage shape."""
entries = _read_raw(booth)
marks = [_hydrate(e) for e in entries]
marks = [_hydrate_safe(e) for e in entries]
# (created, id) rather than created alone: two marks written in the same
# second would otherwise order by however json listed them.
marks.sort(key=lambda m: (m.created, m.id))
@@ -342,6 +556,34 @@ def open_marks(marks: Sequence[Mark]) -> list[Mark]:
return [m for m in marks if _is_open(m)]
def hold_read(booth: Path) -> tuple[list[Mark], str | None]:
"""ONE read of `.marks.json`, answering both questions the LIFETIME rule asks:
what is still open, and whether the file could be read at all.
U4 decides whether a booth may be SWEPT from those two facts. Asking them
with two calls — `marks_for` then `read_error` — reads the file twice, and
two reads of one file are not one read of one state: a write or a repair
landing between them yields a pair that never described the booth at any
instant. The losing pair is `([], None)` — no marks, no error — which is
exactly the one that deletes. Cross-frontier review (2026-09-22) found it;
that is why this exists rather than the obvious two calls.
When the file reads clean the marks are byte-identical to `marks_for`'s:
`_read_raw_strict` raises rather than dropping an entry, so a non-raising
strict read returns the same entries the lenient read would, hydrated and
sorted the same way. The caller can therefore use this ONE read for the
display too, and fall back to `marks_for` only on the error path, where
leniency is the point.
"""
try:
entries = _read_raw_strict(booth, blank_is_corrupt=True)
except MarksCorrupt as exc:
return [], str(exc)
marks = [_hydrate_safe(e) for e in entries]
marks.sort(key=lambda m: (m.created, m.id))
return marks, None
def marks_for_target(marks: Sequence[Mark], rel: str | None) -> list[Mark]:
"""The marks attached to one item, or to the booth itself for None."""
return [m for m in marks if m.target == rel]
@@ -357,16 +599,25 @@ def as_dict(mark: Mark) -> dict:
# ---- write ------------------------------------------------------------------
def declare_pick(booth: Path, mark_id: str, doc: dict) -> Mark:
"""A session poses a pick.
def declare_pick(booth: Path, mark_id: str, doc: dict, target: str | None = None) -> Mark:
"""A session poses a pick, about the booth or about ONE item in it.
Validated through `normalize_ask` BEFORE anything is written, so a session
cannot land a question the renderer would refuse. Re-declaring an existing
id replaces the declaration and CLEARS its answer: the question changed, so
the old judgment is not an answer to it.
the old judgment is not an answer to it — and it may move the target, since
a re-declaration is a new question.
`target` is an `Item.rel`, or None for the booth. It exists because the
2026-09-09 ruling is that a question belongs WITH the artifact it is about: a
four-voice audition wants the radio group under that voice. The record and
the renderer both supported it before this parameter did, which meant a
session could not actually produce one.
"""
if not valid_stem(mark_id):
raise AskError("bad mark id: letters, digits, . _ - only")
if not _valid_target(target):
raise AskError("a pick's target must be a path inside the booth")
normalize_ask(doc, mark_id) # raises AskError; nothing written yet
with _Locked(booth) as lk:
existing = lk.find(mark_id)
@@ -375,7 +626,7 @@ def declare_pick(booth: Path, mark_id: str, doc: dict) -> Mark:
entry = {
"id": mark_id,
"shape": PICK,
"target": existing.get("target") if existing else None,
"target": target,
"created": existing.get("created") if existing else now_stamp(),
"declaration": doc,
"answer": None,
@@ -537,7 +788,8 @@ def import_legacy_asks(booth: Path) -> list[Mark]:
continue
try:
decl = json.loads(p.read_text(encoding="utf-8"))
except (OSError, ValueError, UnicodeDecodeError) as exc:
except (OSError, ValueError, UnicodeDecodeError,
RecursionError, MemoryError) as exc:
found.append((mtime, stem, None, f"unreadable ask: {exc}"))
continue
if not isinstance(decl, dict):
@@ -551,23 +803,49 @@ def import_legacy_asks(booth: Path) -> list[Mark]:
created: list[dict] = []
with _Locked(booth) as lk:
have = {e.get("id") for e in lk.entries}
by_id = {e.get("id"): e for e in lk.entries}
for mtime, stem, decl, err in found:
if stem in have:
continue
answer = None
ap = booth / f"{stem}{ANSWER_SUFFIX}"
try:
loaded = json.loads(ap.read_text(encoding="utf-8"))
if isinstance(loaded, dict):
answer = loaded
except (OSError, ValueError, UnicodeDecodeError):
except (OSError, ValueError, UnicodeDecodeError,
RecursionError, MemoryError):
pass
prior = by_id.get(stem)
if prior is not None:
# The stem is already a mark, so the DECLARATION is not imported
# — that is the idempotence rule, and a mark declared since the
# sidecar outranks it. But a legacy ANSWER must not be stranded:
# if the existing mark is an unanswered pick and the sidecar
# holds the operator's choice, adopt it. Ordinary reads are
# forbidden from looking at sidecars, so a skip here would lose
# that judgment permanently.
if (answer is not None
and prior.get("shape") == PICK
and prior.get("answer") is None):
prior["answer"] = answer
created.append(prior)
continue
entry = {
"id": stem,
"shape": PICK,
"target": None,
"created": datetime.fromtimestamp(mtime).astimezone().isoformat(timespec="seconds"),
# MICROSECONDS, not seconds. `found` is ordered by fractional
# mtime and `marks_for` re-sorts on this string, so truncating
# to the whole second threw away the only thing distinguishing
# two sidecars written in the same second — and the `(created,
# id)` tie-break then silently re-sorted them alphabetically,
# reversing the order the importer had just established. The
# ROADMAP states this import's order is `(mtime, name)`; an
# order that is stated and not kept is worse than one never
# claimed.
"created": datetime.fromtimestamp(mtime).astimezone().isoformat(
timespec="microseconds"),
"declaration": decl,
"answer": answer,
}
@@ -578,4 +856,6 @@ def import_legacy_asks(booth: Path) -> list[Mark]:
# Hydrated AFTER the lock so a broken declaration surfaces as `error` here
# exactly as it does on a normal read, rather than through a second path.
return [_hydrate(e) for e in created]
# `_hydrate_safe`, not `_hydrate`: this is the one path that reads entries
# it did not write, and it was the one without the guard.
return [_hydrate_safe(e) for e in created]
+28
View File
@@ -0,0 +1,28 @@
{# The lifetime line, defined ONCE and called from four surfaces: the index
card (both lanes), the booth header (both branches) and the marks page.
U4: a booth's lifetime is derived from its own state, and a booth that is
not counting down must always SAY WHY — an invisible rule that silently
stopped the clock would be strictly worse than the `.forever` boolean it
replaces, because that one was at least visible as a lane.
`hold` is the REASON, straight off `hold_reason()`, not a bool beside a
string that can disagree with it. Kept wins over a hold because a kept booth
is exempt either way, and showing two reasons for one EXEMPTION is the
two-representations-of-one-state trap.
Unreadable marks are the exception and ride along even on a kept board:
damaged judgment is not a second exemption, it is a thing somebody has to go
and fix, and the kept lane holds the durable boards — the ones where losing
the operator's marks costs most. #}
{% macro lifetime(kept, hold, expires_in) -%}
{%- if kept -%}
kept{% if hold == "unreadable" %} · <span class="held held-broken" title="a mark in this booth cannot be read">marks unreadable</span>{% endif %}
{%- elif hold == "unreadable" -%}
<span class="held held-broken" title="a mark in this booth cannot be read, so the sweeper will not take it">held · marks unreadable</span>
{%- elif hold == "open" -%}
<span class="held" title="an unanswered question holds this booth open">held until answered</span>
{%- else -%}
expires in {{ expires_in|dur }}
{%- endif -%}
{%- endmacro %}
+26 -2
View File
@@ -13,11 +13,35 @@
Works with JS off — plain form POST, every shape. An answered pick shows the
recorded judgment and a collapsed "change" form, because the mark is the
CURRENT judgment and not a log. #}
{# A mark carrying `error` is sorted out FIRST, whatever shape it claims. A
pick keeps its own ⚠ broken rendering below (richer — it has a declaration to
show); a broken note would otherwise render as an empty <pre> with a withdraw
button, indistinguishable from a note the operator wrote and then cleared,
and a broken flag would link to a target that is not there. Unreadable state
is visible state — the rule `_hydrate` states for picks, applied to all
three. #}
{% set broken = marks | selectattr('error') | rejectattr('shape', 'equalto', 'pick') | list %}
{% set picks = marks | selectattr('shape', 'equalto', 'pick') | list %}
{% set notes = marks | selectattr('shape', 'equalto', 'note') | list %}
{% set flags = marks | selectattr('shape', 'equalto', 'flag') | list %}
{% set notes = marks | selectattr('shape', 'equalto', 'note') | rejectattr('error') | list %}
{% set flags = marks | selectattr('shape', 'equalto', 'flag') | rejectattr('error') | list %}
<section class="marks">
{% for a in broken %}
<article class="mark mark-note is-broken" id="mark-{{ a.id }}">
<header class="mark-head">
<span class="mark-state">⚠ broken</span>
<span class="mark-id"><code>{{ a.id }}</code></span>
<span class="board-spacer"></span>
<form class="mark-undo" method="post" action="/b/{{ name_url }}/unmark">
<input type="hidden" name="mark" value="{{ a.id }}">
{% if marks_page %}<input type="hidden" name="back" value="marks">{% endif %}
<button type="submit" class="mark-x" title="withdraw this mark">×</button>
</form>
</header>
<p class="mark-error">This mark could not be read: {{ a.error }}</p>
</article>
{% endfor %}
{% for a in picks %}
<article class="mark mark-pick{% if a.answer and a.answer.complete %} is-answered{% elif a.answer %} is-partial{% elif a.error %} is-broken{% endif %}" id="mark-{{ a.id }}">
<header class="mark-head">
+17
View File
@@ -0,0 +1,17 @@
{# THE ANNOUNCEMENT — who posted this booth and why. Defined ONCE and called
from both index lanes and the booth page header: the kept lane is a separate
block, and patching only the ephemeral one would leave the durable,
most-looked-at boards with exactly the defect this closes.
Four states, and `unannounced` is distinct from `unreadable` on purpose —
folding "cannot be read" into "never said" hides the one case somebody has to
go and fix. The classes are the test hooks; the words are for the operator. #}
{% macro provenance(m) -%}
{% if m is none %}
<div class="prov prov-none">unannounced</div>
{% elif m.error %}
<div class="prov prov-broken" title="{{ m.error }}">unreadable</div>
{% else %}
<div class="prov"><span class="prov-who">{{ m.handle }}</span>{% if m.why %} · <span class="prov-why">{{ m.why }}</span>{% endif %}</div>
{% endif %}
{%- endmacro %}
+28
View File
@@ -255,6 +255,22 @@
.card .name:hover{text-decoration:none;color:var(--aus-bright-cyan)}
.card .sub{color:var(--fg-3);font-size:.72rem;font-family:var(--font-mono);letter-spacing:.03em;margin-top:.3rem}
/* THE ANNOUNCEMENT — who posted this booth and why (U5). Same size and
rhythm as .sub above it, because it is the same class of information: a
second line of card metadata, not a heading. The handle carries the only
colour, so a scan down the index reads as a column of posters. */
.prov{margin-top:.28rem;font-size:.72rem;font-family:var(--font-mono);
letter-spacing:.03em;color:var(--fg-3);line-height:1.45;
overflow-wrap:anywhere}
.prov-who{color:var(--fg-2)}
.prov-why{color:var(--fg-3)}
/* Quiet on purpose. 26 booths arrived before this convention existed and
rsync keeps making more, so the marker has to be visible-if-you-look and
never a badge shouting 26 times. `unreadable` gets the warning tint
because, unlike `unannounced`, it is something somebody has to fix. */
.prov-none{color:var(--fg-muted);font-style:italic}
.prov-broken{color:var(--aus-bright-yellow,#e8c547);font-style:italic;cursor:help}
.wipe{position:absolute;top:.5rem;right:.5rem;margin:0}
/* ★ keep, mirroring .wipe on the other shoulder of the card. Same
hover-to-reveal language as .release in the kept lane. */
@@ -314,6 +330,11 @@
use for state), green check once answered; the accent is a TOP edge, per
Australis, never a coloured left border. */
.badge-mark{background:var(--aus-bright-yellow);color:var(--fg-on-accent)}
/* U4: the lifetime line's HELD states. Marked rather than styled into
invisibility — the whole safety argument for an unbounded hold is that
a booth which stopped counting down says so where the countdown was. */
.held{color:var(--aus-bright-yellow)}
.held-broken{color:var(--fg-3);text-decoration:underline dotted}
.thumb .badge+.badge-mark{top:2.2rem}
.marks{display:flex;flex-direction:column;gap:.9rem;margin:.2rem 0 1.4rem}
.mark{border:1px solid var(--border-subtle);border-top:2px solid var(--aus-bright-yellow);
@@ -408,6 +429,13 @@
.boothhead h1{margin:0;font-family:var(--font-display);font-weight:600;font-size:1.5rem;
letter-spacing:-.01em;word-break:break-word;flex:1 1 auto;color:var(--fg-0)}
.boothhead .sub{color:var(--fg-3);font-size:.74rem;font-family:var(--font-mono);letter-spacing:.06em}
/* Its own row under the title, not another chip in the flex line — a `why`
can run to WHY_MAX and would otherwise shove the zip link around. */
.boothhead .prov{flex:0 0 100%;margin-top:-.35rem}
/* The directory name beside a manifest title: quieter than the title, but
never absent — it is what the URL says and what "the third one" refers to. */
.h1-slug{font-family:var(--font-mono);font-size:.62em;font-weight:400;
letter-spacing:.06em;color:var(--fg-3);margin-left:.5rem;white-space:nowrap}
.wipe-lg{position:static}
/* red-outline danger button — legible on the dark canvas, fills on hover */
.wipe-lg button{width:auto;height:auto;padding:.42rem .85rem;border-radius:var(--radius-md);
+24 -3
View File
@@ -1,4 +1,6 @@
{% extends "base.html" %}
{% from "_provenance.html" import provenance %}
{% from "_lifetime.html" import lifetime %}
{# The blur toggle, defined ONCE. There are three item branches in this file
(doc / media / other) and the first cut of this feature patched only one of
them, so docs rendered with no control at all. A macro makes "patched two of
@@ -55,9 +57,18 @@
{% block content %}
<div class="boothhead">
<a class="back" href="/">‹ all booths</a>
<h1>{{ name }}</h1>
<span class="sub">{% if uploaded %}<span class="badge">⬆ pickup</span> {% endif %}{% if board %}{{ board|length }} link{{ '' if board|length == 1 else 's' }}{% if items %} · {{ items|length }} file{{ '' if items|length == 1 else 's' }}{% endif %}{% else %}{% if marks_open %}<span class="badge badge-mark">{{ marks_open }} open</span> · {% endif %}{{ items|length }} item{{ '' if items|length == 1 else 's' }} · expires in {{ expires_in|dur }}{% endif %}</span>
{# The manifest's TITLE is the display name; the directory name stays visible
beside it because that is the identity the operator navigates by and refers
to positionally, and losing it would be losing the thing the URL says.
Index cards keep the directory name alone for the same reason. #}
{% if manifest and not manifest.error and manifest.title and manifest.title != name %}
<h1>{{ manifest.title }} <span class="h1-slug">{{ name }}</span></h1>
{% else %}
<h1>{{ name }}</h1>
{% endif %}
<span class="sub">{% if uploaded %}<span class="badge">⬆ pickup</span> {% endif %}{% if board %}{{ board|length }} link{{ '' if board|length == 1 else 's' }}{% if items %} · {{ items|length }} file{{ '' if items|length == 1 else 's' }}{% endif %} · {{ lifetime(kept, hold, expires_in) }}{% else %}{% if marks_open %}<span class="badge badge-mark">{{ marks_open }} open</span> · {% endif %}{{ items|length }} item{{ '' if items|length == 1 else 's' }} · {{ lifetime(kept, hold, expires_in) }}{% endif %}</span>
{% if items %}<a class="dl-link" href="/b/{{ name_url }}/?download=1" title="download this booth as a zip">⬇ zip</a>{% endif %}
{{ provenance(manifest) }}
{# A durable multi-writer board gets no one-click wipe — same rule as the
kept lane on the index. Remove rows with the per-row ×, or release the
board from the index and wipe it from there. #}
@@ -94,7 +105,11 @@
back to the flagged items. Always rendered on a gallery booth — the add-note
field is a control, not a result, so it has to be there before the first
mark exists. #}
{% if not board %}
{# `marks or not board`: the standing link board renders as a board rather than
a gallery, and the add-note control would be noise on it — but the
suppression was unconditional, so a pick declared on a booth that happens to
carry a links.md had no form to answer it and nothing said so. #}
{% if marks or not board %}
{% include "_marks.html" %}
{% endif %}
@@ -187,6 +202,12 @@
{% else %}
<pre class="textview doc-body">{{ it.rendered }}</pre>
{% endif %}
{# The doc branch had `markcontrols` and not `marknotes`, so the
operator could point at a report and not write down why — on the
one item kind whose whole content is prose. Exactly the
"patched two of three" failure the blurtoggle macro above was
written to prevent, recurring on the macro written to prevent it. #}
{{ marknotes(name_url, it, item_marks.get(it.name, [])) }}
</details>
</figure>
{% else %}
+13 -3
View File
@@ -33,8 +33,18 @@
white-space:pre-wrap}
</style>
<script>
document.addEventListener('keydown', function (e) {
if (e.key === 'Escape') window.location.href = {{ ('/b/' ~ name_url ~ '/')|tojson }};
});
(function () {
/* Escape leaves the page, so it must not fire from inside a field someone
is typing in — the same guard the image viewer carries, stated in both
places because the handler is on `document` in both. */
function isEditable(el) {
return !!(el && (el.isContentEditable ||
/^(input|textarea|select)$/i.test(el.tagName || '')));
}
document.addEventListener('keydown', function (e) {
if (isEditable(e.target)) return;
if (e.key === 'Escape') window.location.href = {{ ('/b/' ~ name_url ~ '/')|tojson }};
});
})();
</script>
{% endblock %}
+43 -5
View File
@@ -1,4 +1,6 @@
{% extends "base.html" %}
{% from "_provenance.html" import provenance %}
{% from "_lifetime.html" import lifetime %}
{% block content %}
<form class="uploader" method="post" action="/upload" enctype="multipart/form-data">
<label class="drop" for="booth-files">
@@ -38,7 +40,8 @@
</a>
<div class="meta">
<a class="name" href="/b/{{ b.name_url }}/">{{ b.name }}</a>
<div class="sub">{{ b.count }} item{{ '' if b.count == 1 else 's' }} · kept · <a class="dl-link" href="/b/{{ b.name_url }}/?download=1" title="download this booth as a zip">⬇ zip</a></div>
<div class="sub">{{ b.count }} item{{ '' if b.count == 1 else 's' }} · {{ lifetime(true, b.hold, b.expires_in) }} · <a class="dl-link" href="/b/{{ b.name_url }}/?download=1" title="download this booth as a zip">⬇ zip</a></div>
{{ provenance(b.manifest) }}
</div>
{# There IS a × here now (operator, 2026-09-21). The old rule was
release-then-find-it-in-the-other-lane, on the theory that two
@@ -67,11 +70,11 @@
a label changes width. #}
<div class="kept-actions">
<form class="release" method="post" action="/b/{{ b.name_url }}/unkeep"
onsubmit="return confirm('Release \u201c{{ b.name }}\u201d?\n\nIt moves to the ephemeral lane so you can wipe it from there. Nothing is deleted by this step.')">
data-booth="{{ b.name }}" data-confirm="release">
<button title="release this board so it can be wiped">release</button>
</form>
<form class="wipe wipe-kept" method="post" action="/b/{{ b.name_url }}/delete"
onsubmit="return confirm('WIPE the KEPT booth \u201c{{ b.name }}\u201d?\n\nThis deletes it and its files immediately. Kept booths are the ones nothing else will clean up, so nobody else is going to do this for you — and nothing brings it back.')">
data-booth="{{ b.name }}" data-confirm="wipe-kept">
<button title="wipe this KEPT booth now" aria-label="wipe kept booth">×</button>
</form>
</div>
@@ -109,7 +112,8 @@
</a>
<div class="meta">
<a class="name" href="/b/{{ b.name_url }}/">{{ b.name }}</a>
<div class="sub">{{ b.count }} item{{ '' if b.count == 1 else 's' }} · expires in {{ b.expires_in|dur }} · <a class="dl-link" href="/b/{{ b.name_url }}/?download=1" title="download this booth as a zip">⬇ zip</a></div>
<div class="sub">{{ b.count }} item{{ '' if b.count == 1 else 's' }} · {{ lifetime(false, b.hold, b.expires_in) }} · <a class="dl-link" href="/b/{{ b.name_url }}/?download=1" title="download this booth as a zip">⬇ zip</a></div>
{{ provenance(b.manifest) }}
</div>
{# Promote to the kept lane. The /keep route and the `booth keep` CLI verb
both predate this button; until 2026-09-19 the UI could only RELEASE a
@@ -120,7 +124,7 @@
<button title="keep — exempt from the {{ ttl_hours }}h sweep" aria-label="keep booth">★</button>
</form>
<form class="wipe" method="post" action="/b/{{ b.name_url }}/delete"
onsubmit="return confirm('Wipe booth “{{ b.name }}”?')">
data-booth="{{ b.name }}" data-confirm="wipe">
<button title="wipe now" aria-label="wipe booth">×</button>
</form>
</article>
@@ -157,5 +161,39 @@
}
});
})();
/* Destructive-action confirmation, delegated and DATA-DRIVEN.
These were an inline onsubmit calling confirm() with the booth NAME
interpolated straight into the JS string literal. Jinja's autoescape is
HTML-attribute escaping, not JS-string escaping: the browser decodes the
entity back to a quote before the JS parser ever sees it, so a booth name
crafted to close that string executed on submit. Booth names are
agent-authored — making a folder under the data dir is the whole API — so
that is a live path, not a theoretical one.
The name now travels as a DATA ATTRIBUTE, where escaping is escaping, and
never reaches a JS string literal. Same pattern the board controls already
use. With JS off the form submits without a prompt, which is what every
no-JS browser here already did. */
(function () {
var WORDS = {
release: function (n) {
return 'Release \u201c' + n + '\u201d?\n\nIt moves to the ephemeral lane so you '
+ 'can wipe it from there. Nothing is deleted by this step.';
},
'wipe-kept': function (n) {
return 'WIPE the KEPT booth \u201c' + n + '\u201d?\n\nThis deletes it and its files '
+ 'immediately. Kept booths are the ones nothing else will clean up, so nobody '
+ 'else is going to do this for you \u2014 and nothing brings it back.';
},
wipe: function (n) { return 'Wipe booth \u201c' + n + '\u201d?'; }
};
document.addEventListener('submit', function (ev) {
var form = ev.target.closest ? ev.target.closest('form[data-confirm]') : null;
if (!form) return;
var word = WORDS[form.getAttribute('data-confirm')];
if (word && !confirm(word(form.getAttribute('data-booth') || ''))) ev.preventDefault();
}, true);
})();
</script>
{% endblock %}
+2 -1
View File
@@ -1,4 +1,5 @@
{% extends "base.html" %}
{% from "_lifetime.html" import lifetime %}
{% block title %}{{ name }} · marks · The Booth{% endblock %}
{% block content %}
{# The marks page for a booth whose own index.html is served VERBATIM. That page
@@ -11,7 +12,7 @@
{# `marks_open` comes from open_marks() — the ONE openness predicate (INV-2).
This used to re-derive it in Jinja as `selectattr('answer', 'none')`, which
read a half-answered pick as closed. #}
<span class="sub">{% if marks_open %}<span class="badge badge-mark">{{ marks_open }} open</span> · {% endif %}{{ marks|length }} mark{{ '' if marks|length == 1 else 's' }}</span>
<span class="sub">{% if marks_open %}<span class="badge badge-mark">{{ marks_open }} open</span> · {% endif %}{{ marks|length }} mark{{ '' if marks|length == 1 else 's' }} · {{ lifetime(kept, hold, expires_in) }}</span>
</div>
{% if marks %}
{% include "_marks.html" %}
+10
View File
@@ -99,7 +99,17 @@
img.addEventListener('load', evaluate);
window.addEventListener('resize', evaluate);
if (img.complete) evaluate();
/* An arrow key inside the note field is a CARET move, not a navigation.
The handler is on `document` and the note textarea shipped into this same
page, so typing a note and reaching for ← threw the draft away; Escape
did it in one keystroke. Anything editable keeps its own keys. */
function isEditable(el) {
return !!(el && (el.isContentEditable ||
/^(input|textarea|select)$/i.test(el.tagName || '')));
}
document.addEventListener('keydown', function (e) {
if (isEditable(e.target)) return;
if (e.key === 'Escape') window.location.href = BACK;
else if (e.key === 'ArrowLeft' && PREV) window.location.href = PREV;
else if (e.key === 'ArrowRight' && NEXT) window.location.href = NEXT;
+67 -4
View File
@@ -208,10 +208,13 @@ def import_legacy_asks(booth: Path) -> list[Mark]:
`asks.html:11` — and three in Python — `app.py:271` (the index badge),
`app.py:750` and `app.py:751` (the verbatim-booth chip).
- **INV-3 — the judgment travels, like the caption.** U1's rule, extended:
every surface that renders an item renders that item's marks. Gallery tile,
zoom view, doc view. *Falsifiable:* fetch `/b/<n>/view?f=<img>` for a flagged
item carrying a note and assert both the flag state and the note text are in
the served HTML.
every surface that renders an item renders that item's marks. *Falsifiable,
once per surface* — the first draft named three surfaces and checked one, which
all four panel arms flagged as the document's strongest ambiguity: (a) the
**gallery tile** shows the flag control in its current state and the item's
notes; (b) the **zoom view** `/b/<n>/view?f=<img>` carries the flag state and
the note text; (c) the **doc view** `/b/<n>/view?f=<doc>` carries the note
text. Three tests, not one.
- **INV-4 — the pick semantics are byte-identical.** `build_answer` produces,
for every input, the document `write_answer` produced. *Falsifiable:* the
existing `test_asks.py` answer assertions pass against `build_answer` with
@@ -275,6 +278,66 @@ out here because it is a visible change to what the index shows, it is the kind
of thing that looks like a bug when it lands, and the operator should get to
veto it rather than discover it.
## Cross-frontier contract panel — 2026-09-22, four arms, artifact-only
`/heid-contract-review` panel (Gróa / Hulda / Regin / Kimi), thread
`01M33VSNFER4N1554G0Y0VC9C8`, dispatched before implementation and triaged after
it. Every quoted passage was verified verbatim by Heid; no arm fabricated an
identifier. Triaged per the five-category rule — what follows is the disposition,
not the reply.
**Three of these were defects in shipped code, not ambiguities in prose.** v0.2.0
was already tagged and announced to 15 handles when they landed.
| finding | arms | category | disposition |
|---|---|---|---|
| **A write over a corrupt `.marks.json` silently replaced every mark in the booth.** The read path is deliberately lenient (unparseable → `[]` so the page loads); the write path inherited that through the same reader, so one flag click appended to an empty list and atomically replaced the file. | Kimi F2, Hulda F3 | **1 — genuine add** | **FIXED.** `MarksCorrupt`, raised by a strict `_read_raw_strict` used only by the write path. Read stays lenient, write goes strict; the damaged bytes are left on disk. Routes return 409, not 500. Reproduced first, then fixed. |
| **`declare_pick` had no `target`**, so a pick could not be attached to an item — though `Mark.target` carried one, `marks_for_target` retrieved it, and `_marks.html` already rendered "on \<item\>". | Hulda F1, Regin | **1 — genuine add** | **FIXED.** `declare_pick(..., target=None)`, validated like every other target. A re-declaration may move it. |
| **The importer stranded a legacy answer.** A stem already present as a mark was skipped wholesale, so a re-declared-but-unanswered pick with the operator's choice sitting in `<stem>.answer.json` lost that choice permanently — reads are forbidden from looking at sidecars. | Gróa F10 | **1 — genuine add** | **FIXED.** The declaration is still skipped (idempotence), but a legacy answer is ADOPTED when the existing mark is an unanswered pick. An answer made through marks is never overwritten. |
| **INV-3 names "doc view" as a protected surface; nothing tested it.** Shipping the doc view unmarked would have passed. | 4/4 — the panel's strongest convergence | **1 — genuine add** | **TEST ADDED.** The behaviour was already implemented; the gate caught that nothing held it. INV-3's falsifiable below now covers all three surfaces. |
| **The broken-declaration path is three different doors and none is written:** validate-before-write, stored-raw-with-read-time-error, and unparseable-file-yields-`[]`. | 4/4 | **1 — genuine add, prose only** | **PINNED below.** All three are real and distinct cases; the code always handled them separately. The contract conflated them. |
| **"What counts as open" is defined three ways** across assumptions, the signature comment and a test row. | Gróa F1, Regin F5, Hulda F4 | **1 — genuine add, prose only** | **PINNED below.** Code and tests were already correct (partial = open). |
| **INV-2 and INV-5's checks comply in letter:** openness can be re-derived as `(answer or {}).get("complete")` with the grep still green; `importlib` inside a function defeats the AST walk. | Gróa F5/F7, Kimi F4/F5 | **4 — out of place** | Accepted as true and NOT closed. Both describe a future careless change, and the honest statement is that these checks raise the cost of drifting rather than making it impossible. Recorded rather than papered over. |
| **INV-6 named `_with_marks(booth)`; the code has `_Locked`.** And its falsifiable makes the mandated helper unimplementable, since the helper must itself call `os.replace`. | Gróa F8, Kimi F8 | **2 — sharpening** | **FIXED below** — the name and the exemption. |
| `set_flag`'s annotation forbids a booth-level flag; never stated as a decision. | Regin F6 | **2 — sharpening** | It IS a decision: a flag means *this one*, so it needs an item. Stated in the signature. |
| "Cleaning" note text is defined by example only. | Kimi F6, Hulda F6 | **2 — sharpening** | `_clean_text` is CRLF-normalize, strip, truncate at `TEXT_MAX`. Documented at the function. |
| INV-1 self-conflict: the rule allows one function, the check and assumptions exempt `booth_items`' name check. | Gróa F6 | **3 — settled prior** | Already resolved by the seam review (SR-9): the exemption is a NAME check, never a content read. |
**The methodology note the arms volunteered, which is worth more than any single
flag:** this contract's own frontmatter carries a plain-language narrative, so the
paraphrase half was partly re-reading the author's framing back to him. Regin and
Kimi both said the stronger shape for a narrative-heavy contract is the ambiguity
pass with the paraphrase cut to a drift-check. That is a finding about the
*mechanism*, not this document, and it belongs in the skill rather than here.
### The three pinnings
**Openness, normatively, once.** A mark is open when `shape == "pick"` **and** it
has no `error` **and** (`answer is None` **or** `answer["complete"]` is false). A
**partially answered pick is OPEN.** Every other sentence in this document about
openness is descriptive; this one governs, and `open_marks` is its only
implementation.
**A broken declaration, normatively — three distinct cases, not one.**
1. `declare_pick` validates through `normalize_ask` and **raises `AskError`
before writing anything.** A session cannot land a refused question. The
function never returns an invalid mark.
2. A declaration that is invalid **in the stored file** — reachable via the
importer, or a hand-edit — is hydrated with `error` set and is rendered, so a
question the session believes it posted is never silently hidden. It is not
open (it can never be answered), and it cannot be answered: `answer_pick`
re-validates and raises.
3. A **whole file** that cannot be parsed is not a broken declaration. `marks_for`
returns `[]` so the page loads; every WRITE refuses with `MarksCorrupt`.
**INV-6, corrected.** Every writer goes through the one `_Locked(booth)` context
manager, which holds an exclusive flock on `<booth>/.marks.lock` across read,
mutate and atomic replace. *Falsifiable:* no function outside `_Locked` calls
`_write_raw` or `os.replace` on the mark file. (The first draft named a
`_with_marks` helper that does not exist, and forbade the very calls the helper
must make.)
## Slices
Vertical, each one shippable and green before the next starts.
@@ -0,0 +1,606 @@
---
contract_version: "1.0"
module: "booth.app (lifetime)"
purpose: "A booth's lifetime stops being a boolean somebody remembered to press and becomes a fact derived from the booth's own state. Today there is ONE lifetime (24h from the newest mtime in the tree) and ONE escape hatch (`.forever`), and the measurement says the escape hatch is carrying the main load: 17 of 24 live booths (70%) hold the sentinel, up from the 13 of 24 (54%) counted on 2026-09-21. That is not `ephemeral with an exception`; it is two lifetimes wearing one lifetime's clothes, with the operator doing the sorting by hand. This unit adds the two facts the sweeper was missing -- a booth the operator still owes an answer to is HELD, and looking at a booth is ACTIVITY -- so the cases that were pressing `.forever` for `not yet` stop needing it, and `keep` is left meaning only what it says: this is durable."
depends_on:
- "booth.marks (`hold_read` -- ADDED BY THIS UNIT, the one-read pair the lifetime rule needs; and `open_marks` -- THE openness predicate, built for this unit and saying so in its own docstring: `Open is the reading that makes U4 correct: a lifetime rule that unpinned a booth on the first radio click would sweep a review in flight.` U4 CALLS it and does not re-derive it. Also `marks_for` (lenient read, never raises) and `read_error` (strict read, total -- it catches its own `MarksCorrupt` and returns a string). Verified against booth/marks.py, not against U2's contract prose: `marks_for` is `_read_raw` + `_hydrate_safe` + sort at marks.py:514; `read_error` is `_read_raw_strict` in a try/except at marks.py:262 and has no raising path.)"
- "booth.items (the dotfile skip in `booth_items` at items.py -- `.viewed` is excluded from tiles, counts and zips by the EXISTING `p.name.startswith('.')` rule, exactly as `.marks.json`, `.booth.json` and `.forever` are. No new exclusion is added or needed.)"
language: "python"
complexity: "medium"
estimated_loc: 130
used_by:
- "booth.app.sweep_once (gains the hold check beside the keep check -- the one place reaper policy lives)"
- "booth.app.list_booths (the index card gains `held` and `marks_error`, so the card can say WHY it is not counting down)"
- "booth.app.booth_view / booth_view_file / booth_marks_page (each records a view; `/b/<n>/asks` is a 308 redirect into the last of these and so needs no call of its own)"
- "booth.app.booth_unkeep (release is activity -- stated, where it used to be an accident of directory mtime)"
- "booth/templates/index.html, booth/templates/booth.html (the lifetime line: `expires in X` / `held until answered` / `kept`)"
touches:
- "booth/app.py (VIEW_MARKER, record_view, is_held; sweep_once, list_booths, booth_view, booth_view_file, booth_marks_page, booth_unkeep; the module docstring's lifetime paragraph)"
- "booth/templates/index.html (the ephemeral card's sub-line becomes a three-state lifetime line)"
- "booth/templates/booth.html (the same three-state line in the boothhead)"
- "booth/templates/_lifetime.html (new -- the lifetime macro, defined ONCE and called from three surfaces. Not in the first draft of this inventory: a four-state conditional repeated three times is the blurtoggle lesson, and U5 had already established the partial as the house answer.)"
- "booth/templates/marks.html (INV-4's third surface. A verbatim booth has no Booth-rendered header, so without this the booths most likely to be HELD -- a report that asks something -- would be the ones that never say so. Found by looking at the live service, not by the suite.)"
- "booth/templates/base.html (one CSS rule for the held state)"
- "scripts/booth (the header's `THE 24h RULE AND ITS ONE EXCEPTION` block, which states the old doctrine as the whole doctrine, and the `DO NOT unkeep and let it expire` block. The WARNING STAYS AND STAYS TRUE -- release still buys a full TTL, so unkeep-and-wait is still a delay rather than a delete. What changes is that it stops being phrased as a surprise about directory metadata and starts being phrased as the rule it now is. The paraphrase panel read the touches line as possibly meaning the advice was being retired; it is not.)"
- "README.md (the TTL paragraph)"
- "CLAUDE.md (invariant 2's dotfile list gains `.viewed`)"
- "tests/test_lifetime.py (new)"
- "tests/test_booth.py (ONE cross-reference comment. The draft said the two release-clock tests would gain an assertion that the marker is written; implementation showed they must not. `test_releasing_a_board_RESETS_its_ttl_clock` unlinks the sentinel BY HAND, not through the route, so it is a test of the mtime mechanism and asserting a route side-effect in it would be testing the wrong thing. The route behaviour is `test_releasing_a_board_RECORDS_A_VIEW` in the new file; the comment points at it. No existing assertion is touched.)"
assumptions:
- "A VIEW IS RECORDED AS A DOTFILE, AND THE EXISTING AGE RULE READS IT. `.viewed` is a dotfile but NOT a `.lock` dotfile, so `_newest_mtime` already counts it (app.py:220 excludes only `.<name>.lock`). There is therefore NO new arithmetic in `booth_age_seconds`, `is_expired` or `expires_in`: `age = now - newest mtime in the tree` is unchanged, and a view is simply one more thing in the tree. One mechanism, not two. This is the same reason `.booth.json` needed no integration work in U5."
- "THE LOCK EXEMPTION IS WHY THIS IS SAFE. `_newest_mtime` excludes `.<name>.lock` because those are created by a READ-MODIFY-WRITE path, including one that changes nothing -- machinery, not activity. `.viewed` is the opposite: it is written only by a deliberate GET of a booth's own page. The exemption's rule (`machinery does not count, deliberate acts do`) is unchanged and this lands on the counted side of it."
- "RECORDING A VIEW MUST NEVER FAIL THE REQUEST. `record_view` swallows `OSError` -- a read-only mount, a booth owned by another uid, a full disk. The same posture `marks._Locked.__enter__` takes on its `os.utime` and for the same reason, stated there: `Not putting the clock back is a cost this module can absorb; not answering the request is not.` A booth that cannot record a view simply expires on its content mtime, which is today's behaviour."
- "HOLD IS FAIL-SAFE, WHERE READS ARE FAIL-OPEN. `marks_for` is lenient by design -- a damaged `.marks.json` reads as no marks, because a review surface that will not render is worse than one that has lost an annotation. That trade is right for a RENDER and wrong for a DELETE: the same leniency on the sweep path would wipe the booth whose judgment we had just failed to read, artifacts and all. So `is_held` treats an unreadable marks file as held. Reads lenient, deletes strict -- the same asymmetry U2 established between `marks_for` and `_Locked`, extended to the reaper. `THE REAPER` IS THE WHOLE SCOPE OF `deletes strict`, and the paraphrase panel ranked the ambiguity here first by stake: a HAND delete is never strict. `booth rm`, `POST /b/<n>/delete` and `DELETE /b/<n>` take a booth held by unreadable marks exactly as they take a kept one, which is what gives that hold -- the one nothing releases on its own -- an exit at all. Strictness is a property of the TIMER, never of the operator."
- "THE LIFETIME DECISION COMES FROM ONE READ, and that is a correction to this contract's first draft. The draft specified `is_held(marks_for(child), read_error(child))` -- two reads, presented as one answer. They are not: a write or a repair landing between them yields a pair that described the booth at no instant, and the losing pair is `([], None)` -- no marks and no error -- which is exactly the pair that DELETES. Hulda found it on the paraphrase round (2026-09-22) and it is the finding that changed code rather than prose. `booth.marks.hold_read(booth) -> (marks, error)` is the fix: one strict read answering both questions the lifetime rule asks, so `sweep_once` now does ONE read per booth per tick rather than two. And because `_read_raw_strict` RAISES rather than dropping an entry, a non-raising strict read returns exactly what the lenient read would -- so the index uses that same one read for its badge too, falling back to `marks_for` only on the error path, where leniency is the point."
- "AN OPEN PICK HOLDS; A NOTE OR A FLAG DOES NOT. `_is_open` returns False for every shape but `pick`, and False for a pick carrying `error`. That is already correct for U4 and is NOT changed here: a note is the operator's output, not an owed answer, and a pick that hydrated broken can never be answered, so holding a booth on one would be holding it forever for nothing (the CLI already spells that case as exit code 4). A PARTIALLY-answered pick IS open and DOES hold -- operator-settled 2026-09-21, and the reason `open_marks` exists rather than an `answer is None` test."
- "THE HOLD IS UNBOUNDED, AND THAT IS THE POINT -- BUT IT MUST BE VISIBLE. A booth with an unanswered pick is never swept, however old. This is a new way for a booth to become immortal, and it is deliberate: unanswered is unfinished. What makes it safe is not a bound, it is VISIBILITY plus TWO exits that already exist. The card and the booth header say `held until answered` in place of the countdown, so a booth that is not counting down always says why; and `booth rm` / the UI `x` delete a held booth exactly as before -- `sweep_once` is the only caller that honours a hold, precisely as it is the only caller that honours `is_kept`."
- "`keep` IS UNCHANGED AND KEEPS ITS LANE. `.forever` still exempts, still renders in the kept lane, still round-trips through `booth keep` / `booth unkeep` and the UI. U4 does not deprecate it, narrow it or add a reason field to it. The prediction is that its RATE falls because the `not yet` cases stop needing it -- and a prediction is falsified by measuring, not by removing the thing being measured."
- "THE MTIME-RESTORE RACE IN `marks._Locked.__enter__` IS EXPLICITLY CONSIDERED AND LEFT OPEN. The bug-hunt panel flagged it and it was held for U4 because closing it means changing TTL doctrine. U4's answer is that the doctrine stands: the clean fix (ignore a booth directory's own mtime whenever the booth holds anything) would close a two-syscall window that opens ONCE per booth ever, and would in exchange break every `rsync -a` populated booth -- which preserves source mtimes and so has ONLY the directory's freshness to look alive by, and which is the documented path for every host that is not nh3-dev. That is a larger hole than the one being closed. Decided, not deferred; see the OUT OF SCOPE section."
open_questions:
- "Whether `booth ls` should mark held booths the way it marks kept ones with a star. Cheap, and it would need `is_held` (or a stdlib-only sibling) reachable from the CLI. Sessions already have `booth marks`, which answers the same question about their own booth, so this is convenience rather than capability. Parked, not designed."
- "Whether a booth held ONLY by an unreadable `.marks.json` should surface on the index as something to repair, beyond the `marks unreadable` label. It is a held booth that nothing will release, which is the one case where the unbounded hold has no natural exit. The label makes it visible; a repair affordance is a different unit."
---
# U4 — derived lifetime
## The defect, stated precisely
> **One lifetime (24h from last touch) and one shape (a folder), serving five
> jobs with different lifetimes.** — `docs/design/information-architecture.md`
`.forever` is the escape hatch for that mismatch, and the measurement says it is
no longer an exception:
| date | booths carrying `.forever` | rate |
|---|---|---|
| 2026-09-21 (IA doc) | 13 of 24 | 54% |
| 2026-09-21 (re-count) | 14 of 25 | 56% |
| 2026-09-22 | **17 of 24** | **70%** |
Both the rate and the absolute count rose, so this is not the denominator
shrinking as the sweeper ran. A boolean that 70% of the population sets is not
an exception, it is the default with extra steps.
The reason it gets pressed is that it is the only way to say any of these:
| what the operator means | what he has to press |
|---|---|
| "this is a durable reference" | `.forever` |
| "I have not answered the question yet" | `.forever` |
| "I am still looking at this" | `.forever` |
Only the first is what `keep` means. The other two are facts the service already
holds and does not consult: **there is an open pick in `.marks.json`**, and
**somebody just loaded the page**. U4 consults them.
### The diagnosis has a live positive control
Counted 2026-09-22 against `~/booth-data`. A census of the whole population, not
a sample, and every value is a deterministic file fact (existence, mtime) — so
one observation per booth is the measurement, not an anecdote. The population
churns (26 -> 24 over the previous session); re-count rather than trusting these.
| | |
|---|---|
| live booths | 24 |
| carrying `.forever` | 17 (70%) |
| carrying `.marks.json` at all | 4 |
| of those, with an open pick | **4 of 4** |
| **open pick AND `.forever`** | **3** |
Three of the four booths in the fleet that are waiting on an answer have ALSO
been pinned by hand. That is the "not yet" case, caught in the act: the operator
pressed the durable-reference sentinel because there was no other way to say
"do not take this, I have not answered it". U4 makes those three stop needing it.
The staleness distribution says the same thing from the other side. Of the 17
kept booths, **10 are under ONE day old** — younger than the TTL, so the
sentinel has bought them nothing yet and was pressed pre-emptively. (An earlier
draft of this paragraph said "12 under 1.5 days" and called that younger than
the TTL; 1.5 days is not younger than 24 hours, and the claim only holds at the
one-day line. Caught by the cross-frontier paraphrase panel, 2026-09-22 — the
measurement was right and the sentence was not.) Only 4 are old enough
(2.4-4.6 days) that `keep` is the only reason they still exist. A
sentinel pressed on a booth that was in no danger is not a durability decision;
it is "not yet", written in the only vocabulary available.
⚠ The hold's live blast radius is SMALL today — 4 booths have marks at all. The
17-to-something prediction therefore rests on both halves of this unit, and on
the sentinel becoming unnecessary rather than becoming forbidden. If the rate
does not move, the honest readings are: the diagnosis was wrong, OR the habit
outlived the need, and the fortnight re-count cannot tell those apart on its
own. The three open-pick-plus-`.forever` booths are the ones to watch, because
for them the mechanism is now unambiguous.
## The record
A booth is in exactly one lifetime state, decided in this order:
```
KEPT .forever present never swept (unchanged)
HELD an open pick, or a never swept while (new)
.marks.json we cannot read that holds
EPHEMERAL otherwise swept when
age > ttl (unchanged)
```
`age` is unchanged: `now - _newest_mtime(booth)`, the newest mtime in the tree
excluding `.<name>.lock`. **Viewing is folded in through that existing rule**,
not beside it — a view writes `.viewed`, which is a dotfile and not a lock
dotfile, so the age function already counts it. There is no new arithmetic.
### What counts as a view
One line, because CLAUDE.md invariant 6's test applies to rules as well as
orders: **a deliberately-requested response FROM a booth's own page route is a
view; a machine read, an asset fetch, and a request that does not resolve are
not.**
Three words in that rule are load-bearing and the first draft said "HTML page",
which was wrong twice. `?download=1` is a zip served by the booth-page route and
IS a view — the operator asking for the whole booth is as deliberate as looking
at it. And a `/view?f=<missing>` that 404s is NOT one: the route matters, but so
does whether anything was served, or a crawler walking dead zoom URLs holds a
booth open forever. `record_view` therefore sits below the zoom route's file
validation and above the booth route's verbatim/zip fork.
| route | view? | why |
|---|---|---|
| `GET /b/<n>/` | **yes** | the booth page — gallery, verbatim report, or `?download=1` zip |
| `GET /b/<n>/view?f=…` | **yes** | the zoom / doc page; a bookmarked zoom URL is somebody looking |
| `GET /b/<n>/marks` | **yes** | the standalone judgment page — for a verbatim booth this IS the booth page |
| `GET /b/<n>/marks.json` | no | a session polling. An agent must not be able to hold its own booth open |
| `GET /b/<n>/<file>` | no | issued BY the page. A hotlinked image would otherwise keep a booth alive |
| `GET /` | no | the IA's rule: "deliberate act, so it cannot be triggered by browsing the index" |
| `GET /healthz` | no | a monitor is not a viewer |
⚠ Named rather than hidden: `scripts/layout-probe.py` sweeps every booth page,
so running it resets every booth's clock. That is the correct reading of the
rule (it is a GET of every booth page), it is recoverable (one extra TTL), and
it is a dev tool. A note goes in the probe.
⚠ A browser that speculatively prefetches a hovered link records a view the
operator did not quite take. Accepted: the failure mode is a booth living one
extra day because he nearly opened it, and the alternative is sniffing
`Sec-Fetch-*` headers, which is a fragile rule pretending to be a crisp one.
**Checked, because it would have been silent:** nothing in the fleet polls a
booth *page*. Homepage's `siteMonitor` for the Booth is
`http://10.100.10.50:8090/healthz`, which is on the not-a-view list; there is no
cron entry and no systemd timer touching `/b/…`. Had Homepage been pointed at a
booth URL instead, every booth would have become immortal on deploy and nothing
would have reported it.
### Release is activity, on purpose
Removing `.forever` bumps the booth directory's mtime, so a released board
survives another full TTL. Today that is an **accident** of directory metadata
that `app.py` documents as "not intuitive" and `scripts/booth` warns against.
U4 does not change the behaviour and does not retire the test that pins it. It
changes the behaviour's *reason*: `booth_unkeep` calls `record_view`, so a
released board gets one full TTL because **releasing a board is somebody
touching it**, which is a rule, and no longer because of which syscall happened
to write a directory entry, which is not.
The existing tests (`test_releasing_a_board_RESETS_its_ttl_clock`,
`test_released_board_is_sweepable_once_it_ages_again`) are untouched, and that
is a correction to this contract's first draft, which said they would each gain
an assertion that the marker is present. They must not: the first one unlinks
the sentinel **by hand**, not through the route, so it is a test of the mtime
mechanism and a route side-effect does not belong in it. The route behaviour
gets its own test in the new file, and the old test gains a comment pointing at
it.
**The marker's mtime must be NOW**, which `Path.touch()` gives and which the
contract's first draft left unsaid. An implementation that wrote the file with
any older timestamp would satisfy "the marker is there" while the extra TTL
still came from the directory-mtime accident this section exists to replace —
the new reason would be decoration over the old mechanism. Flagged by the
paraphrase panel, 2026-09-22.
## Signatures
```python
# booth/app.py
VIEW_MARKER = ".viewed"
"""Records the last deliberate look at a booth. A dotfile, so `booth_items`
skips it and it costs nothing in counts, galleries or zips — and NOT a `.lock`
dotfile, so `_newest_mtime` counts it and the existing age rule picks up the
view with no new arithmetic."""
def record_view(booth: Path) -> None:
"""Note that somebody deliberately looked at this booth.
Touches VIEW_MARKER; `_newest_mtime` does the rest. NEVER raises: a
read-only mount, a booth we do not own or a full disk cost the timestamp,
not the page. A booth whose view cannot be recorded simply ages on its
content mtime, which is today's behaviour for every booth.
"""
def is_held(marks: Sequence[Mark], error: str | None) -> bool:
"""True if this booth still owes the operator an answer and must not be swept.
PURE — it takes the result of a read and does none of its own, so the index
card and the sweeper cannot answer differently about the same booth. That
is U1's rule (one resolver, every surface reads the record) applied to
lifetime.
FAIL-SAFE on `error`. `marks_for` is lenient because a review page that
will not render is worse than one missing an annotation; the same leniency
on the DELETE path would wipe the booth whose judgment we had just failed
to read. Reads lenient, deletes strict.
Openness itself is `open_marks` and nothing else (U2 INV-2).
"""
return error is not None or bool(open_marks(marks))
```
`is_expired` is **unchanged** and stays a pure age question — the existing
separation ("expiry arithmetic and reaper policy are kept apart so they cannot
drift into each other") is the reason `is_kept` is not consulted there either.
`sweep_once` remains the only caller that honours a pin, and now honours two.
```python
# booth/marks.py — stdlib only, like the rest of that module
def hold_read(booth: Path) -> tuple[list[Mark], str | None]:
"""ONE read of `.marks.json`, answering BOTH questions the lifetime rule
asks: what is still open, and whether the file could be read at all.
Two calls would read the file twice, and two reads of one file are not one
read of one state — the pair that loses the race is `([], None)`, which is
the pair that deletes.
On a clean file the marks are what `marks_for` would return, because
`_read_raw_strict` raises rather than dropping an entry. So one read serves
the badge too, and the lenient reader comes back only on the error path.
"""
```
```python
def sweep_once(data_dir, ttl_seconds, now=None) -> list[str]:
...
if is_kept(child):
continue
if is_held(*hold_read(child)): # NEW — ONE read
continue
if is_expired(child, ttl_seconds, now):
shutil.rmtree(child)
```
```python
def list_booths(data_dir, ttl_seconds, now=None) -> list[dict]:
...
marks, marks_error = hold_read(child) # NEW — one read, both facts
if marks_error is not None:
marks = marks_for(child) # lenient, for the panel
booths.append({
...
"marks_error": marks_error, # NEW — the card says so
"held": is_held(marks, marks_error), # NEW — the same predicate
})
```
## What renders
The lifetime line, on the ephemeral index card and in the booth header. Three
states, one of which is new:
| state | line | why |
|---|---|---|
| ephemeral | `12 items · expires in 3h 20m` | unchanged |
| held, open pick | `12 items · held until answered` | says what holds it AND what releases it |
| held, unreadable | `12 items · held · marks unreadable` | the one hold nothing will release on its own |
| kept | `12 items · kept` | unchanged, kept lane |
**The hold REPLACES the countdown at every age, not only once the booth is
old.** A held booth that is four hours old shows `held until answered`, not
`expires in 20h`. `expires_in` is still computed and still correct (INV-1);
it is simply not what the surface says, because a number counting down to a
deletion that will not happen is the silent-stopped-clock failure in its other
costume — the screen announcing an expiry the sweeper will never carry out.
Flagged as readable-two-ways by the paraphrase panel, 2026-09-22; settled here.
A booth that is not counting down **always says why**. That is the whole safety
argument for an unbounded hold: `.forever` was at least visible as a lane; an
invisible rule that silently stops the clock would be strictly worse than the
boolean it replaces.
**Three surfaces, not two**, and the third was found by looking at the live
service rather than by the suite. A verbatim booth's own `index.html` is served
untouched by design, so it has no Booth-rendered header for the line to live in
— and a report that ASKS the operator something is the archetype of a held
booth. `GET /b/<n>/marks` is the only other page whose chrome the Booth owns, so
the line goes there too. Without it, the booths most likely to be held would be
exactly the ones that never said they were. (U3 is the unit that gives a
verbatim booth real chrome; until then, this is the honest coverage.)
Kept beats held in the display, because a kept booth is in the kept lane and is
exempt either way — showing two reasons for one exemption is the
two-representations-of-one-state trap `flag_id`'s docstring names.
**An unreadable marks file is the exception, and it rides along even on a kept
board**: `kept · marks unreadable`. Damaged judgment is not a second exemption,
it is a thing somebody has to go and fix, and the kept lane holds the durable
boards — the ones where losing the operator's marks costs most. A kept card that
said only `kept` would hide the single case that needs a human. The card's
`held` and `marks_error` are therefore RAW FACTS, true regardless of keep, and
only the display has a precedence. The paraphrase panel found the two readings
of "exactly one lifetime state" that this settles.
## Scope — the blast-radius pass
`graphify explain` on `sweep_once`, `is_kept`, `list_booths`, `_newest_mtime`,
`booth_age_seconds`, `open_marks`, `KEEP_MARKER`, cross-checked with grep.
Graphify reported the call structure and, as expected, **missed both route
callers of `list_booths`** (`index()` and `healthz()`, now at app.py:761 and :774) —
they are function-local inside `create_app`, which is the known AST blind spot.
Grep caught them. Neither tool alone was sufficient; this is the third unit in
a row where that has been true.
**Production, 7 files:** `booth/app.py`, `booth/marks.py` (`hold_read`, added),
`booth/templates/_lifetime.html` (new), `booth/templates/index.html`,
`booth/templates/booth.html`, `booth/templates/marks.html`,
`booth/templates/base.html`.
**Docs/CLI, 4 files:** `scripts/booth`, `scripts/layout-probe.py`, `README.md`,
`CLAUDE.md`.
**Tests, 2 files:** `tests/test_lifetime.py` (new), `tests/test_booth.py`.
⚠ This census said "Production, 4 files" in the first draft and omitted
`_lifetime.html`, `marks.html`, `marks.py` and `layout-probe.py` — three of
which the body text elsewhere required, which is the contradiction both
Gróa and Hulda flagged independently. An inventory that disagrees with the
prose next to it is worse than no inventory: it reads as a closed set.
Not touched, and checked rather than assumed: `booth/items.py`,
`booth/manifest.py`, `booth/links.py`, `booth/asks.py`, `booth/inline.py`.
## The three cross-frontier panels, and what they changed
All three ran on 2026-09-22 and all three earned their place — and each found
a class the other two could not. Triaged per the cross-frontier discipline
rather than adopted.
**Paraphrase panel** (`01M34VX0SH23Y3VC92E7GM4S70`, four arms). Seven flags.
Five folded into the prose above: the hold replacing the countdown at every
age, `deletes strict` scoping to the reaper alone, the zip and the 404 in the
view rule, the marker's mtime, and the CLI warning staying true. Two changed
more than wording:
- **Hulda — the two reads are not one state.** The only finding on this round
that changed CODE. See the `hold_read` assumption in the frontmatter.
- **Gróa and Hulda, independently — the blast-radius census contradicted the
prose beside it.** It named four production files while the body required
three more. An inventory that disagrees with its own document is worse than
none, because it reads as a closed set.
Hulda also caught a number: this contract claimed 12 kept booths were "under
1.5 days old — younger than the TTL". One and a half days is not younger than
twenty-four hours. The measurement was right, the sentence was not, and it is
the one place the diagnosis overstated itself.
**Code-vs-contract panel** (`01M34WAFJC3RTERFYBBZJN1SVG`, four arms). **All
four arms found the same drift** — the strongest signal either panel produced
on this unit. The booth header's sub-line forks on `{% if board %}`, and the
lifetime macro sat only in the `{% else %}`, so a booth carrying `links.md`
rendered a link count and nothing at all about its lifetime. INV-4 says the
templates have no path that renders neither; that was a path, reachable by the
release button or by a hand-made board.
Regin and Kimi recommended amending INV-4 to carve the board header out, on the
grounds that board-header layout belongs to U7. **Declined; the code is fixed
instead.** Cutting an invariant down to fit an implementation gap is the wrong
direction when the fix is one template edit, and U7 owns navigation and section
layout — not whether a header states a lifetime. Gróa's "fix it" was right.
The same panel showed that **most of the INV falsifier tests did not
discriminate**, which is the more useful half of the round. The header test
never rendered a board. The kept-beats-held test only rendered the index, where
kept cards took a hardcoded string and never reached the macro at all. The
INV-5 test called `record_view` directly instead of GETting the routes the
invariant is about. The INV-7 tests asserted the marker's absence rather than
the age, so a handler writing any other non-dot file would have passed. INV-6's
had no doomed sibling, so "spare everything" would have passed. Each is now
written to fail under the change that defeats it, and the board-header pair was
verified RED against the pre-fix template rather than assumed.
**Bug-hunt panel** (`01M34Y2R0RAJRSN36Q8K4KAB36`, four arms). The round that
changed the most code, and the one that found a class the other two could not
see by construction: **a read that FAILED still resolving to "no hold", and
therefore to a delete.** That is the invariant this unit declared to the panel,
and the panel found **four independent paths through it. No single arm found
all four.**
1. **An entry-level hydration error lost its hold.** `.marks.json` parses, one
mark fails normalization, `_hydrate_safe` returns a `Mark` carrying `error`,
and `_is_open` returns False for an errored pick — on purpose, because a
broken pick can never be answered. So the booth read as not-held and swept,
while the panel beside it rendered the broken mark in full. The fail-safe was
built for FILE-level damage and missed ENTRY-level. This is the strongest
finding of all three rounds.
2. **A present-but-blank `.marks.json` swept.** `_read_raw_strict` early-returns
for whitespace-only content — right for the write path it was written for,
wrong for the delete path. Our writer never produces a blank marks document,
so a blank one that exists is something that went wrong.
3. **`_newest_mtime` returned 0.0 when the booth's own stat failed**, making it
maximally ancient and therefore the FIRST thing the sweeper takes. Pre-dates
U4; U4 is what turned the age read into a life-or-death read.
4. **`is_kept` collapsed a stat failure into not-kept.** `Path.exists()` maps
ELOOP and EACCES to False, so a kept booth whose sentinel could not be
stat'd became sweepable.
**`is_held` is gone; `hold_reason` replaced it.** A boolean plus a separate
error string is two representations of one state, and Regin independently
flagged that the display could not distinguish the two holds. One function now
returns the REASON — `"open"`, `"unreadable"`, or None — and every surface reads
it off the same value the sweeper acts on. That closes findings 1 and Regin's
together, which is why it is a rewrite rather than an extra clause.
**Convergent, 3-of-4: `record_view` followed a planted symlink.** `Path.touch()`
follows an existing link, so a booth carrying `.viewed -> /anywhere` turned every
page view into an mtime write at an arbitrary path under the service uid — and
any fleet session can write into a booth, because making a folder is the whole
API. Now an `O_NOFOLLOW` create plus `os.utime(fd)`, so a planted link raises
ELOOP into the existing swallow and view-recording quietly stops for that booth.
The `utime` is also what makes the marker read as NOW, which this contract
already required and `O_CREAT` alone does not do.
**Two more the panel found in code this unit touched:**
- **`?f=.marks.lock` held a booth open.** The zoom route recorded a view for any
path that stats inside the booth, including a lock file the service created
itself. `record_view` now sits below `find_item` and fires only for a real
item — which also makes the comment beside it true, where before it claimed
more than the code did.
- **Releasing an ALREADY-released booth refreshed its TTL forever.** The
unconditional `record_view` on `unkeep` contradicted that route's own no-op
promise and diverged from the CLI, which removes the sentinel without
recording anything. Now gated on something actually having been released. The
same edit fixes a pre-existing 500: a `.forever` that is a DIRECTORY raised
`IsADirectoryError` straight through the route, which made the card's release
button permanently dead for that booth.
**Also fixed: a docstring this unit's own fix made stale.** `sweep_once` still
claimed "one lenient read plus one strict read" after `hold_read` reduced it to
one. Kimi's framing is the right reason to care — a maintainer "optimizes" back
to two calls on the comment's authority, and rebuilds the seam the function
exists to kill.
**Re-declared as parked, not adopted:** Regin distinguished a stale-DECISION
window (hold checked, then rmtree) from the torn-FILE race already parked at
`park/booth-sweeper-rename-then-delete-to-close-the`. The distinction is real
and the fix is the same rename-then-delete, so it parks with its sibling.
**Five pre-existing defects the panel surfaced in touched files** — a booth name
reaching a JS string context, an unguarded `links.md` read, an index sort with
no tie-breaker, `marks.json` reporting damage as empty success — are fixed in
their own commit rather than smuggled into this unit's. See that commit.
⚠ **The capture tooling failed silently and the panel caught it, not us.** The
`files/` tree shipped to the arms was EMPTY: the snapshot loop iterated `for f
in $IN` over a multi-line variable, and **zsh does not word-split unquoted
parameter expansions** the way bash does, so it ran once against a path that was
the whole list. jekyll recovered by re-applying the bundled diff to HEAD and
verified every file byte-identical, so the round is sound — but the failure mode
is the dangerous one: an empty bundle reads exactly like a clean result.
## Seam review — against the real module surface
Checked against `booth/marks.py` itself, not against U2's contract prose.
| borrowed | real surface | verdict |
|---|---|---|
| `open_marks(marks)` | `marks.py:542`, takes `Sequence[Mark]`, returns `list[Mark]` | matches |
| `marks_for(booth)` | `marks.py:514`, `_read_raw` + `_hydrate_safe` + sort; total | matches |
| `read_error(booth)` | `marks.py:262`, returns `str \| None`, catches its own `MarksCorrupt` | matches — **and it is total**, which `is_held`'s fail-safe branch depends on |
| `_is_open` semantics | `marks.py:525`: `pick` only, `error is None`, partial counts open | matches the assumption above |
| `_newest_mtime` lock rule | `app.py:220`: skips `p.name.startswith(".") and p.name.endswith(".lock")` | `.viewed` is counted — confirmed at the source, not inferred |
| `Mark` import in app.py | app.py:145-160 imports `open_marks`, `marks_for`, `marks_for_target`, `as_dict` — **not `Mark`** | `is_held`'s annotation needs `Mark` added to that import list |
| `read_error` import in app.py | **not imported either** — U2 left it to the CLI, which is its only caller today | must be added to the same block; U4 is its first in-service consumer |
| `zip_booth` dotfile skip | `app.py:445`ff: `p.is_file() and not p.name.startswith(".")` | `.viewed` never reaches a zip — confirmed, not inferred from `booth_items` |
| `booth_items` dotfile skip | `items.py:182`: `not p.is_file() or p.name.startswith(".")` | `.viewed` is not an item |
| `GET /b/<n>/asks` | `app.py:1105`, a **308 redirect** to `/marks`, not its own render | records a view through the `/marks` handler. No separate call, and adding one would double-count |
| route concurrency | `booth_view`, `booth_view_file`, `booth_marks_page` are all `def`, not `async def` | FastAPI runs them in a threadpool, so `record_view`'s write cannot block the event loop |
Three rows of that table are the kind of thing only this pass finds: the cold
panel reads one contract, and a signature that is fine in isolation says nothing
about whether the name it needs is in scope at the call site.
**SR-1 — why `read_error` is safe to call per booth per index load, which the
signatures alone do not say.** `_read_raw_strict` checks `S_ISREG` *before* it
calls `read_text` (marks.py:236). That ordering is the v0.2.2 fix: `st_size` is
0 for a FIFO and 0 for a symlink to `/dev/zero`, so a size cap alone lets both
through and `read_text` then either blocks with no EOF or allocates until the
kernel intervenes — across every booth, on `GET /`, which is a service-wide
hang rather than one bad card. U4's decision to spend a second read on the hot
path depends on that guard already being there. It is; checked at the source.
## Out of scope
- **A bound on the hold.** An abandoned pick holds its booth forever. Detecting
"abandoned" needs state the Booth does not have (is any session still
polling?), and the honest alternative — an arbitrary N-day cap — trades a
visible immortal booth for a silent deletion of an open question. Visibility
plus `booth rm` is the answer for v1.
- **A reason string on `.forever`.** "keep survives as an explicit, reasoned
pin" is read here as *a pin the operator reasoned about*, not *a pin carrying
a recorded reason*. A `why` on keep does not close the measured defect — the
70% is people using keep for things that are not keep, and this unit gives
those things their own mechanism. Parked per the anti-creep gate.
- **A third index lane for held booths.** A booth waiting on the operator is the
most actionable thing on the index, and it already carries the `? N open`
badge. Lane structure and ordering are U7's, and adding a lane here would set
an ordering rule that U7 then has to live with.
- **Closing the `marks._Locked.__enter__` mtime-restore race.** See the
assumption above: the clean fix costs every `rsync -a` populated booth. The
comment there stays, and stays accurate.
- **`booth ls` marking held booths.** Open question, parked.
- **Closing the view-during-sweep race, which U4 WIDENS.** `sweep_once` calls
`shutil.rmtree` without holding anything, so a write landing inside that call
can make it raise partway and leave a stump directory. The race is
pre-existing — every write route has always had it — but U4 widens it,
because `record_view` fires on every booth-page GET and the case that
collides is precisely "the first look at a booth that has been silent for 24
hours", which is the state the sweeper acts on.
The fix is known and small: `os.rename` the booth to `.sweeping-<name>` first
(atomic, and a dotfolder the scan already skips), then `rmtree` the renamed
path, plus a cleanup of leftovers at the top of each tick for the
crash-between-the-two case. It is NOT done here, per the anti-creep gate:
both "in" and "park" are defensible, so it parks. The arithmetic is that the
collision needs a GET inside a ~10 ms `rmtree` on a booth nobody has opened in
a day, the sweeper ticks every 15 minutes, and the consequence is a stump that
survives one more TTL — against which a sweeper rewrite is not a v1-path
trade. Named here so it is a decision and not an oversight, and parked on
the henge at `park/booth-sweeper-rename-then-delete-to-close-the` (id 83)
so it has a home rather than only a paragraph.
## Invariants
**INV-1 — Age arithmetic is unchanged.** `booth_age_seconds`, `is_expired` and
the `expires_in` values on both surfaces are computed exactly as before. A view
enters through `_newest_mtime` as a file in the tree, not as a term in a new
formula. *Falsifiable:* a booth with a `.viewed` and a booth with any other
non-lock dotfile of the same mtime report the same age.
**INV-2 — `sweep_once` is the only caller that honours a hold.** `is_expired`
stays a pure age question; `booth rm`, `POST /b/<n>/delete` and
`DELETE /b/<n>` delete a held booth exactly as they delete a kept one.
*Falsifiable:* a held booth is still reported expired by `is_expired` and is
still deleted by the delete routes.
**INV-3 — One predicate, ONE READ, one answer.** The index card's `held`, the
booth header's, the marks page's and the sweeper's exemption all come from the
same pure `is_held`, and each call's two inputs come from a SINGLE read of
`.marks.json` via `hold_read` — never from two reads stitched together, which
is a pair that described the booth at no instant. *Falsifiable:* for any booth, what
`list_booths` reports as `held` and what `sweep_once` refuses to take agree —
tested directly rather than by inspection, because that is the falsifiable form.
(`open_marks` is still called directly for the `N open` COUNT. A count is not a
lifetime decision, and the first draft of this invariant forbade it by accident
— the rule is that no *exemption* and no *held label* is derived except through
`is_held`.)
**INV-4 — A booth that is not counting down says why.** Every non-kept booth
renders either a countdown or a named hold on **every surface whose chrome the
Booth owns**: the index card, the booth header, and the marks page (which is
the only one of the three a verbatim booth has). *Falsifiable:* the templates
have no path that renders neither, and the one line is a single macro rather
than three conditionals that can drift.
**INV-5 — Recording a view cannot fail a request.** `record_view` swallows
`OSError`. *Falsifiable:* a booth whose directory is read-only still returns 200
for its page, its zoom page and its marks page.
**INV-6 — An unreadable `.marks.json` holds its booth.** The reaper never
deletes judgment it could not read. *Falsifiable:* a booth with a corrupt
`.marks.json`, aged past the TTL, survives `sweep_once`.
**INV-7 — Machine reads do not hold a booth open.** `GET /b/<n>/marks.json` and
`GET /b/<n>/<file>` do not write `VIEW_MARKER`. *Falsifiable:* polling either,
repeatedly, leaves the booth's age untouched.
@@ -0,0 +1,405 @@
---
contract_version: "1.0"
module: "booth.manifest"
purpose: "A booth that says what it IS and who posted it. Today the index card shows a name, an item count and a countdown -- nothing about provenance or purpose -- so an agent that wants the operator to look at something has no way to make the booth say so, and posts a URL to the link board instead. That is job 5 (`Announce`), the job nobody named, and its absence is the measured cause of 145 dead link rows (69% of the board pointing at booths that no longer exist). This unit gives job 5 a home: each booth carries `.booth.json` -- `{handle, title, why, created}`, written by the CLI from `$ALTHING_HANDLE` -- and the index card and the booth page header render it. Enforcing the link rule WITHOUT giving job 5 a home first just makes it homeless; this is the home."
depends_on:
- "booth.items (the dotfile skip in `booth_items` -- `.booth.json` is excluded from tiles, counts and zips by the EXISTING `p.name.startswith('.')` rule at items.py:182, exactly as `.marks.json` is. No new exclusion rule is added or needed. Verified, not assumed: `test_a_manifest_is_not_an_item` asserts it.)"
- "booth.marks (the `_write_raw` shape only -- temp file + os.replace, per CLAUDE.md invariant 5. Copied as a pattern, NOT imported: manifest.py must not depend on marks.py, because the CLI imports each module on its own.)"
language: "python"
complexity: "low"
estimated_loc: 150
confidence: 0.85
used_by:
- "booth.app.list_booths (the index card gains `manifest` -- one file read per booth, alongside the `marks_for` read already there)"
- "booth.app.booth_view (the booth page header gains the same provenance line; a booth URL handed to the operator lands HERE, not on the index, and job 5 is literally 'operator, look at this')"
- "booth.app.upload (a pickup booth announces itself as the Booth's own)"
- "scripts/booth (`new` and `add` gain `--why` / `--title`; `link` announces the standing board)"
touches:
- "booth/manifest.py (new -- the record, the write, the lenient read)"
- "booth/app.py (list_booths gains one key; booth_view gains one key; the /upload path writes a manifest. It also adds MANIFEST_FILE to the `used` dedupe set -- CONSISTENCY, not a fix: SR-1 established the collision is unreachable because `safe_upload_name` strips leading dots, which is equally true of the `UPLOAD_MARKER` entry that has sat in that set since before this unit.)"
- "booth/templates/_provenance.html (new -- the provenance macro, defined ONCE and called from both index lanes and the booth header. Not in the first draft of this inventory: the implementation added the partial rather than repeating the four-state conditional three times, which is SR-6 plus the blurtoggle lesson, and the inventory lagged the decision.)"
- "booth/templates/index.html (the provenance line on both lanes' cards -- kept AND ephemeral, or the kept lane silently keeps the old defect)"
- "booth/templates/booth.html (the provenance line in the boothhead, and the h1 renders `title` with the directory name beside it)"
- "booth/templates/base.html (the .prov-* CSS)"
- "scripts/booth (`new` / `add` flag parse; `link` board announcement; usage string; the header doc block)"
- "tests/test_manifest.py (new)"
- "tests/test_marks.py (test_stdlib_only's parametrize list gains `manifest`)"
assumptions:
- "THE MANIFEST IS A DOTFILE, and that is the whole integration story. `booth_items` skips `name.startswith('.')` (items.py:182), `zip_booth` skips it (app.py:351), and the legacy ask scan skips it (marks.py:656). So `.booth.json` costs nothing in item counts, galleries, zips or migration, and needs no new exclusion anywhere. This is the same reason `.marks.json` needed none. Settled -- do not re-derive it."
- "WRITING A MANIFEST IS ACTIVITY. `.booth.json` is a dotfile but NOT a `.lock` dotfile, so `_newest_mtime` counts it (app.py:192 excludes only `.<name>.lock`). Creating or re-announcing a booth resets its TTL, which is correct: both are somebody touching it. The lock exemption exists for machinery that a READ path creates; this is a deliberate write."
- "THE READ IS LENIENT AND THE FAILURE IS VISIBLE. `list_booths` reads every booth on every index load, so a manifest that cannot be parsed must never raise -- that is the v0.2.2 lesson, learned when a poisoned `.marks.json` returned 500 for `/` and `/healthz` across all 25 booths. `read_manifest` returns None for absent and a `Manifest` carrying `error` for damaged, and the card distinguishes them (`unannounced` vs `unreadable`). Silently treating damaged as absent would hide the one case somebody has to fix."
- "THE WRITE IS ATOMIC (CLAUDE.md invariant 5, NOT this unit's INV-5). Temp file + os.replace onto a name no other writer derives, because the CLI writes it in one process while the browser reads it in another -- and because two `booth add` calls on one booth would otherwise share a scratch name, which the atomic-write promise says nothing about: it promises readers never see a partial file, not that writers never race. The pattern is copied from `marks._write_raw` rather than imported: `scripts/booth` imports each module directly under the system python3, and a cross-import between two stdlib-only modules is a second way for INV-1 to break."
- "`booth/manifest.py` IS STDLIB-ONLY and joins the CLAUDE.md invariant 1 list. `scripts/booth` imports it through a `python3 -c` heredoc with no venv, exactly as it imports `marks`, `asks` and `links`. `test_stdlib_only` is parametrized and gains `manifest`; that test is the only thing standing between a casual third-party import and `booth new` breaking on every fleet host."
- "A MISSING MANIFEST IS NORMAL, NOT AN ERROR. All 26 live booths have none, and `rsync -a ./out/ nh3-dev:booth-data/my-run/` -- the documented path for every host that is not nh3-dev -- never runs the CLI at all, so unannounced booths keep arriving after this lands. The card marks them quietly and nothing refuses to render, expire, zip or sweep."
- "THE BOOTH ANNOUNCES ITS OWN BOOTHS rather than exempting them. A pickup booth and the standing link board are created BY the service, so they are written with `handle: booth` -- which is true, not manufactured. The alternative was a pile of exemptions from the unannounced marker; this way there is one rule (a booth with no manifest is unannounced) and no special cases. `handle` therefore names an agent handle OR the service, and the field's docstring says so."
- "NOTHING NEW IS ORDERED, so CLAUDE.md invariant 6 (every ordered collection has a stated, deterministic rule) does not bind here -- there is no new collection for it to bind to. The manifest is one flat record per booth. The index keeps its stated rule -- kept lane first, then ephemeral newest-first by `_newest_mtime` -- and U5 does NOT add a second ordering keyed on `created` (operator, 2026-09-22). A what-landed feed ordered by announcement time is a genuinely different surface: it needs its own stated rule, it competes with the existing order for what 'the third one' means, and it has nothing to sort the 26 manifest-less booths by. Parked for v1.1."
open_questions:
- "Whether `why` should also reach the zip manifest or a `booth ls` column. Both are one-liners over the same record and neither is on the v1 path; deferred rather than designed."
---
# U5 — self-announcing booths
## The defect, stated precisely
The index card is the only thing an agent can put in front of the operator, and
it carries no information the agent chose. Name, item count, countdown, a
thumbnail. Everything about *why this exists* has to travel some other way.
So it travelled some other way. `booth link` exists because a session with
something to show had no way to make the booth itself say "look at this", and
the link board absorbed job 5 until **145 of its 210 rows (69%) pointed at
booths that had already been swept**. The rot is not a link-board bug. The board
was doing a job it was never shaped for, because the shaped thing did not exist.
The lesson the measurement carries, and the reason this unit comes before any
link-board enforcement: **enforcing the link rule without giving job 5 a home
just makes it homeless.**
## The record
```python
@dataclass(frozen=True)
class Manifest:
handle: str # an althing handle, or "booth" for one the service made
title: str # display name; falls back to the directory name
why: str # ONE line: what the operator is looking at and why
created: str # ISO-8601 with offset, from the FIRST announcement
error: str | None = None # a read-time verdict; never stored
```
`.booth.json` on disk is the same four fields, no `error`.
**Every field on an error-carrying record has a stated value**, because the
templates render the record and a careless fill would re-raise the outage in
the renderer: `handle` and `why` and `created` are `""`, `title` is the
normalized directory name, and `error` says which of the six refusals fired.
`created` being `""` is what makes `write_manifest` treat a damaged prior as
having no stamp to preserve (INV-3).
Caps, all applied at the write and again at the read: `handle` 64, `title` 120,
`why` 200, `created` 64. Each is a **display budget**, not a storage limit —
they exist because these strings land in a card's sub-line.
## Signatures
```python
MANIFEST_FILE = ".booth.json"
HANDLE_MAX, TITLE_MAX, WHY_MAX = 64, 120, 200
MANIFEST_MAX_BYTES = 64 * 1024
QUARANTINE_FILE = ".booth.json.broken"
def read_manifest(booth: Path) -> Manifest | None:
"""This booth's announcement, or None if it never made one.
LENIENT, and never raises. `list_booths` calls this once per booth on every
index page load, so a damaged file must cost that booth's provenance and
nothing else — the same posture `marks_for` takes, for the reason v0.2.2
made expensive: a read that can raise, called in a loop over every booth,
is a service-wide outage wearing a single-booth bug's clothes.
"NEVER RAISES" IS BOUNDED, NOT MERELY CAUGHT. An earlier draft of this
contract named a 4 GB file as a tested case and constrained only the RETURN
— which is letter-compliant and purpose-defeating: reading four gigabytes
per booth per index load recreates the same outage in slow motion. The size
is checked by `stat` BEFORE the bytes are touched, and the two exception
classes that are neither `OSError` nor `ValueError` — `MemoryError` from a
huge document, `RecursionError` from a deeply nested one — are caught as
well, so that raising the bound one day cannot quietly re-open the hole.
REGULAR-FILE FIRST, THEN SIZE — and the order is the whole point. `st_size`
is 0 for a FIFO and 0 for a symlink to `/dev/zero`, so both sail under any
byte cap and then the read either blocks forever with no EOF or allocates
until the kernel intervenes. The bound is what made this reachable: a cap
that trusts `st_size` inherits everything `st_size` does not mean. One such
file stalls every `GET /` and `/healthz`, with no error and no recovery
short of a restart.
Absent -> None. Present but too large, unreadable, unparseable, not an
object, or missing `handle` -> a Manifest carrying `error`, so the card can
say `unreadable` rather than quietly showing the same thing as a booth that
never announced.
"""
def write_manifest(booth: Path, handle: str, *, title: str | None = None,
why: str | None = None) -> Manifest:
"""Announce a booth. Atomic per CLAUDE.md invariant 5: temp file +
os.replace, onto a temp name no other writer will pick.
OMITTED MEANS UNCHANGED; `""` MEANS CLEAR. `title` and `why` default to
None. The ordinary sequence is `booth new x --why "..."` then
`booth add x out/*.png`, and while omission meant `""` the second command
silently erased the sentence the first one existed to record. The shell
carries the distinction by leaving the environment variable UNSET rather
than empty.
Re-announcing PRESERVES the original `created` — `created` is when the
booth appeared, and saying something more about it later is not a second
appearance. A prior record carrying `error`, or one whose `created` is
`""`, is treated as having no stamp to preserve and gets `now()`: a stamp
that is silently wrong is worse than one that is silently new.
A WRITE THAT CHANGES NOTHING IS NOT ACTIVITY and does not touch the file,
so it cannot reset the booth's TTL — the rule marks learned in v0.2.0,
needed here because `booth link` re-announces the standing board on every
single post to it.
BYTES THAT COULD NOT BE READ ARE KEPT, not replaced. See INV-6.
A FAILED WRITE LEAVES NOTHING BEHIND. The temp name carries a random suffix
so two writers cannot share it — which also means nothing ever overwrites an
orphan, and `.booth.json.<hex>.tmp` is not a `.lock`, so `_newest_mtime`
counts it and a leak would keep a dead booth alive forever. Cleaned up on
every exit path.
`title` falls back to the directory name, THROUGH the same normalizer the
explicit value gets — a directory name may legally carry a newline on POSIX
and may run to 255 bytes, and the fallback used to hand either straight
into a card's sub-line.
Every stored string is collapsed to a single line — all runs of whitespace,
not only newlines, because a tab or a forty-space indent renders as badly
in a sub-line as a newline does — and truncated to its cap.
An empty `handle` becomes `"booth"` rather than being refused: a manifest
naming no handle does not read back at all, and an unreadable file is the
worse outcome. Unreachable from the CLI, whose fallback chain always yields
something; a direct caller should pass a real one.
"""
```
## What renders
One line, on both surfaces, driven by the same record. The example booth below
is the directory `r18-ab`, announced by the handle `booth-dev`:
| state | the provenance line, on an index card AND on the booth header |
|---|---|
| announced, with a why | `booth-dev · pick the winning denoiser` |
| announced, no why | `booth-dev` |
| no manifest | `unannounced` (muted) |
| damaged manifest | `unreadable` (muted, warning tint, `title=` carries the reason) |
**`title` renders too, and on exactly one surface.** An earlier draft stored it,
surfaced a `--title` flag for it, and rendered it nowhere — a promise of a
display name with no display, caught 4-of-4 and ranked first independently by
every arm. It lands on the **booth page heading**, where there is room:
`<h1>R18 A/B <span class=h1-slug>r18-ab</span></h1>`. The **index card keeps
the directory name alone**, because that is the identity the operator navigates
by and refers to positionally, and CLAUDE.md invariant 6 is about exactly that
kind of reference surviving a re-render. When `title` equals the directory name
— the default — the heading is unchanged from today.
**Both index lanes get it.** The kept lane renders first and is a separate block
in `index.html`; patching only the ephemeral lane would leave the 15 kept booths
— the durable, most-looked-at ones — with exactly the defect this closes. This
is the `blurtoggle` lesson (three item branches, one macro) applied to two lanes.
**The booth page header gets it too**, and that is deliberate scope, not creep:
a booth URL handed to the operator lands on the booth page, never on the index.
Job 5 is "operator, look at this", and the page he actually opens is where the
answer has to be.
## The CLI surface
Operator decision, 2026-09-22 — flags on the existing verbs, not a second verb:
```sh
booth new r18-ab --why "pick the winning denoiser"
booth add r18-ab out/*.png --why "second pass, sharper" --title "R18 A/B"
booth new scratch # still legal — handle + created, no why
```
`handle` comes from `$ALTHING_HANDLE`, falling back to `$BOOTH_SOURCE` then
`hostname -s` — the same resolution `booth link` already uses for its rows, so
provenance means the same thing on the board and on the card.
**Nothing existing breaks.** A bare `booth new x` / `booth add x f.png` keeps
working; the flags are optional and may sit on either side of the file
arguments, because a glob is usually last and a flag usually after it and
nothing enforces that. The alternative — a separate `booth announce` verb — was
rejected because a second step is the step that gets forgotten, which is the
69% rot's own mechanism.
**A bare re-announce does not wipe what the last one said.** On a booth that has
never announced, a bare `new`/`add` writes `{handle, created}` with no `why`. On
one that HAS, an omitted flag leaves the stored value alone and only a supplied
one overwrites — `--why ""` still clears, which is a different intention. This
distinction is load-bearing rather than polite: `booth new x --why "…"` followed
by `booth add x out/*.png` is the ordinary sequence, and the naive reading
erases the sentence on the second command.
**The handle is the CLI's three-step chain**, not `$ALTHING_HANDLE` alone:
`${ALTHING_HANDLE:-${BOOTH_SOURCE:-$(hostname -s)}}`, identical to the one
`booth link` already uses for its rows, so provenance means the same thing on
the board and on the card. A session with no handle set still announces, as its
host.
## Scope — the blast-radius pass
Graphify + grep, both run, because neither is sufficient alone (graphify is
blind to function-local and DI-injected imports; grep misses transitive reach).
**Every site that creates a booth directory:**
| site | gets a manifest? |
|---|---|
| `scripts/booth new` (line 97) | yes — `$ALTHING_HANDLE` |
| `scripts/booth add` (line 103) | yes — `$ALTHING_HANDLE` |
| `scripts/booth link` (line 178) | yes — `handle: booth`, the standing board |
| `app.upload` (app.py:1087) | yes — `handle: booth`, a pickup booth |
| `marks._Locked.__enter__` (marks.py:267) | **no** — `mkdir(exist_ok=True)` on the write path; a mark written to a booth that does not exist is not an announcement, and manifest.py must not be imported by marks.py (INV-1 cross-import) |
| `rsync` from another host | **no** — no CLI runs; this is why `unannounced` exists |
**Every reader of a booth's facts:** `list_booths` (app.py:251) and `booth_view`
— confirmed by `graphify explain list_booths` (15 edges, 4 test consumers) and
by grep for `data_dir.iterdir` (two sites, both in app.py, both enumerating
booths for exactly these two surfaces).
**Sites that already exclude the new file and need no change**, each verified
rather than assumed: `items.booth_items` (items.py:182), `app.zip_booth`
(app.py:351), `marks.import_legacy_asks` (marks.py:656).
**One site the first draft of this contract got WRONG, corrected by the seam
review** (SR-1, below): the upload path's `used: set = {UPLOAD_MARKER}` filename
dedupe set does **not** need to gain `MANIFEST_FILE`. The implementation adds it
anyway, as consistency with the equally-unreachable entry already there, and
says so in a comment rather than claiming it prevents anything.
⚠ **Line numbers in this section are the PRE-CHANGE coordinates** the
blast-radius pass was run against, kept because that is what makes the pass
auditable. They have moved; `grep` the symbol, do not trust the number.
## Seam review — what the real sibling surfaces said
Caller-side pass against the actual modules, not against their prose. Run after
the cold contract panel was dispatched and before any code.
**SR-1 — the upload-collision change is unnecessary, and so is the one already
there.** `safe_upload_name` (app.py) does `base = base.lstrip(".")` with the
comment "a leading dot would hide the file from every listing", so an uploaded
file can never be named `.booth.json` — or `.uploaded`, which means the existing
`UPLOAD_MARKER` entry in that set has never been able to matter either. Adding
`MANIFEST_FILE` alongside it is consistency with a redundant guard, not a fix
for a reachable collision. Do it or don't; what the contract may not do is claim
it prevents something. **This is the exact class the seam review exists for: a
scope item the contract asserted from its own reasoning and the sibling's real
surface refutes.**
**SR-2 — the atomic-write pattern transfers cleanly to a dotfile, verified not
assumed.** `marks._write_raw` derives its temp name as
`path.with_suffix(path.suffix + ".tmp")`. For a dotfile with an extension that
is not obviously safe — `Path(".booth.json").stem` is `".booth"`, which looks
alarming — but `.suffix` is `".json"` and the result is `.booth.json.tmp`.
Checked against the interpreter. The temp file is itself a dotfile, so
`booth_items` and `zip_booth` skip it and no reader can see it mid-write.
**SR-3 — the dotfile skips are on `p.name`, and all three use `rglob` or
`iterdir` over the booth.** `items.booth_items` (items.py:182), `app.zip_booth`
(app.py:351) and `marks.import_legacy_asks` (marks.py:656) each test
`p.name.startswith(".")`. A manifest at the booth root is skipped by every one
of them. Confirmed by reading the three loops, not by trusting the claim.
**SR-4 — `test_stdlib_only` is parametrized `["marks", "asks", "links"]`**
(tests/test_marks.py:279) and gains `"manifest"` as a fourth entry. The test's
docstring calls this INV-5 while `CLAUDE.md` calls it invariant 1; that
inconsistency predates this unit and is left alone.
**SR-5 — `.booth.json` is reachable over HTTP at `/b/<name>/.booth.json`.**
`booth_file` refuses only path escapes and non-files, not dotfiles, so a remote
session with no filesystem access can read a booth's announcement the same way
it already polls `/b/<n>/marks.json`. That is a feature and it is now written
down; there is no secret in a manifest, and the Booth has no auth by design.
**SR-6 — `list_booths` returns plain dicts and the templates read them by key.**
`b.manifest` resolves through Jinja's getitem fallback. A None manifest must be
guarded with an explicit `{% if %}` rather than relying on `b.manifest.handle`
rendering as Undefined, because the two lanes' cards differ and a silent
Undefined in one of them is how the kept lane would quietly keep the old defect.
## Out of scope
Deliberately deferred or never. Divergence here is not drift.
- **A second index ordering keyed on `created`** — a "what landed" feed. Operator
decision, 2026-09-22: parked for v1.1. It is a new ordered collection needing
its own stated rule, it competes with the existing order for what "the third
one" means, and it has nothing to sort the 26 manifest-less booths by.
- **`why` in the zip manifest, or a `booth ls` column.** One-liners over the
same record, neither on the v1 path.
- **Enforcing that a booth MUST announce itself.** `rsync` is the documented
path for every host that is not nh3-dev and never runs the CLI, so a refusal
would break the documented workflow. The marker is the whole mechanism.
- **Deleting, expiring or migrating anything based on the manifest.** U4 owns
lifetime; this unit only describes.
- **Any change to how items, marks, blur, keep or the link board work.** The
manifest is a dotfile and every existing listing already skips it.
- **Auth, or treating a manifest as trusted.** Standing non-goal; the Booth is
LAN-internal and a hand-written `.booth.json` is a supported input.
- **Provenance ON a verbatim-`index.html` booth's own page.** Five live booths
serve the author's HTML raw, and the Booth owns no header there to put a line
into — it currently reaches those pages through six regexes injected into
arbitrary markup, which is precisely the defect U3 exists to fix. Their INDEX
cards carry provenance like everything else; the page itself waits for U3's
declared embed seam. Verified on `pewpew-ui-brief`: page renders 200, card
reads `unannounced`.
## Invariants
Numbered INV-1..5 and local to this unit. Where a repo-wide rule is meant it is
named in words — "CLAUDE.md invariant 5", "CLAUDE.md invariant 6" — never by a
bare number, because an earlier draft used `INV-5` for both the repo's
atomic-write rule and this unit's render rule and the collision was caught
3-of-4.
**INV-1 — one module knows the filename.** `booth/manifest.py` is the only
module that names `MANIFEST_FILE`. No route body, template or CLI verb opens or
parses `.booth.json`; `write_manifest` reads it back inside that module, which
is what INV-3 requires and is not an exception to this rule. Falsifiable and
tested: no other file under `booth/` contains the literal `.booth.json`.
**INV-2 — the read cannot raise, AND cannot cost the caller unboundedly.**
`read_manifest` returns for every input: an absent directory, a `.booth.json`
that is a list, a string, `null`, empty, not UTF-8, wrong-typed, missing its
handle, nested deeply enough to overflow the parser's stack, and one larger
than `MANIFEST_MAX_BYTES` — which is refused by `stat` before a byte is read,
because a bound that only constrains the RETURN recreates the outage in slow
motion. Tested per case, the size and depth cases included.
**INV-3 — `created` survives re-announcement.** A second `write_manifest` on the
same booth preserves the first `created`. A prior record carrying `error`, or
one whose `created` is `""`, has no stamp to preserve and gets `now()`. Tested
against a stamp that could not have come from `now()` — `_now()` is whole-second
resolution, so back-to-back writes share a timestamp and a naive test passes
against an implementation that regenerates it every time.
**INV-4 — stdlib-only, and sibling-free** (this is CLAUDE.md invariant 1
extended by one clause). `booth/manifest.py` imports nothing outside the
standard library and nothing from `booth.*` — a cross-import between two
stdlib-only modules is a second way for the repo rule to break. Relative
imports count; the AST walk sees them.
**INV-6 — bytes that could not be read are never destroyed.** When
`write_manifest` replaces a manifest whose read returned `error`, the old bytes
move to `QUARANTINE_FILE` first. This is the doctrine marks made explicit in
v0.2.1 — reads lenient, writes strict, damaged bytes stay on disk — and this
unit contradicted it by replacing outright, so a file that failed on ONE field
lost the others with it, including a `why` the re-announcer may never have kept
anywhere.
It diverges from marks in HOW it honours the rule, and the divergence is the
interesting part. Marks REFUSE the write and answer 409, because the operator's
judgment is not restatable. A manifest QUARANTINES and proceeds, because
refusing would fail `booth add` and lose the files it was mid-way through
copying — and a booth's own description is something its poster can say again.
One fixed quarantine name rather than a timestamped series: nothing prunes a
booth but the sweep, and the most recent damage is the only copy anyone opens.
**INV-5 — unannounced and unreadable render DIFFERENT TEXT.** Not merely
different styling: the words differ (`unannounced` / `unreadable`), so the
distinction survives a stylesheet change and a reader who cannot see colour. A
one-pixel difference would satisfy a looser wording and encode nothing, and the
point is that one of the two states is something somebody has to go and fix.
@@ -0,0 +1,11 @@
# Every code-changing finding came from the AMBIGUITY pass
_2026-09-21 · booth_
**Every one of the panel's code-changing findings came from the
AMBIGUITY pass, none from a paraphrase divergence** — and two arms independently
proposed cutting the paraphrase to a drift-check for narrative-heavy contracts,
because this contract's own frontmatter carries a plain-language narrative and the
paraphrase was partly reading my framing back to me. That is a finding about the
`/heid-contract-review` **skill**, not about this repo, and it was reported back
to heid. Recorded here only so a future session does not rediscover it.
@@ -0,0 +1,11 @@
# A boolean escape hatch as the lifetime mechanism
_2026-09-21 · booth_
**A boolean escape hatch as the lifetime mechanism.**
`.forever` was added because a 24h TTL genuinely did not fit some booths —
and then 56% of live booths ended up on it, which means it is not "ephemeral
with an exception", it is two lifetimes wearing one lifetime's clothes, with
the operator doing the sorting by hand. Replaced at U4 by lifetime derived
from state (an open mark pins; viewing is activity; `keep` survives as an
explicit reasoned pin rather than the only way to say "not yet").
@@ -0,0 +1,16 @@
# Deterministic order is a cross-cutting v1 invariant
_2026-09-21 · booth_
**Deterministic order is a cross-cutting v1 invariant** —
operator directive, mid-implementation. Every ordered collection the Booth
renders must have a *stated* rule producing the same sequence on every render
of the same state; the rule can be anything defensible (byte order, time, an
explicit number, an arbitrary-but-recorded sequence), but no rule at all is
forbidden. It binds harder here than elsewhere because the Booth's job is
**comparison** — the operator judges tile 47 against tile 47 and refers to
artifacts positionally, so an order that moves between renders misfiles a flag
or a note rather than crashing. Recorded as `ROADMAP.md` § "Cross-cutting
invariant" (with the per-collection table) and `CLAUDE.md` invariant 6, and
tested. Still undecided and must be settled before those units ship: **U7's
section ordering and compare pairing**, and **U6's bench listing**.
@@ -0,0 +1,7 @@
# Extracted from `eshpfi` into its own repo
_2026-09-21 · booth_
**Extracted from `eshpfi` into its own repo.** The accreted
service came over whole, tests included, so `tests/test_booth.py` (1581 lines)
is the regression net the v1 rewrite is checked against.
@@ -0,0 +1,13 @@
# Five mechanisms to get one question beside one artifact
_2026-09-21 · booth_
**Five separate mechanisms to get one question next to one
artifact** — `.forever`, the link board, `inline.py`'s placeholder DSL,
`wrap_verbatim_html`'s six regexes, and the floating amber asks chip plus
`/b/<n>/asks`. Every one is a *correct local fix* to the same global
mismatch, which is exactly why they accumulated without anyone making a bad
call. **The foot-gun is the sixth one:** the next "just add a small thing for
this case" reads as reasonable and is the pattern. The git log carries the
signature — every feature ships, then takes 2–5 patches for cases the single
shape did not anticipate. Check the ROADMAP gate before adding a mechanism.
@@ -0,0 +1,10 @@
# The `.forever` diagnosis is a falsifiable prediction
_2026-09-21 · booth_
**The `.forever` diagnosis is a stated, falsifiable
prediction.** U4 (derived lifetime) predicts the kept-rate falls to the
genuinely-durable booths. Re-measured today: **14 of 25 booths kept (56%)**,
against the 54% the IA doc recorded. **Re-count a fortnight after U4 lands.**
If it does not move, the diagnosis was wrong and the boolean was doing
something else. Tracked in the IA doc's Booth section and by this entry.
@@ -0,0 +1,11 @@
# The information architecture and the v1 gate landed
_2026-09-21 · booth_
**The information architecture and the v1 gate landed**
(`726822b`): `docs/design/information-architecture.md` names the single
defect — *one lifetime (24h from last touch) and one shape (a folder),
serving five jobs with different lifetimes and different shapes* — and
`ROADMAP.md` gates v1 on seven units, each closing a **measured** defect
rather than a wish. Both were written after a measurement pass over the live
service, and the measurements are the load-bearing part.
@@ -0,0 +1,24 @@
# Letting Jinja hot-reload templates in the deployment root
_2026-09-21 · booth_
**Letting Jinja hot-reload templates while the repo is the
deployment root** — the cause of a live outage the same day U2 landed, and the
sharpest foot-gun in the repo. `booth.service` sets `WorkingDirectory` to this
repo, so the running service imports these files with no build step and no
staging copy. Python is read once at process start; Jinja's `FileSystemLoader`
re-reads a template **on every render**. Editing `booth.html` therefore
deployed it instantly against Python from 22:03 that knew nothing about
`item_marks`, and **19 of 25 live booths returned 500** with
`UndefinedError: 'item_marks' is undefined`. Neither the old code nor the new
code was broken — the service was running both at once.
**The lesson that generalises:** a skew between a process and the disk under it
is invisible to the test suite by construction, so no amount of green tests
would have caught it; the operator found it. Fixed at the source rather than
with a reminder — the `Environment` is hand-built with `auto_reload=False`, so
there is now ONE staleness rule (nothing takes effect until you restart) and
the running process is always a coherent snapshot of one commit. Asserted by
`test_templates_do_not_hot_reload_from_disk`. Watch the second-order risk the
fix introduces: a hand-built `Environment` does not inherit `autoescape` from
the `Jinja2Templates` constructor, and booth names, item names and mark text
are all agent-authored strings landing in HTML.
@@ -0,0 +1,12 @@
# Letting the link board absorb the announce job
_2026-09-21 · booth_
**Letting the link board absorb the announce job.** `booth
link` is an `O_APPEND` write with no identity and no stated rule, so
re-announcing a bench appends a row instead of updating one, and a booth URL
rots the moment its booth is swept — **145 of 211 rows (69%) pointed at
nothing**, and 22 were the same target re-posted (talk 5×, peedlar 4×). The
rot is **structural, not drift**. The lesson that cost the most: enforcing
the link rule without first giving the announce job a home (`.booth.json`
provenance on the index, U5) just makes it homeless.
@@ -0,0 +1,21 @@
# Marks are one `.marks.json` per booth
_2026-09-21 · booth_
**Marks are stored as one `.marks.json` per booth**, atomic
temp-file + `os.replace`, `fcntl` lock on the read-modify-write — operator
decision, this session. Two alternatives were weighed and lost: a sidecar
per item (`<rel>.marks.json`) and extending the existing `<stem>.ask.json`
shape. Rationale, and the reason it is not `links.md`-shaped: **(a)** U4
makes *"does this booth owe an answer?"* a hot question — the sweep asks it
per booth per tick and the index asks it per card per page load, so per-item
sidecars turn it into a full walk of all 25 booths, one of which holds 270
files; **(b)** `links.md` is an `O_APPEND` content-hash log because **17
agent handles write it concurrently**, whereas marks have exactly one writer
(the operator, in one browser) and many readers — a different problem that
must not inherit the append-log design; **(c)** `.blurred` / `.pins` /
`.forever` already establish the per-booth dotfile as the house shape for
operator state, and `booth_items()`'s dotfile skip means it costs nothing in
counts, galleries or zips. Accepted cost: a corrupt `.marks.json` loses that
booth's marks rather than one item's. Implementation deferred to U2 —
tracked at `ROADMAP.md` U2 and by this entry.
@@ -0,0 +1,15 @@
# A write over a damaged `.marks.json` wiped the booth
_2026-09-21 · booth_
**A write over a damaged `.marks.json` was wiping every mark in
the booth.** Shipped in `v0.2.0`, found by the panel (Kimi, converged with
Hulda), fixed in `v0.2.1`. `marks_for` is deliberately lenient — unparseable
reads as `[]` so a review page still loads — and the write path inherited that
leniency through the same reader, so one flag click appended to an empty list and
atomically replaced the file. The fix is an **asymmetry**, which is the reusable
part: reads stay lenient, writes go strict (`MarksCorrupt`), damaged bytes stay
on disk, routes answer 409 not 500. A page that renders without an annotation is
recoverable; a file that overwrote the operator's judgment is not. Kimi also
named the class correctly — "an author steeped in the design conversation would
likely read past" it — and that was accurate.
@@ -0,0 +1,10 @@
# A partially-answered pick counts as OPEN
_2026-09-21 · booth_
**A partially-answered pick now counts as OPEN** — declared, not
smuggled. The old index badge tested `answer is None`, so a half-answered
four-question ask read as closed on the index while the panel beside it
rendered `◐ partial`: the two disagreed about the same booth. Open is the
reading that makes U4 correct — a lifetime rule that unpinned a booth on the
first radio click would sweep a review in flight.
@@ -0,0 +1,14 @@
# Regex-injecting chrome into arbitrary author HTML
_2026-09-21 · booth_
**Regex-injecting chrome into arbitrary author HTML**
(`wrap_verbatim_html` + `_HEAD_CLOSE_RE`, `_HTML_OPEN_RE`, `_DOCTYPE_RE`,
`_BODY_CLOSE_RE`, `_HTML_CLOSE_RE`, `_ICON_RE`, and the doctype/charset
ordering constraints they thread). It works today and is **still live** —
but it is the single most fragile thing in the service and it is load-bearing
for the operator's most important workflow. Slated for deletion at U3 in
favour of a declared seam (`/_booth/embed.js`, mounted through a real DOM
API), which costs an author one line and removes the whole class. Do not
extend the regex set in the meantime; if a verbatim page breaks, that is an
argument for U3, not for a seventh pattern.
@@ -0,0 +1,8 @@
# `sindra-finalists` is U2's flag motivation, caught live
_2026-09-21 · booth_
**`sindra-finalists` is U2's `flag` motivation caught in the
act** — 86 items, every one captioned, and the booth's entire name is "the
ones the operator picked." That loop currently runs through chat, which is
the defect `flag` closes. Evidence, not argument.
@@ -0,0 +1,12 @@
# Tagging a release while a review gate was in flight
_2026-09-21 · booth_
**Tagging a release while a review gate was still in flight.**
`v0.2.0` was cut and announced to 15 consuming handles; the
`/heid-contract-review` panel — dispatched BEFORE implementation, as the
discipline says — replied afterwards with three defects in the code that had just
shipped, one of them silent data loss. Nothing about the tier decision was wrong;
the *timing* was. **If a gate is outstanding on the work being released, the tag
waits for it.** The cost was a same-hour `v0.2.1` and a correction note to peers
who had already verified against the broken version.
@@ -0,0 +1,9 @@
# Letting the write path share the read path's leniency
_2026-09-21 · booth_
**Letting the write path share the read path's leniency.** See the
`MarksCorrupt` decision above. The general shape, worth carrying beyond marks:
a tolerant reader and a tolerant writer over the same state are not the same
decision, and pointing both at one function silently makes them one. Tolerate on
read so the surface still renders; refuse on write so nothing is destroyed.
@@ -0,0 +1,13 @@
# Seam review and cold panel had zero overlap, twice
_2026-09-21 · booth_
**The two review gates are complementary, measured on one unit.**
The caller-side **seam review** (nine findings, against the real sibling module
surfaces) and the cold **`/heid-contract-review` panel** (four arms,
artifact-only) had **zero overlap in both directions** on U2. The seam review
found a scope miss the panel structurally could not see: the contract omitted
`inline.py`, whose `place()` indexes by subscript, which a frozen dataclass
refuses. The panel found three code defects and a missing test the seam review
had no lens for. Matches heid's kvasir zero-overlap result on the
conformance-versus-hunt axis. **Run both; neither substitutes.**
@@ -0,0 +1,16 @@
# U2 (marks) landed — one primitive for three mechanisms
_2026-09-21 · booth_
**U2 (marks) landed.** One primitive replacing three
mechanisms. `pick` / `note` / `flag` in one `.marks.json` per booth, one read
path (`marks_for`), one openness predicate (`open_marks`), rendered beside the
artifact on the tile, at full size in the zoom, and in the panel. `flag` and
`note` had no write path at all before this — the selection loop
(`golden-candidates`, `sindra-finalists`, the `pancake-*` ladders) was running
through chat. 242 tests. Details worth carrying: `asks.py` kept `normalize_ask`
and gained `build_answer` (the 2026-09-09 partial-answer semantics preserved by
moving, not rewriting) and LOST its five sidecar-storage functions;
`GET /b/<n>/marks.json` was added because remote sessions polled
`<stem>.answer.json` over HTTP and the sidecar's removal would have taken that
capability with it; `/b/<n>/asks` 308s to `/marks`.
@@ -0,0 +1,17 @@
# The U2 seam review earned its place, and how
_2026-09-21 · booth_
**The U2 seam review earned its place, and the record should
say how.** Nine findings against the real `booth.asks` / `booth.items` /
`booth.inline` surfaces, two of which changed scope or behaviour: `inline.py`
was missing from `touches` entirely (its `place()` indexes asks by
**subscript**, which a frozen dataclass refuses — nothing else in the service
does that), and the partial-answer inconsistency above. The cold
`/heid-contract-review` pass is artifact-only by design and structurally
cannot see a sibling module, so neither it nor a same-model self-review would
have found either. Two more surfaced later and are worth the same note: a
SECOND subscript in `inline.place` the seam review undercounted, and a
regression in my own legacy importer that a retargeted test caught — a
malformed sidecar that renders `⚠ broken` today would have silently vanished
on migration.
@@ -0,0 +1,19 @@
# U7's section premise is half wrong
_2026-09-21 · booth_
**U7's section premise is half wrong, and it is the half that
matters** — found by re-measuring `~/booth-data` rather than trusting the IA
doc. The IA says sections come from subfolders that already exist on disk;
true, but **every booth that actually needs navigation is flat**:
`pancake-v3-full` (270 items, 0 subfolders), `pancake-v4-full` (270, 0),
`sindra20-engines` (98 items + 99 caption sidecars, 0), `sindra-finalists`
(86 + 87, 0). Subfolders exist on exactly two booths — `pewpew-ui-brief` (7,
nested to `_ds/powerpellet-design-system-<uuid>/preview`) and `dfa-concepts`
(1) — and **both are reports**, the job where grid navigation matters least.
So sections stay worth shipping and `Item.section` stays right, but they are
**not** "most of the navigation fix": the rail, the filters and grid keyboard
are all of it. Worth noting for whoever writes U7: `sindra20-engines` encodes
its structure in the **filename prefix** (`b2-s1-<subject>-<seed>`), which is
where a grouping heuristic would actually pay. The IA doc's claim about what
sections buy needs a line struck — not yet edited.
@@ -0,0 +1,14 @@
# v0.2.0 was tagged while a gate was in flight
_2026-09-21 · booth_
**v0.2.0 cut and announced; v0.2.1 fixed what the announcement
was already wrong about.** Operator approved the minor (a v1 unit closed plus a
CLI surface change for 17 consuming handles clears the release-note bar). The
note went to 15 handles — the 17 link-board posters minus `nh3-dev`, a host
label, and `heid`, an oracle that does not script these verbs. Then the
cross-frontier contract panel landed and found **three defects in the code I had
just released**, so `v0.2.1` shipped within the hour. Sequence worth remembering:
the release was correct by the tier bar and still premature by the discipline —
the panel had been dispatched BEFORE implementation and its reply arrived AFTER
the tag. **If a gate is in flight, the tag can wait for it.**
@@ -0,0 +1,79 @@
# The U2 bug-hunt panel — full triage
**Date:** 2026-09-22 · **Thread:** `01M33XEC1H0298C0D968FWBN7A` ·
**Reply:** `01M33YZZ1VYGZ04JGNXNTBXDKS` · **Shipped as:** `v0.2.2`
`/heid-bug-hunt` on U2's diff (+2251/−632, 20 sections, 18 post-change
snapshots). Four arms — Gróa (Grok), Hulda (Codex), Regin (GLM-5.2), Kimi
(kimi-k3) — artifact-only, 4/4 clean transport. Heid adjudicated **9 findings
(6 bug / 3 robustness)**. Staleness was disclosed at build: `app.py` was edited
after the 06:38:52Z capture.
## Triage, five-category
### Category 1 — genuine add (8 taken, all shipped)
| # | finding | where | why it was real |
|---|---|---|---|
| 1 | Lock-inode split on the no-op unlink (**4/4 convergent**) | `marks._Locked` | `flock` binds to an inode; unlinking under a waiter destroys mutual exclusion silently |
| 2 | No-op lock churn resets the TTL via **directory** mtime | `marks._Locked` + `app._newest_mtime` | the guard's own comment reasons about the lock FILE's mtime; the directory is what the sweeper reads |
| 3 | Non-string `text` / `created` raise out of the read path | `marks._clean_text`, `marks_for` sort | `list_booths` reads every booth per page load → one bad file 500s `/` and `/healthz` |
| 4 | Legacy import stamped `created` at whole-second resolution | `marks.import_legacy_asks` | same-second sidecars re-sorted alphabetically, reversing the order the importer had just set — violates the stated `(mtime, name)` rule |
| 5 | `/answer` 500s on a non-string `notes` form value | `app.booth_answer` | the sibling `/note` guards it; same parser, same class of value, two answers |
| 6 | All five mark-write routes hold a blocking `flock` on the event loop | `app.py` | a contended lock freezes every route, not just the one request |
| 7 | CLI conflates a reader crash with "open" / "unanswered" | `scripts/booth` | `marks` printed a traceback and exited 0; `answer --wait` spun the full hour on a damaged file |
| 8 | The inline-doc tile had `markcontrols` and not `marknotes` | `booth.html` | flag a report, cannot say why — on the one item kind that is prose |
Two more taken on the same sweep, found while fixing the above rather than by
the panel: a broken mark of any shape now renders **⚠ broken** instead of as an
empty note (the rule `_hydrate` states for picks, applied to all three shapes),
and the marks panel is no longer suppressed on a booth that carries a
`links.md` *and* has marks.
### Category 3 — restatement of a settled prior (1, no change)
**Corrupt read → filtered writeback → silent deletion** (hulda F2, kimi F3,
gróa F4; Heid ranked it #3). **Already fixed in `v0.2.1`** by
`_read_raw_strict` + `MarksCorrupt` — reads lenient, writes strict. The panel
reviewed the pre-fix capture and the staleness was disclosed up front. Verified
against the current source before declining, not assumed.
This is the exact case the cross-frontier triage discipline warns about: a
confident, well-argued, four-arm-corroborated finding against code that no
longer exists. **Check what the peer actually read before treating an omission
or a defect claim as new.**
### Category 4 — out of place, parked (2)
- **Note-id recycling** (`note-1` reused after a withdrawal) lets a stale tab
delete a newer note. Real mechanism; needs two tabs and an interleaving, and
the Booth has one viewer. Non-reused ids are a schema change, not a patch.
- **Unvalidated flag / note targets** accumulate orphan marks. Targets come
from rendered items; the operator is the only writer through the browser.
### Category 5 — wrong-grounding (1)
**`delete_mark` can remove a pick, not only a note.** Framed as an
access-control divergence. There is no auth by design, and restricting it would
remove the only way to withdraw a pick that hydrates broken. Declined; the
docstring is the thing that was imprecise, not the behaviour.
## What the round is worth remembering for
1. **The two review gates stayed complementary a second time.** The contract
panel (2026-09-21) found three defects; this bug-hunt found eight more, with
**no overlap**. Both ran on the same unit. Neither substitutes.
2. **The panel beat the code's own comments three times.** The bundle's comments
are unusually honest and still wrong about what protected the TTL, and
"written atomically" sat next to a filter-then-replace. **A comment is a
claim, and a claim can be tested.**
3. **The headline bug class shipped with zero guard coverage, and both mutation
tables said so.** `test_a_no_op_write_does_not_touch_the_booth` asserted only
that `.marks.json` was absent — so removing the lock unlink, removing the
whole lock lifecycle, or bumping the directory clock all **SURVIVED** it. The
test asserted an artifact of the property instead of the property. The
replacement asserts `booth_age_seconds` directly, with a positive control (a
real mark still resets the clock) so the fix cannot overshoot into "marking
is never activity".
4. **`scripts/booth` had no tests at all** and two findings lived there. It has
five now, running the real script under the system `python3`.
@@ -0,0 +1,14 @@
# `booth marks` / `booth answer` got real exit codes
_2026-09-22 · booth_
**`booth marks` / `booth answer` got real exit codes**, because
a read that CRASHED was indistinguishable from a read that said no. `marks`
printed a traceback and exited 0 (a caller's `jq` saw success and got
nothing); `answer --wait` read a damaged file as "not yet" and spun for the
full hour before blaming the operator. Now `0 ok · 1 unanswered/timed-out ·
2 no such pick · 3 unreadable`, and `read_error()` was added to `marks.py` so
the CLI can ask the question the browser must not: the page stays lenient, the
machine consumer gets the truth. Also `--wait` now prints ONCE — it was
emitting a whole JSON document per poll, so a captured `--wait` held several
concatenated values and parsed as none of them.
@@ -0,0 +1,13 @@
# An existing test stopped me retiring documented behaviour
_2026-09-22 · booth_
**An existing test stopped me retiring documented behaviour
while fixing a race.** The mtime-restore race is real, and the clean fix —
ignoring a booth directory's own mtime whenever the booth holds anything —
would also have silently retired the rule that RELEASING a kept board resets
its clock, which the CLI header, the README and a deliberately-written test
all pin. That is a TTL doctrine change, not a bug fix. Fixed the concrete half
(a failing `os.utime` used to escape and 500 the route), left the race stated
in the code. **A fix that changes a documented rule is a proposal, not a
patch.**
@@ -0,0 +1,44 @@
# The `.forever` diagnosis got a live positive control
_2026-09-22 · booth_
The U4 diagnosis was that `.forever` is the only way to say three different
things — "this is durable", "I have not answered yet", "I am still looking" —
and that only the first is what keep means. That was an argument. **On
2026-09-22 it stopped being one.**
Census of `~/booth-data`, whole population, every value a deterministic file
fact:
| | |
|---|---|
| live booths | 24 |
| carrying `.forever` | 17 (70%, up from 54% on 2026-09-21) |
| carrying `.marks.json` at all | 4 |
| of those, with an open pick | **4 of 4** |
| **open pick AND `.forever`** | **3** |
**Three of the four booths in the entire fleet that were waiting on an answer
had also been pinned by hand.** That is the "not yet" case caught in the act,
not inferred from a rate.
The staleness distribution says it from the other side: **10 of the 17 kept
booths were under one day old** — younger than the TTL, so the sentinel had
bought them nothing and was pressed pre-emptively. Only 4 were old enough
(2.4-4.6 days) that keep is the reason they still existed.
⚠ **A number I got wrong, caught by a cross-frontier arm, kept here because the
class repeats.** The contract first said "12 are under 1.5 days old — younger
than the TTL". The TTL is 24 hours. 1.5 days is not younger than 24 hours. The
measurement was sound and the sentence was not; the claim only holds at the
one-day line, where it is 10 rather than 12. Nobody on the Claude side caught
it, including the author twice.
⚠ **The hold's live blast radius is SMALL** — only 4 booths have marks at all —
so the `.forever` re-count prediction rests on BOTH halves of U4 and on the
sentinel becoming unnecessary rather than forbidden. **RE-COUNT A FORTNIGHT
AFTER U4 LANDS**, i.e. on or after **2026-10-06**. If the rate does not move,
the honest readings are "the diagnosis was wrong" OR "the habit outlived the
need", and a bare re-count cannot tell those apart. **The three
open-pick-plus-`.forever` booths are the ones to watch**, because for them the
mechanism is now unambiguous.
@@ -0,0 +1,57 @@
# Four independent paths to one fail-open delete
_2026-09-22 · booth_
The U4 bug-hunt panel declared invariant was **"a deletion decision must never
be made from a read that failed"**. The panel found **four independent paths
through it, and no single arm found all four.** That is the strongest argument
yet for running the panel rather than one arm.
1. **An entry-level hydration error lost its hold** (the round's best finding).
`.marks.json` parses; one mark fails normalization; `_hydrate_safe` returns a
`Mark` carrying `error`; `_is_open` returns False for an errored pick — **on
purpose**, because a broken pick can never be answered. So the booth read as
not-held and **swept**, while the panel beside it rendered the broken mark in
full. The fail-safe had been built for FILE-level damage and missed
ENTRY-level. A mark we cannot read is judgment we cannot see; deleting the
booth it belongs to is the one thing we must not do with it.
2. **A present-but-blank `.marks.json` swept.** `_read_raw_strict` early-returns
for whitespace-only content — correct for the WRITE path it was written for
(a blank file is safe to overwrite), wrong for the DELETE path. Fixed with a
`blank_is_corrupt=True` flag used only by `hold_read`. ⚠ The near-regression
worth remembering: a **valid document with an empty `marks` list** is what
deleting the last mark leaves behind, and holding on THAT would make every
finished booth immortal. Blank bytes are damage; an empty list is an answer.
3. **`_newest_mtime` returned 0.0 when the booth's own stat failed**, which made
it maximally ancient and therefore the FIRST thing the sweeper takes — a
permissions problem resolving to a deletion. Now returns `now`: not knowing a
booth's age is a reason to leave it alone. ⚠ Per-entry `FileNotFoundError`
stays a skip, because a dangling symlink raises it and has no mtime worth
counting; only OTHER stat errors mean "something is here we cannot read".
4. **`is_kept` collapsed a stat failure into not-kept.** `Path.exists()` maps
ELOOP and EACCES to False. Now `lstat`, with any non-ENOENT error reading as
KEPT, and a `.forever` symlink counting dangling or not.
**`is_held` was replaced by `hold_reason`, which returns the REASON** —
`"open"`, `"unreadable"`, or None — rather than a bool beside a separate error
string. Two representations of one state drift; Regin independently flagged that
the display could not tell the two holds apart. One value, read by the sweeper
and by all four rendering surfaces.
**Convergent 3-of-4, and the one with teeth beyond lifetime:** `record_view`
used `Path.touch()`, which FOLLOWS an existing symlink. A booth carrying a
planted `.viewed -> /anywhere` turned every page view into an mtime write at an
arbitrary path under the service uid — and **any fleet session can write into a
booth, because making a folder is the whole API.** Now `os.open(..., O_NOFOLLOW)`
plus `os.utime(fd)`; a planted link raises ELOOP into the existing swallow.
⚠ **THE CAPTURE TOOLING FAILED SILENTLY AND THE PEER CAUGHT IT, NOT US.** The
snapshot `files/` tree shipped to the arms was EMPTY. The loop was
`for f in $IN` over a multi-line variable — and **zsh does not word-split
unquoted parameter expansions the way bash does**, so it iterated once against a
path that was the entire list. jekyll recovered by re-applying the bundled diff
to HEAD and verified every file byte-identical, so the round was sound. **The
failure mode is the dangerous one: an empty bundle reads exactly like a clean
result.** Quote-and-split explicitly (`print -r -- $IN | while read f`) or build
the list as a real array. Same family as `[[2026-09-22-vacuous-falsifiers]]` —
an instrument that cannot fail loudly will fail quietly.
@@ -0,0 +1,15 @@
# The lenient reader's blast radius was the whole service
_2026-09-22 · booth_
**The lenient reader's blast radius was the whole service, not
one booth.** `_clean_text` did `(text or "").replace(...)` and `marks_for`
sorts on `(created, id)`, so a stored `text` that was a dict or a `created`
that was a number raised out of the READ path — and `list_booths` reads every
booth's marks on every index load. One hand-edited file 500'd `/` and
`/healthz` for all 25 booths. Fixed in two layers, matching the house posture:
a named type check (`_entry_type_error`) plus a `_hydrate_safe` backstop that
cannot raise, and the panel now RENDERS an unreadable mark as ⚠ broken instead
of as an empty note. **The general shape: a lenient reader is only lenient if
the leniency is bounded by where it runs.** `marks_for` was written for one
booth's page and is called in a loop over every booth.
@@ -0,0 +1,11 @@
# `scripts/booth` went from zero tests to five
_2026-09-22 · booth_
**`scripts/booth` had zero tests and now has five**
(`tests/test_cli.py`). The panel's guard-strength tables returned UNVERIFIED
for every CLI claim because nothing in the suite executed the script — two of
the round's findings lived in exactly that gap. The new tests run the real
script under the system `python3`, which makes them a live check on INV-1
(stdlib-only) as a side effect: a third-party import in `marks.py` now fails
in the suite the same way it would fail on a fleet host.
@@ -0,0 +1,23 @@
# The size cap opened a service-wide hang
_2026-09-22 · booth_
**The U5 bug-hunt panel found a service-wide hang that the
SIZE CAP ITSELF opened — two hours after I added the cap.** `stat` reports
size 0 for a FIFO and 0 for a symlink to `/dev/zero`, so both sail under a
byte cap and then `read_text` blocks with no EOF or allocates until the kernel
intervenes. `list_booths` reads every booth on every `GET /`, so ONE such file
stalls the front page for the whole service with no error and no recovery
short of a restart. Reproduced (`timeout` returned 124), fixed with an
`S_ISREG` check BEFORE the size check in both modules, verified live: the
index answered 200 in 36 ms with two FIFOs planted. **The reusable shape:
`st_size` answers a different question than "can this be read", and a bound
that trusts it inherits everything it does not mean — a hardening fix opened
a worse hole than the one it closed.** Also adopted: the upload path wrote the
manifest ABOVE its own cleanup guard (4/4), so a failure orphaned a half-booth
whose uniquely-named leaked temp then kept it alive forever; replace-over-
damaged destroyed recoverable bytes (4/4, now QUARANTINED rather than refused
— marks refuse because judgment is not restatable, a booth's description is);
and `booth answer` spelled out its own openness test, disagreeing with
`booth marks` about a partially-answered pick, which is a direct violation of
U2's INV-2. Full triage in `persistent-memory.d/2026-09-22-u5-panels.md`.
@@ -0,0 +1,42 @@
# The third one-branch template miss — this repo's recurring blind spot
_2026-09-22 · booth_
**All four arms of the U4 code-review panel found the same drift, independently.**
That is the strongest convergence either panel has produced here.
The booth header's sub-line forks on `{% if board %}`, and the U4 lifetime macro
had been added only to the `{% else %}`. So **a booth carrying `links.md`
rendered a link count and nothing at all about its lifetime** — no countdown, no
hold — while INV-4 said the templates have no path that renders neither. The
standing board being kept by construction (`booth link` drops `.forever` on
first use) is what hid it; a **released** board or a hand-made `links.md` booth
is a live non-kept booth on that path, and both are reachable from the UI.
**This is the third of the same shape in this repo's short history:**
1. `blurtoggle` — the blur only patched the image/video `<figure>`; inline docs
render through their OWN branch and shipped unblurred. Suite green; a live
look caught it.
2. verbatim chrome — a verbatim booth's own `index.html` is served untouched, so
the inline marks panel never renders there. Found by looking at the live
service during U4, not by the suite.
3. the board branch — this one.
**The pattern: the suite renders the surface the author was thinking about.**
Every one of these was a second branch of a conditional the author had already
satisfied once and stopped reading. A cold reader with no idea which branch was
"the real one" finds them; the author does not, and neither does a test the
author wrote.
**Practical consequence for this repo.** When a template gains a fact, grep the
template for `{% if %}` in the block you edited and render EVERY branch in a
test — one test per branch, each rendering only its own surface, or the passing
test on branch A will mask the omission on branch B. U4 now has one per surface
(index card, booth header, board header, marks page) for exactly this reason.
Declined, and worth recording: Regin and Kimi both recommended amending INV-4 to
carve the board header out, on the grounds that board layout belongs to U7.
**Cutting an invariant down to fit an implementation gap is the wrong direction
when the fix is one template edit**, and U7 owns navigation and section layout —
not whether a header states a lifetime.
@@ -0,0 +1,35 @@
# Two reads of one file are not one read of one state
_2026-09-22 · booth_
**The one finding across both U4 panels that changed code rather than prose,
and it came from Hulda (Codex) on the CONTRACT-paraphrase round — before any
code existed.**
The contract specified the hold check as:
is_held(marks_for(child), read_error(child))
Two reads of `.marks.json`, presented as one answer. They are not. A write or a
repair landing between them yields a pair that described the booth at **no
instant**, and the losing pair is `([], None)` — no marks, no error — which is
**exactly the pair that deletes**. A lenient reader plus a strict reader, each
correct on its own, compose into a fail-open delete.
The fix is `booth.marks.hold_read(booth) -> (marks, error)`: ONE strict read
answering both questions. `sweep_once` now does one read per booth per tick
instead of two. And because `_read_raw_strict` **raises rather than dropping an
entry**, a non-raising strict read returns exactly what the lenient read would —
so the index uses that same one read for its badge too and falls back to
`marks_for` only on the error path, where leniency is the point. Better than the
original in both correctness and cost.
**The generalisable class, in heid's words: a two-read seam presented as one
answer is a TOCTOU race even when nothing on the page looks concurrent.** Worth
looking for anywhere two reader functions with different strictness feed one
decision — especially when that decision ends in `rmtree`.
Related: `[[2026-09-21-marks-write-wiped-judgment]]` is the same
reads-lenient/writes-strict asymmetry; U4 extends it to the reaper with
"deletes strict", whose scope is **the sweeper only** — a hand delete is never
strict, which is what gives an unreadable-marks hold an exit at all.
@@ -0,0 +1,19 @@
# The U2 bug-hunt panel was not ceremony
_2026-09-22 · booth_
**The U2 bug-hunt panel landed and it was not ceremony —
`v0.2.2`.** Nine adopted findings across four arms; eight were real against
live code and one was already fixed. The headline was **4/4 convergent from
four different angles**: `_Locked.__exit__` unlinked `.marks.lock` on the no-op
path, and `flock` binds to an INODE — so a writer blocked on the old inode
proceeds while the next writer creates a fresh lock file and takes it at once.
Two processes then run the read-modify-write concurrently and the later
`os.replace` drops a mark, with both of them obeying the protocol. **The
cleanup existed to protect the booth's TTL and it was failing at that too**:
creating and removing a directory entry bumps the DIRECTORY's mtime, which is
what `_newest_mtime` actually seeds from, so a no-op reset the clock it was
written to leave alone. Same code region, two defects, one fix — never unlink
the lock, exempt `.<name>.lock` dotfiles from `_newest_mtime`, and put the
directory's mtime back after creating one. Full triage in
`persistent-memory.d/2026-09-22-bug-hunt-panel.md`.
@@ -0,0 +1,50 @@
# U4 landed — lifetime is derived, not declared
_2026-09-22 · booth_
**A booth's lifetime stopped being a boolean somebody remembered to press.**
Three states now, and `sweep_once` is the only thing that honours the first two:
KEPT `.forever` present never swept (unchanged)
HELD an open pick, or marks we cannot read never swept (new)
EPHEMERAL everything else 24h (unchanged)
Plus **viewing is activity**: a deliberately-served response from a booth's own
page route writes `.viewed`. That dotfile is not a `.lock` dotfile, so
`_newest_mtime` already counts it — **there is no new arithmetic anywhere**.
`booth_age_seconds`, `is_expired` and `expires_in` are byte-for-byte what they
were. A view is one more thing in the tree, which is the same trick `.booth.json`
used in U5.
**What counts as a view, and why the exclusions matter more than the inclusions.**
`/b/<n>/` (gallery, verbatim report, `?download=1` zip), `/b/<n>/view` and
`/b/<n>/marks` count. `/b/<n>/marks.json`, asset GETs, `/`, `/healthz` and a
zoom URL that 404s do NOT. The marks.json exclusion is load-bearing: **an agent
must not be able to hold its own booth open by polling for the answer it is
waiting on.** `/b/<n>/asks` is a 308 into `/marks` and records through it — one
call, not two.
Checked because it would have been silent: **nothing in the fleet polls a booth
page.** Homepage's `siteMonitor` for the Booth is `/healthz`, which is on the
not-a-view list. Had it been pointed at a booth URL, every booth would have
become immortal on deploy and nothing would have reported it.
**The hold is unbounded and that is the point** — unanswered is unfinished. What
makes it safe is visibility plus two exits that already existed: the card and
every Booth-owned header say `held until answered` where the countdown was, and
`booth rm` / the UI x / `DELETE /b/<n>` take a held booth exactly as they take a
kept one. **A hold is protection from the timer, never from the operator.**
**Release is activity, stated rather than accidental.** Releasing a kept board
still buys a full TTL — unchanged — but now because `booth_unkeep` calls
`record_view`, which is a rule, and no longer because unlinking a file happened
to bump a directory's mtime, which is not. The CLI warning against
"unkeep and let it expire" stays and stays true.
⚠ **Running `scripts/layout-probe.py` over booth pages resets every booth's
clock**, because a GET of a booth page is a view and the probe is not exempt
from its own rule. Harmless, recoverable, and noted in the probe so nobody
debugs it later as a sweeper that stopped working.
Contract: `docs/contracts/u4_derived_lifetime.contract.md`. Both heid panels ran
and the bug hunt after them; see the sibling entries.
@@ -0,0 +1,23 @@
# U5's adoption prediction split in two
_2026-09-22 · booth_
**U5's adoption prediction, SPLIT IN TWO within an hour of
landing — and the split is the interesting part.** The baseline was recorded as
0 of 26. Fifty minutes after the deploy, `comfy-dev` created `muse-clothed-repro`
and it announced itself: `{handle: comfy-dev, why: "", created: ...}`. That peer
was told nothing. **The HANDLE propagates for free** — it rides on `booth new`
and `booth add`, so every existing CLI caller starts announcing without learning
anything, which is the flags-on-existing-verbs decision paying off on day zero.
**The WHY does not** — it needs someone to know the flag exists, and this first
one is empty.
So re-measure BOTH on **2026-09-29**, because they answer different questions:
find ~/booth-data -maxdepth 2 -name .booth.json | wc -l # free
grep -l '"why": "[^"]' ~/booth-data/*/.booth.json 2>/dev/null | wc -l # learned
A high first count and a near-zero second is the predicted shape of "nobody was
told", and it is the case the operator's no-announcement decision was designed
to be able to see. Do not read the n=1 above as a rate — it is a code-path
observation (every CLI caller writes a handle), not a sample.
@@ -0,0 +1,21 @@
# Two U5 panels, and prose reached a released outage
_2026-09-22 · booth_
**Two cross-frontier panels on U5, and a paraphrase panel reached
a production outage two modules away.** 3-of-4 flagged the contract's "4 GB"
case as letter-compliant but purpose-defeating; the conformance round found that
unbounded read live in U5's code; walking it to the sibling found the SAME hole
**live in released `v0.2.2`** — `marks._read_raw` catches `(OSError, ValueError,
UnicodeDecodeError)` and `json.loads` on deep nesting raises **RecursionError**,
which is none of them, so 400 KB of brackets in one booth returned 500 for `/`
and `/healthz` across all 26. The v0.2.2 round HAD flagged it and I closed half:
**a finding with two call sites is not closed when one is.** The reusable
instruction — **walk a conformance finding to the sibling module even when the
sibling is out of scope.** Five of ten conformance findings were tests of mine
that pass on the regression they exist to catch, three of them asserting an
ARTIFACT of the property rather than the property; that is three nights running
on the same shape. Two real bugs neither my tests nor I could see: a bare
`booth add` wiped the `why` on the one sequence the feature exists for, and
`--title` was write-only. Full triage in
`persistent-memory.d/2026-09-22-u5-panels.md`.
+102
View File
@@ -0,0 +1,102 @@
# U5's two cross-frontier panels — full triage
**Date:** 2026-09-22 · **Paraphrase:** thread `01M340PNVRS21HPASZT38PXQPN` ·
**Conformance:** thread `01M341E9XAPZEFBSPK9HPGAM0S` · **Shipped as:** `v0.3.0`
Two four-arm artifact-only rounds, dispatched ~30 minutes apart and correctly
firewalled: the paraphrase ran the **pre-seam-review** capture (073612), the
conformance round the **SR-amended** one (074901). Heid diffed the two at
intake and said so.
The conformance round's honest headline is Kimi's: **zero drift in the strict
sense — the code is a clause-for-clause implementation of the contract.** Both
rounds' weight landed one layer down, in test strength and contract finish.
## The result worth keeping
**A paraphrase panel reading nothing but prose reached a production outage two
modules away.** 3-of-4 flagged INV-2's "4 GB" case as *letter-compliant but
purpose-defeating* — the invariant constrained the RETURN, not the cost, so an
unbounded read "recreates the outage in slow motion". The conformance round then
found that exact unbounded read live in U5's shipped code. Walking it to the
sibling module found the same hole **live in released `v0.2.2`**: `marks.py`'s
`_read_raw` catches `(OSError, ValueError, UnicodeDecodeError)`, and
`json.loads` on a deeply nested document raises **RecursionError**, which is
none of them. A 400 KB file of nothing but brackets in any ONE booth returned
500 for `/` and `/healthz` across all 26.
**The v0.2.2 round had flagged this and I closed half of it.** Kimi's R5(c)
named RecursionError explicitly; I adopted "wrap `_hydrate` per-entry" and left
the `json.loads` above it unguarded. **A finding with two call sites is not
closed when one is.**
**The reusable instruction: walk a conformance finding to the sibling module
even when the sibling is formally out of scope.** Heid captured it as its own
lesson.
## The densest class was tests that could not fail
Five of ten adopted conformance findings were tests of mine that pass on the
regression they exist to catch. Three shared one shape — **asserting an
ARTIFACT of the property instead of the property**:
| test | asserted | should have asserted |
|---|---|---|
| `test_the_write_is_atomic` | no `*.tmp` survived | the inode changes (`write_text` leaves no temp file either) |
| INV-3 preservation | a stamp survived a window shorter than the stamp's own resolution | a stamp from 2019 |
| `test_announcing_is_activity` | age via the directory mtime, which the write bumps either way | the file's own mtime, directory clock restored |
That is the same shape as the marks round's guard-strength finding the night
before — **three nights running**. Proposed to heid as a standing
"green-tests-prove-nothing" direction for the skill; routed to the operator
alongside two other methodology proposals from the same night.
⚠ **My first replacement for the atomicity test was ALSO vacuous.** It spied on
`os.open` to prove the published path was never written directly — which passes
trivially, because `Path.write_text` reaches the syscall through `io.open` in C
and never touches the Python-level `os.open`. The dead end is recorded in the
test's own docstring rather than deleted.
## Two real bugs the tests were structurally blind to
**`booth new x --why "…"` then `booth add x out/*.png` erased the why.** Omitted
flags meant empty strings; empty strings overwrote. Two arms predicted it *from
the contract's wording alone* — "gains a manifest with no `why`" does not
distinguish a first write from a re-announce with the flags omitted. Every test
written for this module passed `--why` on both calls, so none could see it.
Omitted means unchanged now; `--why ""` still clears. The shell carries the
distinction by leaving the variable UNSET, not empty.
**`--title` was write-only** — stored, flag-surfaced, rendered nowhere. 4-of-4,
independently top-ranked by every arm of the paraphrase round. It renders on the
booth page heading with the directory name kept beside it, because the directory
name is the identity the operator navigates by and refers to positionally.
## Contract-finish, and why it mattered
**INV-1 contradicted its own falsifiable criterion** (4/4) — "the only place
`.booth.json` is opened" versus INV-3's read-back, which forces `write_manifest`
to open it. One half was already false of a correct implementation. Restated as
*one module knows the filename*, which is true, falsifiable and now tested.
**INV-5 named two different promises** (3/4) — the repo's atomic-write rule and
this unit's render rule. Repo-wide rules are named in words now, never by a bare
number that can collide with a local one.
Regin's meta-observation is the round's methodology keeper and was borne out:
**flags cluster where the same rule is re-voiced per signature**, and four of
eleven contract edits were reconciling a docstring against a prose section
saying the same thing slightly differently. A table-vs-signature consistency
pass would beat the format's prose bias.
## Declined / parked
- **Custom booth pages skip provenance** (hulda, solo, verified) — settled
independently as U3's seam ~20 minutes before the reply landed. Convergence,
not an adoption.
- **Empty-handle coercion misattributes to the service** — kept, documented. A
manifest naming no handle does not read back at all, and an unreadable file is
the worse outcome. Unreachable from the CLI.
- **`used`-set: `touches` versus SR-1 unreconciled** — the code adds the entry
as consistency with the equally-unreachable `UPLOAD_MARKER` entry that
predates this unit, and says so rather than claiming it prevents anything.
@@ -0,0 +1,40 @@
# Five of seven INV falsifiers did not falsify anything
_2026-09-22 · booth_
The U4 contract carried seven invariants, each with a *Falsifiable:* line, and
each had a test. **The code-review panel showed that five of the seven tests
would still pass under a change that defeats the invariant they name.** Gróa's
"per INV entry, what would still pass" section is the single most useful thing
either panel produced on this unit.
| INV | what the test asserted | what still passed |
|---|---|---|
| 1 (no new arithmetic) | the clock moved after a view | special-casing `.viewed` inside `_newest_mtime` — the exact new arithmetic INV-1 forbids |
| 3 (`is_held` is pure) | the right answer, once | `is_held` doing I/O, or `return True` unconditionally |
| 4 (every surface says why) | a substring on `GET /` | dropping the line from the booth header, the marks page, or the board branch |
| 5 (a view cannot fail a request) | `record_view` did not raise | a second `touch` outside the guard, 500ing all three routes |
| 6 (unreadable marks hold) | the corrupt booth survived | a sweeper that deletes nothing at all (no doomed sibling in the fixture) |
| 7 (machine reads do not hold) | `.viewed` was absent | a handler writing any other non-dot file, holding the booth open just as well |
**The shape of the error is the same every time: the test asserted the OUTCOME
the author was thinking about, not the DISCRIMINATOR the invariant names.** A
green test proved the happy path and nothing about the invariant. Writing the
falsifiable line in the contract did not produce a falsifying test — it produced
a test that *cited* one.
Fixed by rewriting each to fail under the change that defeats it: same-mtime
equivalence with an arbitrary non-lock dotfile (plus a `.lock` that must NOT
count); `is_held` called with marks belonging to a booth that does not exist on
disk; one test per rendered surface, each rendering only its own; the three
routes GET against a chmod'd booth; a doomed sibling; the AGE asserted rather
than the marker. **The board-header pair was verified RED against the pre-fix
template rather than assumed** — which is the step that makes "fixed, not
amended" trustworthy.
**The method to keep: for each invariant, name a change that defeats it and ask
whether the test goes red.** If you cannot name one, the invariant is not
falsifiable yet. Regin and Kimi independently proposed this as a contract-time
"vacuity pass"; heid rates this round the strongest evidence for it so far, and
it is a `/heid*` skill proposal sitting with the operator, not a change to this
repo.
+77 -181
View File
@@ -1,6 +1,6 @@
# Persistent memory — booth
_Last updated: 2026-09-21_
_Last updated: 2026-09-22_
> **Always check for `/tmp/booth-dev-handoff.md`** — if it exists and its
> `Written:` stamp is under 8 hours old, read it (it carries the in-flight
@@ -17,190 +17,86 @@ loop it turned out to actually be.
## Current state / in-flight
_As of 2026-09-21:_
_As of 2026-09-22:_
- **v1 is gated on seven units** in `ROADMAP.md`, ordered by dependency:
**U1 → U2 → {U3, U4, U5} → U7**, with **U6 independent** of all of them.
- **U1 (one item record) has landed** at `ce598b3` and is verified against its
own invariants, not just its commit message: INV-1 holds (no `classify` /
`doc_kind` / `read_blurred` / `render_doc` call survives in a route body),
the zoom and doc templates render the caption they now receive, the
re-exports are asserted by a test. 192 tests green, `0.1.15`.
- **U2 (marks) has landed** — `booth/marks.py`, contract at
`docs/contracts/u2_marks.contract.md`, 242 tests green. Not yet deployed.
- **U2 is DEPLOYED and the migration is done.** The service was restarted
2026-09-21 23:41 and again after the `auto_reload` fix; all four legacy
sidecars imported (`dfa-concepts/dfa`, `run07-decisions/decisions`,
`sc-iso-spread/spread`, `sindra-voice-1/anchor`, all still open) with the
sidecars left on disk. Verified live: index + 25 booths x {booth page, marks
page, marks.json} all 200, plus zoom views on five booths.
- **Still needs the operator: the release tier.** U2 changes the CLI surface for
17 consuming handles (`booth asks` -> `booth marks`, new `marks-import`) and is
a v1 unit, so it reads minor-worthy — which needs explicit approval per the
SemVer rule. Nothing is bumped or tagged; the work is committed as SHAs.
- **`/heid-contract-review` on the U2 contract is still in flight** (panel mode,
posted 2026-09-21, redacted copy at
`/tmp/heid-contract-review/booth-20260922-061015/`). Triage it when it lands —
the code is written, so findings land as follow-up fixes rather than contract
edits. The seam review ran in-session and its nine findings are already folded
into the contract and the code.
- **Open, operator's call:** whether U6 (benches) runs in parallel with U2 or
strictly after it. Nothing blocks on the answer; U6 touches different storage
and a different surface, so it cannot be broken by U2.
- Live service is `active` on `:8090` (systemd `--user`), 25 booths.
- **v1 is gated on seven units** in `ROADMAP.md`, dependency-ordered
**U1 → U2 → {U3, U4, U5} → U7**, with **U6 independent**.
- **U1, U2, U4 and U5 are landed.** U1 `ce598b3`; U2 `c7f9437` → `v0.2.0`,
`5e41108` → `v0.2.1`, `026a1fc` → `v0.2.2`; U5 `c015a91` + `95beede` →
`v0.3.0`. **U4 landed 2026-09-22** — 396 tests green (341 → 396), deployed and
verified live, 24/24 booth pages 200, layout probe clean.
- **U4 released as `v0.4.0`** (operator approved the minor on 2026-09-22).
`c3a97c1` is the unit; the release commit carries the pre-existing fixes the
bug-hunt panel surfaced in touched files. The tag waited for the last gate to
close, per the `v0.2.0` lesson — see Tried and abandoned.
- ⚠ **The 17 consuming handles have NOT been told** that `keep` no longer means
"waiting on an answer". That is the one coordination this release genuinely
warrants, and a fleetwide post needs operator approval before it is sent.
- **THE NEXT UNIT IS THE OPERATOR'S CALL.** U3 (declared embed seam) and U6
(benches) are both unblocked; U7 waits on the rest. U6 is independent of
everything and was conceptually unblocked by U5 giving job 5 a home; U3 is
where verbatim-booth provenance was deferred to, and U4 added a fourth reason
to want it — a verbatim booth has no Booth-rendered header, so its lifetime
line lives only on the index card and the marks page.
- **No gate is outstanding.** All three ran on U4 and were folded in: the
`/heid-contract-review` panel (`01M34VX0SH23Y3VC92E7GM4S70`), the
`/heid-code-review` panel (`01M34WAFJC3RTERFYBBZJN1SVG`) and the
`/heid-bug-hunt` (`01M34Y2R0RAJRSN36Q8K4KAB36`). All loops closed with heid.
The U5 round's three are also closed (`01M340PNVRS21HPASZT38PXQPN`,
`01M341E9XAPZEFBSPK9HPGAM0S`, `01M343SXX27Z47C3STXXRC7M42`).
- **Two dated predictions are pending and must not be forgotten.** U5's adoption
re-measure on **2026-09-29** (two counts, see its entry — already at 3 of 24
announced and 2 with a `why`, all from peers told nothing), and the `.forever`
re-count **on or after 2026-10-06**, a fortnight after U4 landed, which is
U4's success criterion. ⚠ Only 4 booths carry marks at all, so the hold's live
blast radius is small and the prediction rests on both halves of U4 — see its
entry for what a null result would and would not mean.
- **Three methodology proposals from this session sit with the operator**, routed
by heid rather than decided unilaterally: reshaping the paraphrase gate toward
a drift-check for narrative-heavy contracts, a standing
"green-tests-prove-nothing" direction for the code-review gate, and regin's
table-vs-signature consistency pass. They are changes to the `/heid*` skills,
not to this repo.
- The booth set churns hard: 26 → 24 during this session as the sweeper ran.
Re-count rather than trusting any number written here.
## Recent decisions
- `[2026-09-21]` **Deterministic order is a cross-cutting v1 invariant** —
operator directive, mid-implementation. Every ordered collection the Booth
renders must have a *stated* rule producing the same sequence on every render
of the same state; the rule can be anything defensible (byte order, time, an
explicit number, an arbitrary-but-recorded sequence), but no rule at all is
forbidden. It binds harder here than elsewhere because the Booth's job is
**comparison** — the operator judges tile 47 against tile 47 and refers to
artifacts positionally, so an order that moves between renders misfiles a flag
or a note rather than crashing. Recorded as `ROADMAP.md` § "Cross-cutting
invariant" (with the per-collection table) and `CLAUDE.md` invariant 6, and
tested. Still undecided and must be settled before those units ship: **U7's
section ordering and compare pairing**, and **U6's bench listing**.
- `[2026-09-21]` **U2 (marks) landed.** One primitive replacing three
mechanisms. `pick` / `note` / `flag` in one `.marks.json` per booth, one read
path (`marks_for`), one openness predicate (`open_marks`), rendered beside the
artifact on the tile, at full size in the zoom, and in the panel. `flag` and
`note` had no write path at all before this — the selection loop
(`golden-candidates`, `sindra-finalists`, the `pancake-*` ladders) was running
through chat. 242 tests. Details worth carrying: `asks.py` kept `normalize_ask`
and gained `build_answer` (the 2026-09-09 partial-answer semantics preserved by
moving, not rewriting) and LOST its five sidecar-storage functions;
`GET /b/<n>/marks.json` was added because remote sessions polled
`<stem>.answer.json` over HTTP and the sidecar's removal would have taken that
capability with it; `/b/<n>/asks` 308s to `/marks`.
- `[2026-09-21]` **A partially-answered pick now counts as OPEN** — declared, not
smuggled. The old index badge tested `answer is None`, so a half-answered
four-question ask read as closed on the index while the panel beside it
rendered `◐ partial`: the two disagreed about the same booth. Open is the
reading that makes U4 correct — a lifetime rule that unpinned a booth on the
first radio click would sweep a review in flight.
- `[2026-09-21]` **The U2 seam review earned its place, and the record should
say how.** Nine findings against the real `booth.asks` / `booth.items` /
`booth.inline` surfaces, two of which changed scope or behaviour: `inline.py`
was missing from `touches` entirely (its `place()` indexes asks by
**subscript**, which a frozen dataclass refuses — nothing else in the service
does that), and the partial-answer inconsistency above. The cold
`/heid-contract-review` pass is artifact-only by design and structurally
cannot see a sibling module, so neither it nor a same-model self-review would
have found either. Two more surfaced later and are worth the same note: a
SECOND subscript in `inline.place` the seam review undercounted, and a
regression in my own legacy importer that a retargeted test caught — a
malformed sidecar that renders `⚠ broken` today would have silently vanished
on migration.
- `[2026-09-21]` **Marks are stored as one `.marks.json` per booth**, atomic
temp-file + `os.replace`, `fcntl` lock on the read-modify-write — operator
decision, this session. Two alternatives were weighed and lost: a sidecar
per item (`<rel>.marks.json`) and extending the existing `<stem>.ask.json`
shape. Rationale, and the reason it is not `links.md`-shaped: **(a)** U4
makes *"does this booth owe an answer?"* a hot question — the sweep asks it
per booth per tick and the index asks it per card per page load, so per-item
sidecars turn it into a full walk of all 25 booths, one of which holds 270
files; **(b)** `links.md` is an `O_APPEND` content-hash log because **17
agent handles write it concurrently**, whereas marks have exactly one writer
(the operator, in one browser) and many readers — a different problem that
must not inherit the append-log design; **(c)** `.blurred` / `.pins` /
`.forever` already establish the per-booth dotfile as the house shape for
operator state, and `booth_items()`'s dotfile skip means it costs nothing in
counts, galleries or zips. Accepted cost: a corrupt `.marks.json` loses that
booth's marks rather than one item's. Implementation deferred to U2 —
tracked at `ROADMAP.md` U2 and by this entry.
- `[2026-09-21]` **U7's section premise is half wrong, and it is the half that
matters** — found by re-measuring `~/booth-data` rather than trusting the IA
doc. The IA says sections come from subfolders that already exist on disk;
true, but **every booth that actually needs navigation is flat**:
`pancake-v3-full` (270 items, 0 subfolders), `pancake-v4-full` (270, 0),
`sindra20-engines` (98 items + 99 caption sidecars, 0), `sindra-finalists`
(86 + 87, 0). Subfolders exist on exactly two booths — `pewpew-ui-brief` (7,
nested to `_ds/powerpellet-design-system-<uuid>/preview`) and `dfa-concepts`
(1) — and **both are reports**, the job where grid navigation matters least.
So sections stay worth shipping and `Item.section` stays right, but they are
**not** "most of the navigation fix": the rail, the filters and grid keyboard
are all of it. Worth noting for whoever writes U7: `sindra20-engines` encodes
its structure in the **filename prefix** (`b2-s1-<subject>-<seed>`), which is
where a grouping heuristic would actually pay. The IA doc's claim about what
sections buy needs a line struck — not yet edited.
- `[2026-09-21]` **`sindra-finalists` is U2's `flag` motivation caught in the
act** — 86 items, every one captioned, and the booth's entire name is "the
ones the operator picked." That loop currently runs through chat, which is
the defect `flag` closes. Evidence, not argument.
- `[2026-09-21]` **The information architecture and the v1 gate landed**
(`726822b`): `docs/design/information-architecture.md` names the single
defect — *one lifetime (24h from last touch) and one shape (a folder),
serving five jobs with different lifetimes and different shapes* — and
`ROADMAP.md` gates v1 on seven units, each closing a **measured** defect
rather than a wish. Both were written after a measurement pass over the live
service, and the measurements are the load-bearing part.
- `[2026-09-21]` **The `.forever` diagnosis is a stated, falsifiable
prediction.** U4 (derived lifetime) predicts the kept-rate falls to the
genuinely-durable booths. Re-measured today: **14 of 25 booths kept (56%)**,
against the 54% the IA doc recorded. **Re-count a fortnight after U4 lands.**
If it does not move, the diagnosis was wrong and the boolean was doing
something else. Tracked in the IA doc's Booth section and by this entry.
- `[2026-09-21]` **Extracted from `eshpfi` into its own repo.** The accreted
service came over whole, tests included, so `tests/test_booth.py` (1581 lines)
is the regression net the v1 rewrite is checked against.
- `[2026-09-22]` **U4 landed — lifetime is derived, not declared** — three states, viewing is activity, and no new arithmetic anywhere → `persistent-memory.d/2026-09-22-u4-derived-lifetime-landed.md`
- `[2026-09-22]` **The `.forever` diagnosis got a live positive control** — 3 of the 4 booths awaiting an answer were ALSO hand-pinned — RE-COUNT 2026-10-06 → `persistent-memory.d/2026-09-22-forever-had-a-live-positive-control.md`
- `[2026-09-22]` **Four independent paths to one fail-open delete** — the bug-hunt panel's class, and the zsh word-splitting trap that shipped an empty bundle → `persistent-memory.d/2026-09-22-four-paths-to-one-fail-open-delete.md`
- `[2026-09-22]` **Two reads of one file are not one read of one state** — a TOCTOU seam that composes two correct readers into a fail-open delete → `persistent-memory.d/2026-09-22-two-reads-are-not-one-state.md`
- `[2026-09-22]` **Five of seven INV falsifiers did not falsify anything** — read before writing a *Falsifiable:* line; a green test cited one rather than being one → `persistent-memory.d/2026-09-22-vacuous-falsifiers.md`
- `[2026-09-22]` **The third one-branch template miss** — this repo's recurring blind spot; read before adding a fact to any template → `persistent-memory.d/2026-09-22-third-one-branch-template-miss.md`
- `[2026-09-22]` **The size cap opened a service-wide hang** — a FIFO has st_size 0; a bound that trusts it inherits what it does not mean → `persistent-memory.d/2026-09-22-size-cap-opened-a-hang.md`
- `[2026-09-22]` **An existing test stopped me retiring documented behaviour** — the clean fix for the mtime race would have silently changed TTL doctrine → `persistent-memory.d/2026-09-22-doctrine-not-defect.md`
- `[2026-09-22]` **Two U5 panels, and prose reached a released outage** — read the detail before assuming a conformance finding stops at its own module → `persistent-memory.d/2026-09-22-u5-panels-reached-a-released-bug.md`
- `[2026-09-22]` **U5's adoption prediction split in two** — the handle rides for free, the why must be learned — RE-MEASURE 2026-09-29 → `persistent-memory.d/2026-09-22-u5-adoption-split-in-two.md`
- `[2026-09-22]` **The U2 bug-hunt panel was not ceremony** — the lock-unlink race and the TTL guard that was failing at its own job → `persistent-memory.d/2026-09-22-u2-bug-hunt-panel.md`
- `[2026-09-22]` **The lenient reader's blast radius was the whole service** — marks_for runs per booth per index load; a raise there is an outage → `persistent-memory.d/2026-09-22-lenient-reader-blast-radius.md`
- `[2026-09-22]` **`booth marks` / `booth answer` got real exit codes** — read it before changing anything the 17 consuming handles call → `persistent-memory.d/2026-09-22-cli-exit-codes.md`
- `[2026-09-22]` **`scripts/booth` went from zero tests to five** — they run the real script under system python3, so they also check INV-1 → `persistent-memory.d/2026-09-22-scripts-booth-got-tests.md`
- `[2026-09-21]` **v0.2.0 was tagged while a gate was in flight** — the sequencing lesson: if a gate is outstanding, the tag waits → `persistent-memory.d/2026-09-21-v020-tagged-with-a-gate-in-flight.md`
- `[2026-09-21]` **A write over a damaged `.marks.json` wiped the booth** — the reads-lenient / writes-strict asymmetry, and why it exists → `persistent-memory.d/2026-09-21-marks-write-wiped-judgment.md`
- `[2026-09-21]` **Seam review and cold panel had zero overlap, twice** — evidence for running both; neither substitutes for the other → `persistent-memory.d/2026-09-21-two-gates-are-complementary.md`
- `[2026-09-21]` **Every code-changing finding came from the AMBIGUITY pass** — a finding about the /heid-contract-review skill, not about this repo → `persistent-memory.d/2026-09-21-ambiguity-pass-did-the-work.md`
- `[2026-09-21]` **Deterministic order is a cross-cutting v1 invariant** — operator directive; read before adding ANY ordered surface → `persistent-memory.d/2026-09-21-deterministic-order-invariant.md`
- `[2026-09-21]` **U2 (marks) landed — one primitive for three mechanisms** — what moved where, and the HTTP mirror remote sessions poll → `persistent-memory.d/2026-09-21-u2-marks-landed.md`
- `[2026-09-21]` **A partially-answered pick counts as OPEN** — declared, not smuggled; it is the reading that makes U4 correct → `persistent-memory.d/2026-09-21-partial-answer-counts-as-open.md`
- `[2026-09-21]` **The U2 seam review earned its place, and how** — inline.place indexes by subscript — the miss a cold panel cannot see → `persistent-memory.d/2026-09-21-u2-seam-review-earned-it.md`
- `[2026-09-21]` **Marks are one `.marks.json` per booth** — operator decision with two rejected alternatives; read before restructuring → `persistent-memory.d/2026-09-21-marks-storage-decision.md`
- `[2026-09-21]` **U7's section premise is half wrong** — every booth that needs navigation is FLAT — read before starting U7 → `persistent-memory.d/2026-09-21-u7-section-premise-half-wrong.md`
- `[2026-09-21]` **`sindra-finalists` is U2's flag motivation, caught live** — evidence, not argument → `persistent-memory.d/2026-09-21-sindra-finalists-is-the-motivation.md`
- `[2026-09-21]` **The information architecture and the v1 gate landed** — the single defect the seven units decompose → `persistent-memory.d/2026-09-21-ia-and-v1-gate-landed.md`
- `[2026-09-21]` **The `.forever` diagnosis is a falsifiable prediction** — U4's success criterion — re-count a fortnight AFTER U4 lands → `persistent-memory.d/2026-09-21-forever-diagnosis-is-a-prediction.md`
- `[2026-09-21]` **Extracted from `eshpfi` into its own repo** — test_booth.py is the regression net the v1 rewrite is checked against → `persistent-memory.d/2026-09-21-extracted-from-eshpfi.md`
## Tried and abandoned
- `[2026-09-21]` **Letting Jinja hot-reload templates while the repo is the
deployment root** — the cause of a live outage the same day U2 landed, and the
sharpest foot-gun in the repo. `booth.service` sets `WorkingDirectory` to this
repo, so the running service imports these files with no build step and no
staging copy. Python is read once at process start; Jinja's `FileSystemLoader`
re-reads a template **on every render**. Editing `booth.html` therefore
deployed it instantly against Python from 22:03 that knew nothing about
`item_marks`, and **19 of 25 live booths returned 500** with
`UndefinedError: 'item_marks' is undefined`. Neither the old code nor the new
code was broken — the service was running both at once.
**The lesson that generalises:** a skew between a process and the disk under it
is invisible to the test suite by construction, so no amount of green tests
would have caught it; the operator found it. Fixed at the source rather than
with a reminder — the `Environment` is hand-built with `auto_reload=False`, so
there is now ONE staleness rule (nothing takes effect until you restart) and
the running process is always a coherent snapshot of one commit. Asserted by
`test_templates_do_not_hot_reload_from_disk`. Watch the second-order risk the
fix introduces: a hand-built `Environment` does not inherit `autoescape` from
the `Jinja2Templates` constructor, and booth names, item names and mark text
are all agent-authored strings landing in HTML.
- `[2026-09-21]` **Five separate mechanisms to get one question next to one
artifact** — `.forever`, the link board, `inline.py`'s placeholder DSL,
`wrap_verbatim_html`'s six regexes, and the floating amber asks chip plus
`/b/<n>/asks`. Every one is a *correct local fix* to the same global
mismatch, which is exactly why they accumulated without anyone making a bad
call. **The foot-gun is the sixth one:** the next "just add a small thing for
this case" reads as reasonable and is the pattern. The git log carries the
signature — every feature ships, then takes 2–5 patches for cases the single
shape did not anticipate. Check the ROADMAP gate before adding a mechanism.
- `[2026-09-21]` **Regex-injecting chrome into arbitrary author HTML**
(`wrap_verbatim_html` + `_HEAD_CLOSE_RE`, `_HTML_OPEN_RE`, `_DOCTYPE_RE`,
`_BODY_CLOSE_RE`, `_HTML_CLOSE_RE`, `_ICON_RE`, and the doctype/charset
ordering constraints they thread). It works today and is **still live** —
but it is the single most fragile thing in the service and it is load-bearing
for the operator's most important workflow. Slated for deletion at U3 in
favour of a declared seam (`/_booth/embed.js`, mounted through a real DOM
API), which costs an author one line and removes the whole class. Do not
extend the regex set in the meantime; if a verbatim page breaks, that is an
argument for U3, not for a seventh pattern.
- `[2026-09-21]` **A boolean escape hatch as the lifetime mechanism.**
`.forever` was added because a 24h TTL genuinely did not fit some booths —
and then 56% of live booths ended up on it, which means it is not "ephemeral
with an exception", it is two lifetimes wearing one lifetime's clothes, with
the operator doing the sorting by hand. Replaced at U4 by lifetime derived
from state (an open mark pins; viewing is activity; `keep` survives as an
explicit reasoned pin rather than the only way to say "not yet").
- `[2026-09-21]` **Letting the link board absorb the announce job.** `booth
link` is an `O_APPEND` write with no identity and no stated rule, so
re-announcing a bench appends a row instead of updating one, and a booth URL
rots the moment its booth is swept — **145 of 211 rows (69%) pointed at
nothing**, and 22 were the same target re-posted (talk 5×, peedlar 4×). The
rot is **structural, not drift**. The lesson that cost the most: enforcing
the link rule without first giving the announce job a home (`.booth.json`
provenance on the index, U5) just makes it homeless.
- `[2026-09-21]` **Tagging a release while a review gate was in flight** — cost a same-hour v0.2.1 and a correction to 15 handles → `persistent-memory.d/2026-09-21-tagging-with-a-gate-in-flight.md`
- `[2026-09-21]` **Letting the write path share the read path's leniency** — a tolerant reader and a tolerant writer are not the same decision → `persistent-memory.d/2026-09-21-tolerant-writer-over-tolerant-reader.md`
- `[2026-09-21]` **Letting Jinja hot-reload templates in the deployment root** — caused a live outage: 19 of 25 booths at 500. Why auto_reload=False → `persistent-memory.d/2026-09-21-jinja-hot-reload-outage.md`
- `[2026-09-21]` **Five mechanisms to get one question beside one artifact** — the accretion signature this whole v1 rewrite is undoing → `persistent-memory.d/2026-09-21-five-mechanisms-one-job.md`
- `[2026-09-21]` **Regex-injecting chrome into arbitrary author HTML** — the defect U3 exists to close → `persistent-memory.d/2026-09-21-regex-injecting-chrome.md`
- `[2026-09-21]` **A boolean escape hatch as the lifetime mechanism** — why `.forever` is a symptom; the defect U4 exists to close → `persistent-memory.d/2026-09-21-boolean-escape-hatch-as-lifetime.md`
- `[2026-09-21]` **Letting the link board absorb the announce job** — 69% rot; U5 gave the job a home, which is what unblocks U6 → `persistent-memory.d/2026-09-21-link-board-absorbing-announce.md`
+1 -1
View File
@@ -1,6 +1,6 @@
[project]
name = "booth"
version = "0.2.0"
version = "0.4.0"
description = "The Booth — a dead-simple standing web server that scans a data dir of drop-folders and renders each as an ephemeral media 'booth' (image/webm/audio auto-gallery, or a folder's own index.html verbatim). Also accepts browser/curl uploads for pickup under a human-readable id. 24h TTL, then the folder is wiped. Fleet tool for CC sessions to surface A/B and smoke results to the operator."
requires-python = ">=3.11"
dependencies = [
+216 -31
View File
@@ -3,13 +3,17 @@
# folder under $BOOTH_DATA_DIR; this is sugar over mkdir/cp so you get the URL
# back.
#
# booth new <name> make an empty booth, print its URL
# booth add <name> <file>... copy files into a booth (creates it), print URL
# booth new <name> [--why W] [--title T]
# make an empty booth, print its URL
# booth add <name> <file>... [--why W] [--title T]
# copy files into a booth (creates it), print URL
# booth url <name> print a booth's URL
# booth ls list booths (kept ones marked ★)
# booth rm <name> wipe a booth now (TTL would eventually anyway)
#
# booth keep <name> exempt a booth from the 24h sweep, forever
# (NOT for "waiting on an answer" — an open
# pick holds its own booth, see below)
# booth unkeep <name> hand it back to the sweeper
# booth link <url> [description] append a link to the standing link board
# booth links list the board, numbered, with entry ids
@@ -22,6 +26,17 @@
# booth answer <name> <id> [--wait [SECS]]
# print ONE pick's answer (exit 1 if unanswered);
# --wait polls until it lands (default 3600 s)
#
# EXIT CODES for the two reading verbs. A read that FAILED gets its own code so
# a caller can tell "not yet" from "the file is damaged" — conflating them is
# how a broken `.marks.json` used to look like an unanswered question and wait
# out the full hour.
# marks 0 read ok · 1 --wait timed out with picks open · 3 unreadable
# answer 0 answered · 1 unanswered · 2 no such pick · 3 unreadable ·
# 4 the pick hydrated broken and can never be answered
#
# `answer` and `marks` use the SAME openness predicate. A partially-answered
# pick is still open to both; a broken one is closed to both.
# booth marks-import <name> import legacy *.ask.json into .marks.json
# booth asks <name> alias for `marks` (deprecated)
#
@@ -46,12 +61,28 @@
# access, so they poll the HTTP mirror instead:
# http://10.100.10.50:8090/b/<name>/marks.json
#
# THE 24h RULE AND ITS ONE EXCEPTION. Every booth is wiped 24h after its last
# THE 24h RULE AND ITS THREE STATES. Every booth is wiped 24h after its last
# activity — that is the contract, and it is why nobody has to clean up after
# themselves. `keep` drops a `.forever` sentinel that exempts one booth from the
# sweep and moves it into its own lane at the top of the index. Use it for
# durable operator-facing boards, not for run output. `unkeep` is just `rm` of
# the sentinel, so putting a board back under the sweeper costs nothing.
# themselves. Two things exempt a booth, and only the first is a button:
#
# KEPT `keep` drops a `.forever` sentinel that exempts one booth from the
# sweep and moves it into its own lane at the top of the index. Use it
# for durable operator-facing boards, not for run output. `unkeep` is
# just `rm` of the sentinel, so putting a board back costs nothing.
# HELD a booth with an UNANSWERED pick is never swept, automatically, for as
# long as the question is open. You do not press anything: `booth ask`
# is what holds it, and the operator answering is what releases it. A
# partially-answered pick still counts as open, so a review in flight
# cannot be swept out from under him.
#
# So: DO NOT `keep` a booth just because you are waiting on an answer. That was
# the old workaround, it is what made 70% of live booths "durable", and it is
# no longer needed. `keep` means durable. The question holds its own booth.
#
# VIEWING IS ACTIVITY TOO. The operator opening a booth page resets its clock —
# if he is still looking at it, it is still alive. Your polling does NOT: `booth
# marks --wait` and the `marks.json` endpoint are machine reads and deliberately
# do not count, so a session cannot hold its own booth open by waiting on it.
#
# DELETING A KEPT BOARD: `booth rm <name>` works on kept boards too and deletes
# NOW — it announces that the board was kept, so wiping something durable is
@@ -59,15 +90,29 @@
# card drops the sentinel, the card moves to the ephemeral lane, and the × wipes
# it from there.
#
# DO NOT "unkeep and let it expire". Removing the sentinel BUMPS the booth
# directory's mtime, and a booth's age is the newest mtime in its tree — so a
# released board's clock RESETS and it survives another full 24h. Unkeep-and-wait
# is a delay, not a delete. Use `rm` (or the UI ×) when you mean now.
# DO NOT "unkeep and let it expire". RELEASING A BOARD IS ACTIVITY — you just
# touched it — so a released board's clock resets and it survives another full
# 24h. Unkeep-and-wait is a delay, not a delete. Use `rm` (or the UI ×) when you
# mean now. (This was true before U4 as an accident of directory mtime; it is
# now the stated rule, which is why it no longer needs a warning shaped like a
# surprise.)
#
# `link` is the reason the exception exists: agent sessions hand the operator
# URLs that then drown in terminal scrollback. They go on a standing kept board
# instead, with provenance, so they outlive the session that produced them.
#
# ANNOUNCE YOUR BOOTH. `--why` is one line saying what the operator is looking
# at and why he should care; it lands on the index card and on the booth page
# beside your handle, taken from $ALTHING_HANDLE. It is optional and nothing
# breaks without it — but a booth that cannot say what it is has no way to ask
# for attention except by posting its URL somewhere, which is exactly how the
# link board came to be 69% dead rows. The booth is the place to say it.
#
# booth add r18-ab out/*.png --why "pick the denoiser, left column is v3"
#
# Re-announcing (a second `new` or `add` on the same booth) updates the why and
# KEEPS the original creation stamp: the booth appeared once.
#
# On a host that is NOT nh3-dev, rsync into the data dir instead, e.g.:
# rsync -a ./out/ nh3-dev:booth-data/my-run/
set -euo pipefail
@@ -78,23 +123,90 @@ KEEP=".forever" # must match KEEP_MARKER in b
BLUR=".blurred" # one booth-relative item path per line; see `blur` below
LINKS_BOARD="${BOOTH_LINKS_BOARD:-links}"
# `--why` / `--title` for `new` and `add`. Pulled out of "$@" wherever they
# appear, so `booth add b *.png --why "..."` and `booth add b --why "..." *.png`
# both work — a glob is usually last and a flag usually after it, but nothing
# enforces that and a session should not have to care.
# OMITTED IS NOT EMPTY. `booth new x --why "..."` then `booth add x out/*.png`
# is the ordinary sequence, and while an omitted flag meant "" the second
# command silently erased the sentence the first one existed to record. So the
# shell tracks WHETHER the flag was given, and only passes it on when it was —
# an explicit `--why ""` still clears, which is a different intention.
WHY=""; TITLE=""; WHY_SET=0; TITLE_SET=0; ARGS=()
strip_announce_flags() {
ARGS=(); WHY_SET=0; TITLE_SET=0
while [ $# -gt 0 ]; do
case "$1" in
--why) [ $# -ge 2 ] || usage; WHY="$2"; WHY_SET=1; shift 2 ;;
--title) [ $# -ge 2 ] || usage; TITLE="$2"; TITLE_SET=1; shift 2 ;;
--why=*) WHY="${1#--why=}"; WHY_SET=1; shift ;;
--title=*) TITLE="${1#--title=}"; TITLE_SET=1; shift ;;
*) ARGS+=("$1"); shift ;;
esac
done
}
# Announce a booth. Goes through booth/manifest.py rather than printf-ing JSON
# from the shell, because a why containing a quote, a backslash or a newline is
# not an edge case — it is a sentence somebody wrote.
# announce <dir> <handle> [title] [why] — the trailing two are passed as
# environment variables that are UNSET when the flag was not given, because
# that is the only way the shell can say "leave it alone" rather than "".
announce() {
local -a envs
envs=( "BOOTH_SRC=$(cd "$(dirname -- "$(readlink -f -- "$0")")/.." && pwd)"
"BOOTH_ANN_DIR=$1" "BOOTH_ANN_HANDLE=$2" )
[ "${TITLE_SET:-0}" = 1 ] && envs+=( "BOOTH_ANN_TITLE=${3:-}" )
[ "${WHY_SET:-0}" = 1 ] && envs+=( "BOOTH_ANN_WHY=${4:-}" )
env "${envs[@]}" python3 -c '
import os, pathlib, sys
sys.path.insert(0, os.environ["BOOTH_SRC"])
try:
from booth.manifest import write_manifest
kw = {}
# Absent means the flag was omitted; present-and-empty means it was given
# as "" and the poster meant to take the line back.
if "BOOTH_ANN_TITLE" in os.environ: kw["title"] = os.environ["BOOTH_ANN_TITLE"]
if "BOOTH_ANN_WHY" in os.environ: kw["why"] = os.environ["BOOTH_ANN_WHY"]
write_manifest(pathlib.Path(os.environ["BOOTH_ANN_DIR"]),
os.environ["BOOTH_ANN_HANDLE"], **kw)
except Exception as exc:
# A booth that could not announce itself is still a booth. Say so on stderr
# and carry on: failing `booth add` over its metadata would lose the files
# the session just copied, which is a far worse trade.
print(f"booth: could not write the announcement: {exc}", file=sys.stderr)
'
}
# Who is posting. The same chain `link` uses for its rows, so provenance means
# the same thing on the board and on the card.
whoami_handle() {
echo "${ALTHING_HANDLE:-${BOOTH_SOURCE:-$(hostname -s 2>/dev/null || echo unknown)}}"
}
usage() {
echo "usage: booth {new <name>|add <name> <file>...|url <name>|ls|rm <name>|keep <name>|unkeep <name>|blur <name> <file>...|unblur <name> <file>...|link <url> [description]|links|unlink <id|index>|ask <name> <id> <prompt> <option>... [--no-notes]|marks <name> [--wait [SECS]]|answer <name> <id> [--wait [SECS]]|marks-import <name>}" >&2
echo "usage: booth {new <name> [--why W] [--title T]|add <name> <file>... [--why W] [--title T]|url <name>|ls|rm <name>|keep <name>|unkeep <name>|blur <name> <file>...|unblur <name> <file>...|link <url> [description]|links|unlink <id|index>|ask <name> <id> <prompt> <option>... [--no-notes]|marks <name> [--wait [SECS]]|asks <name> (deprecated alias for marks)|answer <name> <id> [--wait [SECS]]|marks-import <name>}" >&2
exit 2
}
cmd="${1:-}"; shift || true
case "$cmd" in
new)
strip_announce_flags "$@"
set -- ${ARGS+"${ARGS[@]}"}
[ $# -ge 1 ] || usage
mkdir -p -- "$DATA/$1"
announce "$DATA/$1" "$(whoami_handle)" "$TITLE" "$WHY"
echo "$URL/b/$1/"
;;
add)
strip_announce_flags "$@"
set -- ${ARGS+"${ARGS[@]}"}
[ $# -ge 2 ] || usage
name="$1"; shift
mkdir -p -- "$DATA/$name"
cp -- "$@" "$DATA/$name/"
announce "$DATA/$name" "$(whoami_handle)" "$TITLE" "$WHY"
echo "$URL/b/$name/"
;;
url)
@@ -170,6 +282,11 @@ case "$cmd" in
board="$DATA/$LINKS_BOARD"
mkdir -p -- "$board"
: > "$board/$KEEP" # the board is durable by definition
# The board announces itself as the SERVICE's, not as any one agent's:
# seventeen handles post to it, so no handle owns it. Idempotent — a second
# link keeps the original creation stamp.
TITLE_SET=1 WHY_SET=1 announce "$board" "booth" "$LINKS_BOARD" \
"the standing link board — every agent session posts here"
# Provenance, because a bare URL is unreadable three days later: who posted
# it, from where, and when.
who="${ALTHING_HANDLE:-${BOOTH_SOURCE:-$(hostname -s 2>/dev/null || echo unknown)}}"
@@ -255,7 +372,7 @@ print("removed: %s %s" % (removed["desc"], removed["url"]))
import os, pathlib, sys
sys.path.insert(0, os.environ["BOOTH_SRC"])
from booth.asks import AskError
from booth.marks import declare_pick
from booth.marks import MarksCorrupt, declare_pick
booth, mid, prompt, *opts = sys.argv[1:]
try:
declare_pick(pathlib.Path(booth), mid,
@@ -263,11 +380,22 @@ try:
"notes": os.environ["ASK_NOTES"] == "1"})
except AskError as exc:
sys.exit("bad pick: %s" % exc)
except MarksCorrupt as exc:
sys.exit("this booth'"'"'s .marks.json is damaged, so nothing was written: %s" % exc)
' "$DATA/$name" "$mid" "$prompt" "${opts[@]}"
echo "$URL/b/$name/#mark-$mid"
;;
marks|asks)
# booth marks <name> [--wait [SECS]] (`asks` is the deprecated alias)
#
# EXIT CODES. 0 = the read succeeded and the document is on stdout; 1 =
# --wait gave up with picks still open (the document is still printed); 3 =
# the marks could not be read at all. A reader that CRASHED must never look
# like an answer — the old shape printed a traceback and exited 0, so a
# caller piping to `jq` saw success and got nothing.
#
# Whether anything is still open is in the payload's `open` list. The read
# verb does not encode it in its status: a successful read is a success.
[ $# -ge 1 ] || usage
name="$1"; shift
wait_s=0
@@ -276,19 +404,41 @@ except AskError as exc:
# os.replace, and a 2 s cadence is plenty for a human clicking a radio.
deadline=$(( $(date +%s) + wait_s ))
while :; do
BOOTH_SRC="$(cd "$(dirname -- "$(readlink -f -- "$0")")/.." && pwd)" python3 -c '
# CAPTURED, not streamed. Printing inside the loop wrote one whole JSON
# document per poll, so `booth marks b --wait | jq` got several values
# concatenated and could parse none of them. The wait is a wait; the
# print is the result, and it happens once.
rc=0
out="$(BOOTH_SRC="$(cd "$(dirname -- "$(readlink -f -- "$0")")/.." && pwd)" python3 -c '
import json, os, pathlib, sys
sys.path.insert(0, os.environ["BOOTH_SRC"])
from booth.marks import as_dict, marks_for, open_marks
marks = marks_for(pathlib.Path(sys.argv[1]))
print(json.dumps({"marks": [as_dict(m) for m in marks],
"open": [m.id for m in open_marks(marks)]},
ensure_ascii=False, indent=2))
sys.exit(1 if open_marks(marks) else 0)
' "$DATA/$name" && exit 0
# exit 1 from the reader means at least one pick is still open
if [ "$wait_s" -eq 0 ]; then exit 0; fi
try:
from booth.marks import as_dict, marks_for, open_marks, read_error
booth = pathlib.Path(sys.argv[1])
# Ask FIRST whether the file is readable. `marks_for` answers "no marks"
# for a damaged file, which is the right answer for a page and the wrong
# one for a session that wants to know whether its question survived.
broken = read_error(booth)
if broken:
print(f"booth: {broken}", file=sys.stderr)
sys.exit(3)
marks = marks_for(booth)
doc = json.dumps({"marks": [as_dict(m) for m in marks],
"open": [m.id for m in open_marks(marks)]},
ensure_ascii=False, indent=2)
except Exception as exc:
print(f"booth: cannot read marks: {exc}", file=sys.stderr)
sys.exit(3)
print(doc)
sys.exit(2 if open_marks(marks) else 0)
' "$DATA/$name")" || rc=$?
case "$rc" in
0) printf '%s\n' "$out"; exit 0 ;; # read ok, nothing open
2) if [ "$wait_s" -eq 0 ]; then printf '%s\n' "$out"; exit 0; fi ;;
*) echo "cannot read marks in $name" >&2; exit 3 ;;
esac
if [ "$(date +%s)" -ge "$deadline" ]; then
printf '%s\n' "$out"
echo "timed out after ${wait_s}s with marks still open in $name" >&2; exit 1
fi
sleep 2
@@ -302,20 +452,55 @@ sys.exit(1 if open_marks(marks) else 0)
if [ "${1:-}" = "--wait" ]; then wait_s="${2:-3600}"; fi
deadline=$(( $(date +%s) + wait_s ))
while :; do
BOOTH_SRC="$(cd "$(dirname -- "$(readlink -f -- "$0")")/.." && pwd)" python3 -c '
rc=0
out="$(BOOTH_SRC="$(cd "$(dirname -- "$(readlink -f -- "$0")")/.." && pwd)" python3 -c '
import json, os, pathlib, sys
sys.path.insert(0, os.environ["BOOTH_SRC"])
from booth.marks import marks_for
booth, mid = sys.argv[1:3]
m = next((x for x in marks_for(pathlib.Path(booth)) if x.id == mid), None)
try:
from booth.marks import marks_for, open_marks, read_error
booth, mid = sys.argv[1:3]
broken = read_error(pathlib.Path(booth))
if broken:
print(f"booth: {broken}", file=sys.stderr)
sys.exit(3)
# id AND shape, matching the web route. Matching on id alone reported a
# note id as "unanswered" and then polled it for an hour — a question that
# could never be answered because it was never a question.
marks = marks_for(pathlib.Path(booth))
m = next((x for x in marks if x.id == mid and x.shape == "pick"), None)
# THE openness predicate, not a second spelling of it. `answer is None` is
# what this read used to test, and it disagreed with `marks --wait` on a
# PARTIALLY answered pick: one verb returned the half-filled form while the
# other blocked on the same booth at the same instant. U2 put openness in
# one function precisely so the two could not drift.
still_open = m is not None and m in open_marks(marks)
except Exception as exc:
print(f"booth: cannot read marks: {exc}", file=sys.stderr)
sys.exit(3)
if m is None:
sys.exit(2)
if m.answer is None:
if m.error:
# Not open, and never going to be: the web route refuses this form with a
# 400, so waiting on it is waiting on nothing. `marks --wait` already
# returns immediately here; this is the other half of that agreement.
print(f"booth: pick is broken and cannot be answered: {m.error}",
file=sys.stderr)
sys.exit(4)
if still_open:
sys.exit(1)
print(json.dumps(m.answer, ensure_ascii=False, indent=2))
' "$DATA/$name" "$mid" && exit 0
rc=$?
if [ "$rc" -eq 2 ]; then echo "no such pick: $name/$mid" >&2; exit 1; fi
' "$DATA/$name" "$mid")" || rc=$?
case "$rc" in
0) printf '%s\n' "$out"; exit 0 ;;
2) echo "no such pick: $name/$mid" >&2; exit 2 ;;
# A read that FAILED is not "not yet". Conflating them sent --wait
# spinning for the full hour on a broken file and then blamed the
# operator for not answering.
3) echo "cannot read marks in $name" >&2; exit 3 ;;
# A pick that hydrated broken is refused by the web route, so no answer
# can ever land. Waiting on it is waiting on nothing.
4) exit 4 ;;
esac
if [ "$wait_s" -eq 0 ]; then echo "unanswered: $URL/b/$name/#mark-$mid" >&2; exit 1; fi
if [ "$(date +%s)" -ge "$deadline" ]; then
echo "timed out after ${wait_s}s waiting on $name/$mid" >&2; exit 1
+42 -1
View File
@@ -19,7 +19,22 @@ USAGE
<a python with playwright> scripts/layout-probe.py [URL ...]
Exits 0 if every control is hittable, 1 if any is occluded. No arguments
probes the booth index and every booth linked from it.
probes the INDEX ONLY — it does not follow booth links, and the docstring
claimed it did until 2026-09-22. Pass booth URLs explicitly to cover them:
scripts/layout-probe.py http://10.100.10.50:8090/{,b/my-run/}
⚠ PROBING A BOOTH PAGE RESETS THAT BOOTH'S TTL CLOCK (U4). A GET of `/b/<n>/`
is a view, and a view is activity — that is the rule, and this script is not
exempt from it just because it is ours. Sweeping every booth page therefore
buys every booth another full TTL. Harmless and recoverable (nothing is
deleted, things merely live longer), named here so nobody debugs it later as a
sweeper that stopped working. The index-only default does NOT do this: browsing
the index is deliberately not a view.
⚠ In zsh an unquoted `$URLS` does NOT word-split, so a variable holding
several URLs arrives as ONE argument and the probe silently reports
"2 page(s)" while covering two. Use an array and `"${URLS[@]}"`.
"""
import sys
from playwright.sync_api import sync_playwright
@@ -62,6 +77,32 @@ def probe(page, url: str) -> list[str]:
card.hover(timeout=1500)
except Exception:
pass
# ⚠ OPEN EVERY <details> FIRST. A control inside a CLOSED one is laid out
# but sits outside its collapsed parent's box, so `elementFromPoint` at its
# centre returns an ancestor and it reports OCCLUDED — 23 of them on
# `sindra-set`, every one a false positive, because the only way an operator
# reaches that button is by opening the disclosure first. Verified both
# ways: closed -> elementFromPoint returns div.gallery; opened -> the button
# itself, and a real trial click lands on it.
#
# Opening rather than SKIPPING is deliberate. Skipping would make the probe
# quiet by declaring put-away controls out of scope, and the add-note button
# inside `details.item-addnote` is exactly the kind of control this
# instrument exists to check. Open it and ask the real question.
#
# ⚠ ONE evaluate over the whole document, NOT a locator loop. `.all()` hands
# back positional locators that re-resolve against the CURRENT DOM, and
# `details:not([open])` stops matching an element the moment it is opened —
# so opening them one at a time shrinks the set underneath the indices and
# some are never opened at all. That left exactly the closed-<details>
# false positives this block exists to remove: 1 on booth-redesign, 3 on
# cr123a-to-d-sleeve, stable across five runs and invisible as a bug
# because a false positive looks like a finding. Measured both ways at
# 150 ms and 1000 ms settle: the loop reports them at either wait, the
# single pass reports none at either. The variable was the method, not the
# timing.
page.evaluate("document.querySelectorAll('details').forEach(d => d.open = true)")
page.wait_for_timeout(150)
for el in page.locator("button, a.dl-link, a.thumb").all():
try:
# ⚠ elementFromPoint is VIEWPORT-relative. Without scrolling first,
+11 -1
View File
@@ -781,6 +781,10 @@ def test_releasing_a_board_RESETS_its_ttl_clock(tmp_path):
(kept / KEEP_MARKER).unlink()
# Unlinking the sentinel by hand, which is what this test is about: the
# directory-entry change is what moves the clock. Releasing through the
# ROUTE now also records a view, so the behaviour is stated rather than
# incidental — `test_releasing_a_board_RECORDS_A_VIEW` in test_lifetime.py.
assert booth_age_seconds(kept) < 60, "unlink bumped the dir mtime"
assert sweep_once(tmp_path, ttl_seconds=3600) == [], "so it is NOT swept yet"
assert kept.exists()
@@ -1537,7 +1541,13 @@ def test_booth_page_offers_keep_when_ephemeral_and_release_when_kept(client):
_png(d / "x.png")
body = c.get("/b/bo/").text
assert "☆ keep" in body and "release" not in body.split("boothhead")[1][:900]
# Sliced on the ELEMENT, not the bare word: `boothhead` has appeared in the
# stylesheet this page carries since long before this assertion, so
# `split("boothhead")[1]` was reading CSS and passing on luck. It went red
# the first time a new rule landed above the old one (U5's .prov), which is
# the only reason anybody noticed. Same assertion, aimed at the markup.
head = body.split('class="boothhead"')[1][:900]
assert "☆ keep" in body and "release" not in head
c.post("/b/bo/keep", data={"next": "/b/bo/"}, follow_redirects=False)
body = c.get("/b/bo/").text
+348
View File
@@ -0,0 +1,348 @@
"""`scripts/booth` — the surface every fleet session actually calls.
It had no tests at all, which the 2026-09-22 bug-hunt panel found the hard way:
its guard-strength table returned UNVERIFIED for every CLI claim because nothing
in the suite executes the script. Two of that round's findings live in here.
These run the real script under the real system `python3` with no venv, which
also makes them a live check on INV-1 (stdlib-only): a third-party import in
`marks.py` fails here the same way it fails on a fleet host.
"""
import json
import os
import pathlib
import subprocess
import pytest
SCRIPT = pathlib.Path(__file__).parent.parent / "scripts" / "booth"
# Exit codes the verbs promise. 0 is a successful read; a reader that CRASHED
# must never be one of the meaningful codes, or a caller cannot tell "no" from
# "broken" — which is the whole finding.
OK, UNANSWERED, NO_SUCH_PICK, READER_FAILED = 0, 1, 2, 3
def run(data, *args, **kw):
env = {**os.environ, "BOOTH_DATA_DIR": str(data), "BOOTH_URL": "http://booth.invalid"}
return subprocess.run([str(SCRIPT), *args], capture_output=True, text=True,
env=env, timeout=30, **kw)
@pytest.fixture
def booth(tmp_path):
b = tmp_path / "b"
b.mkdir()
return tmp_path, b
def _declare(booth_dir, mark_id="winner"):
import sys
sys.path.insert(0, str(pathlib.Path(__file__).parent.parent))
from booth.marks import declare_pick
declare_pick(booth_dir, mark_id,
{"prompt": "Which one?", "options": ["A", "B"]})
def test_marks_prints_one_json_document(booth):
"""`booth marks <name>` is a read. Its stdout is parsed by the session that
called it, so it has to be ONE document — and exit 0, because the read
succeeded. Whether a pick is open is in the payload's `open` list, which is
where a caller should read it from."""
data, b = booth
_declare(b)
r = run(data, "marks", "b")
assert r.returncode == OK, r.stderr
doc = json.loads(r.stdout)
assert doc["open"] == ["winner"]
def test_marks_wait_prints_once_not_once_per_poll(booth):
"""`--wait` polls every 2 s and printed the whole document on every pass, so
a capture held several concatenated JSON values and `jq` could not read any
of them. The wait is a wait; the print is the result."""
data, b = booth
_declare(b)
import sys
sys.path.insert(0, str(pathlib.Path(__file__).parent.parent))
from booth.marks import answer_pick
# Answer it after the first poll so --wait genuinely loops at least once.
r = subprocess.Popen([str(SCRIPT), "marks", "b", "--wait", "20"],
stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True,
env={**os.environ, "BOOTH_DATA_DIR": str(data),
"BOOTH_URL": "http://booth.invalid"})
import time
time.sleep(3)
answer_pick(b, "winner", "A")
out, err = r.communicate(timeout=30)
assert r.returncode == OK, err
json.loads(out) # ONE document, or this raises
def test_marks_reports_a_reader_failure_instead_of_printing_garbage(booth):
"""A traceback on stdout with exit 0 is the worst of both: the caller's `jq`
sees success and gets nothing. A read that could not happen is its own
answer and gets its own code."""
data, b = booth
(b / ".marks.json").write_bytes(b"\xff\xfe not utf-8 at all")
r = run(data, "marks", "b")
assert r.returncode == READER_FAILED, f"rc={r.returncode} out={r.stdout!r}"
def test_answer_distinguishes_a_crash_from_an_unanswered_pick(booth):
"""`answer` funnelled a reader crash and "not yet answered" through the same
exit 1, so `--wait` spun for the full hour on a broken file and then blamed
the operator for not answering."""
data, b = booth
_declare(b)
r = run(data, "answer", "b", "winner")
assert r.returncode == UNANSWERED
(b / ".marks.json").write_bytes(b"\xff\xfe not utf-8 at all")
r = run(data, "answer", "b", "winner", "--wait", "6")
assert r.returncode == READER_FAILED, (
"a crash was read as 'unanswered' and waited out the timeout"
)
def test_answer_on_a_note_id_says_no_such_pick(booth):
"""`answer` matched on id alone while the web route filters on shape, so a
note id was reported 'unanswered' and polled forever — a question that could
never be answered because it was never a question."""
data, b = booth
import sys
sys.path.insert(0, str(pathlib.Path(__file__).parent.parent))
from booth.marks import write_note
write_note(b, "a.png", "just a note")
r = run(data, "answer", "b", "note-1")
assert r.returncode == NO_SUCH_PICK
assert "no such pick" in r.stderr
# ---- U5: self-announcing booths ---------------------------------------------
def _manifest(booth_dir):
import sys
sys.path.insert(0, str(pathlib.Path(__file__).parent.parent))
from booth.manifest import read_manifest
return read_manifest(booth_dir)
def test_new_announces_the_booth(tmp_path):
"""`$ALTHING_HANDLE` is the whole provenance story: the session already has
it, so the booth can say who made it without anybody typing a name."""
env = {**os.environ, "ALTHING_HANDLE": "shutter-dev"}
r = subprocess.run([str(SCRIPT), "new", "r18-ab", "--why", "pick the winner"],
capture_output=True, text=True, timeout=30,
env={**env, "BOOTH_DATA_DIR": str(tmp_path),
"BOOTH_URL": "http://booth.invalid"})
assert r.returncode == 0, r.stderr
m = _manifest(tmp_path / "r18-ab")
assert m.handle == "shutter-dev"
assert m.why == "pick the winner"
def test_new_without_a_why_is_still_legal(tmp_path):
"""The flags are optional and existing call sites keep working. A booth
that says only who made it is still a booth that said something."""
r = subprocess.run([str(SCRIPT), "new", "scratch"], capture_output=True,
text=True, timeout=30,
env={**os.environ, "ALTHING_HANDLE": "booth-dev",
"BOOTH_DATA_DIR": str(tmp_path),
"BOOTH_URL": "http://booth.invalid"})
assert r.returncode == 0, r.stderr
m = _manifest(tmp_path / "scratch")
assert m.handle == "booth-dev" and m.why == ""
def test_add_announces_and_still_copies_the_files(tmp_path):
"""`add` is the verb most sessions actually use — it creates the booth AND
fills it — so the why has to ride on it or it rides nowhere."""
src = tmp_path / "src"
src.mkdir()
(src / "a.txt").write_text("content")
r = subprocess.run([str(SCRIPT), "add", "r18-ab", str(src / "a.txt"),
"--why", "second pass", "--title", "R18 A/B"],
capture_output=True, text=True, timeout=30,
env={**os.environ, "ALTHING_HANDLE": "booth-dev",
"BOOTH_DATA_DIR": str(tmp_path),
"BOOTH_URL": "http://booth.invalid"})
assert r.returncode == 0, r.stderr
assert (tmp_path / "r18-ab" / "a.txt").read_text() == "content"
m = _manifest(tmp_path / "r18-ab")
assert m.why == "second pass" and m.title == "R18 A/B"
def test_add_re_announcing_keeps_the_original_created(tmp_path):
"""The common shape: `new` opens the booth, `add` drops the second batch and
sharpens the why. The booth appeared once."""
env = {**os.environ, "ALTHING_HANDLE": "booth-dev",
"BOOTH_DATA_DIR": str(tmp_path), "BOOTH_URL": "http://booth.invalid"}
src = tmp_path / "a.txt"
src.write_text("x")
subprocess.run([str(SCRIPT), "new", "b", "--why", "first"], check=True,
capture_output=True, timeout=30, env=env)
first = _manifest(tmp_path / "b").created
subprocess.run([str(SCRIPT), "add", "b", str(src), "--why", "sharper"],
check=True, capture_output=True, timeout=30, env=env)
after = _manifest(tmp_path / "b")
assert after.created == first
assert after.why == "sharper"
def test_the_link_board_announces_itself_as_the_booths_own(tmp_path):
"""No exemption list. The standing board is made by the service and posted
to by seventeen handles, so no single agent owns it — `booth` is the
truthful answer, and it keeps the rule to one line."""
r = subprocess.run([str(SCRIPT), "link", "http://example.invalid", "a thing"],
capture_output=True, text=True, timeout=30,
env={**os.environ, "ALTHING_HANDLE": "booth-dev",
"BOOTH_DATA_DIR": str(tmp_path),
"BOOTH_URL": "http://booth.invalid"})
assert r.returncode == 0, r.stderr
m = _manifest(tmp_path / "links")
assert m is not None and m.handle == "booth"
assert m.why
def test_the_flags_can_sit_on_either_side_of_the_files(tmp_path):
"""`booth add b *.png --why "..."` and `booth add b --why "..." *.png` both
work. A glob is usually last and a flag usually after it, but nothing
enforces that and a session should not have to remember which."""
src = tmp_path / "a.png"
src.write_bytes(b"x")
env = {**os.environ, "ALTHING_HANDLE": "booth-dev",
"BOOTH_DATA_DIR": str(tmp_path), "BOOTH_URL": "http://booth.invalid"}
for name, args in (("after", ["add", "after", str(src), "--why", "w"]),
("before", ["add", "before", "--why", "w", str(src)])):
r = subprocess.run([str(SCRIPT), *args], capture_output=True, text=True,
timeout=30, env=env)
assert r.returncode == 0, r.stderr
assert _manifest(tmp_path / name).why == "w"
assert (tmp_path / name / "a.png").exists(), "the files stopped being copied"
def test_a_why_survives_quotes_and_non_ascii_and_is_flattened(tmp_path):
"""The reason this goes through manifest.py instead of printf-ing JSON from
the shell: a why containing a quote, a backslash or a newline is not an edge
case, it is a sentence somebody wrote. Newlines flatten because the field
renders inside a card's sub-line."""
r = subprocess.run(
[str(SCRIPT), "new", "b", "--why", 'he said "pick v3" — line1\nline2 · ünï'],
capture_output=True, text=True, timeout=30,
env={**os.environ, "ALTHING_HANDLE": "booth-dev",
"BOOTH_DATA_DIR": str(tmp_path), "BOOTH_URL": "http://booth.invalid"})
assert r.returncode == 0, r.stderr
why = _manifest(tmp_path / "b").why
assert why == 'he said "pick v3" — line1 line2 · ünï'
def test_a_flag_with_no_value_does_not_eat_the_booth_name(tmp_path):
"""`booth new b --why` with nothing after it must not consume `b` as the
value and then create a booth called nothing. Usage, and no directory."""
r = subprocess.run([str(SCRIPT), "new", "b", "--why"], capture_output=True,
text=True, timeout=30,
env={**os.environ, "BOOTH_DATA_DIR": str(tmp_path),
"BOOTH_URL": "http://booth.invalid"})
assert r.returncode == 2
assert "usage:" in r.stderr
assert not (tmp_path / "b").exists()
def test_a_bare_add_does_not_wipe_the_why_the_new_set(tmp_path):
"""`booth new x --why "..."` then `booth add x out/*.png` is THE sequence,
and the second call must not erase the first one's sentence. The module
distinguishes omitted from empty; the shell has to carry that distinction
across, which means an UNSET variable, not an empty one."""
env = {**os.environ, "ALTHING_HANDLE": "booth-dev",
"BOOTH_DATA_DIR": str(tmp_path), "BOOTH_URL": "http://booth.invalid"}
src = tmp_path / "a.png"
src.write_bytes(b"x")
subprocess.run([str(SCRIPT), "new", "b", "--why", "pick the denoiser",
"--title", "R18 A/B"],
check=True, capture_output=True, timeout=30, env=env)
subprocess.run([str(SCRIPT), "add", "b", str(src)],
check=True, capture_output=True, timeout=30, env=env)
m = _manifest(tmp_path / "b")
assert m.why == "pick the denoiser", "a bare `booth add` wiped the why"
assert m.title == "R18 A/B"
def test_an_explicitly_empty_why_still_clears_it(tmp_path):
"""Omitted means unchanged; supplied-and-empty means the poster meant to
take it back. Both have to be reachable from the shell."""
env = {**os.environ, "ALTHING_HANDLE": "booth-dev",
"BOOTH_DATA_DIR": str(tmp_path), "BOOTH_URL": "http://booth.invalid"}
subprocess.run([str(SCRIPT), "new", "b", "--why", "wrong"], check=True,
capture_output=True, timeout=30, env=env)
subprocess.run([str(SCRIPT), "new", "b", "--why", ""], check=True,
capture_output=True, timeout=30, env=env)
assert _manifest(tmp_path / "b").why == ""
def test_answer_and_marks_agree_about_what_open_means(tmp_path):
"""U2 made `_is_open` THE openness predicate — "nothing else may spell this
out" — and `booth answer`'s reader spelled it out anyway, as
`if m.answer is None`. So a PARTIALLY answered pick read as done to
`answer` and still-open to `marks --wait`: one verb returns the half-filled
form and the other blocks on the same booth at the same instant.
Found 2/4. The two verbs are the session's whole view of the loop, and a
session that asks both gets two answers.
"""
import sys
sys.path.insert(0, str(pathlib.Path(__file__).parent.parent))
from booth.marks import answer_pick, declare_pick
b = tmp_path / "b"
b.mkdir()
declare_pick(b, "batch", {
"title": "R18",
"questions": [
{"key": "q1", "prompt": "One?", "options": ["keep", "drop"]},
{"key": "q2", "prompt": "Two?", "options": ["keep", "drop"]},
],
})
answer_pick(b, "batch", {"q1": "keep", "q2": None}) # partial
env = {**os.environ, "BOOTH_DATA_DIR": str(tmp_path),
"BOOTH_URL": "http://booth.invalid"}
marks = subprocess.run([str(SCRIPT), "marks", "b"], capture_output=True,
text=True, timeout=30, env=env)
answer = subprocess.run([str(SCRIPT), "answer", "b", "batch"],
capture_output=True, text=True, timeout=30, env=env)
still_open = "batch" in json.loads(marks.stdout)["open"]
assert still_open, "a partial answer stopped counting as open"
assert answer.returncode == UNANSWERED, (
"`answer` called a partially-answered pick done while `marks` called it open"
)
def test_answer_does_not_poll_forever_on_a_pick_that_cannot_be_answered(tmp_path):
"""The mirror failure. A pick whose declaration went bad hydrates with
`error` set, which makes it NOT open — so `marks --wait` returns at once
while `answer --wait` polled the full hour against a form the web route
refuses with a 400. Nothing was ever going to land."""
b = tmp_path / "b"
b.mkdir()
(b / ".marks.json").write_text(json.dumps({
"version": 1,
"marks": [{"id": "broken", "shape": "pick", "declaration": {},
"error": "pick has no declaration",
"created": "2026-09-21T00:00:00.000000+00:00"}],
}))
r = subprocess.run([str(SCRIPT), "answer", "b", "broken", "--wait", "8"],
capture_output=True, text=True, timeout=40,
env={**os.environ, "BOOTH_DATA_DIR": str(tmp_path),
"BOOTH_URL": "http://booth.invalid"})
assert r.returncode != 0
assert "broken" in r.stderr.lower() or "cannot" in r.stderr.lower()
File diff suppressed because it is too large Load Diff
+719
View File
@@ -0,0 +1,719 @@
"""U5 — self-announcing booths.
A booth carries `.booth.json` saying who posted it and why, and the index card
and the booth page render it. Closes job 5 (`Announce`) — the job nobody named,
whose absence is the measured cause of 145 dead link rows.
See docs/contracts/u5_booth_manifest.contract.md.
"""
import ast
import json
import os
import pathlib
import sys
import pytest
from booth.manifest import MANIFEST_FILE, Manifest, read_manifest, write_manifest
# ---- slice 1: the record and its storage ------------------------------------
def test_an_announcement_round_trips(tmp_path):
b = tmp_path / "r18-ab"
b.mkdir()
written = write_manifest(b, "booth-dev", why="pick the winning denoiser")
assert (b / MANIFEST_FILE).is_file()
got = read_manifest(b)
assert got == written
assert got.handle == "booth-dev"
assert got.why == "pick the winning denoiser"
assert got.error is None
def test_the_title_falls_back_to_the_directory_name(tmp_path):
"""A booth always has a display name. `title` is the one the poster chose
when there is one, and the folder name is a perfectly good one when there
is not — an empty heading on a card is worse than a plain one."""
b = tmp_path / "r18-ab"
b.mkdir()
assert write_manifest(b, "booth-dev").title == "r18-ab"
assert write_manifest(b, "booth-dev", title="R18 A/B").title == "R18 A/B"
def test_a_booth_that_never_announced_reads_as_none(tmp_path):
"""The normal case for every booth that predates this unit, and for every
booth that arrives by rsync — the documented path for any host that is not
nh3-dev, which never runs the CLI at all."""
b = tmp_path / "quiet"
b.mkdir()
assert read_manifest(b) is None
assert read_manifest(tmp_path / "does-not-exist") is None
def test_one_line_by_construction_not_by_convention(tmp_path):
"""`why` renders inside a card's sub-line, so a newline in it would break
the card rather than the field. Truncation and newline-stripping happen at
the WRITE, so nothing downstream has to remember."""
b = tmp_path / "b"
b.mkdir()
m = write_manifest(b, "booth-dev", why="first line\nsecond line\r\nthird")
assert "\n" not in m.why and "\r" not in m.why
assert "first line" in m.why and "second line" in m.why
long = write_manifest(b, "booth-dev", why="x" * 5000)
assert len(long.why) <= 200
assert len(write_manifest(b, "y" * 500).handle) <= 64
assert len(write_manifest(b, "booth-dev", title="t" * 500).title) <= 120
# ---- slice 2: the read cannot raise (INV-2) ---------------------------------
@pytest.mark.parametrize(
"payload",
[
b"{truncated", # not JSON at all
b"[]", # JSON, wrong shape
b'"a string"', # JSON, wronger shape
b"null",
b'{"handle": 7}', # right shape, wrong type
b'{"why": "no handle here"}', # the one required field missing
b"\xff\xfe not utf-8",
b"",
],
ids=["truncated", "list", "string", "null", "wrong-type", "no-handle",
"not-utf8", "empty"],
)
def test_a_damaged_manifest_never_raises(tmp_path, payload):
"""INV-2. `list_booths` calls this once per booth on every index page load,
so a read that can raise is a service-wide outage wearing a single-booth
bug's clothes. That is not a hypothetical — a poisoned `.marks.json` did
exactly that to `/` and `/healthz` across all 25 booths, and the fix shipped
in v0.2.2. The same reader posture, applied before the same mistake."""
b = tmp_path / "b"
b.mkdir()
(b / MANIFEST_FILE).write_bytes(payload)
got = read_manifest(b)
assert isinstance(got, Manifest)
assert got.error, "a damaged manifest read clean"
def test_damaged_is_not_the_same_as_absent(tmp_path):
"""INV-5. Silently folding "cannot be read" into "never announced" would
hide the one case somebody has to go and fix."""
absent = tmp_path / "absent"
absent.mkdir()
damaged = tmp_path / "damaged"
damaged.mkdir()
(damaged / MANIFEST_FILE).write_text("{oops")
assert read_manifest(absent) is None
assert read_manifest(damaged).error
def test_a_manifest_the_module_did_not_write_still_reads(tmp_path):
"""Hand-written is a supported input: the file is plain JSON in a folder the
operator owns, and half the point is that a booth is just a directory. Only
`handle` is required; everything else has a default."""
b = tmp_path / "b"
b.mkdir()
(b / MANIFEST_FILE).write_text(json.dumps({"handle": "shutter-dev"}))
got = read_manifest(b)
assert got.handle == "shutter-dev" and got.error is None
assert got.title == "b"
assert got.why == ""
# ---- slice 3: re-announcement (INV-3) ---------------------------------------
def test_re_announcing_preserves_created(tmp_path):
"""INV-3. `created` is when the booth APPEARED. Saying something more about
it later is not a second appearance, and a `booth add` on an existing booth
is the common case — the poster adds the second batch and sharpens the why."""
b = tmp_path / "b"
b.mkdir()
first = write_manifest(b, "booth-dev", why="first pass")
second = write_manifest(b, "booth-dev", why="second pass, sharper")
assert second.created == first.created
assert second.why == "second pass, sharper"
def test_re_announcing_over_a_damaged_file_does_not_inherit_its_created(tmp_path):
"""A `created` that cannot be read back is replaced rather than guessed at.
The alternative is a stamp that is silently wrong, which is worse than one
that is silently new."""
b = tmp_path / "b"
b.mkdir()
(b / MANIFEST_FILE).write_text("{not json")
m = write_manifest(b, "booth-dev", why="rescued")
assert m.created and m.error is None
assert read_manifest(b).why == "rescued"
# ---- slice 4: the write is atomic, and invisible to every listing -----------
def test_the_write_leaves_no_temp_file(tmp_path):
"""Half of the atomic-write promise, and the weaker half — see
`test_the_write_replaces_rather_than_truncating` for the part that actually
discriminates. Kept because a leaked `.tmp` is its own small defect: it
would sit in the booth forever and, unlike the manifest, nothing would ever
overwrite it."""
b = tmp_path / "b"
b.mkdir()
write_manifest(b, "booth-dev", why="x")
assert not list(b.glob("*.tmp")), "a temp file survived the write"
def test_a_manifest_is_not_an_item(tmp_path):
"""The whole integration story: it is a DOTFILE, so the existing
`startswith('.')` skip in `booth_items` already keeps it out of tiles,
counts and zips. No new exclusion rule anywhere. Asserted rather than
assumed, because the claim is load-bearing for the contract's scope."""
from booth.app import zip_booth
from booth.items import booth_items
b = tmp_path / "b"
b.mkdir()
(b / "a.txt").write_text("real content")
write_manifest(b, "booth-dev", why="x")
assert [i.rel for i in booth_items(b)] == ["a.txt"]
assert MANIFEST_FILE not in zip_booth(b).decode("latin-1")
def test_announcing_is_activity(tmp_path):
"""A manifest is a dotfile but not a `.lock` dotfile, so `_newest_mtime`
counts it. Creating or re-announcing a booth resets its TTL, which is right:
both are somebody touching it. The lock exemption added in v0.2.2 is for
machinery a READ path creates; this is a deliberate write."""
from booth.app import booth_age_seconds
b = tmp_path / "b"
b.mkdir()
old = 1_000_000_000
os.utime(b, (old, old))
write_manifest(b, "booth-dev", why="look at this")
assert booth_age_seconds(b, now=old + 90_000) < 86_400
def test_stdlib_only():
"""INV-4, and the reason this module exists separately from anything that
imports a third-party package. `scripts/booth` imports it under the system
python3 with NO venv, through a `python3 -c` heredoc no AST extractor can
see. It must also not import `booth.*`: a cross-import between two
stdlib-only modules is a second way for the invariant to break."""
src = pathlib.Path(__file__).parent.parent / "booth" / "manifest.py"
roots = set()
for node in ast.walk(ast.parse(src.read_text())):
if isinstance(node, ast.Import):
roots.update(a.name.split(".")[0] for a in node.names)
elif isinstance(node, ast.ImportFrom):
# A RELATIVE import (`from . import marks`) carries no module root
# and used to pass this walk unseen — which matters more here than
# in the shared copy, because this module forbids sibling imports
# outright. Recorded as `booth` so the assertion below catches it.
roots.add("booth" if node.level else
(node.module or "").split(".")[0])
assert not (roots - set(sys.stdlib_module_names)), (
f"booth/manifest.py imports outside the stdlib: "
f"{sorted(roots - set(sys.stdlib_module_names))}"
)
# ---- slice 5: what the operator actually sees -------------------------------
@pytest.fixture
def client(tmp_path):
from fastapi.testclient import TestClient
from booth.app import create_app
return TestClient(create_app(tmp_path, ttl_hours=24, start_sweeper=False)), tmp_path
def _booth(data, name, *, kept=False):
b = data / name
b.mkdir()
(b / "a.txt").write_text("content")
if kept:
(b / ".forever").touch()
return b
@pytest.mark.parametrize("kept", [False, True], ids=["ephemeral", "kept"])
def test_the_index_card_carries_the_announcement(client, kept):
"""BOTH LANES. Kept boards render first and are a separate block in
index.html, so patching only the ephemeral lane would leave the 15 kept
booths — the durable, most-looked-at ones — with exactly the defect this
unit closes. Same lesson as the `blurtoggle` macro: three branches, one
definition; here it is two lanes and one rule."""
c, data = client
b = _booth(data, "r18-ab", kept=kept)
write_manifest(b, "booth-dev", why="pick the winning denoiser")
html = c.get("/").text
assert "booth-dev" in html
assert "pick the winning denoiser" in html
@pytest.mark.parametrize("kept", [False, True], ids=["ephemeral", "kept"])
def test_a_booth_that_never_spoke_up_is_marked(client, kept):
"""All 26 live booths are in this state, and rsync keeps making more. The
marker is what makes the convention adoptable at all: the link board rotted
to 69% precisely because nothing ever showed which rows were dead.
ASSERTED ON THE CLASS, not on the word, and the test is named around it.
`pytest`'s `tmp_path` is derived from the TEST NAME and the index renders
`data_dir` in its empty-state hint — so a test called
`test_an_unannounced_booth_says_so` put the literal string "unannounced"
into the page and passed against a template that did not yet exist. A
structural hook cannot be spelled by accident — though it has to be the
rendered ELEMENT and not the bare class, since base.html ships a
`.prov-none{...}` rule into the very same page."""
c, data = client
_booth(data, "quiet", kept=kept)
html = c.get("/").text
assert 'class="prov prov-none"' in html
assert "unannounced" in html
def test_a_damaged_manifest_reads_differently_from_an_absent_one(client):
"""INV-5 on the surface the operator looks at, not just in the reader."""
c, data = client
b = _booth(data, "damaged")
(b / MANIFEST_FILE).write_text("{oops")
html = c.get("/").text
assert 'class="prov prov-broken"' in html
assert "unreadable" in html
assert c.get("/b/damaged/").status_code == 200
def test_an_announced_booth_with_no_why_shows_only_its_handle(client):
"""`booth new x` with no --why is legal and common. The card shows who made
it and does not invent a purpose or leave a dangling separator."""
c, data = client
b = _booth(data, "scratch")
write_manifest(b, "booth-dev")
html = c.get("/").text
assert "booth-dev" in html
assert 'class="prov prov-none"' not in html
def test_the_booth_page_header_carries_it_too(client):
"""Deliberate scope, not creep: a booth URL handed to the operator lands
HERE, never on the index. Job 5 is 'operator, look at this', so the page he
actually opens is where the answer has to be."""
c, data = client
b = _booth(data, "r18-ab")
write_manifest(b, "booth-dev", why="pick the winning denoiser")
html = c.get("/b/r18-ab/").text
assert "booth-dev" in html
assert "pick the winning denoiser" in html
def test_a_poisoned_manifest_cannot_take_down_the_index(client):
"""The v0.2.2 lesson, asserted for the new reader before it can repeat:
`list_booths` touches every booth on every page load, so one bad file must
cost that booth's provenance and nothing else."""
c, data = client
_booth(data, "good")
bad = _booth(data, "bad")
(bad / MANIFEST_FILE).write_bytes(b"\xff\xfe not utf-8 at all")
assert c.get("/").status_code == 200
assert c.get("/healthz").status_code == 200
def test_a_pickup_booth_announces_itself_as_the_booths_own(client):
"""No exemption list. A booth the service made says the service made it,
which is true — and it keeps the rule to one line: a booth with no manifest
is unannounced."""
c, data = client
r = c.post("/upload", files=[("files", ("a.txt", b"hello", "text/plain"))],
follow_redirects=False)
assert r.status_code in (200, 303)
booth = next(p for p in data.iterdir() if p.is_dir())
got = read_manifest(booth)
assert got is not None and got.handle == "booth"
assert 'class="prov prov-none"' not in c.get("/").text
# ---- findings from the cross-frontier CODE-REVIEW panel, 2026-09-22 ----------
#
# Heid panel (thread 01M341E9XAPZEFBSPK9HPGAM0S). Four arms, artifact-only.
# The round found ZERO drift in the strict sense and landed its weight one layer
# down, in test strength: five of the ten adopted findings are tests of mine
# that pass on the regression they exist to catch.
def test_the_read_survives_a_document_no_one_can_parse(tmp_path):
"""INV-2 said "never raises" and named a 4 GB file as a tested case. It was
not tested, and it did not hold: `except ValueError` catches a truncated
document, but `json.loads` on deeply nested input raises RecursionError,
which is not a ValueError and is not an OSError either.
`list_booths` calls this once per booth on every index load, so the one
file costs the whole front page — the exact outage shape the invariant
cites as its reason for existing. Three of four arms reached it
independently; the eight-payload parametrize above has no size or depth
case, so the hole stayed green.
"""
b = tmp_path / "b"
b.mkdir()
(b / MANIFEST_FILE).write_text("[" * 200_000 + "]" * 200_000)
got = read_manifest(b)
assert isinstance(got, Manifest) and got.error
def test_the_read_refuses_a_document_too_large_to_be_a_manifest(tmp_path):
"""The other half of INV-2's named case. A manifest is four short fields;
anything approaching a megabyte is not one, and reading it into memory to
discover that is the wrong order of operations. Bounded BEFORE the read, so
the size is checked by `stat` rather than survived."""
from booth.manifest import MANIFEST_MAX_BYTES
b = tmp_path / "b"
b.mkdir()
(b / MANIFEST_FILE).write_text('{"handle": "x", "why": "' +
"y" * (MANIFEST_MAX_BYTES + 100) + '"}')
got = read_manifest(b)
assert isinstance(got, Manifest) and got.error
assert "too large" in got.error
def test_a_hostile_directory_name_does_not_reach_the_record_raw(tmp_path):
"""`_one_line(title, TITLE_MAX) or booth.name` — the FALLBACK skips the
normalization the explicit value gets. A directory name may legally carry a
newline on POSIX and may be 255 bytes, and either lands in a card's
sub-line. Same shape on the read path's fallback."""
# 200-odd bytes, under the filesystem's own 255 limit but well over
# TITLE_MAX — and a newline, which POSIX permits in a filename.
name = "we" + "i" * 200 + "rd\nname"
b = tmp_path / name
b.mkdir()
m = write_manifest(b, "booth-dev")
assert "\n" not in m.title and len(m.title) <= 120
assert "\n" not in read_manifest(b).title
def test_the_write_replaces_rather_than_truncating(tmp_path):
"""The previous version of this test asserted only that no `*.tmp` file
survived — which a plain `write_text` passes, since it leaves no temp file
either. All four arms said so, and they were right.
THE INODE IS THE DISCRIMINATOR. `os.replace` publishes a different file over
the old name, so the inode changes; truncate-and-rewrite keeps it. That is
also exactly why the promise holds for a concurrent reader: it either has
the old inode, intact, or opens the new one, complete. A test of the
mechanism rather than of its litter.
(An earlier draft spied on `os.open` to prove the published path was never
opened for writing. It passed — vacuously. `Path.write_text` reaches the
syscall through `io.open` in C and never touches the Python-level
`os.open`, so the spy could not have fired either way. Recorded because
writing a second vacuous test while fixing the first is the failure mode
this whole round is about.)
"""
b = tmp_path / "b"
b.mkdir()
published = b / MANIFEST_FILE
write_manifest(b, "booth-dev", why="first")
first_inode = published.stat().st_ino
write_manifest(b, "booth-dev", why="second")
assert published.stat().st_ino != first_inode, (
"the manifest was rewritten in place, not replaced"
)
assert read_manifest(b).why == "second"
def test_the_temp_file_is_not_a_name_two_writers_share(tmp_path):
"""Every writer derived the same `.booth.json.tmp`. Two `booth add` calls on
one booth could then interleave through a stale descriptor into the
published path — the atomic-write promise is that READERS never see a
partial file, and it says nothing about two writers sharing a scratch name.
Marks are protected from this by their flock; the manifest has none."""
b = tmp_path / "b"
b.mkdir()
seen = set()
for i in range(5):
write_manifest(b, "booth-dev", why=f"pass {i}")
seen.update(p.name for p in b.iterdir() if p.name != MANIFEST_FILE)
assert not seen, f"left temp files behind: {sorted(seen)}"
from booth.manifest import _temp_path
names = {_temp_path(b).name for _ in range(20)}
assert len(names) > 1, "every writer derives the same temp name"
def test_a_bare_re_announce_does_not_wipe_the_why(tmp_path):
"""THE WORKFLOW IS `new --why` THEN `add`. Omitted flags meant empty
strings, and empty strings overwrote — so the second command silently
erased the sentence the first one existed to record, on the single most
common sequence this feature has.
Two arms of the paraphrase panel predicted it from the contract's wording
alone ("gains a manifest with no why" does not distinguish a first write
from a re-announce with the flags omitted). Every test I wrote passed
`--why` on both calls, so none of them could see it.
Omitted now means UNCHANGED; only a value that was actually supplied
overwrites, and an explicit empty string still clears.
"""
b = tmp_path / "b"
b.mkdir()
write_manifest(b, "booth-dev", title="R18 A/B", why="pick the denoiser")
write_manifest(b, "booth-dev") # a bare `booth add`
kept = read_manifest(b)
assert kept.why == "pick the denoiser", "a bare re-announce wiped the why"
assert kept.title == "R18 A/B"
write_manifest(b, "booth-dev", why="sharper") # supplied: overwrites
assert read_manifest(b).why == "sharper"
write_manifest(b, "booth-dev", why="") # explicit: clears
assert read_manifest(b).why == ""
def test_re_announcing_preserves_a_created_from_before_this_second(tmp_path):
"""`_now()` is whole-second resolution, so two `write_manifest` calls in a
row share a timestamp and the old preservation test passed even against an
implementation that regenerated `created` every time. Three of four arms
caught it. Seed a stamp that could not have come from now()."""
b = tmp_path / "b"
b.mkdir()
(b / MANIFEST_FILE).write_text(json.dumps({
"handle": "booth-dev", "title": "b", "why": "first",
"created": "2019-03-04T11:22:33-08:00",
}))
assert write_manifest(b, "booth-dev", why="second").created == \
"2019-03-04T11:22:33-08:00"
def test_only_the_manifest_module_opens_the_manifest(tmp_path):
"""INV-1, which had no guard anywhere. One resolver is only one resolver
while nothing else learns the filename."""
root = pathlib.Path(__file__).parent.parent
offenders = []
for src in sorted((root / "booth").glob("*.py")):
if src.name == "manifest.py":
continue
tree = ast.parse(src.read_text())
# STRING CONSTANTS, not raw text. A comment naming the file is prose
# about the design and harms nothing — the first version of this test
# scanned the whole source and went red on a comment explaining why a
# leaked `.booth.json.<hex>.tmp` keeps a booth alive. The invariant is
# about code that knows the filename, so ask the code.
docstrings = set()
for node in ast.walk(tree):
if isinstance(node, (ast.Module, ast.ClassDef,
ast.FunctionDef, ast.AsyncFunctionDef)):
body = getattr(node, "body", None)
if body and isinstance(body[0], ast.Expr) and \
isinstance(body[0].value, ast.Constant):
docstrings.add(id(body[0].value))
for node in ast.walk(tree):
if (isinstance(node, ast.Constant) and isinstance(node.value, str)
and id(node) not in docstrings and ".booth.json" in node.value):
offenders.append(f"{src.name}:{node.lineno}")
assert not offenders, f"{offenders} name the manifest file in code"
def test_announcing_is_activity_via_the_manifest_file_itself(tmp_path):
"""The previous version could not fail. Writing the manifest creates a
directory entry, which bumps the DIRECTORY's mtime, so the booth read as
fresh whether or not `_newest_mtime` counted the manifest at all — a test
of the side effect rather than of the thing.
Put the directory's clock back afterwards, leaving the manifest's own mtime
as the only thing that can keep the booth alive."""
import os
from booth.app import booth_age_seconds
b = tmp_path / "b"
b.mkdir()
old = 1_000_000_000
os.utime(b, (old, old))
write_manifest(b, "booth-dev", why="look at this")
os.utime(b, (old, old)) # only the file can save it now
assert booth_age_seconds(b, now=old + 90_000) < 86_400
def test_the_booth_header_marks_an_unannounced_booth_too(client):
"""The negative states were asserted on `/` only, so a header that rendered
provenance for clean manifests and nothing for the other two would have
passed the whole suite."""
c, data = client
_booth(data, "quiet")
damaged = _booth(data, "damaged")
(damaged / MANIFEST_FILE).write_text("{oops")
assert 'class="prov prov-none"' in c.get("/b/quiet/").text
assert 'class="prov prov-broken"' in c.get("/b/damaged/").text
def test_the_title_reaches_a_surface(client):
"""`--title` promised a display name and nothing rendered it — 4/4 on the
paraphrase panel, independently the top-ranked flag of that round. It lands
on the booth page heading, where there is room for it; the INDEX card keeps
the directory name, because that is the identity the operator navigates and
refers to positionally."""
c, data = client
b = _booth(data, "r18-ab")
write_manifest(b, "booth-dev", title="R18 A/B — denoiser bakeoff", why="w")
page = c.get("/b/r18-ab/").text
assert "R18 A/B — denoiser bakeoff" in page
assert "r18-ab" in page, "the directory name stopped being visible"
# ---- findings from the cross-frontier BUG-HUNT panel, 2026-09-22 -------------
#
# Heid panel (thread 01M343SXX27Z47C3STXXRC7M42). Four arms, artifact-only,
# diff-scoped. The strongest finding is one the SIZE CAP ITSELF opened.
def test_a_reader_never_blocks_on_a_file_that_is_not_a_file(tmp_path):
"""`stat` reports size 0 for a FIFO, so it sails under the byte cap — and
then `read_text` blocks in `read` with no EOF, so the `except` never runs
and the call never returns. `list_booths` reads every booth on every `GET /`
and `/healthz`, so ONE such file stalls the front page for the whole service,
with no error and no recovery short of a restart.
A symlink to `/dev/zero` is the same hole with unbounded allocation instead
of a hang: `st_size` is 0 there too.
Two of four arms reached it independently. The bound added an hour earlier
is what made it reachable — `st_size` answers a different question than
"can this be read", and a cap that trusts it inherits the difference.
"""
import os
import signal
b = tmp_path / "b"
b.mkdir()
os.mkfifo(b / MANIFEST_FILE)
# ⚠ ALARMED. Without this the RED state of this test does not fail, it HANGS
# — which is the defect itself, and is also useless as a signal: a suite that
# stops is indistinguishable from a suite that is slow. Five seconds is a
# thousand times the budget a read of a four-field file should need.
def _timeout(signum, frame):
raise AssertionError("read_manifest blocked on a FIFO and never returned")
old_handler = signal.signal(signal.SIGALRM, _timeout)
signal.alarm(5)
try:
got = read_manifest(b)
finally:
signal.alarm(0)
signal.signal(signal.SIGALRM, old_handler)
assert isinstance(got, Manifest) and got.error
assert "regular file" in got.error
def test_a_damaged_manifest_is_kept_when_it_is_replaced(tmp_path):
"""4/4, and it contradicted this repo's own doctrine. Marks made the rule
explicit in v0.2.1 — reads stay lenient, writes go strict, damaged bytes
STAY ON DISK — and the manifest's write replaced them outright.
The sharpest leg: a file that fails on ONE field still holds the others.
`{"handle": 7, "why": "the thing I wanted you to look at"}` reads as broken
and used to be destroyed whole, taking a `why` the re-announcer may not have
kept anywhere.
Quarantined rather than refused: refusing would fail `booth add` and lose
the files it was copying, which is the worse trade. One fixed-name
quarantine, so this cannot accumulate.
"""
from booth.manifest import QUARANTINE_FILE
b = tmp_path / "b"
b.mkdir()
damaged = json.dumps({"handle": 7, "why": "the thing I wanted you to see"})
(b / MANIFEST_FILE).write_text(damaged)
write_manifest(b, "booth-dev", why="rescued")
assert read_manifest(b).why == "rescued"
assert (b / QUARANTINE_FILE).read_text() == damaged, "the damaged bytes were destroyed"
def test_a_broken_record_normalizes_the_directory_name_too(tmp_path):
"""The third fallback. `write_manifest`'s and `read_manifest`'s were fixed
in the previous round and `_broken`'s was missed — same raw `booth.name`,
same card sub-line, same newline."""
b = tmp_path / ("wei" + "i" * 200 + "rd\nname")
b.mkdir()
(b / MANIFEST_FILE).write_text("{oops")
got = read_manifest(b)
assert got.error and "\n" not in got.title and len(got.title) <= 120
def test_an_identical_re_announce_does_not_touch_the_booth(tmp_path):
"""Marks learned this in v0.2.0: a write that changes nothing is not
activity and must not reset a booth's TTL. The manifest wrote
unconditionally, so `booth add` on an unchanged booth kept a dead one alive
— and `booth link` does it on every single post to the standing board."""
import os
b = tmp_path / "b"
b.mkdir()
write_manifest(b, "booth-dev", why="x")
path = b / MANIFEST_FILE
os.utime(path, (1_000_000_000, 1_000_000_000))
os.utime(b, (1_000_000_000, 1_000_000_000))
before = path.stat().st_mtime
write_manifest(b, "booth-dev", why="x") # identical
assert path.stat().st_mtime == before, "an identical re-announce rewrote the file"
def test_a_failed_write_leaves_no_temp_file_behind(tmp_path):
"""The unique temp name fixed a cross-writer hazard and created a litter
one: a fixed name is overwritten by the next writer, a random one is not.
And `.booth.json.<hex>.tmp` is NOT a `.lock`, so `_newest_mtime` counts it —
an orphaned temp would keep a dead booth alive forever."""
import os
b = tmp_path / "b"
b.mkdir()
real_replace = os.replace
def boom(src, dst, *a, **kw):
raise OSError("no space left on device")
os.replace = boom
try:
with pytest.raises(OSError):
write_manifest(b, "booth-dev", why="x")
finally:
os.replace = real_replace
assert not list(b.glob("*.tmp")), f"orphaned temp: {list(b.glob('*.tmp'))}"
+664 -3
View File
@@ -276,20 +276,32 @@ def test_as_dict_round_trips_through_json(tmp_path):
# ---- the stdlib-only invariant (INV-5) --------------------------------------
@pytest.mark.parametrize("module", ["marks", "asks", "links"])
@pytest.mark.parametrize("module", ["marks", "asks", "links", "manifest"])
def test_stdlib_only(module):
"""INV-5. scripts/booth imports these under the system python3 with NO venv,
through a `python3 -c` heredoc that no AST extractor can see — so nothing
but this test stands between a casual third-party import and `booth ask`
breaking on every fleet host."""
# `manifest` also carries a stricter copy in tests/test_manifest.py, which
# additionally forbids importing `booth.*` — a cross-import between two
# stdlib-only modules is a second way for this invariant to break.
src = pathlib.Path(__file__).parent.parent / "booth" / f"{module}.py"
tree = ast.parse(src.read_text())
roots = set()
for node in ast.walk(tree):
if isinstance(node, ast.Import):
roots.update(a.name.split(".")[0] for a in node.names)
elif isinstance(node, ast.ImportFrom) and node.level == 0 and node.module:
roots.add(node.module.split(".")[0])
elif isinstance(node, ast.ImportFrom):
# `node.level > 0` is a RELATIVE import (`from . import marks`),
# which has no `module` root to inspect and used to slip through
# this walk entirely. It cannot reach outside the package, so it is
# stdlib-safe by construction — but it is recorded rather than
# ignored, because `manifest.py` additionally forbids importing a
# sibling and its own test needs to see one.
if node.level:
roots.add("booth")
elif node.module:
roots.add(node.module.split(".")[0])
outside = {r for r in roots if r != "booth" and r not in sys.stdlib_module_names}
assert not outside, f"booth/{module}.py imports non-stdlib: {sorted(outside)}"
@@ -713,3 +725,652 @@ def test_a_real_write_then_a_no_op_leaves_the_file_alone(tmp_path):
set_flag(booth, "a.png", True) # idempotent: already flagged
assert path.stat().st_mtime == before, "an idempotent flag rewrote the file"
# ---- findings from the cross-frontier contract panel, 2026-09-22 -------------
#
# Heid panel (thread 01M33VSNFER4N1554G0Y0VC9C8). Four arms, artifact-only.
def test_a_write_over_a_corrupt_marks_file_refuses_instead_of_replacing(tmp_path):
"""DATA LOSS, shipped in v0.2.0. Found by Kimi (flag 2), converged with Hulda.
`marks_for` is deliberately lenient — an unparseable file reads as "no marks"
so a review page still loads. The write path inherited that leniency through
the same reader, so the next flag toggle appended one entry to an empty list
and atomically replaced the file: every judgment in that booth gone, from one
click, silently.
The read stays lenient and the WRITE goes strict. That asymmetry is the fix —
a page that renders without an annotation is recoverable, a file that
overwrote the operator's judgment is not, and this repo's standing rule is
that nothing deletes his data.
"""
from booth.marks import MarksCorrupt, set_flag, write_note
booth = tmp_path / "b"
booth.mkdir()
write_note(booth, "a.png", "judgment one")
write_note(booth, "b.png", "judgment two")
raw = (booth / MARKS_FILE).read_text()
(booth / MARKS_FILE).write_text(raw[: len(raw) // 2]) # truncated mid-write
with pytest.raises(MarksCorrupt):
set_flag(booth, "c.png", True)
# The damaged bytes are still on disk — untouched, recoverable by hand.
assert (booth / MARKS_FILE).read_text() == raw[: len(raw) // 2]
# And the read path is still lenient, so the page renders rather than 500s.
assert marks_for(booth) == []
def test_an_absent_or_empty_marks_file_is_not_corrupt(tmp_path):
"""The strict write path must not mistake "nothing yet" for "damaged"."""
from booth.marks import set_flag
booth = tmp_path / "b"
booth.mkdir()
assert set_flag(booth, "a.png", True) is not None # no file at all
(booth / MARKS_FILE).write_text("")
assert set_flag(booth, "b.png", True) is not None # zero bytes
(booth / MARKS_FILE).write_text('{"version": 1, "marks": []}')
assert set_flag(booth, "c.png", True) is not None # valid but empty
def test_a_pick_can_target_one_item(tmp_path):
"""Found by Hulda (flag 1), converged with Regin.
`Mark.target` carries an item rel, `marks_for_target` retrieves by it, and
the panel template already renders "on <item>" for a pick — but
`declare_pick` had no target parameter, so a session could not actually
produce one. A question about ONE artifact is the 2026-09-09 ruling's whole
point; the record supported it and the door was missing.
"""
from booth.marks import marks_for_target
booth = tmp_path / "b"
booth.mkdir()
declare_pick(booth, "which-crop", _single(), target="v3/DSC03389.jpg")
m = marks_for(booth)[0]
assert m.target == "v3/DSC03389.jpg"
assert [x.id for x in marks_for_target(marks_for(booth), "v3/DSC03389.jpg")] == ["which-crop"]
# and it still answers normally
answer_pick(booth, "which-crop", "A — baseline")
assert marks_for(booth)[0].answer["complete"] is True
def test_a_pick_target_cannot_escape_the_booth(tmp_path):
booth = tmp_path / "b"
booth.mkdir()
for bad in ("../outside.png", "/etc/passwd"):
with pytest.raises(AskError):
declare_pick(booth, "p", _single(), target=bad)
def test_redeclaring_a_pick_may_move_its_target(tmp_path):
booth = tmp_path / "b"
booth.mkdir()
declare_pick(booth, "p", _single(), target="a.png")
declare_pick(booth, "p", _single(), target="b.png")
assert marks_for(booth)[0].target == "b.png"
def test_import_adopts_a_legacy_answer_for_an_already_declared_pick(tmp_path):
"""Found by Gróa (flag 10).
The idempotence rule skipped any stem already present as a mark. If a
session had re-declared that stem through marks (so the mark exists, still
unanswered) while the operator's answer sat in the legacy sidecar, the import
skipped and that answer was stranded on disk forever — with the read path
forbidden from looking at sidecars. Adopting the answer preserves both rules:
idempotent, and never clobbers a NEWER judgment.
"""
from booth.asks import ANSWER_SUFFIX, build_answer, normalize_ask
from booth.marks import import_legacy_asks
booth = tmp_path / "b"
_sidecar(booth, "winner", _single())
doc = build_answer(normalize_ask(_single(), "winner"), "B — async", notes="from the sidecar")
(booth / f"winner{ANSWER_SUFFIX}").write_text(json.dumps(doc))
declare_pick(booth, "winner", _single()) # re-declared, unanswered
assert marks_for(booth)[0].answer is None
import_legacy_asks(booth)
got = marks_for(booth)[0]
assert got.answer is not None, "the legacy answer was stranded"
assert got.answer["choice"] == "B — async"
assert open_marks(marks_for(booth)) == []
def test_import_never_overwrites_an_answer_made_through_marks(tmp_path):
"""The other half of the same rule: a judgment recorded SINCE the sidecar
outranks it, and adoption must not reach back over it."""
from booth.asks import ANSWER_SUFFIX, build_answer, normalize_ask
from booth.marks import import_legacy_asks
booth = tmp_path / "b"
_sidecar(booth, "winner", _single())
old = build_answer(normalize_ask(_single(), "winner"), "A — baseline")
(booth / f"winner{ANSWER_SUFFIX}").write_text(json.dumps(old))
declare_pick(booth, "winner", _single())
answer_pick(booth, "winner", "B — async") # the operator changed his mind
import_legacy_asks(booth)
assert marks_for(booth)[0].answer["choice"] == "B — async"
def test_the_doc_view_carries_the_marks(client):
"""INV-3's third surface — flagged 4/4 by the panel as named in the rule but
covered by no test, so shipping it unmarked would have passed."""
from booth.marks import write_note
c, data = client
b = data / "b"
b.mkdir()
(b / "notes.md").write_text("# report\n\nprose here\n")
write_note(b, "notes.md", "this section is wrong")
html = c.get("/b/b/view?f=notes.md").text
assert "prose here" in html
assert "this section is wrong" in html
def test_a_corrupt_marks_file_gives_the_browser_a_409_not_a_500(client):
"""The request was fine and the service is fine — the state on disk is not,
and the refusal is deliberate. A 500 would read as "the Booth is broken" and
send the operator looking for something to restart."""
c, data = client
b = data / "b"
b.mkdir()
from booth.marks import write_note
write_note(b, "a.png", "keep me")
(b / MARKS_FILE).write_text("{truncated")
r = c.post("/b/b/flag", data={"target": "a.png", "on": "1"}, follow_redirects=False)
assert r.status_code == 409
body = r.json()
assert "cannot be read" in body["error"] and body["fix"]
# the page still renders, so the operator can see the booth at all
assert c.get("/b/b/").status_code == 200
assert c.get("/b/b/marks.json").status_code == 200
# ---- findings from the cross-frontier BUG-HUNT panel, 2026-09-22 -------------
#
# Heid panel (thread 01M33XEC1H0298C0D968FWBN7A). Four arms, artifact-only,
# diff-scoped. The headline was 4/4 convergent and none of it had a guard: the
# panel's own mutation tables showed the lock lifecycle SURVIVED every existing
# test, because `test_a_no_op_write_does_not_touch_the_booth` asserts only that
# `.marks.json` is absent and never looks at the lock or at the clock the
# sweeper actually reads.
def test_the_lock_file_is_never_unlinked(tmp_path):
"""The lock must outlive the operation that created it.
`flock` binds to an INODE, not to a path. Unlinking `.marks.lock` while a
second writer is blocked on it leaves that writer holding an exclusive lock
on a deleted inode — and the next writer along creates a FRESH lock file and
takes it immediately. Two processes then run the read-modify-write
concurrently and the later `os.replace` drops the earlier one's mark, with
no error anywhere. Both of them obeyed the protocol.
The cleanup existed to keep a no-op from leaving a lock file as its only
trace. That is a tidiness goal, and it bought a lost-update race.
"""
from booth.marks import MARKS_LOCK, set_flag
booth = tmp_path / "b"
booth.mkdir()
assert set_flag(booth, "ghost.png", False) is None # a no-op
assert (booth / MARKS_LOCK).exists(), "the no-op path unlinked the lock file"
def test_a_no_op_does_not_reset_the_ttl_clock(tmp_path):
"""The property the no-op guard actually exists for, asserted against the
clock the sweeper reads instead of against one file's absence.
Creating or removing a directory entry bumps the DIRECTORY's mtime, and
`_newest_mtime` seeds from exactly that. So `touch` + `unlink` of the lock
reset the booth's age to zero while leaving no trace behind — the comment on
the create-only guard reasons about the lock FILE's mtime and misses that
the directory moved underneath it. Repeated, it kept a dead booth alive
forever, which is the precise outcome the guard was written to prevent.
"""
import os
from booth.app import booth_age_seconds
from booth.marks import delete_mark, set_flag
booth = tmp_path / "b"
booth.mkdir()
old = 1_000_000_000
os.utime(booth, (old, old))
set_flag(booth, "ghost.png", False) # no-op: never flagged
delete_mark(booth, "nothing") # no-op: no such mark
age = booth_age_seconds(booth, now=old + 90_000)
assert age > 86_400, f"a no-op reset the TTL clock (age fell to {age:.0f}s)"
def test_a_real_mark_still_resets_the_ttl_clock(tmp_path):
"""The other half of the same rule, so the fix cannot overshoot into
'marking is never activity'. Marking IS activity and must reset the clock;
only a write that changes nothing must not."""
import os
from booth.app import booth_age_seconds
from booth.marks import set_flag
booth = tmp_path / "b"
booth.mkdir()
old = 1_000_000_000
os.utime(booth, (old, old))
set_flag(booth, "a.png", True) # a real mark
assert booth_age_seconds(booth, now=old + 90_000) < 86_400
def test_a_non_string_note_text_does_not_crash_the_read(tmp_path):
"""`_clean_text` did `(text or "").replace(...)`, so a stored `text` that is
valid JSON but not a string raised AttributeError out of the READ path.
That is not a marks bug, it is an INDEX bug: `list_booths` reads every
booth's marks on every page load, so one poisoned file took down `/` and
`/healthz` for all 25 booths. The module's stated posture is that a mark it
cannot read renders as broken, never as a 500.
"""
booth = tmp_path / "b"
booth.mkdir()
(booth / MARKS_FILE).write_text(json.dumps({
"version": 1,
"marks": [{"id": "n1", "shape": "note", "text": 7,
"created": "2026-09-21T00:00:00+00:00"}],
}))
marks = marks_for(booth)
assert len(marks) == 1
assert marks[0].error, "a poisoned note read clean instead of reading broken"
def test_a_non_string_created_does_not_crash_the_sort(tmp_path):
"""`marks_for` sorts on `(created, id)`. A stored `created` of the wrong type
made that comparison raise TypeError — same blast radius as the note above,
reached through the sort rather than through hydration."""
booth = tmp_path / "b"
booth.mkdir()
(booth / MARKS_FILE).write_text(json.dumps({
"version": 1,
"marks": [
{"id": "a", "shape": "note", "text": "fine",
"created": "2026-09-21T00:00:00+00:00"},
{"id": "b", "shape": "note", "text": "also fine", "created": 17},
],
}))
marks = marks_for(booth)
assert len(marks) == 2
# An unreadable mark loses its `created` and so sorts FIRST — the stated
# rule is `("", id)` against `(created, id)`. A mark nobody can read is the
# one that wants looking at, and the alternative is it landing at an
# arbitrary position in the middle of the panel.
assert [m.id for m in marks] == ["b", "a"]
assert marks[0].error and not marks[1].error
def test_legacy_import_order_survives_same_second_mtimes(tmp_path):
"""ROADMAP states the legacy import's order is `(mtime, name)`. It was
stamping `created` at whole-second resolution, so two sidecars written in
the same second lost the fractional part that distinguished them and
`marks_for`'s `(created, id)` tie-break silently re-sorted them into
alphabetical order — reversing the pair the importer had just ordered.
Deterministic order is a v1 invariant precisely because the operator refers
to things positionally. An order that is stated and not kept is worse than
one that was never claimed.
"""
import os
from booth.marks import import_legacy_asks
booth = tmp_path / "b"
booth.mkdir()
for stem in ("zeta", "alpha"):
(booth / f"{stem}{ASK_SUFFIX}").write_text(json.dumps(_single()))
# Same whole second, different fractions: `zeta` is OLDER and must come first.
os.utime(booth / f"zeta{ASK_SUFFIX}", (1_700_000_000.10, 1_700_000_000.10))
os.utime(booth / f"alpha{ASK_SUFFIX}", (1_700_000_000.90, 1_700_000_000.90))
imported = [m.id for m in import_legacy_asks(booth)]
assert imported == ["zeta", "alpha"], "the importer's own order is wrong"
assert [m.id for m in marks_for(booth)] == imported, (
"the read path re-sorted what the importer ordered"
)
def test_the_index_survives_a_poisoned_marks_file(client):
"""The blast radius, asserted where it actually hurts.
`list_booths` reads every booth's marks on every index load and `/healthz`
does the same. One hand-edited or foreign-written `.marks.json` therefore
took down the front page for all 25 booths — the single-booth failure the
lenient reader exists to contain, escaping the booth it belongs to.
"""
c, data = client
good = data / "good"
good.mkdir()
_png(good / "a.png")
bad = data / "bad"
bad.mkdir()
(bad / MARKS_FILE).write_text(json.dumps({
"version": 1,
"marks": [{"id": "n1", "shape": "note", "text": {"oops": True}, "created": 3}],
}))
assert c.get("/").status_code == 200
assert c.get("/healthz").status_code == 200
assert c.get("/b/bad/").status_code == 200
def test_answer_treats_a_non_string_notes_field_as_no_notes(client):
"""`booth_note` guards `text` with `isinstance(..., str)`; `booth_answer`
passed `notes` straight to `_clean_notes`, which calls `.replace` on it. A
multipart FILE part named `notes` is a str to nobody, so the route 500'd on
hostile-but-legal input where its sibling handled the same class of value.
Both routes now read the field the same way: a value that is not text is no
value. The CHOICE is the judgment and it still lands — throwing the whole
answer away over a junk optional field would be the wrong trade."""
c, data = client
b = data / "b"
b.mkdir()
declare_pick(b, "winner", _single())
r = c.post(
"/b/b/answer",
data={"ask": "winner", "choice": "A — baseline"},
files={"notes": ("n.txt", b"surprise", "text/plain")},
follow_redirects=False,
)
assert r.status_code == 303
mark = next(m for m in marks_for(b) if m.id == "winner")
assert mark.answer["choice"] == "A — baseline"
assert not mark.answer.get("notes")
def test_an_inline_doc_tile_offers_a_note_control(client):
"""Three item branches, two of them call `marknotes`. The doc branch got the
flag button and not the note field, so the operator could point at a report
and not write down why — on the one item kind whose whole purpose is prose.
This is the exact failure the `blurtoggle` macro comment names ("patched two
of three"), recurring on the macro that was written to prevent it.
"""
c, data = client
b = data / "b"
b.mkdir()
(b / "report.md").write_text("# report\n\nprose here\n")
html = c.get("/b/b/").text
assert 'value="report.md"' in html, "the doc tile has no mark controls at all"
# `marknotes`' add-field, which only that macro emits. The booth-level panel
# has its own note form, so the presence of /note on the page proves nothing.
assert 'placeholder="a note on this item"' in html, (
"an inline doc tile has no way to add a note"
)
def test_the_marks_panel_survives_a_booth_that_also_has_a_link_board(client):
"""The board booth renders as a board instead of a gallery, which is right —
but the suppression was unconditional, so a pick declared on a booth that
happens to carry a `links.md` had no form to answer it and no way to say so."""
c, data = client
b = data / "b"
b.mkdir()
(b / "links.md").write_text("- [a thing](http://example.invalid) <sub>· who · when</sub>\n")
declare_pick(b, "winner", _single())
html = c.get("/b/b/").text
assert "Which render wins?" in html, "a pick on a board booth was unanswerable"
def test_the_zoom_view_does_not_navigate_away_from_a_note_being_typed(client):
"""The viewer's arrow keys move between images and Escape goes back. The
note textarea landed in the same page, and the handler is on `document`, so
an arrow key meant for the caret threw away the draft instead of moving it.
Asserted structurally: the handler must bail on events from an editable
target. There is no browser in this suite, and a guard nobody can test is
exactly how this shipped."""
c, data = client
b = data / "b"
b.mkdir()
_png(b / "a.png")
js = c.get("/b/b/view?f=a.png").text
assert "isEditable" in js, "the viewer's key handler has no editing guard"
@pytest.mark.parametrize("route", ["booth_answer", "booth_note", "booth_flag",
"booth_unmark", "booth_import_asks"])
def test_mark_writes_do_not_block_the_event_loop(route):
"""Every mark write takes a blocking `flock` and does synchronous disk I/O.
In an `async def` handler that runs ON the event loop, so a lock held by
another process — the CLI mid-`marks-import`, a second browser tab — freezes
every other request, including the index and `/healthz`.
Structural, like `test_stdlib_only`, and for the same reason: the failure is
a property of where the call runs, which no single-process response
assertion can see. The rule is that an async mark-write handler hands the
locked section to a worker thread and never calls the writer inline.
"""
src = pathlib.Path(__file__).parent.parent / "booth" / "app.py"
fn = next(
n for n in ast.walk(ast.parse(src.read_text()))
if isinstance(n, ast.AsyncFunctionDef) and n.name == route
)
writers = {"answer_pick", "write_note", "set_flag", "delete_mark",
"import_legacy_asks"}
for node in ast.walk(fn):
if not isinstance(node, ast.Call):
continue
name = getattr(node.func, "id", None) or getattr(node.func, "attr", None)
if name in writers:
pytest.fail(f"{route} calls {name}() on the event loop; "
"dispatch it through run_in_threadpool")
def test_an_unreadable_mark_is_visible_on_the_page(client):
"""Surviving the poisoned file is half of it. A note whose stored `text` is
unreadable hydrates with empty text, and the panel rendered that as an empty
`<pre>` with a withdraw button beside it — which looks exactly like a note
the operator wrote and then cleared.
`_hydrate`'s own docstring forbids this for picks ("a broken question the
session believes it posted has to be visible — silently hiding it is the one
outcome nobody can debug"). It is the same argument for every shape."""
c, data = client
b = data / "b"
b.mkdir()
(b / MARKS_FILE).write_text(json.dumps({
"version": 1,
"marks": [{"id": "n1", "shape": "note", "text": {"oops": True},
"created": "2026-09-21T00:00:00+00:00"}],
}))
html = c.get("/b/b/").text
assert "⚠ broken" in html, "an unreadable mark rendered as an empty note"
assert "n1" in html
def test_a_marks_file_no_one_can_parse_does_not_take_down_the_index(tmp_path):
"""The v0.2.2 round adopted the RecursionError finding and closed only half
of it. `_hydrate_safe` guards hydration; `json.loads` runs BEFORE that, in
`_read_raw`, whose `except (OSError, ValueError, UnicodeDecodeError)` does
not cover RecursionError or MemoryError.
So a 400 KB file of nothing but brackets, in any one booth, still returned
500 for `/` and `/healthz` across every booth on the service. Found by the
U5 code-review panel against the sibling module and confirmed by running it.
The read is bounded now and both classes are caught.
"""
booth = tmp_path / "b"
booth.mkdir()
(booth / MARKS_FILE).write_text("[" * 200_000 + "]" * 200_000)
assert marks_for(booth) == []
def test_a_marks_file_too_large_to_be_marks_is_refused_before_it_is_read(tmp_path):
"""Bounded by `stat`, not survived. A booth holds one marks document, and
the index reads every booth's on every page load."""
from booth.marks import MARKS_MAX_BYTES
booth = tmp_path / "b"
booth.mkdir()
(booth / MARKS_FILE).write_text(" " * (MARKS_MAX_BYTES + 10))
assert marks_for(booth) == []
def test_a_write_over_an_unparseable_marks_file_still_refuses(tmp_path):
"""The strict half of the asymmetry has to see the same failures the lenient
half does, or a file that reads as "no marks" gets replaced by a write that
believed it. Same two exception classes, same bound."""
from booth.marks import MarksCorrupt, set_flag
booth = tmp_path / "b"
booth.mkdir()
(booth / MARKS_FILE).write_text("[" * 200_000 + "]" * 200_000)
with pytest.raises(MarksCorrupt):
set_flag(booth, "a.png", True)
# ---- findings from the U5 diff-scoped BUG-HUNT panel, 2026-09-22 ------------
def test_the_marks_reader_never_blocks_on_a_file_that_is_not_a_file(tmp_path):
"""Same hole the size cap opened in the manifest, in the sibling it was
copied from. `st_size` is 0 for a FIFO, so it passes the cap, and then
`read_text` blocks with no EOF. `list_booths` reads every booth's marks on
every `GET /` and `/healthz`."""
import os
import signal
booth = tmp_path / "b"
booth.mkdir()
os.mkfifo(booth / MARKS_FILE)
def _timeout(signum, frame):
raise AssertionError("marks_for blocked on a FIFO and never returned")
old = signal.signal(signal.SIGALRM, _timeout)
signal.alarm(5)
try:
assert marks_for(booth) == []
finally:
signal.alarm(0)
signal.signal(signal.SIGALRM, old)
def test_new_marks_and_imported_marks_share_one_stamp_format(tmp_path):
"""The v0.2.2 fix for the legacy-import ordering opened a NEW ordering bug,
which is the shape worth remembering. `import_legacy_asks` moved to
microsecond precision while `now_stamp` stayed at whole seconds, and `-` is
0x2D against `.` at 0x2E — so `...T10:00:00-07:00` sorts BEFORE
`...T10:00:00.500000-07:00`, putting a LATER mark ahead of an EARLIER
import inside the same second.
Deterministic order is a v1 invariant precisely because the operator refers
to things positionally. One format, or the rule cannot be stated.
"""
from booth.marks import now_stamp
stamp = now_stamp()
assert "." in stamp.split("T")[1], f"now_stamp is not sub-second: {stamp}"
assert len(stamp.split(".")[1].split("+")[0].split("-")[0]) == 6
def test_the_importer_cannot_raise_out_of_a_poisoned_entry(tmp_path):
"""`marks_for` routes every entry through `_hydrate_safe`; the importer's
return still went through the bare `_hydrate`, so the one path that reads
entries it did not write was the one without the guard."""
booth = tmp_path / "b"
booth.mkdir()
(booth / MARKS_FILE).write_text(json.dumps({
"version": 1,
"marks": [{"id": "n1", "shape": "note", "text": {"bad": True},
"created": "2026-09-21T00:00:00+00:00"}],
}))
(booth / f"q1{ASK_SUFFIX}").write_text(json.dumps(_single()))
from booth.marks import import_legacy_asks
out = import_legacy_asks(booth) # must not raise
assert isinstance(out, list)
def test_a_document_that_would_not_read_back_is_refused_at_the_write(tmp_path):
"""The read bound is on the STORED bytes and the write adds `indent=2`, so a
document that fits in memory can land over the limit on disk and then read
back as no marks at all — every mark in the booth gone, silently. Refuse
loudly instead: a write that fails is recoverable.
Asserted against `_write_raw` directly, because no single mark can get
there: `_clean_text` caps a note at TEXT_MAX and a flag is a fixed shape.
The reachable path is accumulation — `_note_id` puts no ceiling on how many
notes one booth may carry — which is thousands of writes, not one. Testing
it through `write_note` would need a fixture nobody could justify, and
would be testing the cap rather than the guard.
"""
from booth.marks import MARKS_MAX_BYTES, MarksCorrupt, _write_raw
booth = tmp_path / "b"
booth.mkdir()
bulk = [{"id": f"note-{i}", "shape": "note", "text": "x" * 500,
"created": "2026-09-21T00:00:00.000000+00:00"}
for i in range(MARKS_MAX_BYTES // 400)]
with pytest.raises(MarksCorrupt):
_write_raw(booth, bulk)
assert not (booth / MARKS_FILE).exists(), "a refused write still landed"
def test_a_clock_restore_that_fails_does_not_take_the_route_down(tmp_path):
"""The concrete half of the mtime-restore finding.
`_Locked.__enter__` puts the booth directory's clock back after creating its
lock, and `os.utime` can fail — a read-only directory, a booth whose owner
we are not. It used to escape into the route and answer 500 for what is
otherwise a perfectly good request. Not putting the clock back is a cost
this module can absorb; not answering is not.
The RACE half of that finding is documented in the code and deliberately not
closed: the alternative fix would silently retire the documented behaviour
that releasing a kept board resets its clock
(`test_releasing_a_board_RESETS_its_ttl_clock` pins that on purpose), which
is a TTL doctrine change rather than a bug fix.
"""
import os
from booth.marks import MARKS_LOCK, set_flag
booth = tmp_path / "b"
booth.mkdir()
real_utime = os.utime
def boom(path, *a, **kw):
if str(path) == str(booth):
raise PermissionError("read-only directory")
return real_utime(path, *a, **kw)
os.utime = boom
try:
assert set_flag(booth, "a.png", True) is not None
finally:
os.utime = real_utime
assert (booth / MARKS_LOCK).exists()
assert [m.target for m in marks_for(booth)] == ["a.png"]