Files
booth/persistent-memory.md
T
Vuong Hoang a48ef83ef5 feat(manifest): U5 — booths that say who posted them and why
The index card showed a name, an item count and a countdown, and nothing
the poster chose. An agent with something to show therefore had no way to
make the booth say "look at this" and posted a URL to the link board
instead — which is why 145 of that board's 210 rows (69%) ended up
pointing at booths that had already been swept. The board was absorbing a
job it was never shaped for. This is the shape.

Each booth carries `.booth.json` — {handle, title, why, created} — written
by the CLI from $ALTHING_HANDLE, and the provenance line renders on both
index lanes and on the booth page header.

WHAT IS WHERE

- booth/manifest.py, stdlib-only and importing nothing from booth.* either:
  scripts/booth imports it under the system python3 with no venv, and a
  cross-import between two stdlib-only modules is a second way for that
  invariant to break. It joins the shared test_stdlib_only list and keeps
  a stricter copy of its own.
- The read is lenient and cannot raise. list_booths touches every booth on
  every index load, so a manifest that cannot be parsed costs that booth's
  provenance and nothing else. That is the v0.2.2 lesson applied before the
  same mistake rather than after it.
- Absent and damaged render differently — `unannounced` and `unreadable`.
  Folding "cannot be read" into "never said" would hide the one case
  somebody has to go and fix.
- Re-announcing preserves `created`. A second `booth add` sharpening the
  why is not a second appearance of the booth.
- The write is atomic (invariant 5); the temp file is itself a dotfile, so
  no listing can see it mid-write.

THREE OPERATOR CALLS, 2026-09-22

Flags on the existing new/add verbs rather than a separate `announce` verb
(a second step is the step that gets forgotten, which is the rot's own
mechanism). Unannounced booths get a quiet marker rather than nothing — the
convention is only adoptable if the gap is visible. U5 adds provenance only
and does NOT add a second index ordering keyed on announcement time; that
is a different surface needing its own stated rule, parked for v1.1.

NO EXEMPTION LIST

A pickup booth and the standing link board are created by the service, so
they announce themselves with handle `booth`, which is true rather than
manufactured. One rule — a booth with no manifest is unannounced — instead
of a growing set of special cases.

ALSO

tests/test_booth.py's keep/release assertion was slicing the page on the
bare word `boothhead`, which has lived in the stylesheet far longer than
the assertion has; it was reading CSS and passing on luck, and went red the
first time a new rule landed above the old one. Same assertion, aimed at
the markup. A U5 test had the mirror-image bug: pytest derives tmp_path
from the test name and the index renders data_dir, so a test named
`test_an_unannounced_booth_says_so` put the needle in the haystack itself
and passed against a template that did not yet exist.

310 tests (304 before this unit's CLI half). Live service restarted, 26/26
booth pages verified 200, end-to-end smoke through the real CLI.

NOT TAGGED. The cold contract-review panel is still in flight and the
code-review and bug-hunt gates have not run. Tagging with a gate
outstanding is what made v0.2.0 premature.
2026-09-22 00:48:41 -07:00

21 KiB
Raw Blame History

Persistent memory — booth

Last updated: 2026-09-22

Always check for /tmp/booth-dev-handoff.md — if it exists and its Written: stamp is under 8 hours old, read it (it carries the in-flight handoff from the previous session), then delete it. Older than 8 hours: stale — delete it unread.

Repo purpose

The Booth is the fleet's operator-review surface: agents post work by making a folder under ~/booth-data, the operator looks at it and judges it in the browser, and the judgment gets back to the agent that posted it. It was built as a file-shuttle and is being converged, unit by unit, onto the review loop it turned out to actually be.

Current state / in-flight

As of 2026-09-22:

  • v1 is gated on seven units in ROADMAP.md, dependency-ordered U1 → U2 → {U3, U4, U5} → U7, with U6 independent.
  • U1 and U2 are landed and released. Current version 0.2.2, deployed to the live service, 275 tests green, tree clean, 25/25 booth pages verified 200 after the deploy. U1 ce598b3; U2 c7f9437 released as v0.2.0, then 5e41108 as v0.2.1 (four contract-panel findings), then v0.2.2 carrying the bug-hunt panel's nine (below).
  • U5 is IMPLEMENTED and unreleased as of 2026-09-22. booth/manifest.py (stdlib-only, INV-1), .booth.json per booth, the provenance line on both index lanes and the booth page header, --why / --title on booth new and booth add, and the link board + pickup booths announcing themselves as the service's own. 310 tests, live service restarted, 26/26 booth pages verified 200 and all 26 rendering unannounced. Deliberately NOT tagged yet: the cold /heid-contract-review panel is still in flight and the code-review and bug-hunt gates have not run. That ordering is the 2026-09-21 lesson applied — a release whose gate is outstanding is premature even when the tier is right. Contract: docs/contracts/u5_booth_manifest.contract.md (carries its own seam-review section).
  • U5's original framing (operator, 2026-09-21): self-announcing booths. .booth.json carrying {handle, title, why, created}, written by the CLI from $ALTHING_HANDLE; the index card gains provenance and a one-line purpose, and the index becomes the "what landed" feed the link board was being used as. It closes job 5 of the five jobs — the one nobody named, and the reason 145 dead link rows existed. Nothing started: no contract, no blast-radius pass.
  • Two things about U5 are already settled and should not be re-derived. (1) .booth.json is a DOTFILE, so booth_items' existing startswith(".") skip already keeps it out of tiles, counts and zips — the same reason .marks.json needed no new exclusion rule. (2) The deterministic-order invariant applies to whatever U5 adds to the index; the index is ordered newest-first by mtime today and that rule must stay stated. Also worth knowing before scoping: enforcing the link rule without giving job 5 a home first just makes it homeless — that is the lesson from the 69% rot, and U5 is the home.
  • No heid dispatch is outstanding. The /heid-bug-hunt on U2's diff landed 2026-09-22 and shipped as v0.2.2; see the dated entry below.
  • Live service active on :8090, 25 booths, verified 25 × 3 page types after the last deploy. The booth set churns: sindra20-engines and sindra-finalists were swept during the session, cr123a-to-d-sleeve and sindra appeared.

Recent decisions

  • [2026-09-22] The U2 bug-hunt panel landed and it was not ceremony — v0.2.2. Nine adopted findings across four arms; eight were real against live code and one was already fixed. The headline was 4/4 convergent from four different angles: _Locked.__exit__ unlinked .marks.lock on the no-op path, and flock binds to an INODE — so a writer blocked on the old inode proceeds while the next writer creates a fresh lock file and takes it at once. Two processes then run the read-modify-write concurrently and the later os.replace drops a mark, with both of them obeying the protocol. The cleanup existed to protect the booth's TTL and it was failing at that too: creating and removing a directory entry bumps the DIRECTORY's mtime, which is what _newest_mtime actually seeds from, so a no-op reset the clock it was written to leave alone. Same code region, two defects, one fix — never unlink the lock, exempt .<name>.lock dotfiles from _newest_mtime, and put the directory's mtime back after creating one. Full triage in persistent-memory.d/2026-09-22-bug-hunt-panel.md.

  • [2026-09-22] The lenient reader's blast radius was the whole service, not one booth. _clean_text did (text or "").replace(...) and marks_for sorts on (created, id), so a stored text that was a dict or a created that was a number raised out of the READ path — and list_booths reads every booth's marks on every index load. One hand-edited file 500'd / and /healthz for all 25 booths. Fixed in two layers, matching the house posture: a named type check (_entry_type_error) plus a _hydrate_safe backstop that cannot raise, and the panel now RENDERS an unreadable mark as ⚠ broken instead of as an empty note. The general shape: a lenient reader is only lenient if the leniency is bounded by where it runs. marks_for was written for one booth's page and is called in a loop over every booth.

  • [2026-09-22] booth marks / booth answer got real exit codes, because a read that CRASHED was indistinguishable from a read that said no. marks printed a traceback and exited 0 (a caller's jq saw success and got nothing); answer --wait read a damaged file as "not yet" and spun for the full hour before blaming the operator. Now 0 ok · 1 unanswered/timed-out · 2 no such pick · 3 unreadable, and read_error() was added to marks.py so the CLI can ask the question the browser must not: the page stays lenient, the machine consumer gets the truth. Also --wait now prints ONCE — it was emitting a whole JSON document per poll, so a captured --wait held several concatenated values and parsed as none of them.

  • [2026-09-22] scripts/booth had zero tests and now has five (tests/test_cli.py). The panel's guard-strength tables returned UNVERIFIED for every CLI claim because nothing in the suite executed the script — two of the round's findings lived in exactly that gap. The new tests run the real script under the system python3, which makes them a live check on INV-1 (stdlib-only) as a side effect: a third-party import in marks.py now fails in the suite the same way it would fail on a fleet host.

  • [2026-09-21] v0.2.0 cut and announced; v0.2.1 fixed what the announcement was already wrong about. Operator approved the minor (a v1 unit closed plus a CLI surface change for 17 consuming handles clears the release-note bar). The note went to 15 handles — the 17 link-board posters minus nh3-dev, a host label, and heid, an oracle that does not script these verbs. Then the cross-frontier contract panel landed and found three defects in the code I had just released, so v0.2.1 shipped within the hour. Sequence worth remembering: the release was correct by the tier bar and still premature by the discipline — the panel had been dispatched BEFORE implementation and its reply arrived AFTER the tag. If a gate is in flight, the tag can wait for it.

  • [2026-09-21] A write over a damaged .marks.json was wiping every mark in the booth. Shipped in v0.2.0, found by the panel (Kimi, converged with Hulda), fixed in v0.2.1. marks_for is deliberately lenient — unparseable reads as [] so a review page still loads — and the write path inherited that leniency through the same reader, so one flag click appended to an empty list and atomically replaced the file. The fix is an asymmetry, which is the reusable part: reads stay lenient, writes go strict (MarksCorrupt), damaged bytes stay on disk, routes answer 409 not 500. A page that renders without an annotation is recoverable; a file that overwrote the operator's judgment is not. Kimi also named the class correctly — "an author steeped in the design conversation would likely read past" it — and that was accurate.

  • [2026-09-21] The two review gates are complementary, measured on one unit. The caller-side seam review (nine findings, against the real sibling module surfaces) and the cold /heid-contract-review panel (four arms, artifact-only) had zero overlap in both directions on U2. The seam review found a scope miss the panel structurally could not see: the contract omitted inline.py, whose place() indexes by subscript, which a frozen dataclass refuses. The panel found three code defects and a missing test the seam review had no lens for. Matches heid's kvasir zero-overlap result on the conformance-versus-hunt axis. Run both; neither substitutes.

  • [2026-09-21] Every one of the panel's code-changing findings came from the AMBIGUITY pass, none from a paraphrase divergence — and two arms independently proposed cutting the paraphrase to a drift-check for narrative-heavy contracts, because this contract's own frontmatter carries a plain-language narrative and the paraphrase was partly reading my framing back to me. That is a finding about the /heid-contract-review skill, not about this repo, and it was reported back to heid. Recorded here only so a future session does not rediscover it.

  • [2026-09-21] Deterministic order is a cross-cutting v1 invariant — operator directive, mid-implementation. Every ordered collection the Booth renders must have a stated rule producing the same sequence on every render of the same state; the rule can be anything defensible (byte order, time, an explicit number, an arbitrary-but-recorded sequence), but no rule at all is forbidden. It binds harder here than elsewhere because the Booth's job is comparison — the operator judges tile 47 against tile 47 and refers to artifacts positionally, so an order that moves between renders misfiles a flag or a note rather than crashing. Recorded as ROADMAP.md § "Cross-cutting invariant" (with the per-collection table) and CLAUDE.md invariant 6, and tested. Still undecided and must be settled before those units ship: U7's section ordering and compare pairing, and U6's bench listing.

  • [2026-09-21] U2 (marks) landed. One primitive replacing three mechanisms. pick / note / flag in one .marks.json per booth, one read path (marks_for), one openness predicate (open_marks), rendered beside the artifact on the tile, at full size in the zoom, and in the panel. flag and note had no write path at all before this — the selection loop (golden-candidates, sindra-finalists, the pancake-* ladders) was running through chat. 242 tests. Details worth carrying: asks.py kept normalize_ask and gained build_answer (the 2026-09-09 partial-answer semantics preserved by moving, not rewriting) and LOST its five sidecar-storage functions; GET /b/<n>/marks.json was added because remote sessions polled <stem>.answer.json over HTTP and the sidecar's removal would have taken that capability with it; /b/<n>/asks 308s to /marks.

  • [2026-09-21] A partially-answered pick now counts as OPEN — declared, not smuggled. The old index badge tested answer is None, so a half-answered four-question ask read as closed on the index while the panel beside it rendered ◐ partial: the two disagreed about the same booth. Open is the reading that makes U4 correct — a lifetime rule that unpinned a booth on the first radio click would sweep a review in flight.

  • [2026-09-21] The U2 seam review earned its place, and the record should say how. Nine findings against the real booth.asks / booth.items / booth.inline surfaces, two of which changed scope or behaviour: inline.py was missing from touches entirely (its place() indexes asks by subscript, which a frozen dataclass refuses — nothing else in the service does that), and the partial-answer inconsistency above. The cold /heid-contract-review pass is artifact-only by design and structurally cannot see a sibling module, so neither it nor a same-model self-review would have found either. Two more surfaced later and are worth the same note: a SECOND subscript in inline.place the seam review undercounted, and a regression in my own legacy importer that a retargeted test caught — a malformed sidecar that renders ⚠ broken today would have silently vanished on migration.

  • [2026-09-21] Marks are stored as one .marks.json per booth, atomic temp-file + os.replace, fcntl lock on the read-modify-write — operator decision, this session. Two alternatives were weighed and lost: a sidecar per item (<rel>.marks.json) and extending the existing <stem>.ask.json shape. Rationale, and the reason it is not links.md-shaped: (a) U4 makes "does this booth owe an answer?" a hot question — the sweep asks it per booth per tick and the index asks it per card per page load, so per-item sidecars turn it into a full walk of all 25 booths, one of which holds 270 files; (b) links.md is an O_APPEND content-hash log because 17 agent handles write it concurrently, whereas marks have exactly one writer (the operator, in one browser) and many readers — a different problem that must not inherit the append-log design; (c) .blurred / .pins / .forever already establish the per-booth dotfile as the house shape for operator state, and booth_items()'s dotfile skip means it costs nothing in counts, galleries or zips. Accepted cost: a corrupt .marks.json loses that booth's marks rather than one item's. Implementation deferred to U2 — tracked at ROADMAP.md U2 and by this entry.

  • [2026-09-21] U7's section premise is half wrong, and it is the half that matters — found by re-measuring ~/booth-data rather than trusting the IA doc. The IA says sections come from subfolders that already exist on disk; true, but every booth that actually needs navigation is flat: pancake-v3-full (270 items, 0 subfolders), pancake-v4-full (270, 0), sindra20-engines (98 items + 99 caption sidecars, 0), sindra-finalists (86 + 87, 0). Subfolders exist on exactly two booths — pewpew-ui-brief (7, nested to _ds/powerpellet-design-system-<uuid>/preview) and dfa-concepts (1) — and both are reports, the job where grid navigation matters least. So sections stay worth shipping and Item.section stays right, but they are not "most of the navigation fix": the rail, the filters and grid keyboard are all of it. Worth noting for whoever writes U7: sindra20-engines encodes its structure in the filename prefix (b2-s1-<subject>-<seed>), which is where a grouping heuristic would actually pay. The IA doc's claim about what sections buy needs a line struck — not yet edited.

  • [2026-09-21] sindra-finalists is U2's flag motivation caught in the act — 86 items, every one captioned, and the booth's entire name is "the ones the operator picked." That loop currently runs through chat, which is the defect flag closes. Evidence, not argument.

  • [2026-09-21] The information architecture and the v1 gate landed (726822b): docs/design/information-architecture.md names the single defect — one lifetime (24h from last touch) and one shape (a folder), serving five jobs with different lifetimes and different shapes — and ROADMAP.md gates v1 on seven units, each closing a measured defect rather than a wish. Both were written after a measurement pass over the live service, and the measurements are the load-bearing part.

  • [2026-09-21] The .forever diagnosis is a stated, falsifiable prediction. U4 (derived lifetime) predicts the kept-rate falls to the genuinely-durable booths. Re-measured today: 14 of 25 booths kept (56%), against the 54% the IA doc recorded. Re-count a fortnight after U4 lands. If it does not move, the diagnosis was wrong and the boolean was doing something else. Tracked in the IA doc's Booth section and by this entry.

  • [2026-09-21] Extracted from eshpfi into its own repo. The accreted service came over whole, tests included, so tests/test_booth.py (1581 lines) is the regression net the v1 rewrite is checked against.

Tried and abandoned

  • [2026-09-21] Tagging a release while a review gate was still in flight. v0.2.0 was cut and announced to 15 consuming handles; the /heid-contract-review panel — dispatched BEFORE implementation, as the discipline says — replied afterwards with three defects in the code that had just shipped, one of them silent data loss. Nothing about the tier decision was wrong; the timing was. If a gate is outstanding on the work being released, the tag waits for it. The cost was a same-hour v0.2.1 and a correction note to peers who had already verified against the broken version.

  • [2026-09-21] Letting the write path share the read path's leniency. See the MarksCorrupt decision above. The general shape, worth carrying beyond marks: a tolerant reader and a tolerant writer over the same state are not the same decision, and pointing both at one function silently makes them one. Tolerate on read so the surface still renders; refuse on write so nothing is destroyed.

  • [2026-09-21] Letting Jinja hot-reload templates while the repo is the deployment root — the cause of a live outage the same day U2 landed, and the sharpest foot-gun in the repo. booth.service sets WorkingDirectory to this repo, so the running service imports these files with no build step and no staging copy. Python is read once at process start; Jinja's FileSystemLoader re-reads a template on every render. Editing booth.html therefore deployed it instantly against Python from 22:03 that knew nothing about item_marks, and 19 of 25 live booths returned 500 with UndefinedError: 'item_marks' is undefined. Neither the old code nor the new code was broken — the service was running both at once. The lesson that generalises: a skew between a process and the disk under it is invisible to the test suite by construction, so no amount of green tests would have caught it; the operator found it. Fixed at the source rather than with a reminder — the Environment is hand-built with auto_reload=False, so there is now ONE staleness rule (nothing takes effect until you restart) and the running process is always a coherent snapshot of one commit. Asserted by test_templates_do_not_hot_reload_from_disk. Watch the second-order risk the fix introduces: a hand-built Environment does not inherit autoescape from the Jinja2Templates constructor, and booth names, item names and mark text are all agent-authored strings landing in HTML.

  • [2026-09-21] Five separate mechanisms to get one question next to one artifact — .forever, the link board, inline.py's placeholder DSL, wrap_verbatim_html's six regexes, and the floating amber asks chip plus /b/<n>/asks. Every one is a correct local fix to the same global mismatch, which is exactly why they accumulated without anyone making a bad call. The foot-gun is the sixth one: the next "just add a small thing for this case" reads as reasonable and is the pattern. The git log carries the signature — every feature ships, then takes 2–5 patches for cases the single shape did not anticipate. Check the ROADMAP gate before adding a mechanism.

  • [2026-09-21] Regex-injecting chrome into arbitrary author HTML (wrap_verbatim_html + _HEAD_CLOSE_RE, _HTML_OPEN_RE, _DOCTYPE_RE, _BODY_CLOSE_RE, _HTML_CLOSE_RE, _ICON_RE, and the doctype/charset ordering constraints they thread). It works today and is still live — but it is the single most fragile thing in the service and it is load-bearing for the operator's most important workflow. Slated for deletion at U3 in favour of a declared seam (/_booth/embed.js, mounted through a real DOM API), which costs an author one line and removes the whole class. Do not extend the regex set in the meantime; if a verbatim page breaks, that is an argument for U3, not for a seventh pattern.

  • [2026-09-21] A boolean escape hatch as the lifetime mechanism. .forever was added because a 24h TTL genuinely did not fit some booths — and then 56% of live booths ended up on it, which means it is not "ephemeral with an exception", it is two lifetimes wearing one lifetime's clothes, with the operator doing the sorting by hand. Replaced at U4 by lifetime derived from state (an open mark pins; viewing is activity; keep survives as an explicit reasoned pin rather than the only way to say "not yet").

  • [2026-09-21] Letting the link board absorb the announce job. booth link is an O_APPEND write with no identity and no stated rule, so re-announcing a bench appends a row instead of updating one, and a booth URL rots the moment its booth is swept — 145 of 211 rows (69%) pointed at nothing, and 22 were the same target re-posted (talk 5×, peedlar 4×). The rot is structural, not drift. The lesson that cost the most: enforcing the link rule without first giving the announce job a home (.booth.json provenance on the index, U5) just makes it homeless.