Commit Graph
97 Commits
Author SHA1 Message Date
vh 051599a30e docs(contract): r2 — the review flow: the Desk, the lightbox, the review
PROPOSED. Ruled by the operator 2026-09-23 (flow: a_b, compare this_arc,
voice plain, emblem no). Compare is not in this contract; it follows as r3.
Seam-reviewed against the live module surfaces before the cross-frontier
contract panel returned. Four findings are folded in: Mark.created is a
string, the board is BOOTH_LINKS_BOARD, the bench-read error state, and an
unreadable marks file counting as needing the operator.
2026-09-23 08:15:32 -07:00
vh bf55364920 fix(theme): at phone width the JS-off rail fallback is the measured worst case
booth-dev suggested this. At or below 480px, .item's scroll-margin
fallback is 205px, the 16-group rail measured at 390px. With JS on,
--rail-h is exact and nothing changes.

Measured on the same 76 jumps:
- JS off: 0 under the rail, previously 19. At 390px, where a short rail
  gets the full fallback, tiles overshoot by at most 74px, and they stay
  visible.
- JS on: unchanged, 0 under.

660 passed; visual order still matches document order on 32 renders.
2026-09-23 08:12:54 -07:00
vh e8e49ceb14 fix(theme): a group jump lands its tile below the sticky rail, not under it
Heid bug-hunt finding (Gróa, relayed by booth-dev). The rail is sticky
and nothing set a scroll margin, so a fragment jump left the target tile,
and the :target reticle that marks it, hidden under the rail.

The rail wraps, so no CSS value can know its height. A small additive
script publishes the measured height as --rail-h, and a ResizeObserver
keeps it current across widths. .item's scroll-margin-top adds 12px to
that. With JS off, a 120px fallback applies.

Also styles the new empty-filter row (397ea89): the filter name in
heading ink, and a gap before the way back.

Measured, 76 group jumps across 2 booths x 4 widths (rail 48-205px):
- JS on: 0 under the rail, minimum clearance 11px.
- JS off: 0 at desktop widths. 19 at 390px, where a wrapped rail is
  143-205px tall and taller than the fallback.
Positive control: the pre-retheme skin fails 74/76.
Merged onto main 1826d19: 660 passed, mutation_check 20/20.
2026-09-23 08:12:54 -07:00
vh 744fa5263e feat(theme): SVOS retheme — concept-round candidate
Re-skins every Booth surface in the SVOS design system (design-systems
palettes/svos @ ed2f8d8). Visual and interaction layer only: no route,
no copy, no ordering and no information-architecture change.

- _svos_tokens.css: SVOS semantic tokens vendored by copy, with the four
  [data-theme] scopes re-scoped onto prefers-color-scheme and
  prefers-contrast (dark, light, dark-hc, light-hc). Included into
  base.html's <style>; cached at startup like every other template.
- base.html: the accreted Australis sheet is rewritten against semantic
  tokens only. It also fixes four undefined variables (--line, --bg,
  --fg, --muted) that the keep/blur/reveal controls had been reading.
  The three SVOS devices each have exactly one job: reticle = selection
  (grid cursor, :target, picked option), hazard = irreversible (Wipe
  now, armed bulk delete), glow = live power (service dot, live bench).
- The flag list renders as wrapped chips, so a large flag set no longer
  pushes the grid below the fold. The list order is unchanged.
- IBM Plex Sans + JetBrains Mono load via Google Fonts with
  display=swap and system fallbacks (approved by booth-dev).
- view.html, doc.html: inline styles moved onto tokens.
- embed.js: fragment palette as custom properties scoped to .bk-ask;
  `.bk-ask-opt:has(input:checked)` still appears exactly once.
- Favicon (base.html + app.FAVICON_HREF, kept in sync): graphite tile
  with reticle corners.

Verified: 642 passed, the same count as the pre-change baseline.
Visual order matches document order on 32 renders (4 booths x 4 widths
x 2 schemes). A positive control, one tile given `order:-1`, is
detected by the same check.
2026-09-23 08:12:54 -07:00
vh b46ac02be2 docs: the four flow rulings, compare unparked, and the beta premise superseded
All four ruled, all four taking design-dev's recommendation, relayed via Miranda
with booth-dev as sole relay. Verbatim copy committed at docs/rulings/ because
the booth holding it will sweep.

Which is the observation worth keeping: answering a pick removes the hold that
was protecting the record. A booth is held while its question is OPEN, so its
lifetime is shortest exactly when it has just become valuable — before the
answer it is a question, after it is the record of a decision, and only the
first state is protected. Both design booths hit this by different routes, one
withdrawn and one answered. Raised to design-dev as a flow question rather than
patched, since flow is his now.

Compare mode leaves the parking lot: our deferral, his overrule, recorded as his
call so nobody re-parks it by reading the older rule.

And v1.0.0b1's 'no new features' promise no longer describes the arc. The tag
stays as written — rewriting a released tag to flatter the present is how a
version stops being evidence — an alpha drop-back is illegal because 1.0.0a2
sorts below 1.0.0b1, and no further pre-release is cut until the arc lands.
2026-09-23 08:12:05 -07:00
vh f87976b54d memory: correct a review point we got wrong, rather than leave it to be re-asserted
We read design-dev's 'SET order' as 'the order they were set in' and told him it
was already (created, id). He meant the SET's order — by tile number — which
genuinely differs: flag #15 then #07 and today's panel lists #15, #07 while his
tray lists #07, #15.

His rule is also cleaner than the one we proposed. The flag set sorted by its
target's position in sorted(rel) is a total order needing no tie-break at all,
because rels are unique. The memory row now says so explicitly and tells the
next session not to re-raise the point.

Round 1's booth is kept; he releases it once the flow ask is ruled.
2026-09-23 07:21:19 -07:00
vh af57933255 memory: round 2 is up, and the ordering review that preceded it
Four rulings with the operator on booth-flow-concepts. design-dev asked for an
invariant-6 check before building, which is the right order and worth recording
as the pattern.

His ordinals rule is an improvement on invariant 6 rather than compliance with
it: an ordinal counting across all items makes a positional reference stable
under filters, where today 'the third one' silently means something different
the moment a filter is on. Nothing on our side had noticed.

Two corrections returned. Flag 'set order' is already (created, id) — set_flag
upserts and unflag removes the entry, so created IS the set time; what he
actually needs is the tie-break, not a new field. And 'last activity' must reuse
_newest_mtime, whose .lock exclusion was paid for: counting our own lock
sidecars made reading through a write path look like activity.
2026-09-23 07:20:28 -07:00
vh 6ba5a83f81 docs: the operator moved the design ownership boundary, and the fence was ours
Round 1 ruled not-as-shown: 'He didn't go far enough, still looks like the
booth. I want him to consider the flow and the requirements — design touches,
layout, usability all belong to him.'

The handoff paragraph that said we were not asking for layout changes driven by
information architecture is void. design-dev's 'class additions only, no
reordering' was that constraint honoured, so the ruling corrects our brief
rather than his round — worth recording that way round, because the next session
reading only the artifact would read it as a design failure.

Flow, layout, usability and the requirements are his now; the IA is no longer
fenced off. What survives is split in two on purpose: correctness invariants
that are not design opinions, and engineering defaults we chose that he may now
argue with, where a dispute goes to the operator rather than being settled
between agents.
2026-09-23 07:11:53 -07:00
vh dfd806aa9f docs(booth.html): name the .rail cross-file contract at the selector that depends on it
The SVOS retheme makes .rail load-bearing in two files owned by two different
agents: this template's grid-cursor start, and base.html's --rail-h measuring
script that publishes the rail's height for scroll-margin-top (the rail wraps,
so no CSS number can know it).

Neither breaks loudly if it is renamed. Ours starts the cursor one tile too
high; theirs falls back to a fixed guess. design-dev's sheet carries the mirror
of this note above the .rail rule, so the coupling is documented from both ends
rather than from whichever side happened to notice.
2026-09-23 06:57:32 -07:00
vh 06d83dfd2f memory: the staged design-dev ref moves — read it, do not trust a SHA written here
He rebases onto our main and rewrites the ref in place; it has already gone
878ed86 -> a99b7bb. Merging a SHA copied out of the memory file would merge a
pre-rebase branch that predates both his scroll-margin fix and our bug-hunt
batch.

Third instance of one class today: a 'PUSHED' row that was stale when written, a
postbox send-note promoted into durable memory, and now a moving ref recorded by
SHA. The file records what was true when written; anything that moves needs a
command, not a value.
2026-09-23 06:53:29 -07:00
vh 1826d19a1f memory: the bug-hunt panel, the raw-first fragment trap, and five vacuous falsifiers
The mechanic worth keeping: browsers match a URL fragment against element ids
RAW first and percent-decoded only second, so a raw rel on both the anchor and
the id is ambiguous rather than merely unencoded — and encoding one side only
relocates the collision.

The count worth keeping: five falsifiers in one unit were green under the exact
change they forbade, three arms finding the same one independently. A
guard-strength pass is the highest-value part of a panel on a diff that is
already well tested, because the findings sit in the gaps the comments are most
confident about.
2026-09-23 00:06:26 -07:00
vh 397ea89795 fix(u7): six defects from the heid bug-hunt panel, and five vacuous falsifiers
Cross-frontier panel (Gróa/Hulda/Regin/Kimi) on U7's diff, thread
01M368G2Y0JMTJ2T7M3JMTXV5Z. Four of the six fixes are for defects no test in
this repo could have caught, and the panel's guard-strength passes found five of
my own falsifiers green under the exact change they forbade.

THE 4-OF-4 FINDING — the group anchor could land on the WRONG artifact.
The anchor was the raw rel spliced into an href fragment while the tile id was
equally raw. A browser matches a fragment against ids RAW FIRST and only then
percent-decoded, so raw-on-both-sides is not merely unencoded, it is AMBIGUOUS:
with `a b.png` and `a%20b.png` in one booth, the first's href resolves to the
fragment `item-a%20b.png` and the raw pass matches the SECOND file's id. That is
the misfiled-judgment failure invariant 6 exists to prevent, arriving through a
path invariant 6 never looked at. Both sides now use `Item.url`
(`quote(rel, safe="/")`), which is injective here and is the convention
booth_flag has always used. The original test asserted the href occurred as SOME
id on the page — true while pointing at the wrong one.

GRÓA'S STRONGEST SOLO — a zero-hit filter removed the way back.
The rail was gated on the FILTERED list, so a valid filter with no matches
removed the rail, the filter links and the route back to `all`, while the
empty-booth branch announced the booth was empty with rail.total still holding
the real count. No recovery without editing the address bar, and it degraded the
same way with JavaScript off, on the surface the operator actually reviews on.
Gated on all_items now, with an explicit no-match row.

HULDA — one unrepresentable filename took out the INDEX, not just its booth.
A non-UTF-8 filename reaches CPython as a surrogate and quote() raises on it,
outside any per-item handler. booth_items feeds list_booths, so one 0xff byte in
one booth's filename 500s every booth's card. Such a file cannot be linked,
served or zipped, so it is skipped like a dotfile.

HULDA — the `f` shortcut has never worked. The selector named `.flagbtn`, which
nothing in this repo emits, so it fell through to the hidden target input;
clicking a hidden input does not submit its form, and the handler called
preventDefault anyway. Now clicks the flag form's real button, verified end to
end in a real browser.

GRÓA — a group jump was undone by the next keypress. The jump scrolls, the
cursor stayed at -1, and the next arrow focused tile 0 and scrolled back. The
cursor now picks up from the viewport, which also fixes the general
scroll-then-arrow case. Asserted on real scroll geometry in Chromium.

HULDA — the caption sidecar was read whole before being truncated, so a
pathological file was a MemoryError the OSError handler does not catch. Bounded
at the read, and deliberately NOT by st_size: a FIFO reports 0.

ACCEPTED KNOWN RISKS, both now documented rather than implied: no cap on rail
row count (1,000 groups of two would render 1,000 rows; the largest live booth
is 66 items and picking a cap without a booth that needs one is invented work),
and Item.group sits mid-dataclass (one construction site, keyword-only, grepped).
The docstring now names the UPPER median explicitly — two arms flagged that
"the middle group" admits both readings for an even count.

FIVE VACUOUS FALSIFIERS, found by the arms and not by me: the anchor test
survived v[0]->v[-1]; the informativeness guard survived sizes[-1]; the group
count survived len(v)+1; the zero-hit filter test used a fixture that HAD hits;
and the escaping test asserted over the whole page, so it went red on a code
comment. All rewritten, all mutation-proved. The table is up to 20 rows and one
drifted when I changed the line under it — reported by the harness, not silently
skipped, which is the behaviour tests/test_mutation_check.py exists to hold.

660 green; 20/20 proved. Deployed; 21/21 booths 200.

Held for design-dev, not fixed here: Gróa's finding that the sticky rail has no
scroll-margin, so a fragment jump tucks the target under it. It is one line in
base.html, the file he is rewriting from scratch.
2026-09-23 00:04:59 -07:00
vh 6042d10bf3 memory: the SVOS concept round is with the operator, and a latent CSS defect it surfaced
Three rulings open on booth-svos-retheme (ship / voice / emblem). The branch is
an inert ref; merge is gated on the rulings. Verified independently: nothing
checked out, main clean, merge-tree clean, merged tree 649 green.

The fixup hold is now partial — booth.html is released because design-dev does
not touch it, so bug-hunt findings there land immediately.

And a real one he caught on our side: base.html reads four CSS custom properties
and defines none of them, 15 uses without a fallback. An undefined var makes the
whole declaration invalid at computed-value time, so those buttons have had no
border at all and a transparent background — not merely default colours. The U7
rail reads the same names with fallbacks, which is why the rail looked
deliberate and the buttons under it never did. Assigned to his rewrite; fixing
it on main would collide with the one file he is rewriting.
2026-09-22 22:18:56 -07:00
vh 33e7149e24 fix(scripts): the mutation harness must not churn source mtimes
It rewrites a tracked file and restores it byte-for-byte — but the restore
bumped the mtime, and in this repo that is not cosmetic. The repo IS the
deployment root and nothing takes effect until the service restarts, so 'is
:8090 stale?' is answered by comparing the service's start time against source
mtimes. A tool that moves those without changing a byte makes that check lie:
it reported the live service 16 minutes stale while it was serving current code.

Restores atime/mtime with os.utime, with a test whose defeating change is
dropping that line. Found by using the staleness check for real, not by review.

649 green; 12/12 U7 falsifiers still proved.
2026-09-22 22:01:35 -07:00
vh c47b3dba7e memory: pushed v1.0.0b1, and a 'PUSHED' row that was stale when written
main and the annotated v1.0.0b1 tag are on origin; ahead 0, behind 0.

The push carried SIX commits, not the five this session produced: 2f85692 from
the previous session was still unpushed while the memory row above it said
PUSHED. A push is a point in time and this file is not, so the row now says to
run git rev-list rather than to believe it — the same class of error as
promoting a postbox send note into durable memory, twice in one day.
2026-09-22 21:58:46 -07:00
vh 2f6a0ee821 test: keep the mutation harness — scripts/mutation_check.py, with its own controls
Promotes the session-scratchpad harness that proved U7's twelve falsifiers into
a repo tool, on the operator's call. No version bump: test tooling and docs, no
production-code change, per the SemVer SKIP list.

A green test is not evidence. A test that has never seen its own defeating
change may pass under it too, forbidding nothing while reading as though it
forbids something. This repo shipped that three times — twice in one session,
and once an hour after writing the persistent-memory entry about it. Prose in a
memory file is not an instrument.

Tables live in tests/mutations/*.toml, one per unit, committed so a unit's
proofs are an artifact rather than terminal scrollback. Adding a unit means
adding a file, never editing the script. u7_navigation.toml was generated from
the harness that proved those twelve, not retyped, and every anchor was verified
against the source before it landed.

THE TOOL GETS ITS OWN POSITIVE AND NEGATIVE CONTROLS, which is the point. It
shipped two defects in one session, each of which made it report a falsifier
PROVED WITHOUT RUNNING IT, and both were found by accident rather than by
anything checking:

  no green baseline — a test that is ALREADY red reports red for every mutation
  thrown at it, so a broken assertion reads as a certified falsifier

  the bytecode cache — `< 2` -> `< 1` is byte-identical in size, and CPython
  validates a .pyc against the source's (mtime, size) at one-second granularity,
  so a mutation landing in the same second as the revert before it runs against
  cached bytecode; the tell was a verdict flipping between consecutive identical
  runs

tests/test_mutation_check.py now carries a control for each, plus the one
usually skipped: a KNOWN-VACUOUS falsifier the tool must catch. An instrument
that only ever sees unknowns cannot tell "nothing wrong here" from "I am blind",
and twelve `proved` lines from a blind instrument are worth nothing.

Also hardens the tool against itself: it writes to tracked source files, so the
restore is verified rather than assumed, and a .mutation-inflight marker makes a
run killed mid-mutation refuse the next start instead of silently measuring a
mutated tree.

648 tests green; 12/12 U7 falsifiers still proved.
2026-09-22 21:58:12 -07:00
vh 82ac7c44e4 docs: design-dev accepted the SVOS retrofit — the /vor-ui brief is declined, and why
The ROADMAP row requiring a /vor-ui brief predates the IA doc. With that doc,
the landed templates and the seven handoff constraints, a /vor-ui pass would
have cost the operator a serial Q&A to re-derive IA already measured. design-dev
made that argument and it is better than the row it overrides.

Also settles: we merge and restart; he works against a copy, never :8090; the
concept round goes to the operator; webfonts by CDN link with display=swap,
because the CDN-free property turned out to be accreted rather than an
invariant (checked CLAUDE.md, the non-goals and the IA doc).

Corrects a memory defect in the same commit: a postbox send note is a
point-in-time snapshot and one was promoted into persistent memory as a durable
fact about a handle's delivery mode. It was wrong within the hour.
2026-09-22 21:51:35 -07:00
vh 8a18dd13ab memory: snapshot — v1.0.0b1 cut, the version that was two copies, and the design-dev handoff 2026-09-22 21:41:56 -07:00
vh 3126deca00 chore(release): 1.0.0b1 — the v1 target, staged as a beta
All seven v1 capabilities are landed (ROADMAP's v1 target is met), so this is
the first release of the 1.x train. Staged as a beta rather than cut final on
the operator's call: per the canonical policy `-beta.N` means feature-complete,
external testing, no new features, focus is on bugs — which is exactly this
state, with a cross-frontier bug-hunt panel outstanding on U7's diff.

The repo learned this sequencing the hard way once: v0.2.0 was tagged and
announced while a contract panel was in flight, the panel found three defects in
the code just released, and v0.2.1 shipped within the hour. A beta is the
designed answer to that, not a workaround for it.

ALSO FIXES A SECOND COPY OF THE VERSION, found while cutting this one.
`booth.__version__` was the literal `0.1.0` and had been wrong through six
releases. It is now read from pyproject.toml — deliberately NOT from
importlib.metadata, which describes a different artifact: this repo has no build
step and no install step (booth.service runs uvicorn with WorkingDirectory set
to the tree), and the venv was carrying a vestigial booth-0.3.0.dist-info with
no package directory behind it. Installed metadata therefore reported 0.3.0 for
a tree at 1.0.0b1 — confidently wrong and varying by environment, which is worse
than a literal that at least fails the same way everywhere.

booth/__init__.py is also, it turns out, effectively stdlib-only: scripts/booth
imports booth.links / booth.marks / booth.manifest under the system python3 with
no venv, and every one of those executes the package root first. Nothing
asserted it. test_stdlib_only now covers __init__, and the no-venv import path
is verified under python3.11 reporting 1.0.0b1.

642 tests green.
v1.0.0b1
2026-09-22 21:40:09 -07:00
vh bf351a26d1 feat(u7): filename groups — the last v1 unit, and a table that did not reproduce
Completes U7 with its fourth component: a jump-to-group rail derived from
filename prefixes, replacing the subfolder sections ROADMAP named. The scope
departure was ratified by the operator 2026-09-22; this commit deletes
test_no_group_rail_is_shipped_yet, the guard that held it back, in the same
change that builds what it guarded against.

All seven v1 capabilities are now landed. The 1.0 cut is a decision, not a
dependency, and it is the operator's — no version bump here, because a commit
is not a release.

THE RULE CHANGED AT IMPLEMENTATION, ON MEASURED GROUNDS. The contract specified
`strip ONE trailing run of digits`; run against the live set that yields 24
groups for sindra-bakeoff's 40 images and 27 for sindra's 30 — a rail with a row
per tile — because it keys on the END of the stem, where the instance number
lives. The contract's own table claimed 5 and 1 for those two booths and neither
reproduces; the numbers are reachable only by two OTHER heuristics, so the table
that justified the design was assembled from more than one rule. Its own worked
example contradicts it in plain sight.

The shipped rule keys on the first separator-delimited segment, where the family
lives, destemming only when the stem has no separator at all — so `ac01` -> `ac`
while `v30-seed8302` and `v35-seed8302` stay apart. Re-measured across all 17
live booths; the table is in the contract.

INV-3 GAINED ITS SECOND DEGENERACY. The contract guarded one group for
everything (sc-iso-spread: DSC0001-DSC0006). The live set's actual failure is
the opposite — pewpew-ui-brief yields 23 groups for 34 items, dfa-concepts 13
for 20 — and the contract as written would have shipped a rail that is a second
copy of the grid. The rail now renders only when grouping is informative: two or
more groups, and the middle group holding more than one item. That predicate
gets all 17 booths right.

Grouping is a VIEW. The grid stays sorted(rel) and the zoom ring stays that
order filtered to images; the group fixture interleaves across subdirectories
precisely so a (group, rel) re-sort goes red. Groups are derived from the
RENDERED list, not the full gallery, so no anchor points at a filtered-out tile.

booth/items.py       _group_of + Item.group, derived in the resolver (INV-1)
booth/app.py         _groups() builds the rail rows; build_gallery carries it
booth/templates/     the rail-groups nav and its CSS
tests/               +16 tests; 639 green

Every new falsifier was proved by running its defeating change (12/12). Three
were vacuous first time out: one fixture's positional order happened to be
alphabetical, one assertion miscounted elements, and the harness itself
certified a broken test twice — no green baseline, and byte-identical mutations
silently defeated by the pyc cache's one-second mtime granularity.
2026-09-22 21:33:54 -07:00
vh 2f85692e95 memory: a standing no-announcements ruling — the send is the operator's, not the agent's to ask about 2026-09-22 21:05:46 -07:00
vh 6938d21085 memory: the operator ruled on all five — U7's departure approved, main pushed
"accept all recs, or make good ones." Four of five executed.

APPROVED: drop subfolder sections for filename-prefix groups. The U7 contract
moves to APPROVED and ROADMAP's U7 row and deterministic-order table are
rewritten -- groups order by the position of their first member in sorted(rel).

SETTLED: `unanswered` means has-an-open-pick, the reading that shipped. The
has-no-mark-at-all reading is a different question and is parked to v1.1 rather
than left pending.

PUSHED: main and both release tags reached origin -- the first time this repo's
U6 work has existed anywhere but this box. Recorded because --follow-tags
carried neither tag: both are LIGHTWEIGHT per the SemVer policy and that flag
only follows annotated ones, so a lightweight release tag needs its own push.

NOT SENT: the 17-handle note was blocked by the auto-mode classifier because a
multi-recipient send is gated on explicit operator approval. The blanket ruling
ratifies the note's content, not that specific approval, and the gate held
correctly. Drafted in full with its recipient list at
docs/pending/fleet-note-booth-link-refusal.md so it survives a context clear.
Not worked around.

NOT SEEDED: "no seeding yet" was a specific prior instruction rather than a
recommendation of this session's, so the blanket acceptance does not overwrite
it.

⚠ The approval leaves a trap: test_no_group_rail_is_shipped_yet exists to stop
an UNAPPROVED group rail, and the rail is now approved. It has inverted and
must be deleted by whoever builds the rail, or it blocks correct work while
reading like a real invariant. Named in the handoff's first step for that
reason.
2026-09-22 19:49:39 -07:00
vh 8bf5343049 memory: snapshot — U7 three-quarters built, blocked on one ruling
U6 shipped as v0.6.0 and a late fix as v0.6.1; U7's three ratified components
(rail, filters, grid keyboard) are landed and the fourth is deliberately not,
because swapping subfolder sections for filename-derived groups is a scope
departure the operator has not ruled on. A test fails if anyone builds it
anyway.

Two new detail files. One decomposes U7 by ratified-versus-not and records the
two decisions taken under stated assumption. The other keeps the mechanism
behind today's misrouted directive: pane_find addresses seats by a ROLLING PANE
TITLE, which is not a stable address, and the failure is silent from the
sender's side -- Miranda had no signal until infra-ops flagged it. The incident
resolved; the mechanism did not.

The generated handoff committed the modality failure its own step-7 read exists
to catch: it listed push, seed and the 17-handle note as imperative Next steps
when all three are explicitly gated. Rewritten as do-nots, Next steps emptied.
Recorded here because it is the second time the generator has needed that
backstop.
2026-09-22 17:38:40 -07:00
vh a306e2dc6d feat(u7): the rail, the filters and the grid keyboard — the ratified three
ROADMAP's U7 row names four components. Three of them -- a sticky rail,
filters, and grid keyboard -- are already ratified there and are implemented
here. The fourth, replacing directory sections with filename-derived groups, is
a scope DEPARTURE the operator has not ruled on and is deliberately not built;
test_no_group_rail_is_shipped_yet fails the moment somebody builds it anyway,
so it cannot arrive by accident while he is away.

Filters are links carrying a query parameter, resolved server-side, so the
gallery keeps working with JavaScript off -- U3 already cost the verbatim path
its no-JS operation and said so, and the gallery is the surface the operator
actually reviews on. An unknown filter falls back to `all` rather than indexing
a dict by a value that arrives from an operator-editable URL.

`unanswered` means HAS AN OPEN PICK, the U4 hold predicate that already exists.
The other reading is a real and different question and stays open on the
contract rather than being guessed at.

Filtering is a VIEW and never reorders. The grid renders `sorted(rel)` with
non-matching items removed, so "the third one" means the same thing with a
filter on as with it off, and the zoom ring is untouched by any filter -- a
ring that changed with the grid would make `next` depend on how the operator
arrived, which is the misfiled-judgment failure invariant 6 exists for.

⚠ The first version of that invariant's test was VACUOUS and the mutation run
caught it: it compared each filtered view against the unfiltered RESPONSE, so a
reversing mutation reversed both sides and it stayed green under the exact
change it forbade. Rewritten against an independent truth -- U1 INV-3 says the
order IS sorted(rel) -- and re-verified RED. Written an hour after the entry
describing this exact failure class, which is worth recording.

611 -> 623 tests.
2026-09-22 14:45:57 -07:00
vh b50f41bb36 docs(u7): a PROPOSED contract for the last unit — scope departs from ROADMAP on measured grounds
Not approved and not implemented. Frontmatter status says so, the body says so
twice, and the one scope-direction call in it is named as the operator's.

ROADMAP's U7 row is sections, rail, filters, grid keyboard. The measurement
recorded in persistent-memory.d/2026-09-22-u7-remeasured-before-scoping.md
kills the first component -- zero of eleven gallery booths have a subdirectory,
and the only two booths that do are reports -- and supplies a replacement:
stripping a trailing digit-run from the filename stem yields 5 to 16 sensible
groups on four of the five large galleries.

The degenerate fifth is carried as a first-class case rather than an edge: one
group must render NO rail, because a navigation affordance that cannot navigate
is worse than none.

Closes ROADMAP's outstanding U7 ordering question: groups order by the position
of their first member in sorted(rel), so the rail reads in the same direction
as the grid. Grouping and filtering are views and never reorder -- INV-2 exists
because sorting by (group, rel) looks right and silently changes what 'the
third one' means, which is the misfiled-judgment failure invariant 6 was
written for.

Blast radius checked before writing: Item gains one field beside the existing
section, build_gallery carries it, and image_chain is explicitly unchanged.
2026-09-22 14:40:37 -07:00
vh e15ee2c4ab memory: U7 re-measured before scoping — pre-work only, no unit started
The standing instruction is to re-count the booths before scoping U7. Done
against the live 19-booth set, so the scope call is a short read rather than an
investigation.

Two findings. Sections are worth zero and it is now measured twice: not one of
the eleven gallery booths has a subdirectory, and the only two booths that do
are both reports, the job where grid navigation matters least. And the grouping
signal is in the filename rather than the tree -- stripping a trailing digit-run
yields 5 to 16 sensible groups on four of the five large galleries and
degenerates to one group on the fifth, while the competing split-on-second-
hyphen heuristic is useless everywhere.

The sizing case has also moved: the unit was scoped against 270-item booths and
the largest gallery is now 81 items / 40 images.

No U7 code and no U7 contract. The scope direction is the operator's call.
2026-09-22 14:38:26 -07:00
vh 400e254da6 memory: reconcile the snapshot to v0.6.1
The snapshot was written at v0.6.0 and the marks-guard fix landed after it.
Updates the in-flight head commit, the test count, and the ahead-of-origin
count so a fresh session is not told a stale number.
2026-09-22 14:35:53 -07:00
vh 1b394dde18 chore(release): v0.6.1 — the wrong-shaped answer no longer 500s
Patch, agent discretion. Bundles the pre-existing render-time 500 on the
gallery and marks pages, closed at the hydration boundary, plus the
`_safe_fragments` handler that could not survive the failure it was handling.

611 tests.
v0.6.1
2026-09-22 14:35:28 -07:00
vh e702be4e1a fix: a wrong-shaped answer no longer 500s the gallery and the marks page
Pre-existing, measured at 42ea67f, so it predates U3. `_hydrate` checked only
that `answer` was a dict and never that `answer["answers"]` was one, so
`marks_for` and `hold_read` both reported the mark healthy with no read error
-- and `_ask_inline.html` then asked a list for `.get`. The v0.2.2 lesson was
half-implemented: that outage was a file that could not be PARSED and the
reader was made lenient, while this one parses perfectly and breaks one layer
further in, at render, where no leniency existed.

Closed at the hydration boundary rather than by a third copy of the guard --
one predicate, one place, every surface inherits it. Only the multi case is
checked, because only the multi case indexes; requiring `answers`
unconditionally would break every single-question pick, and that direction has
its own test. Measured before and after: gallery and marks pages 500 -> 200,
the error visible on the page, the booth's other healthy pick untouched.

The placement was the one open operator question of the session. It was
surfaced three times without a ruling, so it is taken under a stated assumption
and is cheap to move: the whole fix is one condition in one function.

Two things fell out of it worth more than the fix.

`_safe_fragments` no longer has a reachable natural trigger. Probed every wrong
answer shape a .marks.json can carry: `answers` as a list, a string or null all
become hydration errors now, and a wrong-typed value INSIDE `answers` renders
without raising, because Jinja absorbs attribute access on a non-mapping. U3's
guard is a pure backstop, and its test now says so and trips it synthetically
through the shared macro module rather than asserting a path nothing reaches.
A guard tested by an unreachable input is an untested guard.

And that guard's handler could not survive the failure it was handling: it
caught a raising `_pick_fragments` and rebuilt the broken-ask box through the
SAME macro module that had just raised, so whenever `whole` was the broken
thing it re-raised and took the whole report. Found by accident while building
the falsifier. Fixed, with its own test.

Both new falsifiers were verified RED against their defeating change rather
than assumed.

607 -> 611 tests.
2026-09-22 14:34:28 -07:00
vh c5ac49356f memory: snapshot — U6 released at v0.6.0, six of seven v1 units landed
Nothing in flight. The in-flight section is rewritten to the post-release
state and carries the five things a fresh session must not do: push (main is 8
ahead of origin/main), seed the registry, send the 17-handle note, run either
dated prediction early, or start U7 without re-counting the booths first.

Two new detail files: the release itself, and what each of the five review
passes could only see alone -- the strongest evidence this repo has for running
all of them rather than picking one. The earlier U6 entry is reconciled; it was
written while the gates were still out and said NOT TAGGED.

Restored in the rewrite: the warning that the 17 handles were never told `keep`
stopped meaning "waiting on an answer", which is load-bearing for how the
2026-10-06 re-count reads, and the fact that a remote now exists.
2026-09-22 14:24:12 -07:00
vh 3296a868fa chore(release): v0.6.0 — U6, benches
The sixth of seven v1 units. A bench is a running thing, registered: identity
is the normalized URL so re-posting updates the row instead of appending a
fifth, `booth link` refuses the one shape that now has a better home, and the
board marks the rows whose booths are gone without deleting a single one.

Minor rather than patch, approved by the operator. Two capabilities arrived and
one verb changed behaviour for seventeen agent handles, which is the
push-notification bar in the tier test: `booth bench` is new, the board gained
a dead marker, and `booth link` now refuses a booth URL and a credentialed one.

444 -> 607 tests across the unit and its three cold gates. All four review
gates closed: an in-session seam review (three real contract defects, including
one that would have 404'd the whole board page), an in-session adversarial pass
(four defects, one of them this repo's own FIFO lesson recurring in a new
file), and three cold cross-frontier panels -- contract paraphrase, code-vs-
contract, and a diff-scoped bug hunt -- folded in full with exactly one finding
declined and its reasoning recorded.

Measured before contracted, and the measurement changed the unit: the design
doc's headline 69% rot was two defects wearing one number, and U5 had already
closed the larger half. Identity is the FULL normalized URL rather than the
origin because origin identity merges eight distinct gitea repositories, three
unrelated model cards, and the two LRPG surfaces the design doc itself names as
an example of two real benches.
v0.6.0
2026-09-22 14:21:23 -07:00
vh 8cb21193dc fix(u6): fold the cold bug-hunt panel — a div in a span, a symlink split, and an append outside its lock
/heid-bug-hunt panel 01M35CRRK2RTVWWF1BN09AFQG3, diff-scoped against 91fd8bc.
The most severe of the three rounds, and three of its four convergent findings
were already closed by our own adversarial pass before the reply landed. Three
were not.

- The benches panel was nested inside the booth header's <span class="sub">.
  The insertion had matched the first `{% if board %}` in the template rather
  than the block-level one. A div inside a span is invalid HTML: the parser
  closes the span implicitly and hoists the div out, orphaning the rest of the
  sub-line. Nothing 500s, which is precisely why no test in this suite could
  see it. Moved to block level, pinned by an offset assertion, and verified
  with a real HTML parser.

- _booth_exists used a bare is_dir() while resolve_booth resolves and requires
  the parent to BE the data root. They disagreed on a symlink: the marker
  called a booth pointing outside the root alive while the page 404s it, so the
  row rendered healthy and the link was dead. Same containment now, and
  ValueError joins OSError in the guard -- one bad row must never cost the
  other 220.

- The board append opened its fd OUTSIDE the lock. `flock LOCK printf ... >>
  board` reads as locked and is not: the shell opens the append fd while
  parsing, before flock acquires. A concurrent unlink replaces the inode via
  os.replace, the old fd still points at the unlinked one, and the append
  succeeds, reports success, and vanishes. Pre-existing rather than this
  unit's, but it is silent data loss in the file this unit lives in. Proved by
  holding the lock and asserting nothing is written.

- The atomic write used a predictable .tmp.<pid> name; a pre-planted symlink
  there redirects the write straight through the replace. mkstemp with O_EXCL
  in the same directory, and an fsync before the replace -- os.replace orders
  the rename, not the data behind it.

Declined and recorded: on a host where booth.links cannot be imported, `booth
link` now refuses every URL rather than only booth ones. True, and kept. A
guard that fails open is not a guard, and that state is a broken install in
which most of the CLI is equally broken.

The sharpest line in the reply is one three arms found independently: this repo
had ALREADY paid for the RecursionError class in marks.py, and the new module
re-introduced the unguarded parse. Reading the new module in isolation would
never have surfaced that.

604 -> 607 tests.
2026-09-22 14:20:06 -07:00
vh e3853e2692 docs(u6): the CLI usage strings carry the --apply <id> form
The three places scripts/booth documents itself -- the header block and both
usage lines -- still described a bare --apply, which is now refused. A usage
string that names a form the script rejects is worse than none.
2026-09-22 14:13:41 -07:00
vh 32e3ed65e1 fix(u6): fold the cold contract panel — the import selection gap, and a document arguing with itself
/heid-contract-review panel 01M35BWCJ806MT75NA630Y4WFH. The headline arrived
from all four arms independently and it is a missing feature, not a wording
problem.

`bench import --apply` registered every candidate, while the same contract says
roughly 14 of 35 are reference bookmarks that must stay on the board. There was
no selection mechanism between the dry-run report and the write -- so the write
path did the exact thing this unit's rationale calls impossible, tell a bench
from a bookmark by its URL, silently, to rows that belong where they are. The
report existed precisely because the decision is not mechanizable. `--apply`
now takes the ids the operator names; a bare `--apply` is refused and an
unknown id is refused, both writing nothing.

Two solo findings, both real:

- A successful registration could push the registry past the size its own
  reader refuses, so the LAST bench added would make every other bench
  invisible while reporting success. The writer now respects the reader's cap.
- The credential ban covered bench URLs and not `booth link`, the door this
  unit did not touch -- and the board renders on an unauthenticated LAN
  surface. A password can no longer reach it through either door. A small
  deliberate widening, named rather than smuggled.

Cap semantics were readable three ways (refuse / clip-for-display /
truncate-and-store) with a different build behind each, 4-of-4. Now stated per
field: name and owner truncate, url and state are refused at the write and are
DAMAGE at the read. url is not a display budget -- INV-7 promises the click
goes to the posted address byte for byte, and a clipped URL keeps that promise
in the type system while breaking it in the browser. The code had been clipping
it; fixed.

Two passages disagreed about one character: INV-7's specimen named "a trailing
slash on a non-empty path" as something normalization changes, while the rule
list keeps it and INV-6 makes the two spellings two benches. The rule list is
right; the specimen was wrong. Found by 3-of-4.

Also: INV-6's component list was illustrative where it had to be exhaustive and
was short scheme and port; "writes nothing" appeared twice with different
lists; the dead marker's predicate was readable two ways with 221 rows riding
on it; and INV-2's falsifier read as though three callers agreeing pinned
something, when three callers of one wrong predicate agree perfectly -- the
table's expected values are the real check and now say so.

597 -> 604 tests.
2026-09-22 14:12:36 -07:00
vh 8a7af3eb08 fix(u6): fold the cold code-review panel — four-arm convergence on three surface clauses
/heid-code-review panel 01M35CK8YKEKMV7T15JXEF6A8N, verdict NOT drift-zero.
Three findings arrived from all four arms independently, and they share a
shape: a contract clause written as prose and never converted into an
assertion. That is the lens working.

- The panel dropped the added date the contract promised to show.
- `bench ls` printed no ids, and the URL it printed was truncated to 52 columns
  so the line was not pasteable into `bench state|rm`. The test's docstring
  claimed it printed ids and asserted nothing of the kind.
- `bench import` printed the description instead of the raw URL beside each
  normalized id, hiding the collapse the clause exists to expose.
- An IPv6 literal lost its brackets: http://[::1]:8080/a normalized to
  http://::1:8080/a, a broken identity that no re-post can match. Bracketed
  literals are re-wrapped; an unbracketed one is refused rather than guessed.
- A deeply-nested JSON RecursionError escaped read_benches' except pair. The
  byte cap does not help -- 200k open brackets is 200 KB.
- An empty board hid the whole benches panel, registration form included.
- The link refusal classified by captured-text emptiness, which bash can erase;
  it now answers with a B:/N sentinel so no name reads as "not a booth".

INV-4's tie-break falsifier could not fail: _write_all serializes with
sort_keys=True, so both insertion orders came back already id-sorted and
removing the tie-break left the test green. It now calls order_benches
directly. Same class as the five vacuous U4 falsifiers, found by a cold reader
rather than by us.

Also from the arms' per-invariant vacuity pass: INV-6 had no vector pinning a
non-default port as part of the identity; INV-3 asserted only that links/ was
absent; INV-8's hashed sequence omitted a read verb; INV-9's AST walk is
defeated by a string import. All closed.

Contract amended where the code was right: `updated` means last mutation, the
id cap is write-only because the id is the locator controls post back, INV-8's
file list includes the lock sidecar it always mandated. Every line number is
out of the prose -- the panel found two already stale.

565 -> 593 tests. Nothing declined.
2026-09-22 13:50:40 -07:00
vh 0a2bb1d26c fix(u6): the booth check fails closed with a reason, and a dead write leaves no scratch
Two more from the in-session adversarial pass.

`booth link`'s new booth-URL check shells out to booth/links.py. When that
import cannot run, the command substitution under `set -e` aborted the script
with a bare ModuleNotFoundError traceback: the right DIRECTION (no row was
appended — a guard that fails open is not a guard) reached by accident, and
unactionable when it fires. Handled explicitly now: exit 3, and a message
naming what the check needs. The fail-closed direction is stated rather than
inherited from shell semantics, and a test pins it — the defeating change in
either direction goes red.

_write_all's scratch file was stranded beside the registry if the write died
between create and replace. Cleaned up on every exit path. The prior registry
was never at risk either way: os.replace is the only thing that publishes.

Also pins normalization idempotence, which `bench state <id|url>` and
`bench rm <id|url>` both rely on: they normalize whatever they are handed, so
an id that did not normalize to itself would miss the row it names.
2026-09-22 13:33:59 -07:00
vh 8c7f2127eb fix(u6): a FIFO at the registry path hung the render, and unquote leaked control characters
Both found by the in-session adversarial pass while the cold panels were still
out. The first is this repo's own 2026-09-22 lesson recurring in a new file.

_read_bytes bounded the READ and its docstring claimed that closed the
named-pipe hole. It does not: open() blocks on a FIFO with no writer, before
any byte cap can apply. read_benches runs on the board page's render path, so
one FIFO there is a request that never returns and, with enough hits, the
threadpool behind every route. Guarded with S_ISREG before the open, which is
what marks.py has done since it learned the same thing. The bounded read stays
for the case a stat cannot answer: a regular file that grew between the two.

booth_target handed back whatever unquote produced, including NUL and newline.
Neither can name a directory, and unfiltered they reach is_dir() -- which
raises ValueError on an embedded NUL, and ValueError is not an OSError, so it
escapes the dead marker's guard -- plus the refusal message the CLI prints and
the marker the board renders.

Both tests are written to go red under the exact change that defeats them: the
FIFO test blocks rather than fails if the regular-file check is removed, and
the control-character rows need their own case because %2e%2e and %2f stay
green without the clause.
2026-09-22 13:30:22 -07:00
vh 1c3ce5ddb5 feat(u6): benches — a registry with identity, and the rule enforced
The standing link board carried three jobs because only one of them had a
surface. Re-measured before contracting, its 221 rows split into 178 booth
announcements (156 already dead) and 43 non-booth rows, of which 8 are the same
bench re-posted. U5 gave the booth announcement a home; this gives the running
service one, and refuses the one shape that now has somewhere better to go.

- booth/benches.py (new, stdlib-only and sibling-free): the Bench record, URL
  normalization as the identity, a lenient read on the render path and a strict
  read on the write path, atomic replace under an flock, and a stated total
  order (state rank, name casefolded, id).
- links.booth_target: ONE predicate for "is this a booth URL", consumed by the
  CLI refusal, the board's dead marker and bench import. Host-agnostic,
  path-shaped, percent-decoded, never raises.
- booth link refuses a booth URL, names `booth new --why`, and writes nothing —
  not the row, not the board directory, not the announcement.
- The board marks rows whose booth has been swept. Nothing here deletes a row:
  removal stays the operator's two clicks through the existing bulk control.
- booth bench add|ls|state|rm|import. import writes nothing without --apply and
  never edits links.md.
- docs/archive/links-2026-09-22.md: the board archived verbatim into git.

Identity is the FULL normalized URL, not the origin, and that was measured:
origin identity collapses the 43 non-booth rows to 19 groups by merging eight
distinct gitea repositories into one row, three unrelated HuggingFace model
cards into one, and the two LRPG surfaces on 10.100.10.50:8321 — the design
doc's own example of two real benches — into one. Full-URL identity still
collapses both cases that doc names: talk 5 to 1, Peedlar 3 to 1.

booth link is NOT deprecated. Roughly 14 of the 35 distinct non-booth targets
are reference bookmarks for which the board is the right and only home; the
design doc's plan to deprecate it would have evicted a third of its live
content. Corrected there, along with what "normalized URL" means.

The seam review found three real defects in the contract before any code: the
claim that test_stdlib_only already forbids sibling imports (it exempts `booth`
on purpose), naming resolve_booth as the dead marker's existence check (it
raises HTTPException(404), so one swept booth would have 404'd the whole board
page), and silence on percent-encoding (booth links are emitted through
quote(name, safe=""), so a raw comparison marks every encoded booth dead
forever). That both list_booths and sweep_once skip the registry was verified
against the real functions rather than assumed.

444 -> 555 tests. Deployed and verified live: 23/23 booths 200, and the board
renders 156 dead of 221 rows, matching an independent pre-implementation count.

NOT TAGGED: both cold gates are in flight (contract review
01M35BWCJ806MT75NA630Y4WFH, code review 01M35CK8YKEKMV7T15JXEF6A8N) and the
bug-hunt has not run. Per the v0.2.0 lesson, the tag waits for the gates.
2026-09-22 13:25:32 -07:00
vh 91fd8bc69d memory: snapshot — U3 released at v0.5.0, pushed and deployed
First push of this repo's history: main was 26 commits ahead of origin/main, so
v0.2.0 through v0.5.0 all reached the Gitea remote in one motion. A future
session can assume a remote exists, which no earlier one could.

Records the open defect U3 found and deliberately did not fix -- a well-formed
.marks.json with a wrong-shaped answer 500s the gallery and marks pages,
measured at 42ea67f so it predates the unit -- and marks it explicitly as
awaiting an operator decision with no issue filed, rather than letting it sit in
a detail file nobody is routed to.

No recommendation is recorded for the next unit. U6 and U7 are genuinely
independent and close different defects; the last two before a 1.0 cut are a
scope-direction call.
2026-09-22 13:00:34 -07:00
vh 7996fbd597 chore(release): v0.5.0 — U3, the declared embed seam
Minor rather than patch, and the tie-break rule says default to patch, so the
reason is worth stating: a capability arrived AND one left. Report authors gain
a declared public API -- one line, `<script src="/_booth/embed.js" defer>`, plus
the `data-booth-mark` anchor syntax -- and the verbatim path loses
no-JavaScript operation, which it had since it existed.

That asymmetry is what makes it not a tie. Either half alone would have been
defensible as a patch.

Operator approved 2026-09-22.
v0.5.0
2026-09-22 12:58:35 -07:00
vh 5c20e2f4d5 fix(u3): seven defects two cold panels found in the declared seam
The /heid-code-review and /heid-bug-hunt panels, artifact-only over the U3
diff, between them found four real defects and three vacuous falsifiers. Both
snapshots predate the contract-review fixes, so two of their findings were
already closed; the rest are here.

Prototype pollution in the placement maps. A mark id and a question key are
both [A-Za-z0-9][A-Za-z0-9._-]*, so `toString` and `constructor` are legal in
each. Against a plain `{}` an anchor naming NO mark returned an inherited
function, passed the guard meant to reject it, and threw on .questions.length
-- aborting placement before the tail, so one typo in author markup cost the
page every ask. The `placed` set had the mirror bug: inherited
`got.constructor` read as already-placed and silently dropped a question.
Object.create(null), three times. Found independently by both panels.

A declaring page was not served as written. read_text() opens in
universal-newline mode, so a CRLF report came back LF, and errors="replace"
replaced every byte that was not valid UTF-8. That is this unit's headline
promise, broken by the read itself, and the test could not see it because its
fixture was LF-only ASCII. The verbatim branch reads and serves bytes now; the
decoded copy answers only "does it declare the seam?".

A submit anchor inside the author's own <form> lost ours -- the parser drops a
nested form element outright -- while the code still recorded the pick as
submitted, so no fallback was appended. Every control's form= pointed at
nothing and the button did nothing. It counts as submitted only if the form
survived.

A broken pick's diagnostic never rendered from a submit-only anchor: an errored
pick's submit block is empty, and mounting that then marking it placed made the
tail skip the "broken ask" box entirely. The anchor is left alone instead.

An author's own element could hijack the open-ask chip -- id="bk-ask-winner-
background" satisfies any prefix rule, hyphen boundary included. The chip now
searches only elements this script mounted, which is the identity the deleted
bk-ask-<id>-top anchor used to guarantee, and takes the earliest by
compareDocumentPosition.

No error boundary around fragment rendering. A .marks.json that is well-formed
JSON with a wrong-shaped answer hydrates with no error and then raises in the
macro; this endpoint renders every pick on every load of the report, so that
was the whole seam gone while hold_read called the file readable. Reproduced
before building for it. _safe_fragments gives it the per-mark leniency
_hydrate_safe already applies one layer down.

The gallery and marks pages still 500 on that same entry. Measured at 42ea67f
-- it predates this unit, they render the same macro with no guard, and the
gallery is named out of scope in the contract. Recorded, not quietly widened:
persistent-memory.d/2026-09-22-a-wrong-shaped-answer-500s-the-gallery.md

Also corrected: several comments claimed a multi-question pick POSTs a 400
unless every question is answered. It does not -- an empty submission is
refused, a partial one is recorded on purpose. The real reason an unplaced
question must still be appended is that a question which never reaches the page
cannot be answered at all.

Vacuity pass rebuilt around the rule this session learned: the mutation comes
from the invariant's claim, never from the falsifier's example. 21 mutations,
21 caught, unmutated control green. Getting there took three rounds -- it
passed INV-3 with the contract's own mutation, then found its own fix's hole,
then flagged seven stale mutations and one genuinely vacuous fixture whose
sibling-mark arrangement made the right answer also the first answer.

444 tests. Deployed and verified: 23/23 booths 200, and all four live verbatim
reports served at exactly +46 bytes -- len(EMBED_SCRIPT_TAG) -- with the
authors' own wrappers and headings intact and no console errors.
2026-09-22 11:22:37 -07:00
vh 87e2c5364c feat(u3): a verbatim report declares the seam, the Booth mounts into it
A booth that ships its own index.html was served through ten regular
expressions applied to markup the Booth did not write: six in
wrap_verbatim_html hunting for somewhere to hang a favicon and a chip, four
in booth/inline.py substituting rendered ask markup into the author's own
tags. Both worked. Both were the most fragile thing in the service, on the
path the operator uses most.

The whole class is replaced by a declared seam. A report carries one line —
<script src="/_booth/embed.js" defer></script> — and the chrome mounts
through DOM APIs. What the server does to author HTML is now, in full:

    return html if declares_embed(html) else html + EMBED_SCRIPT_TAG

Two substring tests and a concatenation. Both of the old wrapper's hard
constraints stop existing rather than being satisfied more carefully:
nothing can displace a leading doctype into quirks mode and nothing can push
the charset meta out of its detection window, because nothing in front of
them ever moves. A page that declares the seam is served exactly as written.

Fragments are still rendered by the _ask_inline.html macros and handed over
GET /b/<name>/embed.json; embed.js places them and decides nothing. Openness
comes from open_marks, order from (created, id), questions in declaration
order. A single-question pick normalizes to key None, so the payload carries
questions as a list rather than an object — keying by name would serialize
that as the string "null".

Placement is an anchor fill, not a replacement: el.insertAdjacentHTML(
'beforeend'), so an author's wrapper and its contents survive. The regex it
replaces was eating the opening tag of dfa-concepts' styled .ask blocks and
orphaning their headings, live, unreported.

data-booth-mark is canonical; data-booth-ask stays a kept alias because two
live reports use it. The comment placeholders are dropped — no users.

Declared cost: the verbatim path now needs JavaScript. The never-invisible
guarantee holds through the index badge and /b/<name>/marks, both of which
render server-side.

Deleted: booth/inline.py entire, wrap_verbatim_html and its six patterns,
_BACK_CHIP, asks_chip, inject_asks, FAVICON_LINK, the styles() macro.

Tests 410 -> 434. tests/test_embed_browser.py drives a real Chromium: the
placement algorithm and the form= binding of a scattered multi-question form
cannot be observed any other way, and that binding was measured rather than
assumed (N=3 per condition, with a form-first positive control and a
points-at-nothing negative control).

Contract: docs/contracts/u3_declared_embed_seam.contract.md, with the
in-session seam review and the cold contract panel both recorded. Two of the
panel's findings were code fixes: a vacuous INV-3 falsifier that a renamed
regex walked straight through, and a bare-substring seam detection that read
a report merely quoting the path as declaring it and silently served it with
no chrome.
2026-09-22 10:43:41 -07:00
vh 42ea67f33f memory: snapshot — U4 released at v0.4.0, next unit undecided
Current state rewritten for the post-U4 position: 410 tests, v0.4.0 tagged,
tree not pushed, all three gates closed. Carries the session's U3 recommendation
with its three grounds AND its counter-argument, so the operator can take the
call without reloading the unit.

The methodology-proposals row goes from three to four and is now marked
explicitly untracked by operator choice — the new one is the contract-time
vacuity pass, which is the only one of the four with measured evidence behind
it after five of seven U4 falsifiers turned out not to discriminate.
2026-09-22 10:06:30 -07:00
vh 8f81d8f9d0 memory: no fleetwide notice for U4, and the measurement caveat it creates
Operator decision 2026-09-22: no broadcast to the 17 consuming handles. Same
posture as U5 — adoption gets told apart from design because nobody was primed.

The consequence is a measurement one and it needed writing down before it was
lost. U4's two halves have different adoption costs: the hold rides for free
(a session runs `booth ask` and its booth is held, knowing nothing), but NOT
pressing `keep` has to be learned. So a flat `.forever` rate on 2026-10-06 is
exactly what 'the mechanism works and nobody was told' looks like, and reading
it as a falsification would retire a correct diagnosis on an uncontrolled
measurement.

Records the three counts to report instead, and states the sensitivity floor:
only 4 of 24 booths carry marks at all, so the hold can touch at most a sixth
of the fleet and an effect below one or two booths is not resolvable.
2026-09-22 10:04:51 -07:00
vh c75d7a2797 fix: four defects the U4 bug-hunt panel found in code it did not add
All four pre-date U4 and sit in files it touched, which is why a diff-scoped
robustness lens saw them. They are separated from the unit's own commit so the
feature history stays readable; the release tags both.

* A booth name reached a JS string context. The confirm dialogs interpolated
  the name into a string literal inside `onsubmit`. Jinja's autoescape is
  HTML-attribute escaping, not JS-string escaping: the browser decodes the
  entity back to a quote before the JS parser sees it, so a name crafted to
  close the string executed on submit. Booth names are agent-authored — making
  a folder under the data dir is the whole API — so this was a live path, not a
  theoretical one. The name now travels as a data attribute to a delegated
  handler, where escaping is escaping.

* An unreadable `links.md` returned 500 for the whole booth page. `is_file()`
  then an unguarded `read_text()`. The board is one tile on that page, and a
  page that will not load is worse than one missing a tile — the posture
  `read_blurred`, `marks_for` and `read_manifest` already take.

* The index order had no tie-breaker, which violates the deterministic-order
  invariant. Equal-mtime booths fell back to whatever `iterdir()` yielded, and
  two booths landed by one `rsync` batch share an mtime exactly. Now
  `(mtime, name)` reverse: newest first, then name. The operator refers to
  cards positionally, so a sequence that moves between renders misfiles his
  judgment rather than crashing.

* `/b/<n>/marks.json` reported damage as empty success. `booth marks` exits 3
  on an unreadable file precisely so a caller can tell "not yet" from "broken";
  the HTTP mirror — the only reader a remote session has — returned the same
  empty list for both. It now carries `error` and `detail`. The status stays
  200 deliberately: reads are lenient here, and a pinned status code is a
  promise to remote clients this fix has no business breaking.

Each has a regression test. 410 tests.
v0.4.0
2026-09-22 09:51:14 -07:00
vh c3a97c1b64 feat(u4): a booth's lifetime is derived from its state, not from a boolean
`.forever` was the only way to say three different things — "this is durable",
"I have not answered yet", "I am still looking" — and the census said it was
carrying all three: 17 of 24 live booths (70%, up from 54% the day before).
Three of the four booths in the fleet awaiting an answer had been pinned by
hand as well, and 10 of the 17 were younger than the TTL, so the sentinel had
bought them nothing and was pressed pre-emptively.

Only the first meaning is what `keep` means. The other two are facts the
service already held and did not consult.

    KEPT       `.forever` present                      never swept  (unchanged)
    HELD       an open pick, or marks we cannot read   never swept  (new)
    EPHEMERAL  everything else                         24h          (unchanged)

Viewing is activity: a deliberately-served response from a booth's own page
route writes `.viewed`, which is a dotfile and not a `.lock` dotfile, so
`_newest_mtime` already counts it. There is no new arithmetic — `booth_age_seconds`,
`is_expired` and `expires_in` are unchanged. Machine reads are excluded on
purpose: an agent must not be able to hold its own booth open by polling for
the answer it is waiting on.

The hold is unbounded, and what makes that safe is visibility plus two exits
that already existed. Every surface whose chrome the Booth owns says
`held until answered` where the countdown was, and `booth rm` / the UI x /
`DELETE /b/<n>` take a held booth exactly as they take a kept one. A hold is
protection from the timer, never from the operator.

Three cross-frontier panels ran and each found a class the others could not:

  * the paraphrase panel found that two reads of one file are not one read of
    one state — the contract's `is_held(marks_for(c), read_error(c))` could
    resolve to `([], None)`, the pair that deletes. `hold_read` is one read.
  * the code-review panel found, 4-of-4, that the booth header's board branch
    rendered no lifetime at all; and that five of seven invariant tests passed
    under the change that defeats them.
  * the bug-hunt panel found four more paths where a failed read still
    authorized a delete, and a `record_view` that followed a planted symlink.

`is_held` became `hold_reason`, which returns the reason rather than a bool
beside a string that can disagree with it.

Prediction, to re-count on or after 2026-10-06: the `.forever` rate falls to
the booths that are genuinely durable references. Only 4 booths carry marks at
all, so this rests on both halves of the unit; a null result cannot distinguish
a wrong diagnosis from a habit that outlived its need.

406 tests (341 before). Contract: docs/contracts/u4_derived_lifetime.contract.md
2026-09-22 09:44:25 -07:00
Vuong Hoang d37b81ab9f memory: snapshot — U5 released at v0.3.0, and the index goes two-tier
The two dated log sections had never been split, so every one of their 29
entries sat inline and the startup index had grown to 372 lines — which is
the cost the two-tier scheme exists to remove, paid on every session that
reads the file. 27 entries were over threshold. All 29 now have a detail
file under persistent-memory.d/ and a one-line index entry that routes
rather than restates. Index: 372 -> 93 lines.

No archival. The soft cap fired, but every entry in this repo is dated
2026-09-21 or later, so the under-14-days guard held all of them back — and
the split alone took the index well under the target without moving
anything out of the active file.

The in-flight section is rewritten for the post-release state: nothing is
in flight, no gate is outstanding, and the next unit is explicitly recorded
as the operator's undecided call rather than as a plan. The session's
recommendation (U4, on three grounds) is written down so it does not have
to be re-derived, alongside the two alternatives and why they are
alternatives.

Two dated predictions are carried forward with their dates and their
instruments: the U5 adoption re-measure on 2026-09-29, which already reads
3 of 24 announced and 2 with a why from peers told nothing, and the
.forever re-count a fortnight AFTER U4 lands, which is U4's own success
criterion and is destroyed by running it early.
2026-09-22 08:20:51 -07:00
Vuong Hoang 95beede3c3 fix(manifest)!: the size cap opened a service-wide hang; close it
The diff-scoped bug-hunt panel, four arms, artifact-only. Its strongest
finding is one I created two hours earlier while hardening the reader.

`stat` reports size 0 for a FIFO and 0 for a symlink to /dev/zero, so both
sail under the byte cap added for the RecursionError round — and then
`read_text` either blocks in read() with no EOF, so the except never runs,
or allocates until the kernel intervenes. `list_booths` reads every booth
on every GET / and /healthz, so ONE such file stalls the front page for the
whole service, with no error and no recovery short of a restart.
Reproduced before believing it (timeout returned 124). S_ISREG is checked
BEFORE the size in both modules now; verified against the live service with
two FIFOs planted, which answered 200 in 36ms.

The shape worth carrying: st_size answers a different question than "can
this be read", and a bound that trusts it inherits everything it does not
mean. A hardening fix opened a worse hole than the one it closed.

THE UPLOAD PATH WROTE ABOVE ITS OWN CLEANUP GUARD (4/4)

A failed manifest write orphaned a .uploaded half-booth with no files in
it — and because the temp name now carries a random suffix, nothing ever
overwrote the leak, and .booth.json.<hex>.tmp is not a .lock, so
_newest_mtime counted it and kept that empty booth past every sweep. The
uniqueness fix from the previous round is what made the leak permanent.
Both writes moved inside the guard; the temp is removed on every exit path.

DAMAGED BYTES ARE KEPT, NOT REPLACED (4/4, INV-6)

Marks made this explicit in v0.2.1 and this write path contradicted it: a
manifest that failed on ONE field lost the others with it, including a why
the re-announcer may never have kept anywhere. It diverges from marks in
HOW it honours the rule — marks refuse and answer 409 because the
operator's judgment is not restatable; a manifest quarantines and proceeds,
because refusing would fail `booth add` and lose the files it was copying.

ONE OPENNESS PREDICATE, AS U2 SAID (2/4)

`booth answer` spelled out `if m.answer is None` while `booth marks` asked
`open_marks`, so a partially-answered pick read as done to one verb and
open to the other — at the same instant, on the same booth. U2's INV-2 put
openness in one function precisely so they could not drift. The mirror case
is fixed too: a pick that hydrates broken is refused by the web route, so
`answer --wait` polled an hour on a form nothing could ever land.

ALSO

- now_stamp was whole-second while the importer had moved to microseconds,
  and '-' sorts before '.', so a later mark came out ahead of an earlier
  import inside the same second. One format; the previous round's ordering
  fix had opened this one.
- `_broken` was the third of three directory-name fallbacks and the one
  still handing a raw name into a card's sub-line.
- An identical re-announce rewrote the file and reset the TTL. `booth link`
  does this on every post to the standing board.
- The importer's return went through the bare _hydrate, not _hydrate_safe.
- A marks document could be written larger than it can be read back, and
  then read as no marks at all. Refused at the write instead.
- `choice` reached the answer builder raw while `notes` beside it did not.

AND ONE FINDING DELIBERATELY NOT FULLY CLOSED

The mtime-restore race is real. The clean fix — ignore a booth directory's
own mtime whenever the booth holds anything — also silently retires the
documented rule that releasing a kept board resets its clock, which the CLI
header, the README and a deliberately-written test all pin. That is a TTL
doctrine change, not a bug fix, and an existing test caught the attempt.
The concrete half is fixed (a failing os.utime escaped and 500'd the
route); the race is stated in the code where the next reader will meet it.

341 tests. Live service restarted, 24/24 booth pages verified.
v0.3.0
2026-09-22 02:27:18 -07:00
Vuong Hoang f3193fb054 fix(probe): the disclosure-opening loop was manufacturing its own findings
`page.locator("details:not([open])").all()` hands back POSITIONAL locators
that re-resolve against the current DOM, and `:not([open])` stops matching
an element the moment it is opened — so opening them one at a time shrinks
the set underneath the indices and leaves some closed. Those then report
OCCLUDED, which is exactly the false-positive class the block was added to
remove. One on booth-redesign, three on cr123a-to-d-sleeve, one on
denoise-first-run, and invisible as a bug because a false positive is
shaped like a finding.

Measured both hypotheses rather than guessing between them: per-element
loop against a single document-wide evaluate, at 150 ms and 1000 ms settle.
The loop reports them at either wait; the single pass reports none at
either. The variable was the method, not the timing.

One evaluate over the whole document now. All three pages clean.

Also carries the ROADMAP U5 row, the two-panel record in
persistent-memory.d/, and the memory index line for it.
2026-09-22 01:39:43 -07:00
Vuong Hoang c015a917ee fix(manifest): fold in both cross-frontier panels — and a live hole in v0.2.2
Two four-arm artifact-only rounds landed together: the contract paraphrase
(against the pre-seam-review capture) and the code-vs-contract conformance
review (against the amended one), correctly firewalled from each other.
The conformance round found ZERO drift in the strict sense — the code is a
clause-for-clause implementation of the contract — and the weight of both
rounds landed one layer down, in what green tests structurally cannot
report. Full triage in persistent-memory.d/.

A LIVE HOLE IN RELEASED CODE, FOUND ON THE SIBLING MODULE

v0.2.2 adopted the RecursionError finding from the bug-hunt round and
closed half of it: `_hydrate_safe` guards hydration, but `json.loads` runs
above it in `_read_raw`, whose catch list covers neither RecursionError nor
MemoryError. A 400 KB file of nothing but brackets in any ONE booth
therefore still returned 500 for `/` and `/healthz` across every booth on
the service. Confirmed by running it before believing it.

Both modules now bound the read by `stat` before touching the bytes and
catch both classes anyway, so raising a bound later cannot quietly re-open
the hole. The strict half of the marks asymmetry refuses everything the
lenient half tolerates, or a file that reads as "no marks" gets replaced by
a write that believed it.

THE WHY-WIPE

`booth new x --why "..."` then `booth add x out/*.png` erased the sentence
the first command existed to record. Omitted flags meant empty strings and
empty strings overwrote. Two arms predicted it from the contract's wording
alone; every test here passed --why on both calls and so could not see it.
Omitted now means unchanged and an explicit --why "" still clears — the
shell carries the distinction by leaving the variable UNSET, not empty.

--title WAS WRITE-ONLY

Stored, flag-surfaced, rendered nowhere. 4/4, and independently top-ranked
by every arm of the paraphrase round. It lands on the booth page heading
with the directory name beside it, because the directory name is the
identity the operator navigates by and refers to positionally.

THREE TESTS THAT COULD NOT FAIL

- test_the_write_is_atomic asserted no *.tmp survived, which a plain
  write_text passes. It asserts the inode changes now. (The first
  replacement was ALSO vacuous — it spied on os.open, which Path.write_text
  reaches through io.open in C and never touches. Recorded in the test,
  because writing a second vacuous test while fixing the first is exactly
  the failure this round is about.)
- The INV-3 preservation test passed against an implementation that
  regenerated `created` every time, because _now() is whole-second
  resolution and back-to-back writes share a stamp. Seeded from 2019 now.
- test_announcing_is_activity passed whether or not _newest_mtime counted
  the manifest, because writing it bumps the directory mtime either way.
  The directory's clock is put back, leaving the file as the only thing
  that can keep the booth alive.

ALSO

- The title fallback skipped the normalizer the explicit value gets; a
  directory name may legally carry a newline and run to 255 bytes.
- Every writer derived the same .booth.json.tmp. Marks are protected from
  that by their flock; the manifest has none, so uniqueness stands in.
- test_stdlib_only was blind to relative imports in all four modules.
- INV-1 had no guard at all; INV-5 named two different promises; the
  negative render states were asserted on the index only.

Contract amended throughout: the 4 GB case is a stat-checked bound rather
than a return constraint, every field of an error-carrying record has a
stated value, INV-1 no longer contradicts INV-3, repo-wide rules are named
in words instead of by a colliding number, and touches admits the macro
partial the implementation added.

329 tests.
2026-09-22 01:29:27 -07:00