Files
booth/booth/benches.py
T
vh 8cb21193dc fix(u6): fold the cold bug-hunt panel — a div in a span, a symlink split, and an append outside its lock
/heid-bug-hunt panel 01M35CRRK2RTVWWF1BN09AFQG3, diff-scoped against 91fd8bc.
The most severe of the three rounds, and three of its four convergent findings
were already closed by our own adversarial pass before the reply landed. Three
were not.

- The benches panel was nested inside the booth header's <span class="sub">.
  The insertion had matched the first `{% if board %}` in the template rather
  than the block-level one. A div inside a span is invalid HTML: the parser
  closes the span implicitly and hoists the div out, orphaning the rest of the
  sub-line. Nothing 500s, which is precisely why no test in this suite could
  see it. Moved to block level, pinned by an offset assertion, and verified
  with a real HTML parser.

- _booth_exists used a bare is_dir() while resolve_booth resolves and requires
  the parent to BE the data root. They disagreed on a symlink: the marker
  called a booth pointing outside the root alive while the page 404s it, so the
  row rendered healthy and the link was dead. Same containment now, and
  ValueError joins OSError in the guard -- one bad row must never cost the
  other 220.

- The board append opened its fd OUTSIDE the lock. `flock LOCK printf ... >>
  board` reads as locked and is not: the shell opens the append fd while
  parsing, before flock acquires. A concurrent unlink replaces the inode via
  os.replace, the old fd still points at the unlinked one, and the append
  succeeds, reports success, and vanishes. Pre-existing rather than this
  unit's, but it is silent data loss in the file this unit lives in. Proved by
  holding the lock and asserting nothing is written.

- The atomic write used a predictable .tmp.<pid> name; a pre-planted symlink
  there redirects the write straight through the replace. mkstemp with O_EXCL
  in the same directory, and an fsync before the replace -- os.replace orders
  the rename, not the data behind it.

Declined and recorded: on a host where booth.links cannot be imported, `booth
link` now refuses every URL rather than only booth ones. True, and kept. A
guard that fails open is not a guard, and that state is a broken install in
which most of the CLI is equally broken.

The sharpest line in the reply is one three arms found independently: this repo
had ALREADY paid for the RecursionError class in marks.py, and the new module
re-introduced the unguarded parse. Reading the new module in isolation would
never have surfaced that.

604 -> 607 tests.
2026-09-22 14:20:06 -07:00

417 lines
19 KiB
Python

"""Benches: a running thing, registered.
A bench is NOT a booth and NOT a bookmark. It is a durable middle-to-long-term
testing surface — jackdaw's current bench, talk's current bench, the things that
get promoted to Homepage when they are fully deployed. The standing link board
absorbed the job because it was the only surface on offer, and an O_APPEND log
with no identity turns "here is the bench again" into a fifth row rather than an
update: `talk` is on the board five times and Peedlar's root three.
STDLIB ONLY, AND SIBLING-FREE, ON PURPOSE. `scripts/booth` imports this through
a `python3 -c` heredoc under the system python3 with no venv, exactly as it
imports `marks`, `asks`, `links` and `manifest`. A third-party import breaks
`booth bench` on every fleet host; a `from booth.links import ...` breaks it on
any host where both modules are not importable together, which is a second way
for the same invariant to fall. `tests/test_benches.py` forbids both.
SINGLE-WRITER, MANY-READER — the opposite shape from `links.md`. The board is a
multi-writer append log because seventeen agent handles post to it at once. This
is the operator in one browser plus occasional CLI calls, so it is one file,
rewritten whole under a lock, replaced atomically. Inheriting the append-log
design here would be the mistake CLAUDE.md names by name.
"""
from __future__ import annotations
import fcntl
import json
import os
import stat
import tempfile
from dataclasses import dataclass, replace
from datetime import datetime, timezone
from pathlib import Path
from typing import Iterable
from urllib.parse import urlsplit, urlunsplit
# At the DATA ROOT, not inside a booth. A dotfile there is invisible to
# `list_booths` and to `sweep_once` — both skip a child that is not a directory
# AND a child whose name starts with a dot, so the registry fails two guards
# rather than one. Verified against both functions (seam review SR-4, SR-5)
# rather than assumed: had either guard been absent, the sweeper would have
# eaten this file on its first tick.
BENCHES_FILE = ".benches.json"
BENCH_LOCK = ".benches.lock"
# live → promoted (to Homepage) → retired. Order is meaningful: it is the
# first key of the rendered order, so a retired bench sinks.
BENCH_STATES = ("live", "promoted", "retired")
_STATE_RANK = {s: i for i, s in enumerate(BENCH_STATES)}
# Display budgets, not storage limits — these land in a panel row.
NAME_MAX, OWNER_MAX, URL_MAX = 120, 64, 2048
# The read is on the render path, so it is bounded. 256 KiB holds thousands of
# benches; the live board has 43 non-booth rows total.
BENCHES_MAX_BYTES = 256 * 1024
_SCHEMES = ("http", "https")
@dataclass(frozen=True)
class Bench:
"""One registered bench.
`id` and `url` are two fields ON PURPOSE. The identity must be normalized so
that re-posting updates rather than appends; the href must be verbatim so a
server that cares about a trailing slash, a case-sensitive path or a query
still works when the operator clicks it. Collapsing them would make the
registry quietly change where a link goes — a bug that surfaces as "the
bench 404s" and is never traced back here.
"""
id: str # the normalized URL — identity, and the key on disk
url: str # the URL as posted — what a click goes to
name: str
owner: str # an althing handle, or "booth" for the service
state: str
added: str # ISO-8601 with offset, from the FIRST registration
updated: str # ISO-8601 with offset, from the most recent upsert
error: str | None = None # a read-time verdict; never stored
def normalize_bench_url(url: str) -> str:
"""The identity of a bench. Raises ValueError with a reason a human can act on.
THE RULE, in full, because a vague identity is worse than a wrong one:
* surrounding whitespace stripped
* scheme lowercased; anything but http/https refused
* userinfo (`user:pass@host`) REFUSED, never stripped
* host lowercased; an empty host refused
* port dropped when it is the scheme default (80 http, 443 https)
* path kept verbatim, except that a bare "/" becomes ""
* query kept verbatim INCLUDING parameter order (a query is opaque)
* fragment dropped
WHY THE FULL URL AND NOT THE ORIGIN — measured, not chosen. Collapsing the
live board's 43 non-booth rows by origin yields 19 groups; by full URL, 35.
The difference is not duplication: it is eight distinct gitea repositories
merged into one row, three unrelated HuggingFace model cards merged into
one, and the two LRPG surfaces on `10.100.10.50:8321` merged into one —
which are the information-architecture doc's own example of two real
benches. Origin identity destroys more than it deduplicates. Full-URL
identity still collapses both cases that doc names: talk 5 → 1, Peedlar 3 → 1.
WHY THE QUERY IS IN AND THE FRAGMENT IS OUT. Three ShutterChute rows on the
board differ only by `?token=`; they are three genuinely different one-shot
links, and dropping the query would merge them into a bench that is none of
them. A fragment is a position inside a page, never a different resource.
"""
raw = (url or "").strip()
if not raw:
raise ValueError("a bench needs a URL")
if len(raw) > URL_MAX:
raise ValueError(f"URL is longer than {URL_MAX} characters")
try:
parts = urlsplit(raw)
except ValueError as exc: # malformed IPv6 literal, etc.
raise ValueError(f"could not parse that URL: {exc}") from exc
scheme = parts.scheme.lower()
if scheme not in _SCHEMES:
raise ValueError(
f"a bench must be http or https, not {parts.scheme or '(no scheme)'}"
)
if "@" in parts.netloc:
# Refused, NOT stripped. Stripping would register a bench whose URL no
# longer works while telling the poster it succeeded — and would put a
# credential on a board that renders on an unauthenticated LAN surface
# on the way there.
raise ValueError("a bench URL must not carry credentials; strip the user:pass@ and re-post")
try:
host = (parts.hostname or "").lower()
port = parts.port
except ValueError as exc: # a non-numeric port
raise ValueError(f"could not read the host or port: {exc}") from exc
if not host:
raise ValueError("that URL has no host")
# RE-WRAP A BRACKETED IPv6 LITERAL. `urlsplit().hostname` strips the
# brackets, and rebuilding the netloc from it produces `http://::1:8080/a`
# — not a different spelling of the same URL but a BROKEN one, so a re-post
# never matches the row the operator thinks they are updating. The bracket
# is part of the authority's syntax, not decoration. Detected by the colon,
# which cannot appear in a hostname or an IPv4 literal.
if ":" in host:
host = f"[{host}]"
default = {"http": 80, "https": 443}[scheme]
netloc = host if port in (None, default) else f"{host}:{port}"
# A bare "/" is the same resource as no path at all; a trailing slash on a
# REAL path is not, and is left alone.
path = "" if parts.path == "/" else parts.path
return urlunsplit((scheme, netloc, path, parts.query, ""))
# ---- storage ----------------------------------------------------------------
def _now() -> str:
return datetime.now(timezone.utc).isoformat(timespec="seconds")
def _cap(value: object, limit: int, field: str) -> str:
if not isinstance(value, str):
raise ValueError(f"{field} must be text, not {type(value).__name__}")
return value[:limit]
def _bench_from(bench_id: str, row: object) -> Bench:
"""One stored row to a record. Raises ValueError on any shape it cannot
trust — this is the STRICT half, used by the write path and by the read
path's single try/except."""
if not isinstance(row, dict):
raise ValueError(f"{bench_id}: expected an object, found {type(row).__name__}")
state = row.get("state", "live")
if state not in BENCH_STATES:
raise ValueError(f"{bench_id}: unknown state {state!r}")
url = row.get("url", bench_id)
if not isinstance(url, str):
raise ValueError(f"{bench_id}: url must be text, not {type(url).__name__}")
if len(url) > URL_MAX:
# REFUSED, NOT TRUNCATED — unlike `name` and `owner`. Those are display
# budgets and clipping one costs a few characters in a panel row. A
# clipped URL is a DEAD ANCHOR, and INV-7 promises the click goes to the
# posted address byte for byte; silently shortening it keeps the promise
# in the type system and breaks it in the browser. Nothing this code
# writes can get here (normalize refuses over-long input); a hand-edited
# registry can, and it is damage, which is what the reader reports.
raise ValueError(f"{bench_id}: url is longer than {URL_MAX} characters")
return Bench(
id=bench_id,
url=url,
name=_cap(row.get("name", ""), NAME_MAX, "name"),
owner=_cap(row.get("owner", ""), OWNER_MAX, "owner"),
state=state,
added=_cap(row.get("added", ""), 64, "added"),
updated=_cap(row.get("updated", ""), 64, "updated"),
)
def _read_bytes(path: Path) -> bytes:
"""Read at most BENCHES_MAX_BYTES + 1 bytes from a REGULAR FILE.
REGULAR-FILE FIRST, THEN SIZE, THEN A BOUNDED READ — in that order, and the
order is the whole point. A named pipe blocks in `open()`, before any byte
cap can apply: bounding the read does NOT close that hole, and an earlier
draft of this module claimed it did while hanging on the first FIFO put at
this path. `read_benches` is on the board page's render path, so that hang
is a request that never returns and, with enough of them, the threadpool
behind every route. `marks.py` learned this on 2026-09-22 and guards with
`S_ISREG`; this is the same guard, not a new idea.
The bounded read stays, for the case the stat cannot answer: a regular file
that GREW between the stat and the read.
"""
st = os.stat(path)
if not stat.S_ISREG(st.st_mode):
raise ValueError(f"{path.name} is not a regular file")
if st.st_size > BENCHES_MAX_BYTES:
raise ValueError(f"registry is larger than {BENCHES_MAX_BYTES} bytes")
with path.open("rb") as fh:
return fh.read(BENCHES_MAX_BYTES + 1)
def _load_strict(root: Path) -> dict[str, Bench]:
"""Every bench, or ValueError. The write path's reader.
Whole-file, not per-row: a registry with one unreadable row is a registry
somebody has to look at, and quietly dropping the row is how a bench
disappears without anyone being told.
"""
path = Path(root) / BENCHES_FILE
if not path.exists():
return {}
blob = _read_bytes(path)
if len(blob) > BENCHES_MAX_BYTES:
raise ValueError(f"registry is larger than {BENCHES_MAX_BYTES} bytes")
try:
raw = json.loads(blob.decode("utf-8"))
except (UnicodeDecodeError, json.JSONDecodeError) as exc:
raise ValueError(f"registry is not valid JSON: {exc}") from exc
if not isinstance(raw, dict):
raise ValueError(f"registry must be an object keyed by URL, found {type(raw).__name__}")
return {k: _bench_from(k, v) for k, v in raw.items()}
def read_benches(root: Path) -> tuple[list[Bench], str | None]:
"""Every registered bench in the rendered order, plus a read-time error.
NEVER RAISES. This runs on the render path, and the v0.2.2 lesson in this
repo was learned the expensive way: a poisoned `.marks.json` returned 500
for `/` and `/healthz` across all 25 booths. A registry that cannot be read
costs its own panel, never the page.
ABSENT AND DAMAGED ARE DIFFERENT and must render differently — only one of
them needs a human. Absent is `([], None)`; damaged is `([], "why")`.
"""
try:
return order_benches(_load_strict(root).values()), None
except ValueError as exc:
return [], str(exc)
except OSError as exc:
return [], f"registry could not be read: {exc}"
except RecursionError:
# Deeply nested JSON (`[[[[...`) blows the stack inside json.loads, and
# RecursionError is neither ValueError nor OSError — so it escaped the
# pair above and 500'd the page this function exists to protect. The
# byte cap does not help: 200k open brackets is 200 KB.
return [], "registry is nested too deeply to parse"
def _write_all(root: Path, benches: dict[str, Bench]) -> None:
"""Atomic replace. Caller holds the lock.
Temp file + os.replace, so a reader never sees a partial file and a crash
mid-write cannot truncate the registry into a shorter — and therefore
quieter — set of benches. CLAUDE.md invariant 5.
"""
root = Path(root)
path = root / BENCHES_FILE
payload = {
b.id: {"url": b.url, "name": b.name, "owner": b.owner,
"state": b.state, "added": b.added, "updated": b.updated}
# The key IS the id, so the record does not carry it twice — two copies
# of one fact is two things that can disagree.
for b in benches.values()
}
# Per-pid scratch name so two writers cannot share it: the atomic-replace
# promise is that a READER never sees a partial file, not that two writers
# never collide on the way there.
body = json.dumps(payload, indent=2, sort_keys=True) + "\n"
# THE WRITER RESPECTS THE READER'S CAP. Without this, a successful
# registration can push the file past BENCHES_MAX_BYTES and every
# subsequent read fails — so the LAST bench somebody added is the one that
# makes all the others invisible, and the write that did it reported
# success. The reader is lenient about damage; it is not lenient about
# size, and a writer that ignores a limit its own reader enforces is
# manufacturing exactly the state the leniency exists to survive.
if len(body.encode("utf-8")) > BENCHES_MAX_BYTES:
raise ValueError(
f"that registration would push the registry past {BENCHES_MAX_BYTES} "
f"bytes, which its own reader refuses; nothing was written")
# AN UNPREDICTABLE SCRATCH NAME, IN THE SAME DIRECTORY. `.tmp.<pid>` is
# guessable, and a pre-planted symlink there redirects the write straight
# through the atomic replace — the replace is atomic, not safe. mkstemp
# creates with O_EXCL and 0600, so it cannot land on someone else's file.
# Same directory because os.replace is only atomic within a filesystem.
fd, tmpname = tempfile.mkstemp(dir=str(root), prefix=".benches-", suffix=".tmp")
tmp = Path(tmpname)
try:
with os.fdopen(fd, "w", encoding="utf-8") as fh:
fh.write(body)
fh.flush()
# FSYNC BEFORE THE REPLACE. os.replace orders the rename, not the
# DATA behind it: without this, a power loss can publish a name
# pointing at bytes that never reached the disk, which is a
# truncated registry wearing a successful write's clothes.
os.fsync(fh.fileno())
os.chmod(tmp, 0o644) # mkstemp's 0600 is tighter than the rest
os.replace(tmp, path)
except BaseException:
# A write that dies between create and replace would otherwise strand
# the scratch file beside the registry forever. The prior registry is
# untouched either way — os.replace is the only thing that publishes.
tmp.unlink(missing_ok=True)
raise
class _Locked:
"""Exclusive flock over the whole read-modify-write, on a sidecar."""
def __init__(self, root: Path):
self.root = Path(root)
self.root.mkdir(parents=True, exist_ok=True)
self.path = self.root / BENCH_LOCK
def __enter__(self):
self.path.touch(exist_ok=True)
self.fh = self.path.open("r+")
fcntl.flock(self.fh, fcntl.LOCK_EX)
return self
def __exit__(self, *exc):
fcntl.flock(self.fh, fcntl.LOCK_UN)
self.fh.close()
return False
def upsert_bench(root: Path, url: str, name: str, owner: str) -> tuple[Bench, bool]:
"""Register or update by normalized URL. Returns (bench, created).
READS ARE LENIENT, WRITES ARE STRICT — and this is the strict side. A write
over a registry that cannot be parsed RAISES rather than starting a fresh
one: on 2026-09-21 this repo learned that a tolerant writer over a damaged
`.marks.json` wipes the operator's judgment, and a tolerant reader is a
completely different decision from a tolerant writer.
`added` survives an update; `state` survives too, so a promoted bench that
re-announces itself after a deploy is not silently demoted.
"""
bench_id = normalize_bench_url(url)
with _Locked(root):
benches = _load_strict(root) # raises on damaged — deliberate
prior = benches.get(bench_id)
now = _now()
bench = Bench(
id=bench_id,
url=(url or "").strip(),
name=_cap(name or "", NAME_MAX, "name"),
owner=_cap(owner or "", OWNER_MAX, "owner"),
state=prior.state if prior else "live",
added=prior.added if prior else now,
updated=now,
)
benches[bench_id] = bench
_write_all(root, benches)
return bench, prior is None
def set_bench_state(root: Path, bench_id: str, state: str) -> Bench | None:
"""Move a bench between live / promoted / retired. None if no such bench."""
if state not in BENCH_STATES:
raise ValueError(f"state must be one of {', '.join(BENCH_STATES)}, not {state!r}")
with _Locked(root):
benches = _load_strict(root)
prior = benches.get(bench_id)
if prior is None:
return None
moved = replace(prior, state=state, updated=_now())
benches[bench_id] = moved
_write_all(root, benches)
return moved
def remove_bench(root: Path, bench_id: str) -> Bench | None:
"""Drop one bench. Returns the removed record, or None."""
with _Locked(root):
benches = _load_strict(root)
gone = benches.pop(bench_id, None)
if gone is None:
return None
_write_all(root, benches)
return gone
def order_benches(benches: Iterable[Bench]) -> list[Bench]:
"""ORDER: (state rank, name casefolded, id).
live before promoted before retired, then alphabetical, with the id as a
TOTAL tie-break so two benches sharing a name cannot swap between renders.
CLAUDE.md invariant 6 — the Booth's job is comparison, and an order that
moves between page loads files the operator's judgment against the wrong
row. Pure: no I/O, and the input sequence is not mutated.
"""
return sorted(benches, key=lambda b: (_STATE_RANK.get(b.state, len(BENCH_STATES)),
b.name.casefold(), b.id))