fix(u3): seven defects two cold panels found in the declared seam
The /heid-code-review and /heid-bug-hunt panels, artifact-only over the U3
diff, between them found four real defects and three vacuous falsifiers. Both
snapshots predate the contract-review fixes, so two of their findings were
already closed; the rest are here.
Prototype pollution in the placement maps. A mark id and a question key are
both [A-Za-z0-9][A-Za-z0-9._-]*, so `toString` and `constructor` are legal in
each. Against a plain `{}` an anchor naming NO mark returned an inherited
function, passed the guard meant to reject it, and threw on .questions.length
-- aborting placement before the tail, so one typo in author markup cost the
page every ask. The `placed` set had the mirror bug: inherited
`got.constructor` read as already-placed and silently dropped a question.
Object.create(null), three times. Found independently by both panels.
A declaring page was not served as written. read_text() opens in
universal-newline mode, so a CRLF report came back LF, and errors="replace"
replaced every byte that was not valid UTF-8. That is this unit's headline
promise, broken by the read itself, and the test could not see it because its
fixture was LF-only ASCII. The verbatim branch reads and serves bytes now; the
decoded copy answers only "does it declare the seam?".
A submit anchor inside the author's own <form> lost ours -- the parser drops a
nested form element outright -- while the code still recorded the pick as
submitted, so no fallback was appended. Every control's form= pointed at
nothing and the button did nothing. It counts as submitted only if the form
survived.
A broken pick's diagnostic never rendered from a submit-only anchor: an errored
pick's submit block is empty, and mounting that then marking it placed made the
tail skip the "broken ask" box entirely. The anchor is left alone instead.
An author's own element could hijack the open-ask chip -- id="bk-ask-winner-
background" satisfies any prefix rule, hyphen boundary included. The chip now
searches only elements this script mounted, which is the identity the deleted
bk-ask-<id>-top anchor used to guarantee, and takes the earliest by
compareDocumentPosition.
No error boundary around fragment rendering. A .marks.json that is well-formed
JSON with a wrong-shaped answer hydrates with no error and then raises in the
macro; this endpoint renders every pick on every load of the report, so that
was the whole seam gone while hold_read called the file readable. Reproduced
before building for it. _safe_fragments gives it the per-mark leniency
_hydrate_safe already applies one layer down.
The gallery and marks pages still 500 on that same entry. Measured at 42ea67f
-- it predates this unit, they render the same macro with no guard, and the
gallery is named out of scope in the contract. Recorded, not quietly widened:
persistent-memory.d/2026-09-22-a-wrong-shaped-answer-500s-the-gallery.md
Also corrected: several comments claimed a multi-question pick POSTs a 400
unless every question is answered. It does not -- an empty submission is
refused, a partial one is recorded on purpose. The real reason an unplaced
question must still be appended is that a question which never reaches the page
cannot be answered at all.
Vacuity pass rebuilt around the rule this session learned: the mutation comes
from the invariant's claim, never from the falsifier's example. 21 mutations,
21 caught, unmutated control green. Getting there took three rounds -- it
passed INV-3 with the contract's own mutation, then found its own fix's hole,
then flagged seven stale mutations and one genuinely vacuous fixture whose
sibling-mark arrangement made the right answer also the first answer.
444 tests. Deployed and verified: 23/23 booths 200, and all four live verbatim
reports served at exactly +46 bytes -- len(EMBED_SCRIPT_TAG) -- with the
authors' own wrappers and headings intact and no console errors.
This commit is contained in:
+52
-7
@@ -44,6 +44,7 @@ import shutil
|
||||
import time
|
||||
import zipfile
|
||||
from contextlib import asynccontextmanager
|
||||
from dataclasses import replace
|
||||
from pathlib import Path
|
||||
from typing import Sequence
|
||||
from urllib.parse import quote, unquote
|
||||
@@ -589,8 +590,8 @@ def declares_embed(html: str) -> bool:
|
||||
return any(d in html for d in _EMBED_DECLARATIONS)
|
||||
|
||||
|
||||
def embed_verbatim(html: str) -> str:
|
||||
"""The ONLY thing the Booth does to a verbatim report.
|
||||
def embed_verbatim(raw: bytes) -> bytes:
|
||||
"""The ONLY thing the Booth does to a verbatim report. BYTES IN, BYTES OUT.
|
||||
|
||||
Appended, never inserted, and never prepended. That is what retires both of
|
||||
the old wrapper's hard constraints rather than satisfying them more
|
||||
@@ -598,8 +599,22 @@ def embed_verbatim(html: str) -> str:
|
||||
nothing can push the charset <meta> out of its first-1024-byte detection
|
||||
window, because nothing in front of them moves. Content after `</html>` is
|
||||
parsed into the body by every browser, so there is no seam to find.
|
||||
|
||||
⚠ IT TAKES BYTES BECAUSE TEXT WAS QUIETLY EDITING THE DOCUMENT. The first
|
||||
version read the file with `read_text()` and returned a str. That opens in
|
||||
UNIVERSAL-NEWLINE mode, so a report written with CRLF came back with LF —
|
||||
and `errors="replace"` turned any byte that was not valid UTF-8 into U+FFFD.
|
||||
A declaring page was therefore NOT served as its author wrote it, which is
|
||||
this unit's headline promise, and the test could not see it because its
|
||||
fixture was LF-only ASCII. Found by a cross-frontier bug-hunt panel.
|
||||
|
||||
Decoding still happens — `declares_embed` needs a string to look in — but
|
||||
the decoded copy is used ONLY to answer that question. What goes on the wire
|
||||
is the original bytes, plus the tag's bytes when it is appended, so the
|
||||
source is a byte-exact prefix of the response.
|
||||
"""
|
||||
return html if declares_embed(html) else html + EMBED_SCRIPT_TAG
|
||||
text = raw.decode("utf-8", errors="replace")
|
||||
return raw if declares_embed(text) else raw + EMBED_SCRIPT_TAG.encode("utf-8")
|
||||
|
||||
|
||||
def ask_form_id(stem: str) -> str:
|
||||
@@ -860,8 +875,11 @@ def create_app(
|
||||
# serving raw, which costs it the chrome exactly as it did before.
|
||||
try:
|
||||
if own_index.stat().st_size <= WRAP_MAX_BYTES:
|
||||
raw = own_index.read_text(encoding="utf-8", errors="replace")
|
||||
return HTMLResponse(embed_verbatim(raw))
|
||||
# ONE read, and it is a byte read: see embed_verbatim.
|
||||
return Response(
|
||||
content=embed_verbatim(own_index.read_bytes()),
|
||||
media_type="text/html; charset=utf-8",
|
||||
)
|
||||
except OSError:
|
||||
pass
|
||||
return FileResponse(str(own_index), media_type="text/html")
|
||||
@@ -1104,6 +1122,30 @@ def create_app(
|
||||
],
|
||||
}
|
||||
|
||||
def _safe_fragments(name: str, mark) -> dict:
|
||||
"""`_pick_fragments`, with the promise that it cannot raise.
|
||||
|
||||
`marks_for` hydrates an entry whose JSON is well-formed but whose SHAPE
|
||||
is wrong — `{"answer": {"answers": []}}` survives `_hydrate` with no
|
||||
error and then raises `UndefinedError` in the template, because the
|
||||
macro asks a list for `.get`. Verified, not assumed.
|
||||
|
||||
This endpoint renders every pick in the booth on every page load of the
|
||||
operator's report, so one such entry would 500 the whole seam and the
|
||||
report would show no chrome at all — while `hold_read` reported the file
|
||||
as perfectly readable. Same leniency `_hydrate_safe` already applies one
|
||||
layer down, at the layer that actually renders: one unreadable pick
|
||||
costs that pick, never the page.
|
||||
"""
|
||||
try:
|
||||
return _pick_fragments(name, mark)
|
||||
except Exception as exc: # noqa: BLE001 - deliberate
|
||||
broken = replace(mark, error=f"this question could not be rendered: {exc}")
|
||||
return {"id": mark.id, "error": broken.error,
|
||||
"whole": str(_frag.whole(broken, ask_form_id(mark.id),
|
||||
quote(name, safe=""))),
|
||||
"submit": "", "questions": []}
|
||||
|
||||
@app.get("/b/{name}/embed.json")
|
||||
def booth_embed_json(name: str):
|
||||
"""Everything a verbatim report needs to mount the Booth's chrome.
|
||||
@@ -1133,12 +1175,15 @@ def create_app(
|
||||
picks = [m for m in marks if m.shape == "pick"]
|
||||
body = {
|
||||
"booth": name,
|
||||
"home": "/",
|
||||
# No `home`: the way-home chip mounts from a constant BEFORE this
|
||||
# fetch, so that a failed one still leaves the operator a way out.
|
||||
# Carrying the value anyway would put a second representation of it
|
||||
# on the wire for nothing to read.
|
||||
"favicon": FAVICON_HREF,
|
||||
# Picks only. It is also what keeps a flag's `flag:<target>` id —
|
||||
# the one mark id containing the separator an anchor spec splits
|
||||
# on — out of a payload whose specs split on the first colon.
|
||||
"marks": [_pick_fragments(name, m) for m in picks],
|
||||
"marks": [_safe_fragments(name, m) for m in picks],
|
||||
"open": [m.id for m in open_marks(picks)],
|
||||
}
|
||||
if read_err is not None:
|
||||
|
||||
Reference in New Issue
Block a user