fix(manifest): fold in both cross-frontier panels — and a live hole in v0.2.2
Two four-arm artifact-only rounds landed together: the contract paraphrase (against the pre-seam-review capture) and the code-vs-contract conformance review (against the amended one), correctly firewalled from each other. The conformance round found ZERO drift in the strict sense — the code is a clause-for-clause implementation of the contract — and the weight of both rounds landed one layer down, in what green tests structurally cannot report. Full triage in persistent-memory.d/. A LIVE HOLE IN RELEASED CODE, FOUND ON THE SIBLING MODULE v0.2.2 adopted the RecursionError finding from the bug-hunt round and closed half of it: `_hydrate_safe` guards hydration, but `json.loads` runs above it in `_read_raw`, whose catch list covers neither RecursionError nor MemoryError. A 400 KB file of nothing but brackets in any ONE booth therefore still returned 500 for `/` and `/healthz` across every booth on the service. Confirmed by running it before believing it. Both modules now bound the read by `stat` before touching the bytes and catch both classes anyway, so raising a bound later cannot quietly re-open the hole. The strict half of the marks asymmetry refuses everything the lenient half tolerates, or a file that reads as "no marks" gets replaced by a write that believed it. THE WHY-WIPE `booth new x --why "..."` then `booth add x out/*.png` erased the sentence the first command existed to record. Omitted flags meant empty strings and empty strings overwrote. Two arms predicted it from the contract's wording alone; every test here passed --why on both calls and so could not see it. Omitted now means unchanged and an explicit --why "" still clears — the shell carries the distinction by leaving the variable UNSET, not empty. --title WAS WRITE-ONLY Stored, flag-surfaced, rendered nowhere. 4/4, and independently top-ranked by every arm of the paraphrase round. It lands on the booth page heading with the directory name beside it, because the directory name is the identity the operator navigates by and refers to positionally. THREE TESTS THAT COULD NOT FAIL - test_the_write_is_atomic asserted no *.tmp survived, which a plain write_text passes. It asserts the inode changes now. (The first replacement was ALSO vacuous — it spied on os.open, which Path.write_text reaches through io.open in C and never touches. Recorded in the test, because writing a second vacuous test while fixing the first is exactly the failure this round is about.) - The INV-3 preservation test passed against an implementation that regenerated `created` every time, because _now() is whole-second resolution and back-to-back writes share a stamp. Seeded from 2019 now. - test_announcing_is_activity passed whether or not _newest_mtime counted the manifest, because writing it bumps the directory mtime either way. The directory's clock is put back, leaving the file as the only thing that can keep the booth alive. ALSO - The title fallback skipped the normalizer the explicit value gets; a directory name may legally carry a newline and run to 255 bytes. - Every writer derived the same .booth.json.tmp. Marks are protected from that by their flock; the manifest has none, so uniqueness stands in. - test_stdlib_only was blind to relative imports in all four modules. - INV-1 had no guard at all; INV-5 named two different promises; the negative render states were asserted on the index only. Contract amended throughout: the 4 GB case is a stat-checked bound rather than a return constraint, every field of an error-carrying record has a stated value, INV-1 no longer contradicts INV-3, repo-wide rules are named in words instead of by a colliding number, and touches admits the macro partial the implementation added. 329 tests.
This commit is contained in:
+31
-7
@@ -73,6 +73,14 @@ class MarksCorrupt(RuntimeError):
|
||||
"""
|
||||
|
||||
|
||||
# A booth's whole judgment lives in one document, so this is generous — a
|
||||
# 270-item booth flagged throughout, with notes, is far under it. What it rules
|
||||
# out is the case that is not marks at all: an unbounded read raises MemoryError
|
||||
# and a deeply nested one raises RecursionError out of `json.loads`, neither of
|
||||
# which is an OSError or a ValueError, and `list_booths` calls the reader once
|
||||
# per booth on every index load. Bounded by `stat`, before the bytes are read.
|
||||
MARKS_MAX_BYTES = 4 * 1024 * 1024
|
||||
|
||||
MARKS_FILE = ".marks.json"
|
||||
MARKS_LOCK = ".marks.lock"
|
||||
SCHEMA_VERSION = 1
|
||||
@@ -173,9 +181,12 @@ def _read_raw(booth: Path) -> list[dict]:
|
||||
for the same reason: a review surface that will not load is worse than one
|
||||
that has lost an annotation.
|
||||
"""
|
||||
path = Path(booth) / MARKS_FILE
|
||||
try:
|
||||
raw = json.loads((Path(booth) / MARKS_FILE).read_text(encoding="utf-8"))
|
||||
except (OSError, ValueError, UnicodeDecodeError):
|
||||
if path.stat().st_size > MARKS_MAX_BYTES:
|
||||
return []
|
||||
raw = json.loads(path.read_text(encoding="utf-8"))
|
||||
except (OSError, ValueError, UnicodeDecodeError, RecursionError, MemoryError):
|
||||
return []
|
||||
if not isinstance(raw, dict):
|
||||
return []
|
||||
@@ -198,18 +209,29 @@ def _read_raw_strict(booth: Path) -> list[dict]:
|
||||
case where writing would destroy something.
|
||||
"""
|
||||
path = Path(booth) / MARKS_FILE
|
||||
try:
|
||||
size = path.stat().st_size
|
||||
except FileNotFoundError:
|
||||
return []
|
||||
except OSError as exc:
|
||||
raise MarksCorrupt(f"{path} cannot be read: {exc}") from exc
|
||||
# The strict half has to refuse everything the lenient half tolerates, or a
|
||||
# file that reads as "no marks" gets replaced by a write that believed it.
|
||||
if size > MARKS_MAX_BYTES:
|
||||
raise MarksCorrupt(f"{path} is too large to be a marks document ({size} bytes)")
|
||||
try:
|
||||
text = path.read_text(encoding="utf-8")
|
||||
except FileNotFoundError:
|
||||
return []
|
||||
except (OSError, UnicodeDecodeError) as exc:
|
||||
except (OSError, UnicodeDecodeError, MemoryError) as exc:
|
||||
raise MarksCorrupt(f"{path} cannot be read: {exc}") from exc
|
||||
if not text.strip():
|
||||
return []
|
||||
try:
|
||||
raw = json.loads(text)
|
||||
except ValueError as exc:
|
||||
raise MarksCorrupt(f"{path} is not valid JSON: {exc}") from exc
|
||||
except (ValueError, RecursionError, MemoryError) as exc:
|
||||
raise MarksCorrupt(
|
||||
f"{path} is not valid JSON: {type(exc).__name__}") from exc
|
||||
if not isinstance(raw, dict) or not isinstance(raw.get("marks"), list):
|
||||
raise MarksCorrupt(f"{path} is not a marks document")
|
||||
entries = [e for e in raw["marks"] if isinstance(e, dict) and isinstance(e.get("id"), str)]
|
||||
@@ -671,7 +693,8 @@ def import_legacy_asks(booth: Path) -> list[Mark]:
|
||||
continue
|
||||
try:
|
||||
decl = json.loads(p.read_text(encoding="utf-8"))
|
||||
except (OSError, ValueError, UnicodeDecodeError) as exc:
|
||||
except (OSError, ValueError, UnicodeDecodeError,
|
||||
RecursionError, MemoryError) as exc:
|
||||
found.append((mtime, stem, None, f"unreadable ask: {exc}"))
|
||||
continue
|
||||
if not isinstance(decl, dict):
|
||||
@@ -693,7 +716,8 @@ def import_legacy_asks(booth: Path) -> list[Mark]:
|
||||
loaded = json.loads(ap.read_text(encoding="utf-8"))
|
||||
if isinstance(loaded, dict):
|
||||
answer = loaded
|
||||
except (OSError, ValueError, UnicodeDecodeError):
|
||||
except (OSError, ValueError, UnicodeDecodeError,
|
||||
RecursionError, MemoryError):
|
||||
pass
|
||||
|
||||
prior = by_id.get(stem)
|
||||
|
||||
Reference in New Issue
Block a user