Files
booth/scripts/mutation_check.py
T
vh 213071b6ce fix(embed,mutation): SPYRJA fold — report input on every path; an instrument that cannot certify what it did not run
The heid bug-hunt (hulda, with heid's second voice) on 377e652 found eight
issues. Every fixed one has a red-first test.

embed.js:
- H1: the report-input guard covered only the batch path. A lone changed
  answer went out as a native POST, whose 303 navigation took the report's
  typed text with it. Unsaved report input now routes even a lone answer
  in place. With nothing of ours to send, the submit block says so.
- H7 / V1: a bare <select>, and a range or color input with no value
  attribute, read as typed-into by their default attributes, so every clean
  batch refused its reload with a false message. Report controls are now
  measured against how they stood when the Booth mounted. A control added
  later falls back to its defaults, counting a select's first option as its
  default.
- H2 / V2: a contenteditable region counts as report input.

scripts/mutation_check.py:
- H5: any non-zero exit counted as proof, including a collection error
  where the test never ran. Only pytest's "tests failed" (1) proves now.
- H6: the test run has a timeout (300 s). A hang reports "timed out" and
  the source is still restored.
- H3: source is read and restored as bytes, so a CRLF file comes back
  byte-exact.
- H4: one run per tree, enforced by a lock. The in-flight marker lives with
  the tree it guards.
- H8: anchors are counted with overlaps. The check is `matches()`, not
  str.count.

Tool controls +5 (tests/test_mutation_check.py). u3_submit_all +3 rows.
2026-09-28 17:49:13 -07:00

195 lines
8.7 KiB
Python
Executable File

#!/usr/bin/env python3
"""Prove a falsifier falsifies, by running the change it forbids.
A green test is not evidence. A test that has never seen its own DEFEATING
CHANGE is only evidence that the code and the assertion agree today; it may
agree under the mutation too, in which case it forbids nothing and reads as
though it forbids something. This repo has shipped that three times --
persistent-memory.d/2026-09-22-vacuous-falsifiers.md,
persistent-memory.d/2026-09-22-seven-of-seven-falsifiers.md, and once more in
U7 an hour after the second was written.
So: for each declared mutation, apply it to the source, run the one test that
claims to catch it, and require RED. Revert either way.
.venv/bin/python scripts/mutation_check.py # every table
.venv/bin/python scripts/mutation_check.py u7_navigation # one table
Tables live in tests/mutations/*.toml and are committed, so a unit's proofs are
an artifact rather than terminal scrollback. Adding a unit means adding a file,
never editing this script.
⚠ TWO DEFECTS THIS TOOL HAD, both of which made it CERTIFY A FALSIFIER WITHOUT
RUNNING IT. Neither is obvious and both cost real time:
1. NO GREEN BASELINE. A test that is ALREADY red reports red for every mutation
thrown at it, so a broken assertion reads as a proven falsifier. Every run
now checks the test passes unmutated first; a red baseline is a harness
failure, reported as such, never as a proof.
2. THE BYTECODE CACHE. `< 2` -> `< 1` is BYTE-IDENTICAL IN SIZE, and CPython
validates a .pyc against the source's (mtime, size) at ONE-SECOND
granularity -- so a mutation landing in the same second as the revert before
it is invisible and the unmutated code runs. The tell was a verdict that
flipped between consecutive runs with nothing changed. Caches are dropped
and PYTHONDONTWRITEBYTECODE is set for every run. This biases toward exactly
the mutations most worth making: comparison flips, off-by-one constants,
and/or swaps.
"""
from __future__ import annotations
import fcntl
import os
import shutil
import subprocess
import sys
import tomllib
from pathlib import Path
REPO = Path(__file__).resolve().parent.parent
TABLES = REPO / "tests" / "mutations"
# Written before a source file is touched and removed after it is restored. Its
# presence at startup means a previous run died between the two -- a `kill -9`
# mid-mutation leaves a mutated tracked file that looks like authored code.
INFLIGHT = REPO / ".mutation-inflight"
#: pytest's exit code for "the tests ran and at least one FAILED" — the only
#: result that proves a falsifier. Every other non-zero (2 interrupted or a
#: collection error, 3 internal, 4 usage, 5 nothing collected) means the test
#: never got to judge the mutation (SPYRJA H5, 2026-09-28).
TESTS_FAILED = 1
#: Seconds before a test run is killed. A mutation that loops forever held
#: the checker, and the mutated file, until someone noticed (SPYRJA H6).
DEFAULT_TIMEOUT = 300
def run(test: str, repo: Path = REPO, timeout: float = DEFAULT_TIMEOUT) -> int:
"""Exit code of one test, with the bytecode cache defeated. See defect 2.
Raises subprocess.TimeoutExpired past `timeout`, with the child killed."""
for cache in repo.rglob("__pycache__"):
shutil.rmtree(cache, ignore_errors=True)
return subprocess.run(
[sys.executable, "-m", "pytest", test, "-q", "--no-header", "-p", "no:warnings"],
cwd=repo, capture_output=True, text=True, timeout=timeout,
env=dict(os.environ, PYTHONDONTWRITEBYTECODE="1"),
).returncode
def matches(src: bytes, old: bytes) -> int:
"""How many places `old` STARTS in `src`, overlapping ones included.
`bytes.count` is non-overlapping — `b"aaa".count(b"aa") == 1` — and an
anchor that starts twice is as ambiguous as one that appears twice."""
n, i = 0, src.find(old)
while i != -1:
n, i = n + 1, src.find(old, i + 1)
return n
def check(mutation: dict, repo: Path = REPO,
timeout: float = DEFAULT_TIMEOUT) -> tuple[bool, str]:
"""(proved, note) for one mutation. Never leaves the source mutated.
`repo` is a parameter so the harness can be pointed at a throwaway tree and
given KNOWN-vacuous and KNOWN-good falsifiers — see
tests/test_mutation_check.py. An instrument that only ever sees unknowns
cannot tell "nothing wrong here" from "I am blind", which is the whole of
CLAUDE.md's positive-control rule applied to the tool that enforces it."""
test = mutation["test"]
path = repo / mutation["file"]
try:
base = run(test, repo, timeout)
except subprocess.TimeoutExpired:
return False, f"BASELINE TIMED OUT — {test} ran past {timeout}s unmutated"
if base != 0:
return False, f"BASELINE RED — {test} fails BEFORE the mutation"
# BYTES, in and out: text-mode I/O read CRLF as LF, wrote LF back, then
# compared LF with LF and called it restored (SPYRJA H3).
src = path.read_bytes()
old, new = mutation["old"].encode("utf-8"), mutation["new"].encode("utf-8")
hits = matches(src, old)
if hits == 0:
return False, f"anchor not found in {mutation['file']} — the table has drifted"
if hits > 1:
# The replace below takes the FIRST match, so an anchor that matches
# twice proves by where that match happens to fall, not by the line the
# row names. Two r2b rows proved that way until 2026-09-28.
return False, f"anchor is ambiguous — it matches {hits} places in {mutation['file']}"
inflight = repo / INFLIGHT.name # the marker lives with the tree it guards
inflight.write_text(f"{path}\n")
stat = path.stat() # mtime included; see the restore below
try:
path.write_bytes(src.replace(old, new, 1))
try:
rc = run(test, repo, timeout)
except subprocess.TimeoutExpired:
return False, f"TIMED OUT — the mutated run passed {timeout}s; not a proof"
finally:
path.write_bytes(src)
# Verified, not assumed: a restore that silently failed would leave a
# mutation in a tracked file and the next run would measure it.
assert path.read_bytes() == src, f"RESTORE FAILED for {path} — fix by hand"
# ⚠ AND THE MTIME, which matters more here than it would elsewhere.
# This repo IS its own deployment root and nothing takes effect until
# the service restarts, so "is :8090 stale?" is answered by comparing
# the service's start time against source mtimes. A tool that churns
# those mtimes without changing a byte makes that check lie — it
# reported the live service 16 minutes stale when it was current.
os.utime(path, ns=(stat.st_atime_ns, stat.st_mtime_ns))
inflight.unlink(missing_ok=True)
if rc == 0:
return False, "VACUOUS — stayed green under the change it forbids"
if rc != TESTS_FAILED:
return False, (f"the test did not run under the mutation (pytest exit {rc}: "
"a collection, usage or internal error) — not a proof")
return True, ""
def main(argv: list[str]) -> int:
# ONE RUN PER TREE. Two checkers interleaving can each restore the file
# they read, and the second "restore" writes the first one's mutation back
# in (SPYRJA H4). Held for the whole run; released when the process exits.
lock = open(REPO / ".mutation-lock", "a")
try:
fcntl.flock(lock, fcntl.LOCK_EX | fcntl.LOCK_NB)
except BlockingIOError:
print(f"refusing to run: another mutation_check holds {REPO / '.mutation-lock'}.")
return 2
if INFLIGHT.exists():
print(f"refusing to run: {INFLIGHT} exists, so a previous run died mid-mutation.")
print(f"check `git diff {INFLIGHT.read_text().strip()}`, restore it, then delete the marker.")
return 2
wanted = argv[1:] or None
tables = sorted(TABLES.glob("*.toml"))
if wanted:
tables = [t for t in tables if t.stem in wanted]
if not tables:
print(f"no table matching {wanted} in {TABLES}")
return 2
failed = []
for table in tables:
doc = tomllib.loads(table.read_text())
print(f"\n### {table.stem} — {doc.get('unit', '')}")
for m in doc.get("mutation", []):
proved, note = check(m)
print(f"{' proved' if proved else ' NOT PROVED':14s} {m['label']}")
if not proved:
print(f"{'':14s} ^ {note}")
failed.append(m["label"])
total = sum(len(tomllib.loads(t.read_text()).get("mutation", [])) for t in tables)
print(f"\n{total - len(failed)}/{total} falsifiers proved by running the change they forbid")
return 1 if failed else 0
if __name__ == "__main__":
raise SystemExit(main(sys.argv))