Commit Graph
4 Commits
Author SHA1 Message Date
vh 213071b6ce fix(embed,mutation): SPYRJA fold — report input on every path; an instrument that cannot certify what it did not run
The heid bug-hunt (hulda, with heid's second voice) on 377e652 found eight
issues. Every fixed one has a red-first test.

embed.js:
- H1: the report-input guard covered only the batch path. A lone changed
  answer went out as a native POST, whose 303 navigation took the report's
  typed text with it. Unsaved report input now routes even a lone answer
  in place. With nothing of ours to send, the submit block says so.
- H7 / V1: a bare <select>, and a range or color input with no value
  attribute, read as typed-into by their default attributes, so every clean
  batch refused its reload with a false message. Report controls are now
  measured against how they stood when the Booth mounted. A control added
  later falls back to its defaults, counting a select's first option as its
  default.
- H2 / V2: a contenteditable region counts as report input.

scripts/mutation_check.py:
- H5: any non-zero exit counted as proof, including a collection error
  where the test never ran. Only pytest's "tests failed" (1) proves now.
- H6: the test run has a timeout (300 s). A hang reports "timed out" and
  the source is still restored.
- H3: source is read and restored as bytes, so a CRLF file comes back
  byte-exact.
- H4: one run per tree, enforced by a lock. The in-flight marker lives with
  the tree it guards.
- H8: anchors are counted with overlaps. The check is `matches()`, not
  str.count.

Tool controls +5 (tests/test_mutation_check.py). u3_submit_all +3 rows.
2026-09-28 17:49:13 -07:00
vh 377e652670 fix(inplace,embed): input set back mid-flight, report inputs, ambiguous anchors
Four items owed after S5b, reported by design-dev during the anti-slop run:

- carry() measured a sent-then-changed form against its OLD DEFAULTS. An
  answer set back mid-flight to the value the page first showed read as
  untouched, and the swap put the just-saved value over it. A form sent and
  then changed is now measured against its sent snapshot (sentSet.snapOf).
- The embed's clean-batch reload saw only our own forms. A report's own
  inputs lost whatever the operator had typed into them. Unsaved text in
  any control we don't own now holds the reload, and the page says so.
- Two r2b.toml rows ("D3 a stored theme...", "D3 forced light...") matched
  twice, so they proved only by where the first match fell. Both are
  re-anchored, and scripts/mutation_check.py now refuses any anchor that
  matches more than once. A new tool control covers that.
- The r2_flow contract's C3 steps 2 and 4 now say what S5b superseded. U3
  gains the report-input rule.

Mutation rows: u3_submit_all +1, r2_submit_all +1. Four rows were
re-anchored onto the moved lines.
2026-09-28 16:52:42 -07:00
vh 33e7149e24 fix(scripts): the mutation harness must not churn source mtimes
It rewrites a tracked file and restores it byte-for-byte — but the restore
bumped the mtime, and in this repo that is not cosmetic. The repo IS the
deployment root and nothing takes effect until the service restarts, so 'is
:8090 stale?' is answered by comparing the service's start time against source
mtimes. A tool that moves those without changing a byte makes that check lie:
it reported the live service 16 minutes stale while it was serving current code.

Restores atime/mtime with os.utime, with a test whose defeating change is
dropping that line. Found by using the staleness check for real, not by review.

649 green; 12/12 U7 falsifiers still proved.
2026-09-22 22:01:35 -07:00
vh 2f6a0ee821 test: keep the mutation harness — scripts/mutation_check.py, with its own controls
Promotes the session-scratchpad harness that proved U7's twelve falsifiers into
a repo tool, on the operator's call. No version bump: test tooling and docs, no
production-code change, per the SemVer SKIP list.

A green test is not evidence. A test that has never seen its own defeating
change may pass under it too, forbidding nothing while reading as though it
forbids something. This repo shipped that three times — twice in one session,
and once an hour after writing the persistent-memory entry about it. Prose in a
memory file is not an instrument.

Tables live in tests/mutations/*.toml, one per unit, committed so a unit's
proofs are an artifact rather than terminal scrollback. Adding a unit means
adding a file, never editing the script. u7_navigation.toml was generated from
the harness that proved those twelve, not retyped, and every anchor was verified
against the source before it landed.

THE TOOL GETS ITS OWN POSITIVE AND NEGATIVE CONTROLS, which is the point. It
shipped two defects in one session, each of which made it report a falsifier
PROVED WITHOUT RUNNING IT, and both were found by accident rather than by
anything checking:

  no green baseline — a test that is ALREADY red reports red for every mutation
  thrown at it, so a broken assertion reads as a certified falsifier

  the bytecode cache — `< 2` -> `< 1` is byte-identical in size, and CPython
  validates a .pyc against the source's (mtime, size) at one-second granularity,
  so a mutation landing in the same second as the revert before it runs against
  cached bytecode; the tell was a verdict flipping between consecutive identical
  runs

tests/test_mutation_check.py now carries a control for each, plus the one
usually skipped: a KNOWN-VACUOUS falsifier the tool must catch. An instrument
that only ever sees unknowns cannot tell "nothing wrong here" from "I am blind",
and twelve `proved` lines from a blind instrument are worth nothing.

Also hardens the tool against itself: it writes to tracked source files, so the
restore is verified rather than assumed, and a .mutation-inflight marker makes a
run killed mid-mutation refuse the next start instead of silently measuring a
mutated tree.

648 tests green; 12/12 U7 falsifiers still proved.
2026-09-22 21:58:12 -07:00