docs(training-playbook): 4.6.2 — a null result needs a positive control
From the run-2 gate. A memorisation probe reporting 0.00% across all 72 items is the correct output for a model that has not seen the corpus, and is also the exact output of a probe that is not firing. Nothing in the number distinguishes them. brokkr-smithy-dev drove the overlap function with known-answer inputs (identical 100%, half-verbatim 65.38%, unrelated 0%, empty 0%) before trusting the null, which is what converts a suspicious zero into evidence. This is 4.5's inert gate wearing a different face: there a check that could not return 'fail', here a measurement that cannot return non-zero. A clean null is the most reassuring output an instrument produces and the least self-evidencing. Same section records the identical-on-both-arms variant: the diversity battery's rp family froze zero markers, so its attractor hit rate read 0.0 on base AND tuned. That reads as a clean result and means the instrument cannot discriminate on that family. Report as a bounded limitation, never as a delta of zero -- a check returning the same value for every input is not measuring. Checklist gains the line.
This commit is contained in:
@@ -861,6 +861,38 @@ that, because the instrument was dark.
|
|||||||
|
|
||||||
If breaking it changes nothing, you have documentation, not a test.
|
If breaking it changes nothing, you have documentation, not a test.
|
||||||
|
|
||||||
|
#### ⚠ 4.6.2 ⭐⭐ A NULL RESULT needs a positive control before it counts as a null
|
||||||
|
|
||||||
|
`0.00% / floor 0.00% / max_item 0.0%` across all 72 items is the correct output
|
||||||
|
of a memorisation probe on a model that has never seen the corpus. **It is also
|
||||||
|
the exact output of a probe that is not firing at all**, and nothing in the
|
||||||
|
number distinguishes them.
|
||||||
|
|
||||||
|
The move that separates them costs one minute — drive the metric's own function
|
||||||
|
with inputs whose answer you already know:
|
||||||
|
|
||||||
|
identical text 100.00%
|
||||||
|
half-verbatim 65.38%
|
||||||
|
unrelated English 0.00%
|
||||||
|
empty string 0.00%
|
||||||
|
|
||||||
|
Now the zero means something: the instrument *can* go red, and did not.
|
||||||
|
|
||||||
|
⚠ **This is §4.5's inert gate wearing a different face.** There it was a check
|
||||||
|
that could not return "fail"; here it is a measurement that cannot return
|
||||||
|
non-zero. A clean null is the most reassuring output any instrument produces and
|
||||||
|
the least self-evidencing, so it is precisely the one that has to be earned.
|
||||||
|
|
||||||
|
Same trap in a metric that reads **identical on both arms**: a diversity battery
|
||||||
|
whose rp family froze zero markers reported an attractor hit rate of 0.0 for
|
||||||
|
base *and* tuned. That reads as "no attractors, clean result" and means "this
|
||||||
|
instrument cannot discriminate on this family." Report it as a bounded
|
||||||
|
limitation — that family is measured on one axis rather than two — never as a
|
||||||
|
delta of zero. **A check that returns the same value for every input is not
|
||||||
|
measuring.**
|
||||||
|
|
||||||
|
(Both from the run-2 gate, brokkr-smithy-dev, 2026-08-26.)
|
||||||
|
|
||||||
#### ⚠ 4.6.1 …but calibrate the gate against a CORRECT result, not a convenient one
|
#### ⚠ 4.6.1 …but calibrate the gate against a CORRECT result, not a convenient one
|
||||||
|
|
||||||
The opposite failure, and it costs trust rather than correctness. A coherence
|
The opposite failure, and it costs trust rather than correctness. A coherence
|
||||||
@@ -949,6 +981,9 @@ completed, plausible, wrong run above.
|
|||||||
[ ] log rotates on relaunch one file describes one run
|
[ ] log rotates on relaunch one file describes one run
|
||||||
[ ] present-and-null, not absent a run that claims nothing must say so explicitly
|
[ ] present-and-null, not absent a run that claims nothing must say so explicitly
|
||||||
[ ] watchdog tested negative kill something and confirm it fires
|
[ ] watchdog tested negative kill something and confirm it fires
|
||||||
|
[ ] every NULL has a positive drive the metric with known-answer inputs;
|
||||||
|
control a clean zero is the least self-evidencing
|
||||||
|
result any instrument produces
|
||||||
[ ] stale numbers grepped for the figure, repo-wide, not just in view
|
[ ] stale numbers grepped for the figure, repo-wide, not just in view
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user