Commit Graph
38 Commits
Author SHA1 Message Date
ScottW514 800863fb09 forgetest: the cloud tests wait out the service's thinking
wait_quiet took the machine as quiet after 8 s without the start of a
motion, a park, a run or a lens homing. Between a lid image and its next
move the service is working on the image and the log is silent, and on
the bench reference the re-hunt's moves came 9.6 s apart: cloud.mode-switch
switched back to GRBL mode in the middle of the re-hunt, twice, the first
time canceling a motion at 988 of its 1002 steps.

Every line of a service action now counts as activity: the requests, the
image uploads, the action ends, the runs, the parks, the motions and the
lens homing. The quiet is 30 s. Over 1901 motion, lid-image and hunt
requests in the bench reference's logs, half came within 1.6 s of the
line before, 99 percent within 12.8 s, and two after more than 30 s
(32.8 and 62.2 s).

return_head goes: cloud.mode-switch hands its cloud stretch back through
the cloud client's log (homeoff.cloud_mode_return), and nothing else
called it. RETURN_MAX_MM moves to homeoff, its one user.

Proof: tests/test_cloud_suite.py's new case lands a motion, a lid image
0.8 s later and the next move 1.6 s after the motion, with the quiet at
1 s: the quiet comes after the move. With the old activity marks it is
declared after the image, before the move. forgetest 499 OK.

Acceptance: cloud.mode-switch gates the change on a machine, its re-hunt
waited out before the switch back. The fingerprints of every test in
suite/cloud.py and of every module that imports from it move: 29 tests,
the cloud tests, events.button-telemetry, the exthost tests,
homing.cloud-offsets and setup.check-envelope.
2026-09-25 14:22:51 -04:00
ScottW514 cfef6a1b84 forgetest: the camera-home tests hand the head back where they found it
homing.cloud-offsets, setup.check-envelope and cloud.mode-switch let the
service move the head, and ended with it where the service left it: at
the camera home, or, in cloud.mode-switch, under the camera, where the
service's re-hunt had taken it. Each told the baseline the counters had
been re-zeroed at the starting position, so the hand-back saw nothing to
do. No counter reading can say where the head was found: every service
motion zeroes the counters at its start, and the home and every
controller start zero them again.

suite/homeoff.py now holds the helpers for a test that lets the service
move the head:

- session_travel sums a client's own record of each motion's end ("end
  positions (x, y, z)", the counters the motion zeroed at its start, in
  x8 steps, the one mode a service motion runs at). A motion with no end
  on record, or a log rotated under the run, leaves the travel unknown.
- camera_home_return drops the camera home, jogs the head back by the
  session's travel plus what the counters read since the home, starts
  the controller once more so the counters read zero where the head
  began, and tells the baseline. cloud_mode_return does the same for a
  stay in cloud mode, from the cloud client's log.
- A travel that cannot be known fails the run and moves nothing. A
  hand-back that fails while the test is already failing is logged, and
  the test's own failure is the one reported.
- judge_whole_motions fails a homing motion that stopped short: the
  position declared after it is false.

cloud.mode-switch imports the helpers inside its function, so no other
cloud test's fingerprint moves. Its cloud stretch no longer goes through
return_head, which read 0/0 after the controller's start and left the
head under the camera.

Proof: tests/test_camera_home_return.py, 15 cases on the machine's own
log lines (a session stopped short, a whole three-motion one, a cut log,
a refused motion, homed and unhomed counters, a restart since the home,
the first failure winning). tests/test_cloud_suite.py's mode-switch fakes
now zero the counters and remove the anchor at a controller start, write
the anchor at the home, and move the counters on a jog: 10 cases, a
failed hand-back that must not hide the test's failure among them.
forgetest 498 OK. On the bench reference, with the driver and the runner
fixes: homing.cloud-offsets, setup.check-envelope and cloud.mode-switch
PASS, each ending with the head where it was found; cloud.mode-switch
jogged its cloud stretch 246.06/139.01 mm back with the counters across
the jog agreeing, and every baseline was clean.

Acceptance: the three tests are the change; their fingerprints move and
no other test's does.
2026-09-25 14:22:50 -04:00
ScottW514 619414ed2b cloud suite: a print turned away before the button fails at once, with the reason
The cloud client turns a print away before the button wait for four
reasons of its own (machine._safe_to_move): the lid or the interlock, a
machine that is not idle, the coolant above its start ceiling
(THERMAL.max_start_temp), a coolant sensor that reads nothing. It logs the
reason and finishes the print ':cancelled' within a millisecond. The four
tests that wait for the button looked only for the wait, so on the bench
reference, with the loop at 27.4 C against the ceiling of 27,
cloud.dark-print sat out its 120 s and said "the print never reached the
button wait".

wait_button_wait replaces the four waits: it ends on the print's finish
line as well, and fails with the client's own lines, for example "the
client turned the print away before the button wait: INFO
machine:_safe_to_move machine temp is too high, temp: 27.4 (... finished
with event ":cancelled")". The ceiling is the client's and stays where it
is: the remedy for a warm loop is airflow, a run session for a minute or
two.

Proven. test_a_print_turned_away_before_the_button_fails_at_once_with_the_reason
replays the bench reference's lines and fails in under 30 s with that
reason; with the early exit disabled it fails after the full 120 s with the
old words. Each reason is a line the pinned cloud library can log, so the
phrase check passes. The unit suite passes (422) with no undefined name. On
the bench reference, image 20260920211625 with this cloud.py over the
image's, cloud.dark-print passes through the new wait with the loop at
25.8 C.

Acceptance. cloud.dark-print, cloud.verdict-refuse,
cloud.lid-during-button-wait, and every test that starts an offline print
go through the new wait. The helper is module text, so every cloud.py
test's fingerprint moves; no product behavior changes.
2026-09-20 18:02:31 -04:00
ScottW514 0eb764bf75 Added SPDX 2026-09-18 12:14:22 -04:00
ScottW514 db6015dc81 forgetest: cloud.mode-switch opens the lid behind the controller's start, ahead of the hunt
The supervisor holds every controller spawn until the enclosure is
closed (forgectrl 0.1.25), and the test opened the lid before it asked
for the cloud controller: POST /mode answered "waiting, the lid is
open" and the controller never came up. The round trip now switches
with the lid closed, polls /mode five times a second, and opens the
lid the moment the controller is running. The client requests its
connect-time hunt a few seconds after its start, right behind its
session, so the hunt still finds the lid open. The order is recorded
and judged: the hunt's request line must not be in the client's log
when the lid reads open (hunt_before_lid_open), and the test refuses
to start with the lid open. The catalog text tells the operator to
open the lid at once, with a hand ready on it.

Proof. Host: test_cloud_suite drives the round trip with the hunt
landing only once the lid reads open, as on the bench, plus the lost
race (the hunt requested before the lid opened fails the test with
"before the lid was open") and the start with the lid open refused;
8 mode-switch cases green. Bench reference (dev image 20260915001814,
forgectrl 0.1.25): cloud.mode-switch PASS in 118 s, the lid open 4 s
before the client requested its hunt, no refusal before the hunt's
end, the lens homed, the exhaust row unjudged, 5 service motions
after the lid closed, $H under gfhome homed in 48.4 s with 9 motion
windows. No catalog consequence beyond the test itself: its covers
map is unchanged.
2026-09-15 15:41:35 -04:00
ScottW514 f080d5d9cb Catalog: the cloud client's latch at the run, and a hold that never clears
The cloud client changes (python3-gfhardware: the latch unlocks at the
run and nowhere earlier, the warm-up is supervised, a live feed must
land and finish, fire needs a power byte first, the homing runner
always stops). This commit carries the catalog tests that hold the
part the bench can see.

forgetest/forgetest/suite/cloud.py:
- cloud.dark-print (new, kind operator, one press, dark by
  construction) replaces cloud.verdict-hold. The print arms on the
  press, waits for the engine's acknowledgment and runs: the laser
  latch is locked at the button and through the wait, unlocked only
  for the run (immediately before it starts), locked again when the
  job ends, and the print completes. The engine's own warm-up release
  stays proven by cooling.floor-and-warm-up.
- cloud.verdict-refuse (new, kind operator, one press, dark): the
  start gate far above the coolant and cloud_hold_max_s at its minimum
  keep the armed print under the warm-up past the bound. The client
  waits with the latch locked (sampled every two seconds), cancels at
  the bound with its own log line, never runs, closes the armed window,
  and the print ends ':cancelled'. The settings are restored.
- cloud.verdict-hold is retired: its release rode the loop heater at
  the flow-check duty against a gate one degree above the coolant,
  inside the upstream reading's noise band, and its run hovered for
  minutes with the engine's "warm-up stalled" line in the log. The two
  tests above prove the client's contract without the thermal race.
- cloud.oversize-stream already reads cnc/streaming back at both ends
  of the run, which is the readback the client now insists on. A
  forced streaming write failure has no seam on the board (the write
  goes to sysfs as root) and stays a host test.

tests/test_cloud_suite.py follows: excerpts and failure cases for the
two new tests in place of the retired one's.

Proof. Host: the forgetest unit tests. Bench reference:
cloud.dark-print passed (locked at the button, unlocked in the run,
locked after, ':completed'), cloud.verdict-refuse passed (held 60 s,
31 latch samples all locked, ':cancelled', settings restored);
motion.deadman's kill during $H ended the runner on SIGTERM alone; and
the retired cloud.verdict-hold passed once more on the new client
before it went (the latch locked through an 8 minute warm-up hold, the
release ran the print to completion).
2026-09-14 19:45:20 -04:00
ScottW514 2f2a4160af Stream and motion robustness: harness rules, catalog tests, the hand-back's counter scale
The GRBL driver's stream engine and its motion envelope change
(grblHAL-glowforge: the shipper writes outside the lock, a clamp inside
an armed window faults, the X/Y soft limits follow the home, the
machine's settings are pinned, the homing keys are clamped). This commit
carries the host rules and the catalog tests that hold them; the driver
commit follows, because its CI fetches these harnesses unpinned.

scripts/bench/laser_stream_test.py:
- Rule 28: a 300 ms producer stall while armed faults the stream with
  ALARM:17, the kernel sees no step burst (at most the planned steps per
  100 ticks), the stream ends dark, the latch sideband ends on the lock.
  The same stall unarmed is a warning: the move completes with every
  step, the clamp visible as the burst the kernel counts.
- Rule 29: a 300 ms stall of the sink's write leaves the producer on
  pace: no clamp, every step, lit through, dark at the end.
- The stand-in engine takes GFSINK_STALL_MS and GFSINK_WRITE_STALL_MS,
  null-sink only.

scripts/bench/z_envelope_test.py:
- Rule 10: homed (a gfcloud home), a program move past X max, Y max or
  the near edge alarms with ALARM:2 before any motion, a jog past the bed
  is refused with error 15, a move inside the bed runs, and a $20 write
  keeps the limits. The core repeats the last error for the line after a
  refused jog until an empty line clears it, so the rule sends one.

forgetest/forgetest/suite/motion.py:
- motion.soft-limits (kind auto, no emission): homes through the cloud
  suite's gfhome homing when the machine is not homed, then the three
  refusals (ALARM:2, the kernel counters still), the refused jog, the
  inside move, and the return to the corner read at rest.
- motion.deadman phase 2b: a 100 ms SIGSTOP mid-move, inside the
  kernel's queue: no underrun, the controller's log warns of the clamped
  late events, the move completes with every step (read at rest), the
  latch stays locked. The armed clamp is proven on the host (rule 28).

forgetest/forgetest/suite/cloud.py:
- gfhome_homing drains the driver's answer to $H once the session ends:
  it sits behind the status reports and passed for the reply to the
  caller's next command (a setting read as None).

forgetest/forgetest/baseline.py:
- The hand-back reads the position counters at the kernel's own
  microstep mode (cnc/x_mode, read before the sysfs restore puts the
  settings' mode back). At the x8 constant, an x32 machine's 30 mm read
  as 120 mm, beyond the return bound, and the displaced head was left in
  place. The dead band scales the same way. tests/test_baseline.py holds
  both.

Proof. Host: stream rules 1 to 29 and z_envelope rules 1 to 10 against
the null-sink driver, the forgetest unit tests. Bench reference:
motion.soft-limits passed (X 495 and Y 279 refused with ALARM:2 and
0.000 mm of kernel motion, X-1 the same, the jog error:15, the inside
move ran, back at the corner 0.0/0.0 mm); motion.deadman passed with
the new phase (92 late events clamped, max behind 92.3 ms, no underrun,
30.0 mm counted, latch locked) and the hand-back jogged the head back
under the x32 scale.
2026-09-14 17:28:28 -04:00
ScottW514 57ba3454b4 Clear a setting the only way the daemon accepts, everywhere it is cleared
An empty form field never reaches the request, so POST /settings with an
empty value in the body carries no key at all and the daemon answers 400,
"no known setting in request". Only the query-string form gets an empty
value through. Measured on the bench reference:

    POST /settings -d "lid_policy="   -> 400, the value unchanged
    POST /settings?lid_policy=        -> 200, the key cleared

Three places in the suite already knew this and say so in a comment; two
did not.

motion.lid-policy-hold restored with the form body, so on a machine whose
lid_policy has never been set the restore was refused and the test failed
its own check. It now uses the query-string form for the empty value, as
the others do. This is the same test as the commit before: that one
stopped it from writing the word "cancel" where the machine had nothing,
and this one lets it write the nothing.

cloud.mode-switch had the older shape of the same fault: it restored
homing_mode to the literal "none" where the machine had it unset. On the
bench reference the restore branch never ran, because that machine homes
through the cloud, so it has not fired yet; on a machine with homing_mode
unset it would have. It now restores exactly what it found.

The rest of the suite was audited for this and is clean: camera.key-read,
cooling.critical-tier, cooling.tec-drive, cooling.crash-watch-plumbing,
cooling.fire-watch-tiers, cooling.floor-and-warm-up, cloud.verdict-hold,
forgectrl.settings, logs.levels and motion.xy-microsteps all either guard
the empty value or go through the shared Restore helper, which has always
done it correctly. commission.cloud-header-capture uses that helper too.
2026-09-10 10:26:25 -04:00
ScottW514 97287aa6a9 commissioning: the layer, the acceptance tests, the harness rule, the docs, and the bench drills
meta-forgefirm: the forgefirm-users init replays the account at boot;
sshd refuses root and empty passwords and runs only while the panel
turns it on; the release image keeps an empty root password for the
console; the console banner; avahi announces forgefirm.local; https in
libmicrohttpd and ulfius; the panel on 80 and 443; the license bundle on
the rootfs; release.sh checks the root policy on the built rootfs.

forgetest: the commission suites (commission, commission_dark,
commission_sheet: 23 cases); the runner turns cloud mode on with the
typed phrase for a test that declares it; the baseline's motor_lock is
0; the log-export test checks the bundle for the camera key; the record
helpers write bytes as given and join the daemon's paths as POSIX. The
stream harness gains rule 24: a hold verdict is held again after a
resume. Bench drills: lens_travel.py and lens_stop_accel.py.

Docs: BRINGUP carries the present state; CAMPAIGN-LOG carries the dated
record.
2026-09-06 19:56:05 -04:00
ScottW514 88ec984e28 Audit follow-through: runbook, bench tools, recipes, release tooling
BRINGUP describes the present: the 54-test catalog and its seven-test
always core, the tier counts, the shipped low-temperature gates, the
density floor ($35 = 10), the two local core commits, the ffboot env
write, the aa-offset route, the current bench image, and the bench
measurements the audit asks for (pooled into the next session). The
workstation shell notes and every em dash are gone.

forgetest: the takeover waits for the cloud client too (found by its
command line); the unauthenticated /boot probe names the endpoint's
parameter; the UI prose is American English. Recipes: forgetest
fetches its package directory and init script only and drops
__pycache__ at unpack; the dev image no longer re-adds forgectrl; the
release image's remove list drops the gfui-client the BSP no longer
has; the platform identity strips the kernel's local-version hash
from the modules directory name, so a re-patched kernel keeps its
fingerprints. grblhal restart is stop then start. release.sh --dev
packs the dev image. fixture.sh refuses a readable env file.

Bench tools: the live-fire drills measure the lid-IR baseline before
every run and point at the fire-watch thresholds the engine reads;
one thermistor conversion (gfbench.degc) serves every drill; the six
dated measurement records leave the tool directory; feeder.c names the
two sysfs writes its caller makes.

Host tests: forgetest 258 pass; the coverage lint reports no uncovered
path across 54 tests. Acceptance: forgectrl.auth covers the /boot
probe; update.* cover ffboot and the manifest identity; the runbook
and bench-tool changes have no catalog consequence.
2026-09-02 09:51:23 -04:00
ScottW514 44ba07ea94 forgetest: cloud.verdict-hold drills the warm-up wait of an armed print
The start gate set just above the coolant opens the armed session under
the warm-up; the cloud client waits it out after the button, nothing
runs, and the release starts the print, which completes. The mid-run
hold and its bound are host-tested: the gates apply at session open, so
no setting can produce a hold mid-run on the bench. Bench-excerpt unit
tests: the pass, and the failure when a run starts under the hold.
2026-09-02 07:24:35 -04:00
ScottW514 8ff37d188d forgetest: the lid-at-button-wait drill uses a job longer than the ring
A job longer than the ring keeps its feeder alive through the button
wait, so the cancel there has to stop the feeder before the park clears
the ring. The drill now loads such a job, requires the "longer than the
ring" line, and after the cancel reads the program total twice over the
feeder's retry period (zero both times) and cnc/streaming (zero). The
bench-excerpt unit test carries the long-job line and the new evidence.

Covers map unchanged: the drill already names gfhardware/machine.py.
2026-09-02 06:52:09 -04:00
ScottW514 b54b94e8ae forgetest: the bench actuator plugs into the action seam
fixture.py: the bench's /data/forgetest/fixture.json (hostname, key,
optional ip, the channels wired, arm_press), a resolver for
<hostname>.local asked of the network directly (the image has no mDNS
resolver), and the client. The runner probes it before every run and
at most every 30 s otherwise; ctx.act asks it for a channel it covers
and still waits for the machine's own reading, falling back to the
operator's notice when the box fails. A test declares with hands=(...)
what it asks of a person beyond its typed actions; an operator test
with none, whose actions the fixture covers, is routed into the
unattended queue, its Ready gates pass, and a prompt it raises anyway
is a FAIL naming the undeclared step. Live tests never move; their arm
press stays a person's unless the bench opted in, in which case the
fixture presses when the button lights. Whatever the box still holds
after a run is released before the baseline's post pass and recorded.
The page shows what the fixture covers. Contract in ACCEPTANCE.md; the
wiring facts, with the interlock connector left to the bench to settle
(SAFETY.md and the sister map differ), in BRINGUP.

Catalog unchanged in its definitions; the cloud and laser suites'
shared code moved, so their implementation hashes move with it.
2026-08-23 14:15:53 -04:00
ScottW514 53e4fa20e9 forgetest: the cloud tests' coverage maps follow the split
The six job tests covered all three cloud components whole, so a
one-line change anywhere re-required one real print and five attended
tests. The maps now say what each test proves: the protocol test the
web session, the emulator and its fixtures; the offline tests the run
loop, the hardware it drives, the offline dispatch and the pulse path;
cloud.mode-switch the homing path; every one of them the client's
common ground. The one real print keeps the coarse maps as the
integration and the lint's floor. Two entries of the protocol test
named app files without the recipe's subdirectory and selected nothing;
the paths are spelled out now.
2026-08-23 12:54:55 -04:00
ScottW514 969bac6013 forgetest: the service's hunt paid only where it is the subject
A cloud client the tool starts for anything but homing comes up under
the /run/gfcloud-nohunt marker: the real client back after the
emulator, a mode the runner switches to or hands back, a controller it
restarts. The service keeps the head position it has. cloud.mode-switch
and cloud.service-protocol keep their hunts, and so does the one real
print: enter_cloud reuses a running session only when that client has
hunted the machine itself (session_hunted: never the emulator's, never
a no-hunt start), otherwise it restarts the client with the hunt, since
a print placed on a head position the service only believes can run the
gantry into a rail. The markers are one start, taken down by the client
that read them first thing; the tool's own removal stays for a start
that never happened. Catalog unchanged; the cloud tests' shared code
moved, so their implementation hashes move with it.
2026-08-23 11:27:36 -04:00
ScottW514 1c8197faa3 forgetest: cloud.service-protocol, the service answered by the emulator in this machine's identity
The service-protocol half of the cloud catalog on its own test: the cloud
client restarted as gfutilities' emulator under the /run/gfcloud-emulate
marker signs in, passes the firmware check, opens the WebSocket, answers
the connect-time hunt and the image requests with the dev image's canned
frames, and runs a print from the app through the real download path to
':completed' - nothing moves, nothing arms, and only the app has to be
driven, by a person or an agent through the prompt API. The real client
is restarted afterward and its hunt waited out. session_live now knows
the emulator's session is not the machine's, so enter_cloud restarts it
rather than reusing it; restart_client is the one restart the offline
and emulator entries share.

The dev image adds python3-gfutilities-emulator (the fixtures, packaged
on their own in meta-openglow); forgefirm-app moves to 12ad3b1 (gfcloud
--emulate). Catalog: 44 tests, 27 auto / 9 operator / 8 live; the new
test covers the gfutilities service layer and examples/, which step 4
will take off the other cloud tests. Replays over the prompt script;
contract and BRINGUP updated. A layer change (the dev image recipe):
everything re-requires on the next image.
2026-08-22 20:19:46 -04:00
ScottW514 0cb9044e1d forgetest: the offline jobs removed after each test; the record of the first offline campaign
Every offline test now removes the jobs it wrote under /tmp/forgetest
on its way out (the bench rule: nothing left behind in the session that
put it there). CAMPAIGN-LOG gets the 2026-08-22 entry for dev image
20260822232347: the offline service dry-checked, then campaign
c-20260822233344-08de, 13 run and 43 of 43 with the four offline tests in
5.5 minutes, and what the machine said under it. BRINGUP: the offline
service is done and bench-validated.
2026-08-22 19:59:36 -04:00
ScottW514 9cc2e4eef4 forgetest: the machine's print behavior under the offline service
Four cloud tests no longer need the app, an account, a network, or
anything on the bed: cloud.lid-interlock-abort, lid-during-button-wait,
paused-lid-cancel and oversize-stream run under the offline service
(enter_offline restarts the cloud client with the /run marker for that
one start; Offline is the socket; offline_job writes the job). The jobs
come from forgetest/puls.py: the header of a factory print of this
machine type (134 tags, MCsn 0, so the client's limits and settings come
from where a service job's do) over a square traced at a steady feed
with a leading power byte of zero and no LASER bit anywhere - the arm
unlocks the latch, the beam is never commanded, so the tests stay live
and need no scrap. A job longer than the ring (33 MiB of ticks, an hour
of squares) is an 87 kB gzip written in a tenth of a second, in place of
a full-bed raster designed in the app.

session_live reads the offline mark as "no web session"; enter_cloud
restarts an offline client with the service, so cloud.pause-resume (the
one real print left, with cloud.mode-switch the service-protocol half of
the catalog) follows the offline tests without the operator's hand.
Replays over a fake socket; the contract and BRINGUP say how the cloud
catalog splits. Catalog consequence: the four re-ported tests move;
nothing else is invalidated.
2026-08-22 19:19:50 -04:00
ScottW514 296fd686e1 forgetest: the button's commanded level, and one app cancel instead of two
laser.arm-wait-lid failed its first bench run on the witness, not the
machine: the cancel relocked, disarmed and emitted nothing, but the
"button dark" check read the LEDs' brightness the instant after, and
the smooth trigger fades it. The controller writes `target`, so that is
what hw.button_leds() reads now (brightness where no target exists),
and check_button_dark waits a few seconds for the command to land.

The operator's campaign showed cloud.oversize-stream and
cloud.pause-cancel-paths both cancelling a print from the app. A print
longer than the ring has to be ended that way, and the operator is at
the app for it anyway, so the app cancel lives there and is judged in
full (the abort tail: stop, park to the job start, relock, disarm, the
button dark, ':cancelled'); cloud.pause-cancel-paths loses its second
print and becomes cloud.paused-lid-cancel - one print, one arm press.
Replays follow. Catalog consequence: the two cloud tests and the laser
module are re-required; the count stays 43.
2026-08-22 17:47:26 -04:00
ScottW514 de324cc32f forgetest: the operator's part asked for by name, and fewer hands in a campaign
A campaign asked a person for about eighty things: lid, button and
interlock actions, app jobs, and sixteen confirmations by eye, most of
them as popups to read and answer while the head was already moving.
This is the forgetest-only step of cutting that down.

The operator channel. A test asks for its operator's part in four ways:
ctx.ready() pre-announces a timed step and waits for the click that
starts it; ctx.notice() is a standing instruction with no button, the
test watching the machine for the result; ctx.act(channel, state) is a
machine action by name (lid, interlock, button) - a notice for the
operator today, proven done by the switch reading or an `until`
condition, recorded in evidence.actions with who performed it, and the
seam a bench actuator plugs into through runner.fixture; ctx.confirm()
stays for the yes/no the evidence cannot answer. Tests declare
`actions`; a `precheck` refuses a start the machine cannot honor (a
reason, no result, a queue skips it and carries on).

The page shows what you will do before it is asked: the running test's
steps, a queue's attended tests still waiting, or the test whose title
you clicked while idle; notices and prompts sit under it. The campaign
card no longer carries baseline, queue and leftover notes: those go to
the runner journal (daemon.log, syslog as `forgetest`, and the run in
progress), with a Runner journal button in the footer.

The catalog, 43 tests (27 auto, 8 operator, 8 live; was 45: 25/12/8):
cloud.mode-switch absorbs cloud.hunt-lid-open and cloud.gfhome-homing
(the connect made with the lid open, the hunt judged lid-open with its
Z cycle, the re-hunt waited out after the close, the switch back, then
$H judged by gfhome's own "homing complete" line with its motion
windows; precheck homing_mode = gfcloud). kernel.fire-line is auto with
the HV-not-good precheck, camera.snapshot is auto (a second frame with
the lid lamp off differs and is smaller), motion.jog-roundtrip is auto
(the head accelerometer per leg, the supervisor's own witness). The
remaining attended tests use Ready gates and act(); the head's beam
detector and the button LEDs replace the eye, leaving two confirms: the
emission witness's mark and the app's display in cloud.pause-resume.

Proof: tests/test_operator.py (the channel, the precheck, the journal),
the cloud replays re-targeted to the merged round trip over a fake grbl
and the bench excerpts, 196 host tests green, coverage lint 0 uncovered.
Every re-ported attended test is owed one bench run on the next dev
image (BRINGUP). Catalog consequence: the merged and reclassified
tests' implementation hashes move; nothing else is invalidated.
2026-08-22 16:36:03 -04:00
ScottW514 c1591e4f69 The fan floors proven on a pinned image; a fan fault ends with its session
The dated record of dev image 20260822135848: the campaign of every
non-operator, non-live test at 18 of 18 PASS with the measured floors and
the operating-point rule (cooling.fan-gate-trips and the hunt leg of
cloud.mode-switch as recorded), and the unplugged-exhaust-fan drill:
AIRFLOW at the grace plus three ticks with the exhaust dead, the other
fans held, the reason relayed on the Grbl port, the replugged fan ok
inside the next session's grace.

The drill showed the fault riding into idle, where the hold canceled
jogs and would have refused the cloud print that re-proves the fan.
forgectrl pin d51dbdb: the fault ends with its run session, and the next
session judges every fan afresh. cooling.fan-gate-trips checks the
verdict is OK with no hold once the tripped session is over (the unit
fake mirrors it); its covers, and the cooling tests' shared covers, gain
src/coolfmt.* (the tree manifest carries the new files at the bumped pin,
and the lint was right to ask). cloud.mode-switch samples the hunt's gate
rows twice a second: a hunt's run phase is a few seconds long.

Docs: COOLING 3a, BRINGUP item 19, CAMPAIGN-LOG.
2026-08-22 10:46:41 -04:00
ScottW514 adcd1adb9a Fan floors from the measurement; a hunt is measured, not judged
forgectrl pin 47e4256: the airflow floors set from the bench measurement
(exhaust 6400, intake 2290, air assist 6000 rpm, purge current 300, grace
15 s) and the operating-point rule: a fan is judged while the laser is
armed, when a job's profile may raise it but never lower it below the run
duty, or whenever it is commanded at the run duty; the service's hunts,
sent with the extraction fans off, are measured and published unjudged.

Catalog: cloud.mode-switch now waits out the connect-time hunt sampling
/cool/status and requires no AIRFLOW, the exhaust row unjudged and the
exhaust actually off; its covers gain forgectrl src/cool.* and
src/airflow.*. cooling.fan-gate-trips uses an 8 s test grace (the
intakes take 7 s to 90 percent and tripped under the old 2 s at the new
floor) and its off leg waits for the row state as well as gates_off.
fan_floor_measure.py names a reply that is not JSON and calls the purge
readings idle and run (the pump is always on).

Docs: COOLING 3a and the settings table, BRINGUP item 19 and a facts-bank
entry with the measured fan speeds, the bench README, and a CAMPAIGN-LOG
entry for the measurement, the two status-document finds, the hunt find
and the hot-deployed bench runs of both tests.
2026-08-22 09:56:53 -04:00
ScottW514 5249c1ecf8 CAMPAIGN-LOG: the job's limits pass through on a live session
cloud.pause-resume passed on dev image 20260821220926 with the service's
hunt windows (10 to 50 C) ignored as looser and the print's window (33 C,
5 C floor, 116 rpm air-assist floor) matched. The test now quotes the
print's job-limits line rather than the session's first (a hunt's), and
keeps the engine line that carries the header beside the last one.
forgectrl pin moves to e0b41b3 (the "not stricter" notice once per value).
2026-08-21 18:39:43 -04:00
ScottW514 d530f773e2 cloud.pause-resume reads the job's limits through; pin the pass-through
The pulse header's envelope now reaches the cooling engine: the cloud
client (python3-gfhardware c34faa1) derives the coolant window and the
fans' minimum speeds from the header and rides them on every report,
and the engine (forgectrl 57f6064) resolves each as the stricter of its
setting and the job's, never looser and never overruling an off gate.
cloud.pause-resume, which runs a real print, now also reads the client's
"job limits from the header:" line and the engine's "effective limits:"
line from the two logs; two host cases cover the failure paths. The
needle guard lists the engine's phrases as not the app's.

COOLING.md section 2 explains what a cloud job brings with it; BRINGUP
item 19 records the pass-through as landed. Pins: forgectrl 57f6064,
forgefirm-app c34faa1 (0.1.14+git); fetch-verified.
2026-08-21 18:07:41 -04:00
ScottW514 6c1d68f2c3 Judge the cloud resume on the lines the app logs; guard every needle against the pinned app
cloud.pause-resume failed a print that paused, resumed with its laser
lead, completed and parked: the test waited for the single line "button
pressed while paused; resuming", and the app has logged that as two
lines since its feeder work ("button pressed while paused", then
"resuming (laser lead N ticks)" from _resume_retraced). The replay
fixture carried the old wording, so the host test kept passing.

The pause and resume are now judged on PAUSE_LINES + RESUME_LINES
through one checker shared by the three tests that drive a pause
(cloud.pause-resume, the streamed pause, the pause-then-lid test), which
also fails on the app's "resume refused" line with the reason. The
fixture carries the app's two lines.

So the wording cannot drift silently again: tests/test_cloud_needles.py
reads every log phrase the cloud suite greps for out of cloud.py (the
left side of each `x in ln`, every wait_log needle, the mark tuples, and
the phrases it builds) and checks each against the logger calls in the
app sources at the revisions the recipes pin, read from the manifest
cache the tree manifest builds (the sibling checkouts locally),
placeholder-aware under a rule that never lets a placeholder stand for
the phrase itself. CI now builds the tree manifest before the unit
tests so the cache is there. The old needle fails that check.

Replays added: the second press seen but no retraced restart, and a
refused resume. 159 unit tests pass; coverage lint clean. No catalog
consequence beyond the suite module's own hash.
2026-08-21 14:09:51 -04:00
ScottW514 54e1689889 Let a test declare the controller mode it needs; the runner switches to it
The cloud job tests enter cloud mode and stay there, by design, so a
queue (or an operator) that goes on to a motion test reaches it with
gfcloud as the controller and no grblHAL process to find:
motion.step-timing-under-load failed on exactly that, before it touched
the machine. Nothing in the runner put the machine into the mode a test
needed; the baseline only preserved the mode it found.

A test now declares `mode="grbl"` (or "cloud") in @test. The runner's pre
pass, after the leftovers are handled and before the preserved state is
captured, switches through POST /mode, waits for the supervisor to settle
(controller running, motion verified) and for the Grbl port to answer,
and fails the test with the reason when the mode cannot be established.
Capturing after the switch means the post pass keeps the mode the test
asked for, so the machine changes mode only where the next test asks for
it and never between tests of the same mode. The cloud job tests keep
managing their own entry (enter_cloud also waits for the service
session) and declare nothing.

Tagged: every motion.* test but the mode-agnostic liveness probe, the six
laser.* tests, cooling.fans-quiet-after-motion, cloud.mode-switch and
cloud.gfhome-homing (both start in GRBL mode). controller_pid() now says
what mode forgectrl reports when the process is missing. The page shows
the declared mode as a badge; the Grbl port probe moved to hw.

Proof: tests/test_mode.py (switch_mode against the fake forgectrl,
including a refused switch, a controller that never comes up and a port
that never opens; the runner end to end from cloud mode, from grbl mode,
an undeclared test, and a failed switch). 151 unit tests pass; the
coverage lint is clean. No catalog consequence beyond the suite modules'
own source hashes: the change is to how a test is started, not to what
it proves.
2026-08-21 12:34:03 -04:00
ScottW514 aaabfdf9b9 Check the progress a print reports, where a print already runs
Two tests already run a print end to end, and progress is a property of a
running print, so the checks go there rather than into a test of their own
that would cost the operator another job.

cloud.pause-resume takes the job that fits the ring: the client names the
length it is reporting against, and the operator is asked the question only a
person can answer, whether the bar actually moved.

cloud.oversize-stream takes the job that does not fit, which is where a moving
denominator would show: the kernel's program total grows all run long under a
live feed, and the test already samples it growing, so the check is that the
figure progress divides by is larger than that - the job, not the count the
ring had swallowed when the run started.

The forgetest replay plays a captured log from a build that predates the line,
so it carries the line where the current build emits it, as it already does
for the warm-up and the rest.

BRINGUP's cloud item now says a print reports itself again, and what is left
on it is a print watched from the app.
2026-08-20 14:16:27 -04:00
ScottW514 d122a6ff1d Check a print's warm-up and rest where a print already runs
cloud.pause-resume runs a print end to end, which is exactly what the job
lifecycle needs to be seen: a non-zero hold before the first fire, a non-zero
rest after the park, and neither on the connect-time hunt in the same
session. Folding the checks in there costs the operator nothing, where a
test of its own would cost another print.

The forgetest replay noticed first: it plays a real captured log from a build
that predates those lines. Rather than editing what the machine said that
day, the replay carries the two lines where the current build emits them.

BRINGUP's cloud item now names what is actually open on the header keys, the
park and the two periods, and records that the pause constants were looked
for in the wrong place.
2026-08-20 10:05:38 -04:00
ScottW514 942d9dde99 Drill the backtrack boundary, and pause a streamed print
kernel.backtrack-bounds plays a program, stops it, and holds the readback
to the bytes it played less the deceleration tail: a step past the boundary
has to be refused rather than quietly shortened, and the run at the boundary
has to play out and come back idle. Motors locked, latch locked, duty zero,
so nothing moves and nothing fires.

cloud.oversize-stream pauses and resumes the live-fed print it already has
running. That is the pause the kernel change makes possible, and it costs a
minute of a job that is on the bed either way.

BRINGUP's pause bullet, its ring facts and item 18 all said a ring under a
live feed has nothing left to back into. The gap the writer keeps clear says
otherwise; what is still open for GRBL mode is the bookkeeping above the
ring, not the kernel below it.
2026-08-20 07:41:41 -04:00
ScottW514 7bd9c45fd3 Accept a print longer than the ring
cloud.oversize-stream drives a job the ring cannot hold and checks the
signature of a live feed: the load reporting the job as longer than the
ring, the device in live-feed mode, the kernel's program total growing
during the run, no underrun, and a clean cancel afterwards.

read_program_total gives the catalog the counter that growth is read
from.
2026-08-19 21:21:31 -04:00
ScottW514 cd5098b6b3 Acceptance catalog: merge the tests that share a setup (39 -> 35, live 10 -> 7)
A sweep of the whole catalog for the overlap the drill work found. The test
applied was "do these share a SETUP", not "do these share a subsystem":
combining only pays where a human waits - an arm press, scrap, a takeover,
a mode entry, a lid choreography - and it costs failure isolation, because
the campaign inherits per test and a merged test invalidates as a unit. The
17 auto tests were left alone for exactly that reason: they cost no operator
time and separate ids give the coverage map and the domain invalidation
finer teeth.

  kernel.k3-unlock + kernel.fire-abu -> kernel.fire-line
      Four phases behind ONE takeover of the pulse device instead of two:
      A/B/U on the FIRE line, then the mid-ramp unlock. Same HV-not-good
      gate, same zero duty, same safe state on the way out.

  laser.expected-stop + laser.kill-mid-fire -> laser.armed-kill
      Both ways an armed job is killed, on one scrap setup: the supervisor's
      expected stop (with its separate operator-judged restart) and then a
      SIGKILL of the restarted controller. The second phase re-reads the pid
      after the restart, so it kills the process the supervisor just spawned.

  laser.lid-cancel-mid-fire + laser.pause-resume-live -> laser.pause-resume-lid-cancel
      One armed burn in the order the factory uses the machine's controls:
      press (pause - emission stops, latch stays UNLOCKED, armed window
      stays open), press (resume), lid (cancel, reset without alarm, return
      to the job start, button latch SET). One arm press instead of two.

  cloud.lid-abort + cloud.interlock-abort-park -> cloud.lid-interlock-abort
      Two prints in one test: the lid, then the interlock with the lid opened
      during the park. The tail both share - park complete, kernel counters
      back at the job start, ':cancelled', latch locked, armed window closed -
      is one helper now, so both triggers are judged the same way.

Not merged, though they share code: cloud.gfhome-homing and cloud.hunt-lid-open.
Both drive Machine._hunt (gfhome.py and gfcloud.py build the same gfhardware
machine, so the lid ungating is the same lines), but each is about the opposite
value of the same variable - gfhome's homing is camera-corrected and needs the
lid CLOSED in GRBL mode, the hunt test needs it OPEN through a cloud-mode
connect. Merging would put a mid-test mode switch back into the cloud tests.

Shared live-test prologue (the arm cue, the mark job, the wait for the emission
witness) and the post-kill trail/judging are helpers now.

Host proof: 104 unit tests - the merged cloud test replays the machine's own
lid-abort excerpt twice, once with the interlock substituted for the trigger,
and three negative cases hold it honest (loop already open, a park the lid can
interrupt, a stop that is not edge-driven) - and the coverage lint at 0
uncovered across 35 tests. Catalog time 137 -> 125 minutes, 96 of it attended.
2026-08-17 08:39:35 -04:00
ScottW514 0870a835c4 Acceptance catalog: the rest of the lid/button/interlock drills, and a pause/resume chain-timing tool
Catalog 34 -> 39. Every remaining bench drill of the lid/button parity work
is now a test, with drills that exercise the same path combined:

  motion.lid-cancel-home   also cancels from a hold - a job paused on the
                           button is ended by the lid, never resumed - so the
                           armed window and the hardware button latch stay in
                           agreement by construction.
  motion.cancel-abort      also asserts what a sender abort must NOT do: it
                           stops where it stopped and never returns home. The
                           return-to-start belongs to the lid policy alone.
  motion.interlock-cancel-park (new)  the interlock loop cancels like the lid,
                           and the lid opened during the return home does not
                           interrupt it.
  motion.lid-policy-hold (new)  the other policy: Door park, no cancel, no
                           return, and a cycle start finishes the move. The
                           setting is restored on the way out.
  laser.pause-resume-live (new)  the button pause/resume during a live cut:
                           emission stops, the latch stays UNLOCKED and the
                           armed window open (a pause is not a cancel), the
                           next press resumes and the job finishes.
  cloud.interlock-abort-park (new)  the same interlock/park pair in cloud mode.
  cloud.pause-cancel-paths (new)  the two non-finishing ends of a print, each
                           from the state the factory ends it in: paused on the
                           button then cancelled by the lid, and cancelled from
                           the app while running.

The shared cancel tail (reason reported, reset without an alarm, position
kept, head back at the job start with the KERNEL counters confirming it) is
now one helper, so every trigger is judged the same way.

scripts/bench/resume_dark_lead.py: samples LASER_ON, FIRE, HV_ENABLE, the
charge-pump watchdog, the button and the doors off the SoC pads at ~2 kHz
through /dev/mem, with motion dated from the kernel step counters, across a
pause and a resume. Levels are taken at idle and everything after is reported
as a change from that baseline, so no polarity assumption is baked in. Dry by
default, with --auto driving the pause and resume through ! / ~ for an
unattended rehearsal; --run live adds the dark lead between FIRE and LASER_ON,
in milliseconds and in millimeters at the job's feed.

Bench registry: argument specs can name a flag (--feed 600) instead of being
positional, and an optional argument with an empty default is left off the
command line entirely.

Host proof: 104 unit tests (20 in the cloud suite - the two new cloud tests
replay the machine's own lid-abort excerpt, with the interlock and the app
cancel substituted for the trigger, and fail for the right reasons), coverage
lint 0 uncovered across 39 tests.
2026-08-17 07:54:47 -04:00
ScottW514 86ce0419e5 forgetest: cloud tests stay in cloud mode; the page can ignore prerequisites
The cloud job tests (lid-abort, lid-during-button-wait, hunt-lid-open,
pause-resume) run in cloud mode and leave the machine there: enter_cloud
reuses a live session (pid-scoped from the client's own websocket state
lines) and switches once from GRBL mode, declaring the change to the
baseline; nothing switches back. Each test judges the log from its own
window, the print by its own "print [id]: finished" line, and waits the
service's follow-up moves out (wait_quiet) before it ends. hunt-lid-open
restarts the cloud client through the supervisor's stop/start lever for
a fresh connect. The former switch-back is what failed the last bench run
of hunt-lid-open (409 machine is not idle: the service was still
re-finding the head after the lid closed).

The baseline is mode-aware: in cloud mode the client owns the GRBL
controller's init values, the lid lamp, and the position counters; the
mode itself is preserved unless the run declared the change
(Context.mode_changed); controller_mode is never handed back as a bare
setting (a bare write left the persisted mode out of step with the live
one). The undeclared-change restore now uses the captured state, which
the post pass never saw before.

The acceptance page gets an "Ignore prerequisites" switch (remembered by
the browser): POST /start {ignore_requires} starts a test whose requires
are unmet, and the run records the unmet prerequisites in its evidence
and log; they stay required for the release. The cloud tests' requires
no longer chain through cloud.mode-switch.

Proof: tests/test_cloud_suite.py replays the four tests on the bench's
own gfcloud excerpts (fixtures/) and the run loop's emitted pause/resume
lines under the real runner Context against a fake forgectrl; baseline
mode tests and the server override test; the whole suite (98) and the
coverage lint pass. Catalog consequence: cloud.* fingerprints move with
the module; the catalog hash moves with the requires.
2026-08-17 06:42:48 -04:00
ScottW514 e908db7f3a cloud.* lid/button tests: judge only the print's own window
cloud.lid-during-button-wait failed on a healthy machine: it counted
every "starting run" since the session start, and a cloud session runs
the connect-time hunt and several service moves before the print. The
checks now use the print's own window: runs between "waiting for
button" and the print's ":cancelled" (none allowed); the print's run is
the "starting run" after its button wait (lid-abort, pause-resume); the
lid edge timed against the stop is the last edge before the stop line
(an earlier open to place the scrap is not the one); hunt-lid-open
judges the hunt's own terminal line and refusals before it (service
moves after the hunt are rightly refused with the lid open).
2026-08-16 21:41:10 -04:00
ScottW514 7bf8e3d4ed forgefirm-app: bump to 37854f0 (0.1.8+git) - the park clears the ring first; cloud.lid-abort checks the kernel counters
cloud.lid-abort now proves the park with the machine's own counters
(cloud clears them at every job start, so a completed park reads back
at (0,0)); stale ring bytes replayed ahead of the park would not.
2026-08-16 20:59:54 -04:00
ScottW514 c255b265c5 Acceptance: cloud lid-abort, lid-during-button-wait, hunt-lid-open, pause-resume
Four cloud-mode catalog tests driven from the app with the operator,
proven from the client's log (the same record the wire gets), forgectrl
/status and /cool/status, and the kernel latch readback:

- cloud.lid-abort (live): the lid edge reaches the controlled stop
  within 60 ms, the head parks with the lid still open, the latch
  relocks, the armed window closes, the print ends ':cancelled';
- cloud.lid-during-button-wait (operator): the lid at the white-button
  prompt relocks and cancels; no run starts;
- cloud.hunt-lid-open (operator): the connect-time hunt runs and
  completes with the lid open;
- cloud.pause-resume (live): the button pauses (stop + backtrack) and
  resumes; the job completes and parks; nothing relocks or cancels.

Shared helpers enter/leave cloud mode the way cloud.mode-switch does
and return the head afterward.
2026-08-16 19:11:23 -04:00
ScottW514 a80d7aabfb forgetest: cloud tests prove the session from the client's log and hand the head back
cloud.mode-switch took the optional connect-time firmware probe file as
the evidence of a live session and failed on a bench where that check is
off; the evidence is now gfcloud's own authenticate/ws-connect lines in
the unified log after the switch, the probe recorded when present. Cloud
mode's connect clears the kernel position counters at the head's start
and its hunt homes the head to the corner: the test tells the runner the
counters were re-zeroed (Context.counters_rezeroed) and jogs the head
back by the counter-measured displacement, and hands the lid lamp back
at the level it found. cloud.gfhome-homing documents that it leaves the
machine homed at the corner.

Baseline: a displaced head is jogged back along its own path by the
kernel-measured X/Y delta through the GRBL controller (bounded 100 mm,
waits out a controller respawn backoff); Z is never touched.
2026-08-16 14:21:54 -04:00
ScottW514 c0f53a865f forgetest: the release acceptance tool and the bench diagnostics page
A stdlib-only daemon on the dev image (HTTP :8090) that runs the acceptance
catalog against the machine from a self-contained page, keeps the append-only
result log under /data/forgetest, and exports the release artifact the gate
reads. Tests declare kind (auto / operator / live), hardware (api / takeover),
coverage globs, prerequisites, and core membership; a test's domain
fingerprint is the hash of the manifest files its globs select plus the
platform and its own implementation, so a PASS stays valid exactly while
nothing it covers changed. Campaign rules: a FAIL ends the campaign, the core
(image health, kernel latch and drills, one live emission witness) is never
inherited, invalidate-all forces a full campaign, no SKIP. Live tests need the
operator acknowledgment and the physical arm press through the controller;
takeover tests stop forgectrl for the duration with a crash-recoverable
marker; the tool never touches the laser latch.

Catalog v1: 24 tests ported from the proven bench drills with their recorded
pass criteria (image, kernel K1-K3 and fire A/B/U, forgectrl API and logs,
motion incl. dead-man, cooling, live laser, camera, update, cloud). The bench
tab lists every scripts/bench tool and runs the board-side ones as
subprocesses (takeover tools wrapped). 44 host unit tests, including the gate
verification fixtures. Installed only by forgefirm-image-dev, with the bench
scripts under /usr/share/forgetest/bench.
2026-08-15 15:57:11 -04:00