mirror of
https://github.com/openglow-org/forgefirm.git
synced 2026-09-27 16:51:12 -07:00
9f7744b4590faca5be6422cc6fef84255c1d90c8
98
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8dc7c0bc56 |
forgetest: the stale boot reference judged at a held uptime, on both sides of the limit
test_stale_preconfig_reference_is_retaken_on_a_fresh_boot read the host's
uptime after boot_reference() had read it to decide, and expected the
decision its own reading implied. On a CI runner the uptime crossed
BOOT_MAX_AGE_S (600 s) between the two readings: boot_reference retook
the reference and the test expected the stale one (forgetest-ci on
|
||
|
|
0eb0a16452 |
forgetest: forgeext's packages/ exemption goes with the packages
OpenGlow's packages left forgeext for repositories of their own, so the
non-behavioral entry ("forgeext", "packages/**") names nothing. The
manifest test checks the shared package workflow's kit action as
non-behavioral in its place (".github/**" covers it). The manifest tests
pass (26); manifest.py is in no fingerprint.
|
||
|
|
d2307a899e |
forgetest: the cooling tests judge the log line their own run wrote
Six checks in five cooling tests read forgectrl's log tail once for the line the engine writes when it acts: the TEC-on line (cooling.tec-drive), the off gate's line and the run-end temperature line (cooling.gate-off), the warm-up release and the floor's off line (cooling.floor-and-warm-up), the critical line's off line (cooling.critical-tier), and the two put-back lines (cooling.fan-duty-readback). The engine writes to the device first and logs after, and the line reaches the file through rsyslog a moment later still: on the bench reference cooling.tec-drive read thermal/tec_on = 1 and then the tail a few milliseconds before "TEC on: coolant 26.3 C over 24.5 C, airflow up" landed, and failed. A tail read also takes a line an earlier run left behind for this run's. Each check now marks forgectrl's log before the action that makes the line (the settings write, the M8, the duty written behind the engine's back) and waits up to 5 s for the line among those written after the mark. Every line was checked against where cool.c writes it: gates_apply for an off gate, the run's end for the temperatures, the warm-up release, fans_verify for a put-back, the TEC policy for the TEC. The helpers are motion's log offset pair, imported inside each test's function as the fire-watch tests already do; the tests that poll the tail already (_tail_wait) are unchanged. Proof: tests/test_cooling_suite.py's fake engine writes its lines to a scratch forgectrl log as well as to the fake tail, and a new case puts an earlier run's off-gate line in the log with the engine writing none this run: it fails as it must, and against the checks before this it passes. 25 cooling cases pass; forgetest 500 OK. On the bench reference, image 20260925183749 with this file mounted: cooling.tec-drive, cooling.gate-off, cooling.floor-and-warm-up, cooling.critical-tier and cooling.fan-duty-readback PASS, each baseline clean. Acceptance: the five tests are the change; their fingerprints move and no other test's does. |
||
|
|
800863fb09 |
forgetest: the cloud tests wait out the service's thinking
wait_quiet took the machine as quiet after 8 s without the start of a motion, a park, a run or a lens homing. Between a lid image and its next move the service is working on the image and the log is silent, and on the bench reference the re-hunt's moves came 9.6 s apart: cloud.mode-switch switched back to GRBL mode in the middle of the re-hunt, twice, the first time canceling a motion at 988 of its 1002 steps. Every line of a service action now counts as activity: the requests, the image uploads, the action ends, the runs, the parks, the motions and the lens homing. The quiet is 30 s. Over 1901 motion, lid-image and hunt requests in the bench reference's logs, half came within 1.6 s of the line before, 99 percent within 12.8 s, and two after more than 30 s (32.8 and 62.2 s). return_head goes: cloud.mode-switch hands its cloud stretch back through the cloud client's log (homeoff.cloud_mode_return), and nothing else called it. RETURN_MAX_MM moves to homeoff, its one user. Proof: tests/test_cloud_suite.py's new case lands a motion, a lid image 0.8 s later and the next move 1.6 s after the motion, with the quiet at 1 s: the quiet comes after the move. With the old activity marks it is declared after the image, before the move. forgetest 499 OK. Acceptance: cloud.mode-switch gates the change on a machine, its re-hunt waited out before the switch back. The fingerprints of every test in suite/cloud.py and of every module that imports from it move: 29 tests, the cloud tests, events.button-telemetry, the exthost tests, homing.cloud-offsets and setup.check-envelope. |
||
|
|
cfef6a1b84 |
forgetest: the camera-home tests hand the head back where they found it
homing.cloud-offsets, setup.check-envelope and cloud.mode-switch let the
service move the head, and ended with it where the service left it: at
the camera home, or, in cloud.mode-switch, under the camera, where the
service's re-hunt had taken it. Each told the baseline the counters had
been re-zeroed at the starting position, so the hand-back saw nothing to
do. No counter reading can say where the head was found: every service
motion zeroes the counters at its start, and the home and every
controller start zero them again.
suite/homeoff.py now holds the helpers for a test that lets the service
move the head:
- session_travel sums a client's own record of each motion's end ("end
positions (x, y, z)", the counters the motion zeroed at its start, in
x8 steps, the one mode a service motion runs at). A motion with no end
on record, or a log rotated under the run, leaves the travel unknown.
- camera_home_return drops the camera home, jogs the head back by the
session's travel plus what the counters read since the home, starts
the controller once more so the counters read zero where the head
began, and tells the baseline. cloud_mode_return does the same for a
stay in cloud mode, from the cloud client's log.
- A travel that cannot be known fails the run and moves nothing. A
hand-back that fails while the test is already failing is logged, and
the test's own failure is the one reported.
- judge_whole_motions fails a homing motion that stopped short: the
position declared after it is false.
cloud.mode-switch imports the helpers inside its function, so no other
cloud test's fingerprint moves. Its cloud stretch no longer goes through
return_head, which read 0/0 after the controller's start and left the
head under the camera.
Proof: tests/test_camera_home_return.py, 15 cases on the machine's own
log lines (a session stopped short, a whole three-motion one, a cut log,
a refused motion, homed and unhomed counters, a restart since the home,
the first failure winning). tests/test_cloud_suite.py's mode-switch fakes
now zero the counters and remove the anchor at a controller start, write
the anchor at the home, and move the counters on a jog: 10 cases, a
failed hand-back that must not hide the test's failure among them.
forgetest 498 OK. On the bench reference, with the driver and the runner
fixes: homing.cloud-offsets, setup.check-envelope and cloud.mode-switch
PASS, each ending with the head where it was found; cloud.mode-switch
jogged its cloud stretch 246.06/139.01 mm back with the counters across
the jog agreeing, and every baseline was clean.
Acceptance: the three tests are the change; their fingerprints move and
no other test's does.
|
||
|
|
ef2390f7d4 |
forgetest: a takeover hands back the cloud client it found
In cloud mode forgectrl's start at the end of a takeover starts a cloud
client, and what that client is comes only from gfcloud's one-start
markers, which the client that starts reads and takes down. The takeover
started it bare: the online client, with the service's connect-time
hunt. Seen on the bench with exthost.armed-freeze, which runs its print
under the offline service and puts the setup record back under a
takeover: the machine came back online in cloud mode, the hunt homed the
head at 14:21:01, the run's result was written at 14:21:02, and the
service sent a head move at 14:21:08, after the run had ended, where the
next test's switch to GRBL mode would have cut it off. It also broke the
harness's own rule that every cloud client start it makes comes up
without the hunt.
The takeover now reads, on entry, whether the running cloud client is
the offline service: the client itself holds the listening socket at
/run/gfcloud-offline.sock (hw.listens_on reads /proc/net/unix and the
process's own descriptors, so a socket file another process left behind
does not count). forgectrl's start is made under the no-hunt marker in
cloud mode, and under the offline marker as well when the offline
service was found. After the start the client that came up must listen
on the offline socket within 60 s; a miss is recorded on the run's
baseline capture and the post pass turns it into a leftover ("cloud
client at the takeover end", not restored: a cloud test starts the
client it needs), which fails the run. A marker no client read (the
gate, standby, a fault) is taken down, so a later start never comes up
under it. The judge never breaks the exit path.
Proof: forgetest's unit tests pass on the host (481 OK). The new
tests/test_takeover_client.py drives a real enter and exit against the
fake forgectrl, a fake /proc, and an init script that plays gfcloud's
start: the offline service comes back offline, an online client after
it fails the run, a start that never happened takes the markers down,
an online client comes back without the hunt, a stale socket file is not
the offline service, GRBL mode starts under no marker. Negative
controls: without the markers four of them fail, without the judge
three. Bench drill on the bench reference with this package from /tmp:
GRBL mode, the offline service started as enter_offline starts it (pid
4834), a takeover; the client forgectrl brought back (pid 4898) read
both markers, logged "OFFLINE service", listened on the socket, and in
the 10 s after had no web session, no hunt, and no motion. The bench was
then handed back to GRBL mode as found.
Acceptance: runner.py, baseline.py and hw.py are in no test's
fingerprint, so no result moves; the catalog test that exercises the
change is exthost.armed-freeze, whose takeover now fails the run unless
the offline service comes back.
|
||
|
|
b7dd23d539 |
forgetest: a takeover judges where the head stands
forgectrl's start at the end of a takeover is a controller start, which
re-zeroes the step counters wherever the head then stands. After it a
head left out reads as home, so the baseline's position check after the
run could not see a head a test left out before its takeover.
setup.check-envelope did exactly that on 2026-09-23 and passed (fixed by
|
||
|
|
308a033146 |
laser.verdict-cut judges the hold by the gated output; laser.disarm-in-hold presses at the arm
Both tests failed for the operator on the bench reference on 2026-09-24, image 20260923232513, from the harness and not the machine. laser.verdict-cut (00:10:52) needed the kernel's sampled LASER_ON count to read 0 within 2.5 s of the hold. That count latches once a second over the second before, so it reads 0 only once a whole window has closed inside the hold, up to 2 s in, and this hold lasted 1.37 s. Every sample of the hold read bit 0 of interlock_circuit as 1: the gated LASER_ON, active low, dark. The dark judge now reads the gated output itself, cnc/laser_on (1 = on), about 3300 times a second on the bench reference, yielding the CPU between reads (the stream threads run SCHED_FIFO). Its witness is the same reads over the second before the pause, which must see the cut lit (10 reads or more, none failed). The hold is dark when every read that falls wholly inside it, from 0.3 s after the first Hold:0 to the last Hold row, reads off, none failed, and there are at least 500. The 0.3 s is the pause tier's first deceleration, lit on purpose, still playing out of the driver's 200 ms queue and its 10 ms lead when the controller reports Hold:0. The daemon's pause is 3.0 s, as the code already had and the description now says. dark_span() is pure, kept inside the test's function, and tests/test_laser_verdict.py runs it through the function's code object over synthetic trails (8 cases: a dark hold, the lit deceleration inside the drain, emission after it caught, a burst the resume ends, failed reads counted, the drain counted from the first Hold:0, no Hold:0, and a hold too short for the drain). laser.disarm-in-hold (00:06:23, 180 s) pressed the button through ctx.act at once, before the job reached its arm wait, so the press was lost and the move never started. It now uses ctx.arm_press(), which presses when the button lights, and waits for Run. Proof: forgetest's unit tests pass on the host (474 OK, 4 skipped). Only the fingerprints of laser.verdict-cut and laser.disarm-in-hold move. Acceptance: the change is the two catalog tests themselves; both are attended (laser emission, the operator present) and a campaign runs both again. |
||
|
|
dc9bede034 |
forgetest: the first /state parses each suite module once
The first GET /state after forgetest starts hashes every test's implementation, and it took from 81 s to more than 12 minutes on the bench reference. Two causes: - sibling_imports() read and parsed a test's module and every sibling it imports, transitively, once per test: 280 ast.parse calls for 26 suite modules. module_parts() kept its own parse, sibling_imports() kept none. - The page gives up on a poll after 20 s and polls again; the server thread goes on computing. Each new poll started the same cold work beside the first, all of it on the one CPU, so the longer the first took the more copies ran. That is the spread between 81 s and 12 minutes. Each module's text and tree are now read and parsed once and kept, as are its direct sibling imports, and the implementation hash is filled by one thread at a time: a poll that arrives while it is computed waits for it instead of repeating it. catalog.forget(path) drops what is kept about a file, for the unit tests that edit their modules. The hashes do not move: every test's implementation hash and domain fingerprint on the image manifest of 20260923220034 is byte-identical before and after, so no result is invalidated. On the host the cold computation went from 7.00 s to 0.29 s (280 parses to 25). On the bench reference (the file bind-mounted on image 20260923220034, the page open) the first /state answered in 8 to 12 s after a restart, where the same restart earlier in the evening had not answered after 7 minutes. tests/test_responsiveness.py pins both: each suite module parsed at most once while every test's implementation hash is computed, and three threads reading one test's hash compute it once. With the old behavior put back by a patch both fail (280 parses for 26 modules; the hash computed 3 times). forgetest's unit tests 454 OK. forgetest is the dev-only harness, outside the catalog's coverage; its catalog consequence is none, since no fingerprint moves. |
||
|
|
4e47b7fbf1 |
forgetest: exthost.catalog, and the coverage lint no longer allows an advisory away
exthost.catalog (suite/extcat.py, its own module): GET /ext/catalog answers the index the host keeps and the one address it is fetched from. On a scratch root under /tmp, with a throwaway key standing in for the OpenGlow extension key, the machine's own forgeext keeps an index signed with it, and the author key it names for one id makes a package of that id read as community and endorsed, where before it was unverified; the same key on another id counts for nothing. On the machine's own root that index is refused in words, a package handed over as an index is refused by the product gate, and the index kept is left as it was. The relay refuses an id with no such form (400) and one the kept index does not list (404, or 409 with none kept) before anything is fetched, and a refresh from the fixed address keeps OpenGlow's index when one is published there and is 502 in curl's words when none is, the kept index left as it was. Nothing is left in the staging directory. The coverage lint had a gap: coverage_report() let the allowlist's docs/** and **/*.md take out a path the BEHAVIORAL list keeps in every fingerprint, so the four first-run advisory documents were covered by no test and the lint passed. A change to the privacy advisory would have invalidated nothing. A behavioral path is now never allowed away, and setup.advisories-rehash, which accepts every first-run document at its current hash, covers the four. Proof: on the bench reference, with forgectrl 848ccc1 and forgeext a64b933 bind-mounted and the privacy document accepted again at its new hash with the fixture's press, exthost.catalog PASS (the refresh was 502: GitHub answered 404, nothing is published at the address yet), and setup.advisories-rehash PASS with the rest of the campaign. The lint's new unit test reports the uncovered advisory, and the old reading (the override ignored) reports nothing for it, which is the gap. forgetest's unit tests pass (452), and the coverage lint passes with --enforce. |
||
|
|
71b44598eb |
test_cloud_suite: the quiet wait's deadline is no longer the test
The cloud suite's host tests replay a print's log from threads, and the suite's quiet wait gave up after 3 s (QUIET_TIMEOUT_S in setUp). On a loaded host the last replayed lines landed after that, and a test failed with "still running service moves after 3 s" instead of its own finding: about one run of the module in five, a different test each time, alone as well as under the full suite. The deadline is now 15 s. It is a deadline and not a wait: a quiet machine is seen at once, so a passing run is no slower. The one test that waits the deadline out on purpose (a motion that never goes idle) sets 3 s for itself; tearDown puts the module's value back. Proof: before, test_cloud_suite failed 2 of 10 runs by itself, every failure the quiet deadline (test_lid_during_button_wait_on_the_bench_excerpt, test_pause_resume_fails_without_the_retraced_restart, test_mode_switch_fails_when_gfhome_never_saw_the_head_move, test_pause_resume_fails_when_the_kernel_refuses_the_resume). After, 10 of 10. forgetest's whole unit suite: 451 tests OK, 4 skipped, in 492 s against 491 s before. A host-test change: no catalog consequence. |
||
|
|
b7e7c0fb2c |
Manifest: the advisories are behavior, and forgeext's kit and packages are not
forgectrl embeds its advisory documents in the binary, serves them, and records the operator's consent to one by its hash, so an edited advisory is a changed consent. They are Markdown under docs/, which the non-behavioral list takes out of every fingerprint, and that left setup.extensions-consent's covers entry for docs/advisories/extensions.md selecting nothing: the enforced coverage lint fails on it (exit 1 on the dev image's manifest of 20260922225653), and an edited advisory moved no fingerprint at all. A BEHAVIORAL list now keeps forgectrl's docs/advisories/** in, ahead of the non-behavioral one. forgeext's recipe installs the binary and its init script and nothing else, so packages/ (the official packages, which carry their own acceptance artifact), sdk/ (the author's kit), template/ and tools/ are non-behavioral for the image: without that, an edit to the alignment page or the kit would make every exthost test stale on the next image. Proof: test_manifest passes its 25 cases, the new ones among them; the enforced coverage lint on the dev image's manifest of 20260922225653 exits 0 with no empty entry and nothing uncovered, where it exited 1 before. The whole forgetest suite ran its 451 tests; one, test_cloud_suite's test_pause_resume_passes_on_the_machines_lines, errored under the suite's load and passes 16 runs of 16 alone, at HEAD and on this tree alike: it replays a print against timed hooks. |
||
|
|
eb0a1e7838 |
exthost.panel-install: the key goes in through the panel, and the box holds the button
The test added the owner's key by copying the file. It now adds it the way an operator does: POST /ext/key with the machine's button held. Without the button it is 409 and no file lands; a name with a space, a name that is a path, and no key at all are 400; a key that is no key is 409 from the host; with the button held it is added and listed with its id. Removed again, the same archive reads unverified, and removing a key that is not there is 409. The bench actuator's press was 200 ms, the firmware's default, and a request that must reach the machine while the button is down often missed it: the unverified install took three presses on one run and all ten on another. ctx.act() now takes `ms`, the fixture client passes it to the box (the firmware clamps it into 20 to 500), and the two button-held steps ask for 500. Both landed on the first or second press afterward. Proven. The unit suite: 451 tests, 0 undefined names (both stub fixtures take the new argument, and the fixture test pins that the box is asked for the longest press). On the bench reference, image 20260921161446 with the cross-built forgectrl and extension host mounted over the image's: exthost.panel-install PASS, exthost.package-routes PASS, exthost.hold-pause-tier PASS. Against the image's own daemons the install test FAILS, as it should. Acceptance. exthost.panel-install covers forgectrl's src/extpkg.*, src/main.c, src/auth.* and forgeext's src/main.c, src/install.*, src/pkg.*; the runner and the fixture client are harness, outside the suite and outside every fingerprint. |
||
|
|
1a306d5d8d |
forgetest: a campaign's machine is extension-free
An extension package is software the image does not carry, and a result taken beside one is not a result about the image. Two places hold the line. The baseline (_ext_side, in every pre and post pass): a package under the tests' own prefix (org.forgetest.) and an owner key named forgetest-*.pub are what a test made and left behind; they are removed (forgeext remove, the key's file), recorded as restored, and the host stops the service on its next turn. A process that still runs under a pool account after that belongs to the operator's own packages: it is recorded as unrestorable and never touched. An installed package that does not run is nobody's leftover. hw.pool_pids() reads each process's Uid line; hw.ext_packages() lists the package directory. image.health (5b): the extension host is one process (/usr/bin/forgeext run), its start link sorts after forgectrl's and its kill link before it, no package is installed, and nothing runs under a pool account. Proven. test_baseline.py, ExtensionFreeTests, over a stand-in forgeext and a fake package tree: an installed package that does not run leaves nothing; a test's package and key are removed and the operator's package and key stay; a removal that fails says so in the host's own words; running extensions are reported and left alone. The unit suite passes (451 tests, 0 undefined names). The link order image.health asks for is the one the built root filesystems of image 20260921014201 have: S90forgectrl before S91forgeext, K09forgeext before K90forgectrl. Acceptance. image.health gains forgeext's init/** in its covers map; it runs first in every campaign and is the on-image proof of 5b. The baseline is harness, outside the suite and outside every fingerprint. |
||
|
|
619414ed2b |
cloud suite: a print turned away before the button fails at once, with the reason
The cloud client turns a print away before the button wait for four reasons of its own (machine._safe_to_move): the lid or the interlock, a machine that is not idle, the coolant above its start ceiling (THERMAL.max_start_temp), a coolant sensor that reads nothing. It logs the reason and finishes the print ':cancelled' within a millisecond. The four tests that wait for the button looked only for the wait, so on the bench reference, with the loop at 27.4 C against the ceiling of 27, cloud.dark-print sat out its 120 s and said "the print never reached the button wait". wait_button_wait replaces the four waits: it ends on the print's finish line as well, and fails with the client's own lines, for example "the client turned the print away before the button wait: INFO machine:_safe_to_move machine temp is too high, temp: 27.4 (... finished with event ":cancelled")". The ceiling is the client's and stays where it is: the remedy for a warm loop is airflow, a run session for a minute or two. Proven. test_a_print_turned_away_before_the_button_fails_at_once_with_the_reason replays the bench reference's lines and fails in under 30 s with that reason; with the early exit disabled it fails after the full 120 s with the old words. Each reason is a line the pinned cloud library can log, so the phrase check passes. The unit suite passes (422) with no undefined name. On the bench reference, image 20260920211625 with this cloud.py over the image's, cloud.dark-print passes through the new wait with the loop at 25.8 C. Acceptance. cloud.dark-print, cloud.verdict-refuse, cloud.lid-during-button-wait, and every test that starts an offline print go through the new wait. The helper is module text, so every cloud.py test's fingerprint moves; no product behavior changes. |
||
|
|
19ce4d78d2 |
forgetest: the built-in extensions in setup.cloud-disabled-surface
forgectrl gains a table of built-in extensions, with cloud mode as entry one, GET /extensions to serve it, and POST /settings asking it which selections point at the cloud, how to refuse them, and what they fall back to. It also gains an example client under examples/, and its test tokens lose the names of clients nobody is building. setup.cloud-disabled-surface is the gate for "nothing points at the cloud while it is off", so it now holds the list to the settings three times: as found (enabled as cloud_enabled says, the two roles with their providers and fallbacks, each active exactly when its setting selects it), with the cloud off (not enabled, no role active), and as restored. The two refusals are held to the table's words. Its covers name src/builtin.*. The helper lives inside the test's own function, so no other test of the module changes its fingerprint. The unit test's fake daemon serves the route the way builtin.c does, and gains a case with two lists that lie (enabled after the cloud went off; a role that stays active), each of which must fail the test. forgectrl.tokens: the jog token is "forgetest jogger". manifest: forgectrl's examples/** joins the non-behavioral paths. They are clients that run on another computer: outside the image, outside every fingerprint, and outside the coverage lint. Proven. The unit suite passes (421) with no undefined name; the two lying lists fail the test as they should. On the bench reference, forgectrl's registry daemon hot-deployed over image 20260920152153, on a machine with cloud mode on: setup.cloud-disabled-surface PASS through the whole path (off, the sweep, the refusals, the restore), and forgectrl.tokens PASS with the renamed token. Acceptance. setup.cloud-disabled-surface is the gate for forgectrl's built-in table through the settings route. |
||
|
|
888b47abe8 |
forgetest: the hand-back never moves the head across a lost counter frame
Seen on the bench reference, twice in one session: at the end of a passing run the hand-back jogged the head 30 mm into the back-left stop blocks, from a head that had not moved. The baseline compares the kernel's step counters at the end of a run with the start and jogs the head back by the difference. The GRBL controller zeroes those counters at every start (the lens's startup reference) and at every home (home_completed()), and rewrites its anchor, /run/grblhal.homed, each time. Across either event the difference between two counter readings is not a distance the head traveled. It stayed hidden because the counters normally read zero between tests. homing.manual broke that: a manual home zeroes the counters 30 mm out from where the test began, so after the test's own correct return they read -6400, and the next test that restarts the controller (setup.check-flow-verify, then motion.release) ended at 0, "expected -6400", and was "returned" by 30 mm. An operator who jogs the head from a Grbl client and then starts any takeover test from the page would have met the same thing, by whatever distance they had jogged. The baseline's capture() now records the counters' frame, the anchor's inode and mtime. If the frame changed during the run and the test did not vouch for the new one, the hand-back logs that the two readings share no frame, and moves nothing. ctx.counters_rezeroed() is how a test vouches: it now sets rezero_declared beside the position it expects. The start_reads argument it briefly took is gone, since it made the baseline accept counters that nothing after it could live with. homing.manual restarts the controller once more after returning the head, so it ends with the counters at zero where it began, and declares that. events.stream waits for its three places. A stream an earlier test closed keeps its place until the daemon's next write to it (its keep-alive), as documented, so run straight after forgectrl.lease the third stream drew 503. The test now opens the three once they can be opened, and says so in its log. Proven. test_baseline gains test_a_lost_counter_frame_never_moves_the_head: counters at -6400, a new anchor, counters at 0: no jog, no leftover, and the log says why; the same counters with the frame intact are still a displaced head; a declared re-zero is held to the position it declared. With the frame check unable to see the change (the first cut of the test reused an inode inside one clock tick) the case fails with ['position'], which is the old behavior. The unit suite passes. On the bench reference, arranged so a failure would move the head away from the stop: the head jogged to +60 mm (counters 12800), motion.release run, PASS, "the controller re-zeroed its counters during the run ... the head is not moved", and nothing moved. forgectrl.lease then events.stream: two logged waits, PASS. homing.manual then setup.check-flow-verify, the sequence that drove the head into the stop: both PASS with a clean hand-back. |
||
|
|
f306c9983f |
forgetest: the hand-back reads an engine hold again past the engine's next tick
The cooling engine publishes its state once a tick (1 Hz). A /cool/status
read inside the tick after a run ended still shows the run: while a
diagnostic owns the hardware every tick publishes phase "diag" with the
hold set, and a fail tier's hold stands until the tick that ends its
session. The baseline took one read, and by its rule an arm or a hold is
the run's doing, so a test that finished inside that second failed its
hand-back on a hold the engine's next tick cleared.
Seen on the bench reference twice. cooling.aa-offset-calibrate in campaign
c-20260919215024-c402: the diagnostic reported done at 22:07:14 with its
offset measured (15.7 counts, spread 0.7), and the hand-back at 22:07:15
read "cool=diag/armed=False/hold=True ... -> waited" and failed the run;
the test had passed on five images before, the last one earlier the same
day. cooling.fail-tier-stop in c-20260919202934-3d3a, the same way on the
crash fault's hold (
|
||
|
|
dc7170d0d1 |
forgetest: cooling.flow-under-load finds the verdict by offset, not by a tail count
The test counted "heater rise" in the last 300 lines of the forgectrl log before the job, then waited for the count to grow. A tail of fixed length cannot show that: the new verdict line comes in at the bottom as an old one leaves at the top, the count does not move, and a check that verified reads as one that never judged. Seen on the bench reference in campaign c-20260919210959-8780: the engine logged "coolant flow verified (heater rise 11.5 C, dT 9.5 C; laser 1.6 off 13.1)" at 21:37:46, 70 s into the job and inside the wait, and the test failed at 21:39:21 with "the engine published no flow verdict within 120 s". The 21:17:27 verdict of an earlier test sat near the top of the 300-line tail when the test began. Replayed against that log, the old method reads 4 before the job and 4 with the new verdict in the tail; the search from the byte offset finds the 21:37:46 line. The test now takes the log's byte offset before the job (_log_offset, as the fail-tier and liveness tests do) and searches what the file gained since with the verdict expression, stopping only on a line that matches. _log_since reads a file that is now shorter than the offset from its start: a rotation under the test leaves only newer lines. The /logs/tail helper and its line count are gone with their one user. Proven: tests/test_cooling_suite.py FlowVerdictLogTests - the text after the offset alone, the real verdict line through the expression, an old verdict before the offset not taken for the new one, a rotated log read whole, a missing log read as nothing; the cooling host tests pass under Linux, 24 tests. No component source changed, so no pin moves. |
||
|
|
d482e76402 |
installer: a download that resumes and retries, and an install log
A field install failed on "firmware download failed", and worked after a reboot. The download was one bare curl -fL: no retry, no resume, no bound on a stalled transfer, and nothing on the machine recorded what had gone wrong. The download. download_fw makes up to five tries, 5, 15, 30 and 60 seconds apart. Each try resumes the partial file (curl -C -) and is bounded: 20 s to connect, and a transfer below 1 KB/s for 30 s ends the try. The file is written as forgefirm.fw.part and takes its name only when curl finished; the signature check that follows is what vouches for its content. A full disk (curl 23) and a release that is not there (HTTP 404) end the tries at once, because waiting cannot fix them. A partial file the server will not resume (curl 33 or 36, HTTP 416) starts over. The loop is the installer's own rather than curl --retry: the factory curl on the bench reference is 7.69.1, whose --retry does not count a resolver failure or a dropped transfer as retryable, and older factory builds carry older curls. The owner sees the reason in words with each retry, and the final failure says that a re-run goes straight to the download, because the archives are kept. The log. Every run appends to /data/log/forgefirm/install/install.log, in the log tree's own line format (UTC, program "install"): the installer's md5 (which revision ran), the factory version and the slots, the owner's answers, each archive, each download try with curl's exit code, the HTTP code and the reason, the machine's clock at each try (a wrong clock breaks TLS), and after a failed try the address, the default route, the resolver and whether github.com resolves; then the signature and identity checks, the write, the boot selection, and the reason for any failure through die(). The log is appended across runs, so the run that failed is still there after the run that worked. Logging never fails the install. forgectrl's log export carries the directory (forgectrl 0dae758). Proven: tests/test_installer.py runs the installer's own functions under sh against a scripted curl - a clean download, a resolver failure and a dropped transfer that resume to the full file, the tries running out, 404 and a full disk ending them at once, a stale partial file starting over, the TLS reason naming the clock, every log line in the tree format, die() leaving its reason, and an unwritable log not failing the run. The whole host suite, 409 tests, passes under Linux and the coverage lint is clean. Bench: the same functions under the factory firmware's own shell (busybox 1.31.1 ash, the factory slot of the bench reference in a chroot) resumed, retried, ran out of tries and logged exactly as under sh. Acceptance: logs.tree-tail-export now plants a probe file in the install directory and requires it back in the export bundle, its line intact and its MAC and IPv4 address redacted, and requires an install log in the bundle when the machine has one. The installer itself is not on the image: the install page fetches it from master, so it is live with this push. |
||
|
|
8232c8c9fe |
forgetest: image.network-boot - the boot does not wait on the network
A client that waits in ifup's foreground on a server's answer holds the whole machine, because init starts the rest of the boot - sshd, forgectrl, the console login - only after S01networking returns. A field machine sat there forever on a router that refused DHCPv6 (the record is in the meta-openglow commit that drops the DHCPv6 client). Nothing in the catalog looked at the network boot path; this test does. It asserts: wlan0 is in ifupdown's state file and no ifup is running; the console getty is up; udhcpc runs with -b (it leaves ifup after three unanswered discovers) and has been reparented to init; no DHCPv6 client is named in /etc/network/interfaces, running, or on the image, and its hook script is gone; IPv6 on wlan0 is the kernel's own - enabled, router advertisements accepted, a link-local address up. A global address is evidence only: a network whose router advertisement offers no SLAAC prefix gives none. What a hostile server does to a client is a bench drill, not a test. covers is empty by design, as with setup.machine-name: the interfaces file is layer content, in the platform identity of every fingerprint, so a change there already makes every test necessary again. Proven: the host tests for the parsers and the registration (tests/test_image.py), and the whole host suite, 398 tests, under Linux. The test's logic, run read-only on the bench reference against an image that carries the client, fails exactly the four DHCPv6 checks and passes every other one. The test fails on any image built from a meta-openglow that still carries the client, so the kas lock moves to the layer head that drops it. |
||
|
|
cd4c176a87 |
Finish the x32 xy_microsteps default in the baseline test and the stream harness
The x32 default landed in forgetest/baseline.py and the driver, but two callers still judged the machine at x8 and both failed on the host. forgetest/tests/test_baseline.py: setUp seeded the fake machine from the x8 FIXED_SYSFS literals while enforce() compares against fixed_sysfs() of the resolved mode, so x_mode, y_mode, step_freq and ramp_rate read as deviations on a clean machine - 23 failures across BaselineTests and TransientNotLeftoverTests. It seeds from fixed_sysfs() now, and the tick expectations come from it (DEFAULT_TICK) rather than a typed 28160. The xy_mode_of and ref_xy_mode unset/invalid cases expect 32, with an explicit "8" case added that had no coverage. Two reference_preconfig dumps taken on an x8 machine carry xy_microsteps = 8, because the markers are read at the reference's own mode. wait_configured wrote the static CONFIGURED_MARKERS where the function watches configured_markers() of the mode in force, and the held-controller jog typed 221 steps for "4.144 mm", which is 1.036 mm at x32; both derive from the mode now. 69 tests, all pass. scripts/bench/laser_stream_test.py: STEPS_PER_MM was the x8 53.333, so the X-peak check failed at 2133 steps against an expected 533. The whole harness now derives from XY_MICROSTEPS_BASE/DEFAULT the way glowforge.h and baseline.py do, which uncovered five more x8-only expectations behind the first: the machine tick, the fire-gap limit (it grows as sqrt(k), not k - a finer mode shortens the accel interval by sqrt(k) while speeding the tick by k), the rung split in fire_spans, the density period and minimum burst (laser_pulse_ticks is in x8 ticks and the stream scales it, so the config keeps the x8 numbers and the measured lengths scale), and the decel/hold budgets. Run against the null-sink build: all stream emission rules hold. xy_mode_test.py's docstring still described the no-key case as x8 while its own assertions had moved to x32. No behavior change and no acceptance-catalog consequence: these are test expectations and a bench harness, not image component sources. The coverage lint is unchanged at 0 uncovered paths. |
||
|
|
f9f4de31c1 | Fixed artifact exporter | ||
|
|
0eb764bf75 | Added SPDX | ||
|
|
db6015dc81 |
forgetest: cloud.mode-switch opens the lid behind the controller's start, ahead of the hunt
The supervisor holds every controller spawn until the enclosure is closed (forgectrl 0.1.25), and the test opened the lid before it asked for the cloud controller: POST /mode answered "waiting, the lid is open" and the controller never came up. The round trip now switches with the lid closed, polls /mode five times a second, and opens the lid the moment the controller is running. The client requests its connect-time hunt a few seconds after its start, right behind its session, so the hunt still finds the lid open. The order is recorded and judged: the hunt's request line must not be in the client's log when the lid reads open (hunt_before_lid_open), and the test refuses to start with the lid open. The catalog text tells the operator to open the lid at once, with a hand ready on it. Proof. Host: test_cloud_suite drives the round trip with the hunt landing only once the lid reads open, as on the bench, plus the lost race (the hunt requested before the lid opened fails the test with "before the lid was open") and the start with the lid open refused; 8 mode-switch cases green. Bench reference (dev image 20260915001814, forgectrl 0.1.25): cloud.mode-switch PASS in 118 s, the lid open 4 s before the client requested its hunt, no refusal before the hunt's end, the lens homed, the exhaust row unjudged, 5 service motions after the lid closed, $H under gfhome homed in 48.4 s with 9 motion windows. No catalog consequence beyond the test itself: its covers map is unchanged. |
||
|
|
f080d5d9cb |
Catalog: the cloud client's latch at the run, and a hold that never clears
The cloud client changes (python3-gfhardware: the latch unlocks at the run and nowhere earlier, the warm-up is supervised, a live feed must land and finish, fire needs a power byte first, the homing runner always stops). This commit carries the catalog tests that hold the part the bench can see. forgetest/forgetest/suite/cloud.py: - cloud.dark-print (new, kind operator, one press, dark by construction) replaces cloud.verdict-hold. The print arms on the press, waits for the engine's acknowledgment and runs: the laser latch is locked at the button and through the wait, unlocked only for the run (immediately before it starts), locked again when the job ends, and the print completes. The engine's own warm-up release stays proven by cooling.floor-and-warm-up. - cloud.verdict-refuse (new, kind operator, one press, dark): the start gate far above the coolant and cloud_hold_max_s at its minimum keep the armed print under the warm-up past the bound. The client waits with the latch locked (sampled every two seconds), cancels at the bound with its own log line, never runs, closes the armed window, and the print ends ':cancelled'. The settings are restored. - cloud.verdict-hold is retired: its release rode the loop heater at the flow-check duty against a gate one degree above the coolant, inside the upstream reading's noise band, and its run hovered for minutes with the engine's "warm-up stalled" line in the log. The two tests above prove the client's contract without the thermal race. - cloud.oversize-stream already reads cnc/streaming back at both ends of the run, which is the readback the client now insists on. A forced streaming write failure has no seam on the board (the write goes to sysfs as root) and stays a host test. tests/test_cloud_suite.py follows: excerpts and failure cases for the two new tests in place of the retired one's. Proof. Host: the forgetest unit tests. Bench reference: cloud.dark-print passed (locked at the button, unlocked in the run, locked after, ':completed'), cloud.verdict-refuse passed (held 60 s, 31 latch samples all locked, ':cancelled', settings restored); motion.deadman's kill during $H ended the runner on SIGTERM alone; and the retired cloud.verdict-hold passed once more on the new client before it went (the latch locked through an 8 minute warm-up hold, the release ran the print to completion). |
||
|
|
2f2a4160af |
Stream and motion robustness: harness rules, catalog tests, the hand-back's counter scale
The GRBL driver's stream engine and its motion envelope change (grblHAL-glowforge: the shipper writes outside the lock, a clamp inside an armed window faults, the X/Y soft limits follow the home, the machine's settings are pinned, the homing keys are clamped). This commit carries the host rules and the catalog tests that hold them; the driver commit follows, because its CI fetches these harnesses unpinned. scripts/bench/laser_stream_test.py: - Rule 28: a 300 ms producer stall while armed faults the stream with ALARM:17, the kernel sees no step burst (at most the planned steps per 100 ticks), the stream ends dark, the latch sideband ends on the lock. The same stall unarmed is a warning: the move completes with every step, the clamp visible as the burst the kernel counts. - Rule 29: a 300 ms stall of the sink's write leaves the producer on pace: no clamp, every step, lit through, dark at the end. - The stand-in engine takes GFSINK_STALL_MS and GFSINK_WRITE_STALL_MS, null-sink only. scripts/bench/z_envelope_test.py: - Rule 10: homed (a gfcloud home), a program move past X max, Y max or the near edge alarms with ALARM:2 before any motion, a jog past the bed is refused with error 15, a move inside the bed runs, and a $20 write keeps the limits. The core repeats the last error for the line after a refused jog until an empty line clears it, so the rule sends one. forgetest/forgetest/suite/motion.py: - motion.soft-limits (kind auto, no emission): homes through the cloud suite's gfhome homing when the machine is not homed, then the three refusals (ALARM:2, the kernel counters still), the refused jog, the inside move, and the return to the corner read at rest. - motion.deadman phase 2b: a 100 ms SIGSTOP mid-move, inside the kernel's queue: no underrun, the controller's log warns of the clamped late events, the move completes with every step (read at rest), the latch stays locked. The armed clamp is proven on the host (rule 28). forgetest/forgetest/suite/cloud.py: - gfhome_homing drains the driver's answer to $H once the session ends: it sits behind the status reports and passed for the reply to the caller's next command (a setting read as None). forgetest/forgetest/baseline.py: - The hand-back reads the position counters at the kernel's own microstep mode (cnc/x_mode, read before the sysfs restore puts the settings' mode back). At the x8 constant, an x32 machine's 30 mm read as 120 mm, beyond the return bound, and the displaced head was left in place. The dead band scales the same way. tests/test_baseline.py holds both. Proof. Host: stream rules 1 to 29 and z_envelope rules 1 to 10 against the null-sink driver, the forgetest unit tests. Bench reference: motion.soft-limits passed (X 495 and Y 279 refused with ALARM:2 and 0.000 mm of kernel motion, X-1 the same, the jog error:15, the inside move ran, back at the corner 0.0/0.0 mm); motion.deadman passed with the new phase (92 late events clamped, max behind 92.3 ms, no underrun, 30.0 mm counted, latch locked) and the hand-back jogged the head back under the x32 scale. |
||
|
|
7c0412075c |
Normalize line endings to LF
.gitattributes sets text=auto with eol=lf, so every text file is stored and checked out with LF, and a patch keeps its bytes. The files that carried CRLF from a Windows editor are renormalized. No content changes. |
||
|
|
d8b8adaef4 |
forgetest: the setup suite, and the sensors check asks nothing
The commission.* tests are setup.* (subsystem "setup"): suite/setup.py, setup_dark.py, setup_sheet.py, and the host tests test_setup_*.py. The coverage maps name src/setup.* in place of src/commission.*, the record is setup.json, and the bench seed in forgetest.init creates /run/forgefirm/setup-override. setup.check-sensors follows the check as it is now: no question is asked and no setting is written. The test reads the settings before and after and fails on any change, and fails at once if the check opens a prompt. The host test drives the fake daemon with no prompt and proves both outcomes. Proof: test_setup_dark.py, test_setup_sheet.py, and test_setup_suite.py pass (46 tests). The coverage lint names src/commission.c and .h as uncovered until the forgectrl pin moves to the revision that carries the rename. |
||
|
|
7b8f72b632 |
Run the commissioning sheet as a fresh machine, and pin forgectrl 0.1.18 (the lens frame)
forgectrl f8ddb17 puts the lens frame in one place and writes the focus window before the card's controller starts, after two commissioning cards ended in ALARM:2 on a Z the wizard sent from one source while the controller's Z limit stood on another. The acceptance run had passed only because the bench's settings already held the stops from an earlier focus run, so the run is now a fresh machine's. commission.sheet clears the three lens settings inside its Restore before the cards, checks that the frame runs in the fallback window, and, after the focus card, that the settings hold the window the ladder ran in (the stops found or the fallback), that every ladder height lies in that window's reach, and that the program served now agrees with /status. Every served program's Z is checked against the reach /status reports before the card starts, so a stray Z fails the test with nothing burned. The test covers src/lens.*. The host mock in tests/test_commission_sheet.py mirrors the daemon (the /status lens block from the settings, the ladder served from the settings, the window written at the focus start); two regression tests reproduce the defects: a focus result naming a window the settings do not hold, and a served program with a Z beyond the reach. 10/10 green. scripts/bench/z_envelope_test.py, the grblHAL CI harness, adds the referenced-lens cases: with forgectrl's marker and the shared settings, the fallback window and a 14/20 window run to the ends of their reach and two half-steps past either end alarms, and a count of 41 falls back on its side alone. Passed on the null-sink build. The bench page's description of the harness follows. The forgectrl pin moves to f8ddb17 (0.1.18); every test covering forgectrl re-runs. |
||
|
|
b94c990bbc |
Leave the lens to the cloud client in cloud mode
cloud.verdict-hold printed to completion and then failed its hand-back on head/z_mode=0. That job's header carried ZSmd 0, gfhardware's set_mode_from_puls turns that into z_mode 0, and z_mode 0 is full step (lenshome.c): the cloud client had set the lens the way the pulse file it was playing asked. Nothing was left behind - the machine did what the job said. The baseline already skips nine attributes in cloud mode, for the reason written above the list: the cloud client sets its own values for them from every pulse header, and forcing the GRBL values under it would be the baseline configuring another controller's machine. head/z_current and head/z_mode are the same thing and were not on the list. They are now, in cloud mode only; in GRBL mode the pair is still checked against 1 and 1, which is the half of this an exemption could quietly break, so a test holds it. Host-proven: 89 baseline unit tests. |
||
|
|
8d8ca2a8c6 |
Judge a hand-back on what a run left, not on the machine still working
A leftover is dirt a run left behind. Four of the checks were reporting
something else: the machine part-way through work of its own, which it
finishes on its own a few seconds later. Each one failed a run that had
done nothing wrong.
laser.arm-wait-lid found it. The test passed every check it makes and
failed on cool=smoke/armed=False/hold=False. The smoke clear is the
engine's post-job work - run duty for smoke_s after an armed session,
timed, ending by itself (cool.c, Cool_Smoke) - and the machine was idle
fifteen seconds later without being asked. Armed and hold were both
false: nothing had been left.
The four, and what each was catching:
cool a phase alone (the smoke clear, a cooldown). Armed or
holding stays a leftover: the run left a job alive and the
engine is keeping the fans up for it.
controller the supervisor's own start. A takeover ends by starting
forgectrl again, and the respawn runs the liveness probe
and the lens reference before it reports running and
verified. The run had put it back.
state the ring draining to the end of a job. An underrun stays a
leftover and is still acknowledged with cnc/stop: that one
is the run's.
leds read_led read brightness, write_led writes target, and the
smooth trigger fades brightness toward target. An LED the
machine had already released still read lit mid-fade, and
the restore called itself done before the fade had moved.
Judged on target now: a run that left the button lit left
a target standing.
The waiting, the restoring and the reporting of anything a run actually
left are unchanged. What changed is which readings count as dirt.
Host-proven: 87 baseline unit tests, two of them new - a fading button
LED is not a leftover, a lit one is.
|
||
|
|
d3fe1d90b9 |
Hand the machine back, do not describe what is wrong with it
The hand-back is the promise that a run leaves the machine where it found it. Two of its checks reported instead of restoring, and the machine sat in the state the run left it in. The cooling engine: a run that ends without ending its job leaves one alive behind it - the engine armed and holding for a job that is never coming back, the fans at run duty. The check waited two minutes for that to resolve itself, which it cannot, and wrote "failed: still run/armed=True/hold=True". On the bench reference the fans then ran for an hour. The pulse ring: bytes the last job never played sit there, and the next run replays them before its own. The check refused the return jog and wrote "clear the ring (controller restart) before moving" - the instruction, to a log, instead of the action. Both end the same way: stop the controller and let the supervisor bring it back. The job goes, the arm and the hold go with it, and the ring is empty on the way in. stand_down() does that and proves it settled; the cooling check calls it when the engine will not idle on its own, and the return jog calls it when the ring has residue, refusing only if bytes survive a restart. A leftover still reports what it found - the record of what the run did is the point - but it reports having fixed it. Host-proven: 62 baseline unit tests, including one that the hand-back calls the stand-down for ring residue rather than describing it. |
||
|
|
594b6990fd |
Stop a finished run's clock
Run.snapshot() computed elapsed_s from the current time on every call, whether the run had ended or not. The page shows the last run until the next one starts, and it polls, so a finished run's figure went on counting: the result badge said FAIL beside a number still climbing, and the run read as still going. On the bench reference a test that ended after 5082 s was showing 5242 s and rising. The figure a finished run should carry was already recorded next to it: finished["duration_s"], fixed when the result was written. snapshot() now returns that once the run has ended, and the live count only while it is running. Host-proven: two unit tests - a finished run's clock reads its duration and does not move across a poll, a running one's still climbs. |
||
|
|
b2c42f0cc3 |
Do not fail a hand-back on the step the counters round to
The post-run pass compared the kernel position counters exactly, so a test that put the head back within a hundredth of a millimeter failed whenever that distance did not round to the same step. motion.lid-cancel-home found it on the bench reference. The test cancels a job with the lid twice and the controller returns the head to the job start each time, landing 0.038 and 0.037 mm out - the same figures as the run before it, which passed. This time the two returns left four steps on X, 0.075 mm at 53.333 steps per mm, and the hand-back called it a leftover. Twenty-three runs of that test, all passing, on returns of the same accuracy: whether the residue rounds to zero is chance, not a property of the machine or of the test. A difference inside POSITION_DEADBAND_MM (0.1 mm, five steps at x8) on X and Y is now the quantization rather than a leftover. The head is still put back, so nothing accumulates over a campaign - only the failure goes. Z stays exact: the return never moves the lens, so a Z difference is still a leftover and still unrestorable. Host-proven: four new unit tests on the boundary (four steps and a moved Z on either side of it, and an unreadable reading), and the forgetest suite. |
||
|
|
1dd6ad608b |
Press the button when the machine asks, not when its LED is on
The sheet handed the bench actuator every press after the operator's presence press, and the actuator pressed every time - into nothing. On the bench reference all six cards were pressed before the card asked: the frame by 8 s, the focus by 33, the floor by 10, the dose curve by 8, the corner by 9, and the flow-load card by 69. The operator then pressed all six himself, which is the opposite of what the ready gate promises. arm_press() waits for hw.button_lit(), which is true when any button LED is on, and burn() started that wait before run_check had even started the wizard. The button is lit through parts of a card that are not the arm - the lens reference, the program on its way to the controller - so the wait ended at once, the press landed before the job waited for it, and the thread was gone by the time the real cue came. forgectrl uses the same predicate but only inside the job's own sample callback, with the tube still dark, where a lit button does mean the arm. The machine already says when it wants the press: a live check opens a `press` wait prompt at that moment, and run_check sees every prompt. It now presses there, through a new Ctx.press_now() - no LED read, no waiting thread, no timing guess. That retires the per-card lit-timeout column of CARDS, which existed only to give the flow-load card's coolant settle enough room for a wait that was reading the wrong thing. The four grbl-driven arm_press() callers in laser.py and cooling.py are left as they are: they call it with the machine idle and its LEDs dark, so the level read is the edge they mean. The same shape would bite them if that ever stopped being true. Pin: forgectrl 0.1.15 (6e69479, the flow-load tail ending at the coolant peak and the way out of a finished setup page). Host-proven: 357 forgetest unit tests, including two new ones - the actuator presses on the prompt with button_lit stubbed false throughout, so the LED is provably not consulted, and the press falls to the operator without a takeover. |
||
|
|
f0c40e7d4f |
Name every machine after its own MAC address, and drop mDNS
One name for every machine was wrong: an operator with two of them on a network had one forgefirm.local, and mDNS does not work on many networks at all. The machine now calls itself forgefirm-<xxxx>, from the last four hex digits of its WiFi MAC address, and sends that name with its DHCP request, so a network with dynamic DNS publishes it and a router lists the machine by name. The name is the same at every boot, two machines take different names, and no serial number leaves the machine. forgefirm-hostname (new): reads the wlan0 MAC address (eth0 on a machine with no WiFi) at S38 in rcS, after udev has probed the network drivers and before poky's hostname.sh reads the file and before the network starts. The rootfs is read-only, so the name is written through a bind-mounted copy under /run/forgefirm. A bounded wait covers a slow probe. hostname:pn-base-files is "forgefirm": the name before S38, and the fallback when no MAC address can be read. avahi is deleted - the bbappend, the daemon configuration, the service file, the image install and the distro block. The address is the way in that works on every network, and the DHCP name covers the rest. forgefirm-banner: the marker lines are gone. "# ForgeFIRM addresses" and "# end" delimited the address block inside /etc/issue, and getty prints every line of that file, so both markers were on the console. The script now keeps the image's own text in a second copy under /run/forgefirm, captured once per boot before the first write, and renders the whole banner from it. The block is the addresses alone: no mDNS name. forgefirm-image.bb: the ForgeFIRM mark, under the OpenGlow one the base image carries, with the version on the mark's own last line, right-justified to the mark's last column. The mark is written once and rendered per reader, because /etc/issue is parsed by busybox getty (a backslash or a percent sign starts an escape, so the art goes in with every backslash doubled) while /etc/motd is written out as it is. Widths are measured in columns, not bytes: the color sequences take no room on the screen. /etc/issue.net stays unused - the machine tells a client that has not logged in nothing. Acceptance: commission.mdns-announce is replaced by commission.machine-name, which checks the name against the MAC address, the bind-mounted /etc/hostname, the DHCP client's hostname option, the banner's addresses, and that no mDNS responder is on the image; it covers nothing by design, like the test it replaces. forgectrl.auth gains the own-name Host check and its refusal with a domain on it. image.health checks the /etc/hostname mount and the version on the mark's last line in both files. commission.ssh-until-reboot asserts there is no pre-authentication banner. commission_dark's lens coverage widens to src/lenshome.* so src/lenshome.h is covered; the lint is clean at 83 tests. Pins: forgectrl 0.1.14 (9e5330f, the hostname certificate and the Host rule), meta-openglow ced2af2 (the DHCP hostname option and the motd mark) in the kas lock. Proven on the bench reference, hot-deployed and rebooted (image 20260910000208 dev): hostname forgefirm-b00a from MAC 2c:6b:7d:0d:b0:0a, live and in the bind-mounted file; the DHCP client running with -x hostname:forgefirm-b00a; the console banner and the motd carrying both marks with the version aligned to the mark's last column, no marker line and no .local name; forgectrl regenerating its certificate for the new name. Host tests: 357 forgetest unit tests, forgectrl clean under -Werror, tls_test and sanitize_test. |
||
|
|
468e92655a |
Teach the manifest guard test about the version file
The guard keeps scripts/manifest-from-tree.py and forgefirm-image-manifest.bbclass computing the same identity, so it pins the skip list by value and the new suffix failed it. It now expects the version file in both places, and a second case proves the point the change exists for: setting the release number leaves the layer hash where it was. My miss: I proved the identity by hand and pushed without running the unit suite first. |
||
|
|
8fc5250d6d |
XY microstep modes: the baseline, the catalog test and the bench tools
The baseline derives x/y_mode, step_freq, ramp_rate and the configured markers from the xy_microsteps setting; ramp_rate joins the GRBL controller's set (the driver writes it; the cloud client runs at the module's). motion.microstep-modes cycles 8, 16 and 32: the save restarts the idle controller, the kernel reads the mode with its tick and ramp, $100/$101 are the mode's and a typed $100 is overwritten, a 40 mm jog at top speed returns to Idle with the kernel counters over the mode agreeing with the commanded travel and the accelerometer seeing the head move; the setting is put back as found. Bench tools: xy_mode_test.py (the null-sink harness the grblHAL CI runs), raster_dry.py (a top-speed raster per mode), xy_pattern_accel.py (the operator's pattern from home with the machine silent and the head accelerometer listening), arc_tolerance_sweep.py (a $12 ladder on the 9 in circle: the chord rate the protocol loop feeds, about 300 a second, is the ceiling, not the core), and the xymode and xycircle live drills. BRINGUP carries the present state and the facts; CAMPAIGN-LOG the dated record, including the planner-blocks spin (a $398 of 255 or more loops forever at start, a core bug) and its recovery. |
||
|
|
b0fa4ccaf5 |
A release publishes the source of the software it installs
The release build merges kas/source-bundle.yml, which turns on the Yocto archiver: the upstream source of each recipe as upstream publishes it, the patches with their series file, and the recipe with its includes. The overlay adds tasks only, so the image manifest is unchanged and an acceptance result still applies; proven on the build host, where the archiver build and a plain rebuild of the same tree give the same content_sha256. scripts/source-bundle.py packs forgefirm-source-v<version>.tar.gz: the archives, both license manifests, the license texts, the ForgeFIRM layers, the kas configuration, the layer revisions and the build identity of the image. What the bundle must hold comes from the image, not from a list in the script: every recipe of license.manifest and image_license.manifest whose license is in the include list must have an archive, or the release stops with the recipe named. release.sh attaches the bundle and covers it with sha256sums.txt; FORGEFIRM_SOURCE_SKIP=1 bypasses deliberately. On the build host: 68 of the image's 111 recipes carry source, 313.7 MiB, under the 2 GiB limit of a release asset. No acceptance catalog consequence: the change is release tooling on the build host and puts no file and no behavior on the machine. The host-side proof is forgetest/tests/test_source_bundle.py, which holds the license decision, the choice of archive and the refusal. |
||
|
|
a86d66049f | forgetest: the invalidate-all notice ends when a campaign starts after it; its epoch stays | ||
|
|
0df0162c8d | forgetest: the sheet test covers the runner's header; the hollow generator entry is dropped | ||
|
|
97287aa6a9 |
commissioning: the layer, the acceptance tests, the harness rule, the docs, and the bench drills
meta-forgefirm: the forgefirm-users init replays the account at boot; sshd refuses root and empty passwords and runs only while the panel turns it on; the release image keeps an empty root password for the console; the console banner; avahi announces forgefirm.local; https in libmicrohttpd and ulfius; the panel on 80 and 443; the license bundle on the rootfs; release.sh checks the root policy on the built rootfs. forgetest: the commission suites (commission, commission_dark, commission_sheet: 23 cases); the runner turns cloud mode on with the typed phrase for a test that declares it; the baseline's motor_lock is 0; the log-export test checks the bundle for the camera key; the record helpers write bytes as given and join the daemon's paths as POSIX. The stream harness gains rule 24: a hold verdict is held again after a resume. Bench drills: lens_travel.py and lens_stop_accel.py. Docs: BRINGUP carries the present state; CAMPAIGN-LOG carries the dated record. |
||
|
|
d6f648b539 |
forgetest: presence is proved once per test, not once per ready gate
The ready gate lives inside the arm-and-fire helper, and the kill drill calls that helper twice, once for the expected stop and once for the SIGKILL. So the presence gate asked the operator for a second press part way through a test they had already proved themselves present for, with the actuator standing by holding the presses. It is the only test in the catalog with two ready gates. The second gate now returns at once. Its setup line still goes up, because the second half may want the scrap moved, but there is no press to make. |
||
|
|
ba0bf41749 |
forgetest: the operator proves presence at the machine, the bench presses
A live test asked a person to click Ready on a page and then make timing-critical presses in the middle of a burning cut. That is how tonight's pause test became unanswerable: the actuator had dropped off the network, the harness fell back to the operator without saying so, and afterwards nobody could tell a second press from the machine resuming on its own. Where an actuator is up and wired to the button, the ready gate now takes a press on the machine's own button as the presence check, and the actuator performs every press in that test. The button does nothing at Idle, so the press is only a presence check, and the gate waits for the release so it is never read as the arm press. With no actuator the operator does the presses and answers on the page, as before. An actuator lost after that takeover is now said out loud, in the log and in the evidence, instead of quietly becoming a person's press. The live-fire cue was four lines of machine-shaped prose. It is now what a person needs: protection, exhaust, extinguisher, scrap, lid. |
||
|
|
9258dea885 |
forgetest: fans-quiet also proves a daemon restart on a busy machine returns to idle airflow
A forgectrl started while the kernel is not idle takes the cooldown airflow (forgectrl's busy-start rule), and on the bench it kept it: after kernel.fire-line's takeover restarted the daemon with the kernel in the drill's safe state, the exhaust ran at 6200 rpm on an idle machine until the daemon was restarted by hand. cooling.fans-quiet-after-motion gains the case: forgectrl stopped, cnc/disable written, forgectrl started, and within 90 s the controller must be running with the idle duties applied. The host replay stubs the init script and holds both outcomes: the idle duties after the start, and a daemon that keeps the cooldown duties. Bench: with forgectrl 522cdb2, the busy start logged, idle airflow one tick later, the duties idle 15 s after the start, PASS. Catalog consequence: the cooling.* implementation hashes move. |
||
|
|
970f10a9e2 |
forgetest: the dwell-gap latch rule judges the hardware's unlocked window, the live runs keep a trail
laser.emission-witness required the hardware button latch clear in every sample the engine reported armed, and, after a first fix, in every sample up to the last nonzero emission count. Both windows were drawn from lagging signals: the engine's armed flag follows the controller's next report, and the emission counter latches once per second and reads nonzero about two seconds past the relock. Both reached into the tail where the job-end relock sets the button latch by design, and the rule refused three clean runs on image 20260902144848 (all four sides burned; the trail shows the latch clear from the press to the relock, emission through the fourth side, HV_ENABLE's dip in the dwell and its return). The rule now uses the window the hardware defines: from the first emission, in every sample whose readback word shows the laser latch unlocked, the button-latch bit of that same word must be clear. That spans the kernel-run gap of the dwell and ends at the relock, and no lagging flag can misplace it. dwell_gap() is a pure function; tests/test_laser_dwell.py holds the relocked tail, a set inside the gap, and a trail without emission. The recorded trail of the third run replays to a pass (47 unlocked samples, none set). The live runs keep a per-sample trail in the evidence (TRAIL_FIELDS: the readback word, the switches, the lock flag, the controller's state and messages), so a run's timeline can be read back without a rerun. A fourth run then errored on a name the refactor had removed and one later check still used; py_compile does not catch it and a live drill never executes on the host, so the CI job now fails on any undefined name in the harness (pyflakes). Catalog consequence: the laser implementation hashes move. |
||
|
|
7605a90946 |
forgetest: the update drill reads the daemon's reply, and a queue start waits for the fixture probe
update.slots-and-signature's apply section required 200 from
POST /update/apply, where the daemon answers 202 with started, like
every job endpoint, so its first bench run on image 20260902144848
ended before the job did; the cleanup then deleted the staged archive
under the running job. The drill requires 202 and started, and looks
for the daemon's refusal ("archive is not signed with the ForgeFIRM
release key"). Bench: the apply started, the job ended with that
refusal, PASS.
A queue started 4 s after a forgetest restart ran 7 tests instead of
10. The bench page's /state poll had a fixture probe in flight (an mDNS
answer), probe_fixture stamped its time at its start, and the queue
start read the stale fixture, none, so the three operator tests the
fixture runs in the unattended queue were routed to nobody. The probe
now runs under a lock and is stamped when it completes: a caller that
arrives during a probe waits for its answer. tests/test_fixture.py
holds the race with a slow scripted probe; it fails on the old code.
Catalog consequence: the update implementation hash moves; the runner
change is dev-only.
|
||
|
|
1fef9c6f51 |
forgetest: the air-assist offset calibration runs before the heater tools
A queue runs the catalog in registration order among tests with the same prerequisites, and cooling.aa-offset-calibrate followed cooling.flow-verify. A flow-verify trial heats the tube water (a no-flow trial by 17 C on the bench) and the warm slug circulates past the coolant sensors for minutes afterward; the calibration's stationary gate passed 44 s after the trial on image 20260902144848 and the edges read the wave as disagreement. The calibration is registered first now, with the reason beside it, and tests/test_cooling_order.py holds the order. Catalog consequence: the cooling.* implementation hashes move (the suite file changed). |
||
|
|
b3efab9c46 |
forgetest: the latch-unlock gate waits for the safety chain to release
The gate that refuses a latch unlock while the chain may hold HV_ENABLE up (charge_pump_alive or a pulse engine not idle) ran at the start of phases B, U and K3 of kernel.fire-line, within a second of the previous phase's run. A run feeds the charge-pump watchdog every 200 ms and the one-shot holds ALIVE for 0.45 s after the last feed, so the gate read alive=1 and refused: the first bench run of the gate (forgefirm 64f552fc; the laser_pgood gate before it was vacuous) failed phase B on image 20260902144848. wait_hv_off() polls the chain for up to 3 s before it refuses, logs the release when it was not immediate and records every wait in the evidence (hv_release_s). require_hv_off and check_hv_off use it. The bench scripts that copy the gate (fire_test.py per phase, gate_a_kernel_drills.py K3 after K2) get the same wait. Bench: kernel.fire-line PASS on 20260902144848 with the chain released after 0.41 s at each of the three phase boundaries. Host: tests/test_kernel_suite.py covers release inside the window, a chain held past it, and a chain already off. Catalog consequence: the kernel.* implementation hashes move (the suite file changed); the kernel set re-ran and passed. |