An extension package is software the image does not carry, and a result
taken beside one is not a result about the image. Two places hold the line.
The baseline (_ext_side, in every pre and post pass): a package under the
tests' own prefix (org.forgetest.) and an owner key named forgetest-*.pub
are what a test made and left behind; they are removed (forgeext remove,
the key's file), recorded as restored, and the host stops the service on
its next turn. A process that still runs under a pool account after that
belongs to the operator's own packages: it is recorded as unrestorable and
never touched. An installed package that does not run is nobody's
leftover. hw.pool_pids() reads each process's Uid line; hw.ext_packages()
lists the package directory.
image.health (5b): the extension host is one process (/usr/bin/forgeext
run), its start link sorts after forgectrl's and its kill link before it,
no package is installed, and nothing runs under a pool account.
Proven. test_baseline.py, ExtensionFreeTests, over a stand-in forgeext and
a fake package tree: an installed package that does not run leaves
nothing; a test's package and key are removed and the operator's package
and key stay; a removal that fails says so in the host's own words;
running extensions are reported and left alone. The unit suite passes (451
tests, 0 undefined names). The link order image.health asks for is the one
the built root filesystems of image 20260921014201 have: S90forgectrl
before S91forgeext, K09forgeext before K90forgectrl.
Acceptance. image.health gains forgeext's init/** in its covers map; it
runs first in every campaign and is the on-image proof of 5b. The
baseline is harness, outside the suite and outside every fingerprint.
The cloud client turns a print away before the button wait for four
reasons of its own (machine._safe_to_move): the lid or the interlock, a
machine that is not idle, the coolant above its start ceiling
(THERMAL.max_start_temp), a coolant sensor that reads nothing. It logs the
reason and finishes the print ':cancelled' within a millisecond. The four
tests that wait for the button looked only for the wait, so on the bench
reference, with the loop at 27.4 C against the ceiling of 27,
cloud.dark-print sat out its 120 s and said "the print never reached the
button wait".
wait_button_wait replaces the four waits: it ends on the print's finish
line as well, and fails with the client's own lines, for example "the
client turned the print away before the button wait: INFO
machine:_safe_to_move machine temp is too high, temp: 27.4 (... finished
with event ":cancelled")". The ceiling is the client's and stays where it
is: the remedy for a warm loop is airflow, a run session for a minute or
two.
Proven. test_a_print_turned_away_before_the_button_fails_at_once_with_the_reason
replays the bench reference's lines and fails in under 30 s with that
reason; with the early exit disabled it fails after the full 120 s with the
old words. Each reason is a line the pinned cloud library can log, so the
phrase check passes. The unit suite passes (422) with no undefined name. On
the bench reference, image 20260920211625 with this cloud.py over the
image's, cloud.dark-print passes through the new wait with the loop at
25.8 C.
Acceptance. cloud.dark-print, cloud.verdict-refuse,
cloud.lid-during-button-wait, and every test that starts an offline print
go through the new wait. The helper is module text, so every cloud.py
test's fingerprint moves; no product behavior changes.
forgectrl gains a table of built-in extensions, with cloud mode as entry
one, GET /extensions to serve it, and POST /settings asking it which
selections point at the cloud, how to refuse them, and what they fall back
to. It also gains an example client under examples/, and its test tokens
lose the names of clients nobody is building.
setup.cloud-disabled-surface is the gate for "nothing points at the cloud
while it is off", so it now holds the list to the settings three times: as
found (enabled as cloud_enabled says, the two roles with their providers
and fallbacks, each active exactly when its setting selects it), with the
cloud off (not enabled, no role active), and as restored. The two refusals
are held to the table's words. Its covers name src/builtin.*. The helper
lives inside the test's own function, so no other test of the module
changes its fingerprint.
The unit test's fake daemon serves the route the way builtin.c does, and
gains a case with two lists that lie (enabled after the cloud went off; a
role that stays active), each of which must fail the test.
forgectrl.tokens: the jog token is "forgetest jogger".
manifest: forgectrl's examples/** joins the non-behavioral paths. They are
clients that run on another computer: outside the image, outside every
fingerprint, and outside the coverage lint.
Proven. The unit suite passes (421) with no undefined name; the two lying
lists fail the test as they should. On the bench reference, forgectrl's
registry daemon hot-deployed over image 20260920152153, on a machine with
cloud mode on: setup.cloud-disabled-surface PASS through the whole path
(off, the sweep, the refusals, the restore), and forgectrl.tokens PASS with
the renamed token.
Acceptance. setup.cloud-disabled-surface is the gate for forgectrl's
built-in table through the settings route.
Seen on the bench reference, twice in one session: at the end of a passing
run the hand-back jogged the head 30 mm into the back-left stop blocks,
from a head that had not moved.
The baseline compares the kernel's step counters at the end of a run with
the start and jogs the head back by the difference. The GRBL controller
zeroes those counters at every start (the lens's startup reference) and at
every home (home_completed()), and rewrites its anchor, /run/grblhal.homed,
each time. Across either event the difference between two counter readings
is not a distance the head traveled. It stayed hidden because the counters
normally read zero between tests. homing.manual broke that: a manual home
zeroes the counters 30 mm out from where the test began, so after the
test's own correct return they read -6400, and the next test that restarts
the controller (setup.check-flow-verify, then motion.release) ended at 0,
"expected -6400", and was "returned" by 30 mm. An operator who jogs the
head from a Grbl client and then starts any takeover test from the page
would have met the same thing, by whatever distance they had jogged.
The baseline's capture() now records the counters' frame, the anchor's
inode and mtime. If the frame changed during the run and the test did not
vouch for the new one, the hand-back logs that the two readings share no
frame, and moves nothing. ctx.counters_rezeroed() is how a test vouches: it
now sets rezero_declared beside the position it expects. The start_reads
argument it briefly took is gone, since it made the baseline accept
counters that nothing after it could live with.
homing.manual restarts the controller once more after returning the head,
so it ends with the counters at zero where it began, and declares that.
events.stream waits for its three places. A stream an earlier test closed
keeps its place until the daemon's next write to it (its keep-alive), as
documented, so run straight after forgectrl.lease the third stream drew
503. The test now opens the three once they can be opened, and says so in
its log.
Proven. test_baseline gains test_a_lost_counter_frame_never_moves_the_head:
counters at -6400, a new anchor, counters at 0: no jog, no leftover, and the
log says why; the same counters with the frame intact are still a displaced
head; a declared re-zero is held to the position it declared. With the
frame check unable to see the change (the first cut of the test reused an
inode inside one clock tick) the case fails with ['position'], which is the
old behavior. The unit suite passes. On the bench reference, arranged so a
failure would move the head away from the stop: the head jogged to +60 mm
(counters 12800), motion.release run, PASS, "the controller re-zeroed its
counters during the run ... the head is not moved", and nothing moved.
forgectrl.lease then events.stream: two logged waits, PASS. homing.manual
then setup.check-flow-verify, the sequence that drove the head into the
stop: both PASS with a clean hand-back.
The cooling engine publishes its state once a tick (1 Hz). A /cool/status
read inside the tick after a run ended still shows the run: while a
diagnostic owns the hardware every tick publishes phase "diag" with the
hold set, and a fail tier's hold stands until the tick that ends its
session. The baseline took one read, and by its rule an arm or a hold is
the run's doing, so a test that finished inside that second failed its
hand-back on a hold the engine's next tick cleared.
Seen on the bench reference twice. cooling.aa-offset-calibrate in campaign
c-20260919215024-c402: the diagnostic reported done at 22:07:14 with its
offset measured (15.7 counts, spread 0.7), and the hand-back at 22:07:15
read "cool=diag/armed=False/hold=True ... -> waited" and failed the run;
the test had passed on five images before, the last one earlier the same
day. cooling.fail-tier-stop in c-20260919202934-3d3a, the same way on the
crash fault's hold (0ceb4ab made that test wait for its own fault; this is
the general case).
The cooling check moves into Baseline._cool_side. An arm or a hold is read
again for up to COOL_PUBLISH_S (2.5 s: two ticks and a margin) before it is
called the run's doing. One the next tick cleared is logged as the engine's
last word on the run and judged on what the engine reads then: idle is
clean, a cooldown phase is waited out as the engine's own post-job work. A
hold a run did leave stands for a job that is never coming back, so it is
still there after the tick and is recorded, stood down and failed exactly
as before; a hold that clears only later is still a leftover ("waited").
Proven: tests/test_baseline.py CoolPublishTests - a diagnostic's hold the
next tick clears, a fail tier's hold clearing into the smoke phase, a hold
that outlives the tick (recorded, machine stood down), a hold that clears
only later (recorded as waited), an idle engine and a silent daemon; the
baseline, queue, operator, mode, responsiveness and server host tests
pass under Linux. baseline.py is not a suite module, so no test's
fingerprint moves with it.
The test counted "heater rise" in the last 300 lines of the forgectrl log
before the job, then waited for the count to grow. A tail of fixed length
cannot show that: the new verdict line comes in at the bottom as an old one
leaves at the top, the count does not move, and a check that verified reads
as one that never judged.
Seen on the bench reference in campaign c-20260919210959-8780: the engine
logged "coolant flow verified (heater rise 11.5 C, dT 9.5 C; laser 1.6 off
13.1)" at 21:37:46, 70 s into the job and inside the wait, and the test
failed at 21:39:21 with "the engine published no flow verdict within 120 s".
The 21:17:27 verdict of an earlier test sat near the top of the 300-line
tail when the test began. Replayed against that log, the old method reads 4
before the job and 4 with the new verdict in the tail; the search from the
byte offset finds the 21:37:46 line.
The test now takes the log's byte offset before the job (_log_offset, as
the fail-tier and liveness tests do) and searches what the file gained
since with the verdict expression, stopping only on a line that matches.
_log_since reads a file that is now shorter than the offset from its start:
a rotation under the test leaves only newer lines. The /logs/tail helper
and its line count are gone with their one user.
Proven: tests/test_cooling_suite.py FlowVerdictLogTests - the text after the
offset alone, the real verdict line through the expression, an old verdict
before the offset not taken for the new one, a rotated log read whole, a
missing log read as nothing; the cooling host tests pass under Linux, 24
tests. No component source changed, so no pin moves.
A field install failed on "firmware download failed", and worked after a
reboot. The download was one bare curl -fL: no retry, no resume, no bound
on a stalled transfer, and nothing on the machine recorded what had gone
wrong.
The download. download_fw makes up to five tries, 5, 15, 30 and 60 seconds
apart. Each try resumes the partial file (curl -C -) and is bounded: 20 s
to connect, and a transfer below 1 KB/s for 30 s ends the try. The file is
written as forgefirm.fw.part and takes its name only when curl finished;
the signature check that follows is what vouches for its content. A full
disk (curl 23) and a release that is not there (HTTP 404) end the tries at
once, because waiting cannot fix them. A partial file the server will not
resume (curl 33 or 36, HTTP 416) starts over. The loop is the installer's
own rather than curl --retry: the factory curl on the bench reference is
7.69.1, whose --retry does not count a resolver failure or a dropped
transfer as retryable, and older factory builds carry older curls. The
owner sees the reason in words with each retry, and the final failure says
that a re-run goes straight to the download, because the archives are kept.
The log. Every run appends to /data/log/forgefirm/install/install.log, in
the log tree's own line format (UTC, program "install"): the installer's
md5 (which revision ran), the factory version and the slots, the owner's
answers, each archive, each download try with curl's exit code, the HTTP
code and the reason, the machine's clock at each try (a wrong clock breaks
TLS), and after a failed try the address, the default route, the resolver
and whether github.com resolves; then the signature and identity checks,
the write, the boot selection, and the reason for any failure through
die(). The log is appended across runs, so the run that failed is still
there after the run that worked. Logging never fails the install.
forgectrl's log export carries the directory (forgectrl 0dae758).
Proven: tests/test_installer.py runs the installer's own functions under
sh against a scripted curl - a clean download, a resolver failure and a
dropped transfer that resume to the full file, the tries running out, 404
and a full disk ending them at once, a stale partial file starting over,
the TLS reason naming the clock, every log line in the tree format, die()
leaving its reason, and an unwritable log not failing the run. The whole
host suite, 409 tests, passes under Linux and the coverage lint is clean.
Bench: the same functions under the factory firmware's own shell (busybox
1.31.1 ash, the factory slot of the bench reference in a chroot) resumed,
retried, ran out of tries and logged exactly as under sh.
Acceptance: logs.tree-tail-export now plants a probe file in the install
directory and requires it back in the export bundle, its line intact and
its MAC and IPv4 address redacted, and requires an install log in the
bundle when the machine has one. The installer itself is not on the image:
the install page fetches it from master, so it is live with this push.
A client that waits in ifup's foreground on a server's answer holds the
whole machine, because init starts the rest of the boot - sshd, forgectrl,
the console login - only after S01networking returns. A field machine sat
there forever on a router that refused DHCPv6 (the record is in the
meta-openglow commit that drops the DHCPv6 client). Nothing in the catalog
looked at the network boot path; this test does.
It asserts: wlan0 is in ifupdown's state file and no ifup is running; the
console getty is up; udhcpc runs with -b (it leaves ifup after three
unanswered discovers) and has been reparented to init; no DHCPv6 client is
named in /etc/network/interfaces, running, or on the image, and its hook
script is gone; IPv6 on wlan0 is the kernel's own - enabled, router
advertisements accepted, a link-local address up. A global address is
evidence only: a network whose router advertisement offers no SLAAC prefix
gives none. What a hostile server does to a client is a bench drill, not a
test.
covers is empty by design, as with setup.machine-name: the interfaces file
is layer content, in the platform identity of every fingerprint, so a
change there already makes every test necessary again.
Proven: the host tests for the parsers and the registration
(tests/test_image.py), and the whole host suite, 398 tests, under Linux.
The test's logic, run read-only on the bench reference against an image
that carries the client, fails exactly the four DHCPv6 checks and passes
every other one. The test fails on any image built from a meta-openglow
that still carries the client, so the kas lock moves to the layer head
that drops it.
The x32 default landed in forgetest/baseline.py and the driver, but two
callers still judged the machine at x8 and both failed on the host.
forgetest/tests/test_baseline.py: setUp seeded the fake machine from the x8
FIXED_SYSFS literals while enforce() compares against fixed_sysfs() of the
resolved mode, so x_mode, y_mode, step_freq and ramp_rate read as deviations
on a clean machine - 23 failures across BaselineTests and
TransientNotLeftoverTests. It seeds from fixed_sysfs() now, and the tick
expectations come from it (DEFAULT_TICK) rather than a typed 28160. The
xy_mode_of and ref_xy_mode unset/invalid cases expect 32, with an explicit
"8" case added that had no coverage. Two reference_preconfig dumps taken on
an x8 machine carry xy_microsteps = 8, because the markers are read at the
reference's own mode. wait_configured wrote the static CONFIGURED_MARKERS
where the function watches configured_markers() of the mode in force, and
the held-controller jog typed 221 steps for "4.144 mm", which is 1.036 mm at
x32; both derive from the mode now. 69 tests, all pass.
scripts/bench/laser_stream_test.py: STEPS_PER_MM was the x8 53.333, so the
X-peak check failed at 2133 steps against an expected 533. The whole harness
now derives from XY_MICROSTEPS_BASE/DEFAULT the way glowforge.h and
baseline.py do, which uncovered five more x8-only expectations behind the
first: the machine tick, the fire-gap limit (it grows as sqrt(k), not k - a
finer mode shortens the accel interval by sqrt(k) while speeding the tick by
k), the rung split in fire_spans, the density period and minimum burst
(laser_pulse_ticks is in x8 ticks and the stream scales it, so the config
keeps the x8 numbers and the measured lengths scale), and the decel/hold
budgets. Run against the null-sink build: all stream emission rules hold.
xy_mode_test.py's docstring still described the no-key case as x8 while its
own assertions had moved to x32.
No behavior change and no acceptance-catalog consequence: these are test
expectations and a bench harness, not image component sources. The coverage
lint is unchanged at 0 uncovered paths.
The supervisor holds every controller spawn until the enclosure is
closed (forgectrl 0.1.25), and the test opened the lid before it asked
for the cloud controller: POST /mode answered "waiting, the lid is
open" and the controller never came up. The round trip now switches
with the lid closed, polls /mode five times a second, and opens the
lid the moment the controller is running. The client requests its
connect-time hunt a few seconds after its start, right behind its
session, so the hunt still finds the lid open. The order is recorded
and judged: the hunt's request line must not be in the client's log
when the lid reads open (hunt_before_lid_open), and the test refuses
to start with the lid open. The catalog text tells the operator to
open the lid at once, with a hand ready on it.
Proof. Host: test_cloud_suite drives the round trip with the hunt
landing only once the lid reads open, as on the bench, plus the lost
race (the hunt requested before the lid opened fails the test with
"before the lid was open") and the start with the lid open refused;
8 mode-switch cases green. Bench reference (dev image 20260915001814,
forgectrl 0.1.25): cloud.mode-switch PASS in 118 s, the lid open 4 s
before the client requested its hunt, no refusal before the hunt's
end, the lens homed, the exhaust row unjudged, 5 service motions
after the lid closed, $H under gfhome homed in 48.4 s with 9 motion
windows. No catalog consequence beyond the test itself: its covers
map is unchanged.
The cloud client changes (python3-gfhardware: the latch unlocks at the
run and nowhere earlier, the warm-up is supervised, a live feed must
land and finish, fire needs a power byte first, the homing runner
always stops). This commit carries the catalog tests that hold the
part the bench can see.
forgetest/forgetest/suite/cloud.py:
- cloud.dark-print (new, kind operator, one press, dark by
construction) replaces cloud.verdict-hold. The print arms on the
press, waits for the engine's acknowledgment and runs: the laser
latch is locked at the button and through the wait, unlocked only
for the run (immediately before it starts), locked again when the
job ends, and the print completes. The engine's own warm-up release
stays proven by cooling.floor-and-warm-up.
- cloud.verdict-refuse (new, kind operator, one press, dark): the
start gate far above the coolant and cloud_hold_max_s at its minimum
keep the armed print under the warm-up past the bound. The client
waits with the latch locked (sampled every two seconds), cancels at
the bound with its own log line, never runs, closes the armed window,
and the print ends ':cancelled'. The settings are restored.
- cloud.verdict-hold is retired: its release rode the loop heater at
the flow-check duty against a gate one degree above the coolant,
inside the upstream reading's noise band, and its run hovered for
minutes with the engine's "warm-up stalled" line in the log. The two
tests above prove the client's contract without the thermal race.
- cloud.oversize-stream already reads cnc/streaming back at both ends
of the run, which is the readback the client now insists on. A
forced streaming write failure has no seam on the board (the write
goes to sysfs as root) and stays a host test.
tests/test_cloud_suite.py follows: excerpts and failure cases for the
two new tests in place of the retired one's.
Proof. Host: the forgetest unit tests. Bench reference:
cloud.dark-print passed (locked at the button, unlocked in the run,
locked after, ':completed'), cloud.verdict-refuse passed (held 60 s,
31 latch samples all locked, ':cancelled', settings restored);
motion.deadman's kill during $H ended the runner on SIGTERM alone; and
the retired cloud.verdict-hold passed once more on the new client
before it went (the latch locked through an 8 minute warm-up hold, the
release ran the print to completion).
The GRBL driver's stream engine and its motion envelope change
(grblHAL-glowforge: the shipper writes outside the lock, a clamp inside
an armed window faults, the X/Y soft limits follow the home, the
machine's settings are pinned, the homing keys are clamped). This commit
carries the host rules and the catalog tests that hold them; the driver
commit follows, because its CI fetches these harnesses unpinned.
scripts/bench/laser_stream_test.py:
- Rule 28: a 300 ms producer stall while armed faults the stream with
ALARM:17, the kernel sees no step burst (at most the planned steps per
100 ticks), the stream ends dark, the latch sideband ends on the lock.
The same stall unarmed is a warning: the move completes with every
step, the clamp visible as the burst the kernel counts.
- Rule 29: a 300 ms stall of the sink's write leaves the producer on
pace: no clamp, every step, lit through, dark at the end.
- The stand-in engine takes GFSINK_STALL_MS and GFSINK_WRITE_STALL_MS,
null-sink only.
scripts/bench/z_envelope_test.py:
- Rule 10: homed (a gfcloud home), a program move past X max, Y max or
the near edge alarms with ALARM:2 before any motion, a jog past the bed
is refused with error 15, a move inside the bed runs, and a $20 write
keeps the limits. The core repeats the last error for the line after a
refused jog until an empty line clears it, so the rule sends one.
forgetest/forgetest/suite/motion.py:
- motion.soft-limits (kind auto, no emission): homes through the cloud
suite's gfhome homing when the machine is not homed, then the three
refusals (ALARM:2, the kernel counters still), the refused jog, the
inside move, and the return to the corner read at rest.
- motion.deadman phase 2b: a 100 ms SIGSTOP mid-move, inside the
kernel's queue: no underrun, the controller's log warns of the clamped
late events, the move completes with every step (read at rest), the
latch stays locked. The armed clamp is proven on the host (rule 28).
forgetest/forgetest/suite/cloud.py:
- gfhome_homing drains the driver's answer to $H once the session ends:
it sits behind the status reports and passed for the reply to the
caller's next command (a setting read as None).
forgetest/forgetest/baseline.py:
- The hand-back reads the position counters at the kernel's own
microstep mode (cnc/x_mode, read before the sysfs restore puts the
settings' mode back). At the x8 constant, an x32 machine's 30 mm read
as 120 mm, beyond the return bound, and the displaced head was left in
place. The dead band scales the same way. tests/test_baseline.py holds
both.
Proof. Host: stream rules 1 to 29 and z_envelope rules 1 to 10 against
the null-sink driver, the forgetest unit tests. Bench reference:
motion.soft-limits passed (X 495 and Y 279 refused with ALARM:2 and
0.000 mm of kernel motion, X-1 the same, the jog error:15, the inside
move ran, back at the corner 0.0/0.0 mm); motion.deadman passed with
the new phase (92 late events clamped, max behind 92.3 ms, no underrun,
30.0 mm counted, latch locked) and the hand-back jogged the head back
under the x32 scale.
.gitattributes sets text=auto with eol=lf, so every text file is
stored and checked out with LF, and a patch keeps its bytes. The files
that carried CRLF from a Windows editor are renormalized. No content
changes.
The commission.* tests are setup.* (subsystem "setup"): suite/setup.py,
setup_dark.py, setup_sheet.py, and the host tests test_setup_*.py. The
coverage maps name src/setup.* in place of src/commission.*, the record
is setup.json, and the bench seed in forgetest.init creates
/run/forgefirm/setup-override.
setup.check-sensors follows the check as it is now: no question is
asked and no setting is written. The test reads the settings before
and after and fails on any change, and fails at once if the check
opens a prompt. The host test drives the fake daemon with no prompt
and proves both outcomes.
Proof: test_setup_dark.py, test_setup_sheet.py, and
test_setup_suite.py pass (46 tests). The coverage lint names
src/commission.c and .h as uncovered until the forgectrl pin moves to
the revision that carries the rename.
forgectrl f8ddb17 puts the lens frame in one place and writes the focus
window before the card's controller starts, after two commissioning
cards ended in ALARM:2 on a Z the wizard sent from one source while the
controller's Z limit stood on another. The acceptance run had passed
only because the bench's settings already held the stops from an
earlier focus run, so the run is now a fresh machine's.
commission.sheet clears the three lens settings inside its Restore
before the cards, checks that the frame runs in the fallback window,
and, after the focus card, that the settings hold the window the ladder
ran in (the stops found or the fallback), that every ladder height lies
in that window's reach, and that the program served now agrees with
/status. Every served program's Z is checked against the reach /status
reports before the card starts, so a stray Z fails the test with nothing
burned. The test covers src/lens.*. The host mock in
tests/test_commission_sheet.py mirrors the daemon (the /status lens block
from the settings, the ladder served from the settings, the window
written at the focus start); two regression tests reproduce the defects:
a focus result naming a window the settings do not hold, and a served
program with a Z beyond the reach. 10/10 green.
scripts/bench/z_envelope_test.py, the grblHAL CI harness, adds the
referenced-lens cases: with forgectrl's marker and the shared settings,
the fallback window and a 14/20 window run to the ends of their reach
and two half-steps past either end alarms, and a count of 41 falls back
on its side alone. Passed on the null-sink build. The bench page's
description of the harness follows.
The forgectrl pin moves to f8ddb17 (0.1.18); every test covering
forgectrl re-runs.
cloud.verdict-hold printed to completion and then failed its hand-back on
head/z_mode=0. That job's header carried ZSmd 0, gfhardware's
set_mode_from_puls turns that into z_mode 0, and z_mode 0 is full step
(lenshome.c): the cloud client had set the lens the way the pulse file it
was playing asked. Nothing was left behind - the machine did what the job
said.
The baseline already skips nine attributes in cloud mode, for the reason
written above the list: the cloud client sets its own values for them
from every pulse header, and forcing the GRBL values under it would be
the baseline configuring another controller's machine. head/z_current and
head/z_mode are the same thing and were not on the list. They are now, in
cloud mode only; in GRBL mode the pair is still checked against 1 and 1,
which is the half of this an exemption could quietly break, so a test
holds it.
Host-proven: 89 baseline unit tests.
A leftover is dirt a run left behind. Four of the checks were reporting
something else: the machine part-way through work of its own, which it
finishes on its own a few seconds later. Each one failed a run that had
done nothing wrong.
laser.arm-wait-lid found it. The test passed every check it makes and
failed on cool=smoke/armed=False/hold=False. The smoke clear is the
engine's post-job work - run duty for smoke_s after an armed session,
timed, ending by itself (cool.c, Cool_Smoke) - and the machine was idle
fifteen seconds later without being asked. Armed and hold were both
false: nothing had been left.
The four, and what each was catching:
cool a phase alone (the smoke clear, a cooldown). Armed or
holding stays a leftover: the run left a job alive and the
engine is keeping the fans up for it.
controller the supervisor's own start. A takeover ends by starting
forgectrl again, and the respawn runs the liveness probe
and the lens reference before it reports running and
verified. The run had put it back.
state the ring draining to the end of a job. An underrun stays a
leftover and is still acknowledged with cnc/stop: that one
is the run's.
leds read_led read brightness, write_led writes target, and the
smooth trigger fades brightness toward target. An LED the
machine had already released still read lit mid-fade, and
the restore called itself done before the fade had moved.
Judged on target now: a run that left the button lit left
a target standing.
The waiting, the restoring and the reporting of anything a run actually
left are unchanged. What changed is which readings count as dirt.
Host-proven: 87 baseline unit tests, two of them new - a fading button
LED is not a leftover, a lit one is.
The hand-back is the promise that a run leaves the machine where it found
it. Two of its checks reported instead of restoring, and the machine sat
in the state the run left it in.
The cooling engine: a run that ends without ending its job leaves one
alive behind it - the engine armed and holding for a job that is never
coming back, the fans at run duty. The check waited two minutes for that
to resolve itself, which it cannot, and wrote "failed: still
run/armed=True/hold=True". On the bench reference the fans then ran for
an hour.
The pulse ring: bytes the last job never played sit there, and the next
run replays them before its own. The check refused the return jog and
wrote "clear the ring (controller restart) before moving" - the
instruction, to a log, instead of the action.
Both end the same way: stop the controller and let the supervisor bring
it back. The job goes, the arm and the hold go with it, and the ring is
empty on the way in. stand_down() does that and proves it settled;
the cooling check calls it when the engine will not idle on its own, and
the return jog calls it when the ring has residue, refusing only if bytes
survive a restart. A leftover still reports what it found - the record of
what the run did is the point - but it reports having fixed it.
Host-proven: 62 baseline unit tests, including one that the hand-back
calls the stand-down for ring residue rather than describing it.
Run.snapshot() computed elapsed_s from the current time on every call,
whether the run had ended or not. The page shows the last run until the
next one starts, and it polls, so a finished run's figure went on
counting: the result badge said FAIL beside a number still climbing, and
the run read as still going. On the bench reference a test that ended
after 5082 s was showing 5242 s and rising.
The figure a finished run should carry was already recorded next to it:
finished["duration_s"], fixed when the result was written. snapshot()
now returns that once the run has ended, and the live count only while
it is running.
Host-proven: two unit tests - a finished run's clock reads its duration
and does not move across a poll, a running one's still climbs.
The post-run pass compared the kernel position counters exactly, so a
test that put the head back within a hundredth of a millimeter failed
whenever that distance did not round to the same step.
motion.lid-cancel-home found it on the bench reference. The test cancels
a job with the lid twice and the controller returns the head to the job
start each time, landing 0.038 and 0.037 mm out - the same figures as the
run before it, which passed. This time the two returns left four steps on
X, 0.075 mm at 53.333 steps per mm, and the hand-back called it a
leftover. Twenty-three runs of that test, all passing, on returns of the
same accuracy: whether the residue rounds to zero is chance, not a
property of the machine or of the test.
A difference inside POSITION_DEADBAND_MM (0.1 mm, five steps at x8) on X
and Y is now the quantization rather than a leftover. The head is still
put back, so nothing accumulates over a campaign - only the failure goes.
Z stays exact: the return never moves the lens, so a Z difference is
still a leftover and still unrestorable.
Host-proven: four new unit tests on the boundary (four steps and a moved
Z on either side of it, and an unreadable reading), and the forgetest
suite.
The sheet handed the bench actuator every press after the operator's
presence press, and the actuator pressed every time - into nothing. On
the bench reference all six cards were pressed before the card asked:
the frame by 8 s, the focus by 33, the floor by 10, the dose curve by 8,
the corner by 9, and the flow-load card by 69. The operator then pressed
all six himself, which is the opposite of what the ready gate promises.
arm_press() waits for hw.button_lit(), which is true when any button LED
is on, and burn() started that wait before run_check had even started the
wizard. The button is lit through parts of a card that are not the arm -
the lens reference, the program on its way to the controller - so the
wait ended at once, the press landed before the job waited for it, and
the thread was gone by the time the real cue came. forgectrl uses the
same predicate but only inside the job's own sample callback, with the
tube still dark, where a lit button does mean the arm.
The machine already says when it wants the press: a live check opens a
`press` wait prompt at that moment, and run_check sees every prompt. It
now presses there, through a new Ctx.press_now() - no LED read, no
waiting thread, no timing guess. That retires the per-card lit-timeout
column of CARDS, which existed only to give the flow-load card's coolant
settle enough room for a wait that was reading the wrong thing.
The four grbl-driven arm_press() callers in laser.py and cooling.py are
left as they are: they call it with the machine idle and its LEDs dark,
so the level read is the edge they mean. The same shape would bite them
if that ever stopped being true.
Pin: forgectrl 0.1.15 (6e69479, the flow-load tail ending at the coolant
peak and the way out of a finished setup page).
Host-proven: 357 forgetest unit tests, including two new ones - the
actuator presses on the prompt with button_lit stubbed false throughout,
so the LED is provably not consulted, and the press falls to the operator
without a takeover.
One name for every machine was wrong: an operator with two of them on a
network had one forgefirm.local, and mDNS does not work on many networks
at all. The machine now calls itself forgefirm-<xxxx>, from the last four
hex digits of its WiFi MAC address, and sends that name with its DHCP
request, so a network with dynamic DNS publishes it and a router lists
the machine by name. The name is the same at every boot, two machines
take different names, and no serial number leaves the machine.
forgefirm-hostname (new): reads the wlan0 MAC address (eth0 on a machine
with no WiFi) at S38 in rcS, after udev has probed the network drivers
and before poky's hostname.sh reads the file and before the network
starts. The rootfs is read-only, so the name is written through a
bind-mounted copy under /run/forgefirm. A bounded wait covers a slow
probe. hostname:pn-base-files is "forgefirm": the name before S38, and
the fallback when no MAC address can be read.
avahi is deleted - the bbappend, the daemon configuration, the service
file, the image install and the distro block. The address is the way in
that works on every network, and the DHCP name covers the rest.
forgefirm-banner: the marker lines are gone. "# ForgeFIRM addresses" and
"# end" delimited the address block inside /etc/issue, and getty prints
every line of that file, so both markers were on the console. The script
now keeps the image's own text in a second copy under /run/forgefirm,
captured once per boot before the first write, and renders the whole
banner from it. The block is the addresses alone: no mDNS name.
forgefirm-image.bb: the ForgeFIRM mark, under the OpenGlow one the base
image carries, with the version on the mark's own last line,
right-justified to the mark's last column. The mark is written once and
rendered per reader, because /etc/issue is parsed by busybox getty (a
backslash or a percent sign starts an escape, so the art goes in with
every backslash doubled) while /etc/motd is written out as it is. Widths
are measured in columns, not bytes: the color sequences take no room on
the screen. /etc/issue.net stays unused - the machine tells a client that
has not logged in nothing.
Acceptance: commission.mdns-announce is replaced by
commission.machine-name, which checks the name against the MAC address,
the bind-mounted /etc/hostname, the DHCP client's hostname option, the
banner's addresses, and that no mDNS responder is on the image; it covers
nothing by design, like the test it replaces. forgectrl.auth gains the
own-name Host check and its refusal with a domain on it. image.health
checks the /etc/hostname mount and the version on the mark's last line in
both files. commission.ssh-until-reboot asserts there is no
pre-authentication banner. commission_dark's lens coverage widens to
src/lenshome.* so src/lenshome.h is covered; the lint is clean at 83
tests.
Pins: forgectrl 0.1.14 (9e5330f, the hostname certificate and the Host
rule), meta-openglow ced2af2 (the DHCP hostname option and the motd mark)
in the kas lock.
Proven on the bench reference, hot-deployed and rebooted (image
20260910000208 dev): hostname forgefirm-b00a from MAC 2c:6b:7d:0d:b0:0a,
live and in the bind-mounted file; the DHCP client running with
-x hostname:forgefirm-b00a; the console banner and the motd carrying both
marks with the version aligned to the mark's last column, no marker line
and no .local name; forgectrl regenerating its certificate for the new
name. Host tests: 357 forgetest unit tests, forgectrl clean under
-Werror, tls_test and sanitize_test.
The guard keeps scripts/manifest-from-tree.py and forgefirm-image-manifest.bbclass computing the same identity, so it pins the skip list by value and the new suffix failed it. It now expects the version file in both places, and a second case proves the point the change exists for: setting the release number leaves the layer hash where it was. My miss: I proved the identity by hand and pushed without running the unit suite first.
The baseline derives x/y_mode, step_freq, ramp_rate and the configured
markers from the xy_microsteps setting; ramp_rate joins the GRBL
controller's set (the driver writes it; the cloud client runs at the
module's). motion.microstep-modes cycles 8, 16 and 32: the save restarts
the idle controller, the kernel reads the mode with its tick and ramp,
$100/$101 are the mode's and a typed $100 is overwritten, a 40 mm jog at
top speed returns to Idle with the kernel counters over the mode agreeing
with the commanded travel and the accelerometer seeing the head move; the
setting is put back as found.
Bench tools: xy_mode_test.py (the null-sink harness the grblHAL CI runs),
raster_dry.py (a top-speed raster per mode), xy_pattern_accel.py (the
operator's pattern from home with the machine silent and the head
accelerometer listening), arc_tolerance_sweep.py (a $12 ladder on the 9
in circle: the chord rate the protocol loop feeds, about 300 a second, is
the ceiling, not the core), and the xymode and xycircle live drills.
BRINGUP carries the present state and the facts; CAMPAIGN-LOG the dated
record, including the planner-blocks spin (a $398 of 255 or more loops
forever at start, a core bug) and its recovery.
The release build merges kas/source-bundle.yml, which turns on the Yocto
archiver: the upstream source of each recipe as upstream publishes it, the
patches with their series file, and the recipe with its includes. The
overlay adds tasks only, so the image manifest is unchanged and an
acceptance result still applies; proven on the build host, where the
archiver build and a plain rebuild of the same tree give the same
content_sha256.
scripts/source-bundle.py packs forgefirm-source-v<version>.tar.gz: the
archives, both license manifests, the license texts, the ForgeFIRM layers,
the kas configuration, the layer revisions and the build identity of the
image. What the bundle must hold comes from the image, not from a list in
the script: every recipe of license.manifest and image_license.manifest
whose license is in the include list must have an archive, or the release
stops with the recipe named. release.sh attaches the bundle and covers it
with sha256sums.txt; FORGEFIRM_SOURCE_SKIP=1 bypasses deliberately.
On the build host: 68 of the image's 111 recipes carry source, 313.7 MiB,
under the 2 GiB limit of a release asset.
No acceptance catalog consequence: the change is release tooling on the
build host and puts no file and no behavior on the machine. The host-side
proof is forgetest/tests/test_source_bundle.py, which holds the license
decision, the choice of archive and the refusal.
meta-forgefirm: the forgefirm-users init replays the account at boot;
sshd refuses root and empty passwords and runs only while the panel
turns it on; the release image keeps an empty root password for the
console; the console banner; avahi announces forgefirm.local; https in
libmicrohttpd and ulfius; the panel on 80 and 443; the license bundle on
the rootfs; release.sh checks the root policy on the built rootfs.
forgetest: the commission suites (commission, commission_dark,
commission_sheet: 23 cases); the runner turns cloud mode on with the
typed phrase for a test that declares it; the baseline's motor_lock is
0; the log-export test checks the bundle for the camera key; the record
helpers write bytes as given and join the daemon's paths as POSIX. The
stream harness gains rule 24: a hold verdict is held again after a
resume. Bench drills: lens_travel.py and lens_stop_accel.py.
Docs: BRINGUP carries the present state; CAMPAIGN-LOG carries the dated
record.
The ready gate lives inside the arm-and-fire helper, and the kill drill
calls that helper twice, once for the expected stop and once for the
SIGKILL. So the presence gate asked the operator for a second press part
way through a test they had already proved themselves present for, with
the actuator standing by holding the presses. It is the only test in the
catalog with two ready gates.
The second gate now returns at once. Its setup line still goes up,
because the second half may want the scrap moved, but there is no press
to make.
A live test asked a person to click Ready on a page and then make
timing-critical presses in the middle of a burning cut. That is how
tonight's pause test became unanswerable: the actuator had dropped off
the network, the harness fell back to the operator without saying so,
and afterwards nobody could tell a second press from the machine
resuming on its own.
Where an actuator is up and wired to the button, the ready gate now
takes a press on the machine's own button as the presence check, and the
actuator performs every press in that test. The button does nothing at
Idle, so the press is only a presence check, and the gate waits for the
release so it is never read as the arm press. With no actuator the
operator does the presses and answers on the page, as before.
An actuator lost after that takeover is now said out loud, in the log
and in the evidence, instead of quietly becoming a person's press.
The live-fire cue was four lines of machine-shaped prose. It is now what
a person needs: protection, exhaust, extinguisher, scrap, lid.
A forgectrl started while the kernel is not idle takes the cooldown
airflow (forgectrl's busy-start rule), and on the bench it kept it: after
kernel.fire-line's takeover restarted the daemon with the kernel in the
drill's safe state, the exhaust ran at 6200 rpm on an idle machine until
the daemon was restarted by hand. cooling.fans-quiet-after-motion gains
the case: forgectrl stopped, cnc/disable written, forgectrl started, and
within 90 s the controller must be running with the idle duties applied.
The host replay stubs the init script and holds both outcomes: the idle
duties after the start, and a daemon that keeps the cooldown duties.
Bench: with forgectrl 522cdb2, the busy start logged, idle airflow one
tick later, the duties idle 15 s after the start, PASS.
Catalog consequence: the cooling.* implementation hashes move.
laser.emission-witness required the hardware button latch clear in every
sample the engine reported armed, and, after a first fix, in every sample
up to the last nonzero emission count. Both windows were drawn from
lagging signals: the engine's armed flag follows the controller's next
report, and the emission counter latches once per second and reads
nonzero about two seconds past the relock. Both reached into the tail
where the job-end relock sets the button latch by design, and the rule
refused three clean runs on image 20260902144848 (all four sides
burned; the trail shows the latch clear from the press to the relock,
emission through the fourth side, HV_ENABLE's dip in the dwell and its
return).
The rule now uses the window the hardware defines: from the first
emission, in every sample whose readback word shows the laser latch
unlocked, the button-latch bit of that same word must be clear. That
spans the kernel-run gap of the dwell and ends at the relock, and no
lagging flag can misplace it. dwell_gap() is a pure function;
tests/test_laser_dwell.py holds the relocked tail, a set inside the gap,
and a trail without emission. The recorded trail of the third run
replays to a pass (47 unlocked samples, none set).
The live runs keep a per-sample trail in the evidence (TRAIL_FIELDS: the
readback word, the switches, the lock flag, the controller's state and
messages), so a run's timeline can be read back without a rerun.
A fourth run then errored on a name the refactor had removed and one
later check still used; py_compile does not catch it and a live drill
never executes on the host, so the CI job now fails on any undefined
name in the harness (pyflakes).
Catalog consequence: the laser implementation hashes move.
update.slots-and-signature's apply section required 200 from
POST /update/apply, where the daemon answers 202 with started, like
every job endpoint, so its first bench run on image 20260902144848
ended before the job did; the cleanup then deleted the staged archive
under the running job. The drill requires 202 and started, and looks
for the daemon's refusal ("archive is not signed with the ForgeFIRM
release key"). Bench: the apply started, the job ended with that
refusal, PASS.
A queue started 4 s after a forgetest restart ran 7 tests instead of
10. The bench page's /state poll had a fixture probe in flight (an mDNS
answer), probe_fixture stamped its time at its start, and the queue
start read the stale fixture, none, so the three operator tests the
fixture runs in the unattended queue were routed to nobody. The probe
now runs under a lock and is stamped when it completes: a caller that
arrives during a probe waits for its answer. tests/test_fixture.py
holds the race with a slow scripted probe; it fails on the old code.
Catalog consequence: the update implementation hash moves; the runner
change is dev-only.
A queue runs the catalog in registration order among tests with the
same prerequisites, and cooling.aa-offset-calibrate followed
cooling.flow-verify. A flow-verify trial heats the tube water (a no-flow
trial by 17 C on the bench) and the warm slug circulates past the
coolant sensors for minutes afterward; the calibration's stationary gate
passed 44 s after the trial on image 20260902144848 and the edges read
the wave as disagreement.
The calibration is registered first now, with the reason beside it, and
tests/test_cooling_order.py holds the order.
Catalog consequence: the cooling.* implementation hashes move (the suite
file changed).
The gate that refuses a latch unlock while the chain may hold HV_ENABLE
up (charge_pump_alive or a pulse engine not idle) ran at the start of
phases B, U and K3 of kernel.fire-line, within a second of the previous
phase's run. A run feeds the charge-pump watchdog every 200 ms and the
one-shot holds ALIVE for 0.45 s after the last feed, so the gate read
alive=1 and refused: the first bench run of the gate (forgefirm
64f552fc; the laser_pgood gate before it was vacuous) failed phase B on
image 20260902144848.
wait_hv_off() polls the chain for up to 3 s before it refuses, logs the
release when it was not immediate and records every wait in the
evidence (hv_release_s). require_hv_off and check_hv_off use it. The
bench scripts that copy the gate (fire_test.py per phase,
gate_a_kernel_drills.py K3 after K2) get the same wait.
Bench: kernel.fire-line PASS on 20260902144848 with the chain released
after 0.41 s at each of the three phase boundaries. Host:
tests/test_kernel_suite.py covers release inside the window, a chain
held past it, and a chain already off.
Catalog consequence: the kernel.* implementation hashes move (the suite
file changed); the kernel set re-ran and passed.
The manifest lists the modules directory without the kernel's
LOCALVERSION_AUTO hash (the hash does not reproduce across a re-patch of
the same source, so the image manifest strips it). image.health still
compared the full running release against that list and failed on the
first post-flash run of image 20260902144848 with the kernel
6.12.20-fslc-fslc-g72a0b1431a9d against the manifest's 6.12.20-fslc-fslc.
kernel_ident() strips the same suffix from both sides, so a manifest with
or without the hash matches the running kernel, and a different base
release still fails. tests/test_image.py covers both forms.
Catalog consequence: only image.health's own implementation hash moves;
it is an always-run test, so no inherited result is affected.
BRINGUP describes the present: the 54-test catalog and its seven-test
always core, the tier counts, the shipped low-temperature gates, the
density floor ($35 = 10), the two local core commits, the ffboot env
write, the aa-offset route, the current bench image, and the bench
measurements the audit asks for (pooled into the next session). The
workstation shell notes and every em dash are gone.
forgetest: the takeover waits for the cloud client too (found by its
command line); the unauthenticated /boot probe names the endpoint's
parameter; the UI prose is American English. Recipes: forgetest
fetches its package directory and init script only and drops
__pycache__ at unpack; the dev image no longer re-adds forgectrl; the
release image's remove list drops the gfui-client the BSP no longer
has; the platform identity strips the kernel's local-version hash
from the modules directory name, so a re-patched kernel keeps its
fingerprints. grblhal restart is stop then start. release.sh --dev
packs the dev image. fixture.sh refuses a readable env file.
Bench tools: the live-fire drills measure the lid-IR baseline before
every run and point at the fire-watch thresholds the engine reads;
one thermistor conversion (gfbench.degc) serves every drill; the six
dated measurement records leave the tool directory; feeder.c names the
two sysfs writes its caller makes.
Host tests: forgetest 258 pass; the coverage lint reports no uncovered
path across 54 tests. Acceptance: forgectrl.auth covers the /boot
probe; update.* cover ffboot and the manifest identity; the runbook
and bench-tool changes have no catalog consequence.
A test's fingerprint covered its own text and its module's shared text
only, so a judge imported from a sibling suite module (laser.py takes
its motion judges from motion.py) could change without moving the
fingerprints of the tests that call it. The shared text of every sibling
module a module imports now rides along, transitively; unit test.
update.slots-and-signature claimed to refuse a tampered signature but
fed fwup one garbage file. It now makes a throwaway key pair on the
machine, signs a tiny archive, checks that the archive verifies with its
own key and fails against the shipped release key, and asks the update
job to apply it without confirm_unsigned: the job refuses it for its
signature before touching the slot.
The inheritance walk skipped every record that was not a PASS on the
current fingerprint, so a FAIL or ERROR recorded after a PASS on the
same image was stepped over and the older PASS inherited into the next
campaign. The newest record on the fingerprint now decides: a PASS is
inherited, a FAIL or ERROR blocks it (reason failed-since), an ABORTED
run says nothing. Unit tests for all three orders.
ffboot, the tool that rewrites the boot environment on every install and
slot switch, was packaged from scripts/ outside every fingerprint. It
now lives in the recipe's files and the recipe inherits the manifest
class; the tree manifest tool fingerprints file components the same
way, and the update tests cover the component.
forgectrl.settings-bounds fell back to ui_units=mm, which the whitelist
refuses, so the always-required test failed on a fresh machine; the
fallback is metric.
The start gate set just above the coolant opens the armed session under
the warm-up; the cloud client waits it out after the button, nothing
runs, and the release starts the print, which completes. The mid-run
hold and its bound are host-tested: the gates apply at session open, so
no setting can produce a hold mid-run on the bench. Bench-excerpt unit
tests: the pass, and the failure when a run starts under the hold.
A job longer than the ring keeps its feeder alive through the button
wait, so the cancel there has to stop the feeder before the park clears
the ring. The drill now loads such a job, requires the "longer than the
ring" line, and after the cancel reads the program total twice over the
feeder's retry period (zero both times) and cnc/streaming (zero). The
bench-excerpt unit test carries the long-job line and the new evidence.
Covers map unchanged: the drill already names gfhardware/machine.py.
The first campaign with the bench actuator wired failed
motion.button-hold-resume on the tool, not the machine: the second press
was asked while the first 200 ms pulse was still on, the fixture answered
409, the runner handed the step to an operator who was not in the room,
and the post pass could not jog a controller left in Hold.
- fixture.py: a press waits for the last pulse to end (the fixture's
pulse_ms) plus a 300 ms release, so the controller sees the edge; a
409 for a pulse in progress is waited out against button_pulsing and
retried once.
- runner.py: in an unattended run a fixture refusal ends the test at
once as ERROR naming the refusal; the operator fallback stays for
attended runs.
- baseline.py: a controller in Hold or Door gets a soft reset before the
return jog, position kept.
- tests: the fake fixture refuses a press while one is in progress and
reports button_pulsing; FakeGrbl records ^X and can land a reset in a
chosen state; five new tests.
- docs: ACCEPTANCE.md fixture rules, fixture/README.md tool's side.
No catalog consequence: tool-side change, no covers map moves.
Bench: campaign c-20260824174545-0bdc 25/25 with every action by the
fixture; the hold reset proven by a dry drill.
The acceptance page is assembled by page.py from forgetest/forgetest/ui/
(index.html, page.css, help.js, app.js) plus theme.css and the vendored
Bootstrap files, which are byte for byte the ones forgectrl's panel
carries, so the two pages look like one product and share the light and
dark themes (same localStorage key). A plain file is read in a checkout;
on the dev image the recipe installs ui/ gzipped and page.py reads the
.gz sibling, inflating once at first request: the rootfs is raw ext4, so
bytes in the package are bytes on the image. The explanatory prose
(campaign rules, the queues, the campaign actions, the prerequisites
switch, the bench intro) is a "?" popover with a link into the
documentation site; operator steps, prompts, notices and the live-laser
acknowledgment stay in the page, and confirmLive() stays a blocking
dialog. The page's own rules hold: rows, prompt buttons and tool entries
are built once and updated in place, and the popovers sit on static
markup only, so no rebuild orphans one. On a phone the Run pane goes to
the top for the duration of a run.
scripts/check-ui-vendor.py compares the shared files against forgectrl
at its pinned revision (or a local checkout with --forgectrl); it runs
in forgetest-ci.yml, so the copies cannot drift.
Tests: test_page.py (the gzipped install assembles to the same bytes as
a checkout, one self-contained response, the token placeholder once, a
missing marker refused); test_server asserts the served page's
invariants; test_responsiveness keeps its rules with needles pointed at
the new files, its ASCII rule applied to our own sources (Bootstrap's
CSS carries an em dash of its own), and its self-contained rule testing
asset tags rather than the presence of https:// (the documentation links
are meant to be there). forgectrl.panel-serves gains two needles for the
panel's theme attribute and save bar. Proof: the unit suite, and the
page in Chrome against a fake catalog (both themes, popovers, the bench
tab, a full operator run with its prompt, abort).
forgectrl pinned at 9d1f6f2 (the panel on Bootstrap, one save bar, help
popovers, themes, the gzipped page); PV unchanged. The pin moves only
forgectrl's fingerprint. The forgetest changes are the harness's own and
have no catalog consequence.
fixture.py: the bench's /data/forgetest/fixture.json (hostname, key,
optional ip, the channels wired, arm_press), a resolver for
<hostname>.local asked of the network directly (the image has no mDNS
resolver), and the client. The runner probes it before every run and
at most every 30 s otherwise; ctx.act asks it for a channel it covers
and still waits for the machine's own reading, falling back to the
operator's notice when the box fails. A test declares with hands=(...)
what it asks of a person beyond its typed actions; an operator test
with none, whose actions the fixture covers, is routed into the
unattended queue, its Ready gates pass, and a prompt it raises anyway
is a FAIL naming the undeclared step. Live tests never move; their arm
press stays a person's unless the bench opted in, in which case the
fixture presses when the button lights. Whatever the box still holds
after a run is released before the baseline's post pass and recorded.
The page shows what the fixture covers. Contract in ACCEPTANCE.md; the
wiring facts, with the interlock connector left to the bench to settle
(SAFETY.md and the sister map differ), in BRINGUP.
Catalog unchanged in its definitions; the cloud and laser suites'
shared code moved, so their implementation hashes move with it.
The coverage lint already allowed docs, CI, unit tests and licenses to
go uncovered; the same list now keeps them out of every fingerprint,
so a README edit in any component re-requires nothing. The list moves
to the manifest module as NON_BEHAVIORAL, the one place both uses read
it. And a coverage entry that selects no file of its component (a glob
without the recipe's subdirectory, a component the manifest lacks, a
glob naming docs only) fails the lint: such an entry covers nothing and
the test's fingerprint ignores the file it meant. The contract says
both. Every test whose maps reached a doc or a test file gets a new
fingerprint once.
POST /mode, /controller/start and /controller/stop answer only when the
switch is done: the old controller gone, the new one started after any
pending liveness probe, and its first job-state report in (15 s without
one). The client's 10 s timeout read a slow but honest switch as a dead
daemon and errored cloud.service-protocol on the bench; those three
paths now get 120 s. No catalog consequence: the tests and their covers
are unchanged, the client only waits longer.
A cloud client the tool starts for anything but homing comes up under
the /run/gfcloud-nohunt marker: the real client back after the
emulator, a mode the runner switches to or hands back, a controller it
restarts. The service keeps the head position it has. cloud.mode-switch
and cloud.service-protocol keep their hunts, and so does the one real
print: enter_cloud reuses a running session only when that client has
hunted the machine itself (session_hunted: never the emulator's, never
a no-hunt start), otherwise it restarts the client with the hunt, since
a print placed on a head position the service only believes can run the
gantry into a rail. The markers are one start, taken down by the client
that read them first thing; the tool's own removal stays for a start
that never happened. Catalog unchanged; the cloud tests' shared code
moved, so their implementation hashes move with it.
The service-protocol half of the cloud catalog on its own test: the cloud
client restarted as gfutilities' emulator under the /run/gfcloud-emulate
marker signs in, passes the firmware check, opens the WebSocket, answers
the connect-time hunt and the image requests with the dev image's canned
frames, and runs a print from the app through the real download path to
':completed' - nothing moves, nothing arms, and only the app has to be
driven, by a person or an agent through the prompt API. The real client
is restarted afterward and its hunt waited out. session_live now knows
the emulator's session is not the machine's, so enter_cloud restarts it
rather than reusing it; restart_client is the one restart the offline
and emulator entries share.
The dev image adds python3-gfutilities-emulator (the fixtures, packaged
on their own in meta-openglow); forgefirm-app moves to 12ad3b1 (gfcloud
--emulate). Catalog: 44 tests, 27 auto / 9 operator / 8 live; the new
test covers the gfutilities service layer and examples/, which step 4
will take off the other cloud tests. Replays over the prompt script;
contract and BRINGUP updated. A layer change (the dev image recipe):
everything re-requires on the next image.